Repository navigation
add BLAKE3 to hashlib #83479
Description
Activity
From 3/4 of the team that brought you BLAKE2, now comes... BLAKE3!
https://github.com/BLAKE3-team/BLAKE3
BLAKE3 is a brand new hashing function. It's fast, it's paralellizeable, and unlike BLAKE2 there's only one variant.
I've experimented with it a little. On my laptop (2018 Intel i7 64-bit), the portable implementation is kind of middle-of-the-pack, but with AVX2 enabled it's second only to the "Haswell" build of KangarooTwelve. On a 32-bit ARMv7 machine the results are more impressive--the portable implementation is neck-and-neck with MD4, and with NEON enabled it's definitely the fastest hash function I tested. These tests are all single-threaded and eliminate I/O overhead.
The above Github repo has a reference implementation in C which includes Intel and ARM SIMD drivers. Unsurprisingly, the interface looks roughly the same as the BLAKE2 interface(s), so if you took the existing BLAKE2 module and s/blake2b/blake3/ you'd be nearly done. Not quite as close as blake2b and blake2s though ;-)
Reacted by bvd0- added3.9 (EOL)end of lifeend of lifestdlibStandard Library Python modules in the Lib/ directoryStandard Library Python modules in the Lib/ directorytype-featureA feature request or enhancementA feature request or enhancement
on Jan 11, 2020 I've been playing with the new algorithm, too. Pretty impressive!
Let's give the reference implementation a while to stabilize. The code has comments like: "This is only for benchmarking. The guy who wrote this file hasn't touched C since college. Please don't use this code in production."
For what it's worth, I spent some time producing clean benchmarks. All these were run on the same laptop, and all pre-load the same file (406668786 bytes) and run one update() on the whole thing to minimize overhead. K12 and BLAKE3 are using a hand-written C driver, and compiled with both gcc and clang; all the rest of the algorithms are from hashlib.new, python3 configured with --enable-optimizations and compiled with gcc. K12 and BLAKE3 support several SIMD extensions; this laptop only has AVX2 (no AVX512). All these numbers are the best of 3. All tests were run in a single thread.
-----------------+----------+----------+----+-----------------------
hash algorithm|elapsed s |mb/sec |size|hash
-----------------+----------+----------+----+-----------------------
K12-Haswell 0.176949 2298224495 64 24693954fa0dfb059f99...
K12-Haswell-clang 0.181968 2234841926 64 24693954fa0dfb059f99...
BLAKE3-AVX2-clang 0.250482 1623547723 64 30149a073eab69f76583...
BLAKE3-AVX2 0.256845 1583326242 64 30149a073eab69f76583...
md4 0.37684668 1079135924 32 d8a66422a4f0ae430317...
sha1 0.46739069 870083193 40 a7488d7045591450ded9...
K12-clang 0.498058 816509323 64 24693954fa0dfb059f99...
BLAKE3 0.561470 724292378 64 30149a073eab69f76583...
K12 0.569490 714093306 64 24693954fa0dfb059f99...
BLAKE3-clang 0.573743 708800001 64 30149a073eab69f76583...
blake2b 0.58276098 697831191 128 809ca44337af39792f8f...
md5 0.59936016 678504863 32 306d7de4d1622384b976...
sha384 0.64208886 633352818 96 b107ce5d086e9757efa7...
sha512_224 0.66094102 615287556 56 90931762b9e553bd07f3...
sha512_256 0.66465768 611846969 64 27b03aacdfbde1c2628e...
sha512 0.6776549 600111921 128 f0af29e2019a6094365b...
blake2s 0.86828375 468359318 64 02bee0661cd88aa2be15...
sha256 0.97720436 416155312 64 48b5243cfcd90d84cd3f...
sha224 1.0255457 396538907 56 10fb56b87724d59761c6...
shake_128 1.0895037 373260576 32 2ec12727ac9d59c2e842...
md5-sha1 1.1171806 364013470 72 306d7de4d1622384b976...
sha3_224 1.2059123 337229156 56 93eaf083ca3a9b348e14...
shake_256 1.3039152 311882857 64 b92538fd701791db8c1b...
sha3_256 1.3417314 303092540 64 69354bf585f21c567f1e...
ripemd160 1.4846368 273918025 40 30f2fe48fec404990264...
sha3_384 1.7710776 229616579 96 61af0469534633003d3b...
sm3 1.8384831 221198006 64 1075d29c75b06cb0af3e...
sha3_512 2.4839673 163717444 128 c7c250e79844d8dc856e...If I can't have BLAKE3, I'm definitely switching to BLAKE2 ;-)
I'm in the middle of adding some Rust bindings to the C implementation in github.com/BLAKE3-team/BLAKE3, so that
cargo testandcargo benchcan cover both. Once that's done, I'll follow up with benchmark numbers from my laptop (Kaby Lake i5-8250U, also AVX2 with no AVX-512). For benchmark numbers with AVX-512 support, see the Performance section of the BLAKE3 paper (https://github.com/BLAKE3-team/BLAKE3-specs/blob/master/blake3.pdf). Larry, what processor did you run your benchmarks on?Also, is there anything currently in CPython that does dispatch based on runtime CPU feature detection? Is this something that BLAKE3 should do for itself, or is there existing machinery that we'd want to integrate with?
According to my order details it is a "8th Generation Intel Core i7-8650U".
Ok, I've added Rust bindings to the BLAKE3 C implementation, so that I can benchmark it in a vaguely consistent way. My laptop is an i5-8250U, which should be very similar to yours. (Both are "Kaby Lake Refresh".) My end result do look similar to yours with TurboBoost on, but pretty different with TurboBoost off:
with TurboBoost on
------------------
K12 GCC | 2159 MB/s
BLAKE3 Rust | 1787 MB/s
BLAKE3 C Clang | 1588 MB/s
BLAKE3 C GCC | 1453 MB/swith TurboBoost off
-------------------
BLAKE3 Rust | 1288 MB/s
K12 GCC | 1060 MB/s
BLAKE3 C Clang | 1094 MB/s
BLAKE3 C GCC | 943 MB/sThe difference seems to be that with TurboBoost on, the BLAKE3 benchmarks have my CPU sitting around 2.4 GHz, while for the K12 benchmarks it's more like 2.9 GHz. With TurboBoost off, both benchmarks run at 1.6 GHz, and BLAKE3 does better. I'm not sure what causes that frequency difference. Perhaps some high-power instruction that the BLAKE3 implementation is emitting?
To reproduce these numbers you can clone these two repos (the latter is where I happen to have a K12 benchmark):
https://github.com/BLAKE3-team/BLAKE3
https://github.com/oconnor663/blake2_simdThen in both cases checkout the "bench_406668786" branch, where I've put some benchmarks with the same input length you used.
For Rust BLAKE3, at the root of the BLAKE3 repo, run: cargo +nightly bench 406668786
For C BLAKE3, the command is the same, but run it in the "./c/blake3_c_rust_bindings" directory. The build defaults to GCC, and you can "export CC=clang" to switch it.
For my K12 benchmark, at the root of the blake2_simd repo, run: cargo +nightly bench --features=kangarootwelve 406668786
I plan to bring the C code up to speed with the Rust code this week. As part of that, I'll probably remove comments like the one above :) Otherwise, is there anything else we can do on our end to help with this?
53 remaining items
On 23.03.2022 17:53, Larry Hastings wrote:
Ok, I give up.
Sorry to spoil the fun, but there's no need to throw
everything in the bin ;-)A lean and fast blake3 C package would still be a great thing
to have on PyPI, e.g. provide support for platforms, which
Jack's blake3 Rust package doesn't cover, e.g.Raspis:
https://www.piwheels.org/project/blake3/Android (e.g. via termux):
https://wiki.termux.com/wiki/Main_Page
https://wiki.termux.com/wiki/Pythonetc.
The Rust version is already quite "lean". And it can be much faster than the C version, because it supports internal multithreading. Even without multithreading I bet it's at least a hair faster.
Also, Jack has independently written a Python package based around the C version:
https://github.com/oconnor663/blake3-py/tree/master/c_impl
so my making one would be redundant.
I have no interest in building standalone BLAKE3 PyPI packages for Raspberry Pi or Android. My goal was for BLAKE3 to be one of the "included batteries" in Python--which would have meant it would, eventually, be available on the Raspberry Pi and Android builds that way.
With "lean" I meant: doesn't use much code and is easy to compile
and install.I built a wheel from Jack's experimental package and it comes out to
just under 100kB on Linux x64, compared to around the 1.1MB the
Rust wheel needs:Archive: blake3_experimental_c-0.0.1-cp310-cp310-linux_x86_64.whl
Length Date Time Name
--------- ---------- ----- ----
348528 2022-03-23 18:38 blake3.cpython-310-x86_64-linux-gnu.so
3183 2022-03-23 18:38 blake3_experimental_c-0.0.1.dist-info/METADATA
105 2022-03-23 18:38 blake3_experimental_c-0.0.1.dist-info/WHEEL
7 2022-03-23 18:38 blake3_experimental_c-0.0.1.dist-info/top_level.txt
451 2022-03-23 18:38 blake3_experimental_c-0.0.1.dist-info/RECORD
--------- -------
352274 5 filesArchive: blake3-0.3.1-cp310-cp310-manylinux_2_5_x86_64.manylinux1_x86_64.whl
Length Date Time Name
--------- ---------- ----- ----
3800 2022-01-13 01:26 blake3-0.3.1.dist-info/METADATA
133 2022-01-13 01:26 blake3-0.3.1.dist-info/WHEEL
48 2022-01-13 01:26 blake3/init.py
4195392 2022-01-13 01:26 blake3/blake3.cpython-310-x86_64-linux-gnu.so
382 2022-01-13 01:26 blake3-0.3.1.dist-info/RECORD
--------- -------
4199755 5 filesI don't know why there is such a significant difference in size. Perhaps
the Rust version includes multiple variants for different CPU
optimizations ?!I can't answer why the Rust one is so much larger--that's a question for Jack. But the blake3-py you built might (should?) have support for SIMD extensions. See the setup.py for how that works; it appears to at least try to use the SIMD extensions on x86 POSIX (32- and 64-bit), x86_64 Windows, and 64-bit ARM POSIX.
If you were really curious, you could run some quick benchmarks, then hack your local setup.py to not attempt adding support for those (see "portable code only" in setup.py) and do a build, and run your benchmarks again. If BLAKE3 got a lot slower, yup, you (initially) built it with SIMD extension support.
To anyone else who comes along with motivation:
I'm fine with blake3 being in hashlib, but I don't want us to guarantee it by carrying the implementation of the algorithm in the CPython codebase itself unless it gains wide industry standard-like adoption status.
We should feel free to link to both the Rust blake3 and C blake3-py packages from the hashlib docs regardless.
Performance wise... The SHA series have hardware acceleration on modern CPUs and SoCs. External libraries such as OpenSSL are in a position to provide implementations that make use of that. Same with the Linux Kernel CryptoAPI (https://bugs.python.org/issue47102).
Hardware accelerated SHAs are likely faster than blake3 single core. And certainly more efficient in terms of watt-secs/byte.
Here's a wheel which only includes the portable code (I disabled
all the special cases as you suggested).Archive: dist/blake3_experimental_c-0.0.1-cp310-cp310-linux_x86_64.whl
Length Date Time Name
--------- ---------- ----- ----
297680 2022-03-23 19:26 blake3.cpython-310-x86_64-linux-gnu.so
3183 2022-03-23 19:26 blake3_experimental_c-0.0.1.dist-info/METADATA
105 2022-03-23 19:26 blake3_experimental_c-0.0.1.dist-info/WHEEL
7 2022-03-23 19:26 blake3_experimental_c-0.0.1.dist-info/top_level.txt
451 2022-03-23 19:26 blake3_experimental_c-0.0.1.dist-info/RECORD
--------- -------
301426 5 filesI didn't run any benchmarks, but it's clear that the SIMD code was
used in my initial build and this adds some 50kB to the .so file.
This is on a older Linux x64 box with Intel i7-4770k CPU.Could be that the Rust version adds several such SIMD variants and
then branches based on the platform running the code.In any case, the C extension is indeed very easy to build and
install with a standard compiler setup.Rust based anything comes with a baseline level of Rust code overhead. https://stackoverflow.com/questions/29008127/why-are-rust-executables-so-huge
That seems expected.
Performance wise... The SHA series have hardware acceleration on
modern CPUs and SoCs. External libraries such as OpenSSL are in a
position to provide implementations that make use of that. Same with
the Linux Kernel CryptoAPI (https://bugs.python.org/issue47102).Hardware accelerated SHAs are likely faster than blake3 single core.
And certainly more efficient in terms of watt-secs/byte.I don't know if OpenSSL currently uses the Intel SHA1 extensions.
A quick google suggests they added support in 2017. And:- I'm using a recent CPU that AFAICT supports those extensions.
(AMD 5950X) - My Python build with BLAKE3 support is using the OpenSSL implementation
of SHA1 (_hashlib.openssl_sha1), which I believe is using the OpenSSL
provided by the OS. (I haven't built my own OpenSSL or anything.) - I'm using a recent operating system release (Pop!_OS 21.10), which
currently has OpenSSL version 1.1.1l-1ubuntu1.1 installed. - My Python build with BLAKE3 doesn't support multithreaded hashing.
- In that Python build, BLAKE3 is roughly twice as fast as SHA1 for
non-trivial workloads.
- I'm using a recent CPU that AFAICT supports those extensions.
Hardware accelerated SHAs are likely faster than blake3 single core.
Surprisingly, they're not. Here's a quick measurement on my recent ThinkPad laptop (64 KiB of input, single-threaded, TurboBoost left on), which supports both AVX-512 and the SHA extensions:
OpenSSL SHA-256: 1816 MB/s
OpenSSL SHA-1: 2103 MB/s
BLAKE3 SSE2: 2109 MB/s
BLAKE3 SSE4.1: 2474 MB/s
BLAKE3 AVX2: 4898 MB/s
BLAKE3 AVX-512: 8754 MB/sThe main reason SHA-1 and SHA-256 don't do better is that they're fundamentally serial algorithms. Hardware acceleration can speed up a single instance of their compression functions, but there's just no way for it to run more than one instance per message at a time. In contrast, AES-CTR can easily parallelize its blocks, and hardware accelerated AES does beat BLAKE3.
And certainly more efficient in terms of watt-secs/byte.
I don't have any experience measuring power myself, so take this with a grain of salt: I think the difference in throughput shown above is large enough that, even accounting for the famously high power draw of AVX-512, BLAKE3 comes out ahead in terms of energy/byte. Probably not on ARM though.
sha1 should be considered broken anyway and sha256 does not perform well on 64bit systems. Truncated sha512 (sha512-256) typically performs 40% faster than sha256 on X86_64. It should get you close to the performance of BLAKE3 SSE4.1 on your system.
Truncated sha512 (sha512-256) typically performs 40% faster than sha256 on X86_64.
Without hardware acceleration, yes. But because SHA-NI includes only SHA-1 and SHA-256, and not SHA-512, it's no longer a level playing field. OpenSSL's SHA-512 and SHA-512/256 both get about 797 MB/s on my machine.
You missed the key "And certainly more efficient in terms of watt-secs/byte" part.
I did reply to that point above with some baseless speculation, but now I can back up my baseless speculation with unscientific data :)
https://gist.github.com/oconnor663/aed7016c9dbe5507510fc50faceaaa07
According to whatever
powerstat -Rmeasures on my laptop, running hardware-accelerated SHA-256 in a loop for a minute or so takes 26.86 Watts on average. Doing the same with AVX-512 BLAKE3 takes 29.53 Watts, 10% more. Factoring in the 4.69x difference in throughput reported by those loops, the overall energy/byte for BLAKE3 is 4.27x lower than SHA-256. This is my first time running a power benchmark, so if this sounds implausible hopefully someone can catch my mistakes.Reacted by Joakim Soderlund and Lucas ParzianelloReacted by Techcable
Note: these values reflect the state of the issue at the time it was migrated and might not reflect the current state.
Show more details
GitHub fields:
bugs.python.org fields: