Llamafile 0.7 Brings AVX-512 Support: 10x Faster Prompt Eval Times for AMD Zen 4
phoronix.com
phoronix.com
Question: if I already have a previous version .llamafile, is there a way to repack the model blob with the new llamafile executable? That is, say I have mixtral_llamafilev6 promote to mixtral_llamafilvev7?
From the release page
- Prompt evaluation now goes much faster on CPU. For example, f16 weights on Raspberry Pi 5 are now 8x faster. These new optimizations mostly apply to F16, BF16, Q8_0, Q4_0, Q4_0, and F32 weights. Depending on the hardware and weights being used, we've observed llamafile-0.7 going anywhere between 30% to 500% faster than llama.cpp upstream.
...
- Support for AVX512 has been introduced. Owners of CPUs like Zen4 can expect to see 10x faster prompt eval times.https://github.com/Mozilla-Ocho/llamafile?tab=readme-ov-file...
One issue I hope they overcome for Windows however, is being able to run the executable when it’s more than 4GB (a Windows limitation) as all the powerful models are WAY in excess of that. I believe they will figure out a nice workaround in time for Windows users.
“External weights are particularly useful for Windows users because they enable you to work around Windows' 4GB executable file size limit.
For Windows users, here's an example for the Mistral LLM:
curl -L -o llamafile.exe https://github.com/Mozilla-Ocho/llamafile/releases/download/...
curl -L -o mistral.gguf https://huggingface.co/TheBloke/Mistral-7B-Instruct-v0.1-GGU...
./llamafile.exe -m mistral.gguf -ngl 9999
The only exceptions are the loads/stores from/to the L1 cache, which have double throughput on Intel and the floating-point fused-multiply-add units, where the most expensive Xeon SKUs can do 2 FMAs per cycle, while Zen 4 can do only 1 FMA + 1 FADD per cycle.
Zen 4 implements the BF16 instruction set, which is likely to increase the speed many times for any AI/ML workload that uses BF16. It also implements the VNNI instruction set, which will accelerate any inference that uses INT8.
Even when these dedicated instructions are not used, AVX-512 is usually much faster on Zen 4, by eliminating bottlenecks caused by instruction fetch and decoding and by using the better designed AVX-512 instructions.
If you make simple 1-to-1 transition from AVX256 to AVX512, the speedup is usually <2x, unless there is a major bottleneck in instruction fetch and decoding. Also AVX512 in Zen is still double-pumped. Regarding FMA, again, if you compare FMA in AVX256 and AVX512, latencies and throughput are the same[1].
Comparing performance between different datatypes is probably fine, but it should be stated directly. Unless "Owners of CPUs like Zen4 can expect to see 10x faster prompt eval times" means comparison with Skylake.
[1] https://uops.info/table.html?search=vfmadd132ps&cb_lat=on&cb...
And as the sibling comment mentions, AVX512 is not just about the width but also about newer instructions available at 128b and 256b widths as well under AVX512VL.
Intel has only three 256-bit execution units, which are also ganged into two 512-bit execution units, by adding an extra 256-bit unit, which stays idle in 256-bit mode.
So the throughput for most register-register 512-bit operations is the same for Intel and AMD, except that AMD has a single FP64 multiplier vs. two FP64 multipliers on Intel and that the path to the L1 cache has double width on Intel (while AMD does one 512-bit load per cycle + one 512-bit store every other cycle, Intel can do two 512-bit loads from L1 + one store per cycle).