HNHacker News
TopNewBestAskShowJobs

eldar_ciki

19 karma · joined December 15, 2019

submissionscomments
eldar_ciki··on We Ran Over Half a Million Evaluations on Quantized LLMs
The ML community has recently questioned whether quantized LLMs can genuinely compete with their full-precision counterparts. To address this, we conducted over half a million evaluations on quantized Llama-3.1-{8B, 70B, 405B}-Instruct models across FP8, INT8, and INT4 quantization schemes. We looked at various benchmarks, from open-ended challenges like Arena-Hard, to rigorous academic benchmarks such as MMLU, MMLU-Pro, Big Bench Hard, ARC-Challenge, IFEval, GPQA (and others from OpenLLM Leaderboard v1 and v2), and coding tests like HumanEval and HumanEval+.

-> Long story short: when models are carefully quantized, everything looks good. The lowest accuracy recovery we found was 96% relative to unquantized baseline, and this happened only for 8B model at weight-only INT4 quantization mostly because the unquantized baseline model has close to random prediction accuracy on a couple of Leaderboard v2 benchmarks. As one could imagine, "recovering" random accuracy is a bit noisy.

eldar_ciki··on Show HN: Panza: A personal email assistant, trained and running on-device
you might want to open an issue in our git repo, and if there is interest we would be more than happy to add a solution for Apple silicon