We do plan to do larger benchmark suites though!
3,617 karma · joined September 8, 2021
1. Used to work at NVIDIA RAPIDS cuML
2. Discord: https://discord.gg/unsloth
3. Github: https://github.com/danielhanchen
4. Twitter / X: x.com/danielhanchen
5. Email: my handle @ gmail.com
6. Bug fixes for Gemma: https://news.ycombinator.com/item?id=39671146
7. Bug fixes for Gradient Accumulation: https://x.com/danielhanchen/status/1846235913443262891?lang=en
We do plan to do larger benchmark suites though!
As someone in the comments said we made a separate Q4_0 MTP if that's helpful so you can use that.
But I would suggest using UD-IQ3_XXS for 10.9GB for 16GB machines or Q2_K_XL
Sometimes they're just slow and expensive, so we we KLD as a proxy measure and it's very high correlation (95%+)
Also note its best to follow Gemma4's official sampling params since they evaled with it - dry multiplier sometimes works, but it actually screws up reasoning sometimes
It works in [Windows, Linux, Mac, WSL] X [NVIDIA GPUs, AMD GPUs, CPUs]!
If there are any issues / suggestions - we'll be quick to fix / implement them!
95% of it is fully human done - the maths, algos, code snippets, screenshots & benchmarks are done / conducted by us and NVIDIA :)
We did use AI to fix spelling errors + made some nice plots using Chat (ours would look horrible lol)
Update - Just got rid of the spiced up intro
For SSH - we haven't yet done that - for now we have a SHA256 encryption approach, but it's not SSH yet. HTTPS will also sadly have to be the end user's setup process as well - we plan to make it better soon!
In general so this is funny and a quirk of quantization - sometimes 8bit, 4bit models do BETTER on downstream benchmarks (SWE Bench for eg), since sometimes rounding can actually somehow act as a "regularization" method (this is just my hunch).
So KLD isn't that expensive, since we leverage the trick of causal attention - since causal attention is lower triangular, we can do 1 forward pass on the enter text (say 2048 tokens), and you attain logits for the prediction for every token's position - so this is O(N^2).
However coding benchmarking require actual inference, and cannot use the causal attention trick, and it's best to run them 10 times since temperature = 1.0 is not deterministic - and take an average. We plan to maybe do something like https://marginlab.ai/trackers/claude-code/, which takes a random sample and does it over time.
We'll try our best to compress it more, but it's tough
Interesting on diskpart - let me check and get back to you [EDIT] - visual studio build tools, python 3.13, git, cmake, node.js are all msi-based installers - so these are likely the culprits on using diskpart - essentially MSI installers check if there's enough disk space before installing items
Oh yes we added a custom folder button which can pull .gguf files for now from any folder - it supports LM Studio and Ollama ones - but afreed it's still a mess.
One of the goals is to somehow quick search for .gguf folders, and add recommended folders - we currently have folders for Ollama and LM Studio for eg