HNHacker News
TopNewBestAskShowJobs

danielhanchen

3,617 karma · joined September 8, 2021

Unsloth github.com/unslothai/unsloth - finetune Llama 2x faster + use 70% less VRAM

1. Used to work at NVIDIA RAPIDS cuML

2. Discord: https://discord.gg/unsloth

3. Github: https://github.com/danielhanchen

4. Twitter / X: x.com/danielhanchen

5. Email: my handle @ gmail.com

6. Bug fixes for Gemma: https://news.ycombinator.com/item?id=39671146

7. Bug fixes for Gradient Accumulation: https://x.com/danielhanchen/status/1846235913443262891?lang=en

submissionscomments
danielhanchen··on Unsloth Dynamic 3.0 GGUFs
We made something called Divergence-300 @32 (and later @512) which tests actual inference across 32 tokens on a held out test (Terminal Bench, DeepSWE, Math etc)

We do plan to do larger benchmark suites though!

danielhanchen··on Unsloth Dynamic 3.0 GGUFs
Hey we did not remove the MTP for sizes above 8GiB - but yes for small GGUFs under 8 ish GiB, we removed the MTP module (IQ2_XXS and lower), because it's 500MiB to 750MiB in size, and on small 8 GiB machines, even 500MiB is needed.

As someone in the comments said we made a separate Q4_0 MTP if that's helpful so you can use that.

But I would suggest using UD-IQ3_XXS for 10.9GB for 16GB machines or Q2_K_XL

danielhanchen··on Qwen 3.8 27B
Actually we do publish non KLD benchmarks - top-1% is better - for NVFP4 for eg we did MMLU Pro, GPQA, AIME 2025: https://unsloth.ai/docs/models/qwen3.6#nvfp4-benchmarks

Sometimes they're just slow and expensive, so we we KLD as a proxy measure and it's very high correlation (95%+)

danielhanchen··on Qwen 3.8 27B
That wasn't our problem right? Gemma officially updated tool calling which we adopted
danielhanchen··on Qwen 3.8 27B
We also made NVFP4 ones if that helps! https://huggingface.co/unsloth/Qwen3.8-27B-NVFP4
danielhanchen··on Qwen3.8-2.4T
Ok no worries - if there are any future issues - feel free to message / make a HF issue - we'll fix promptly!

Also note its best to follow Gemma4's official sampling params since they evaled with it - dry multiplier sometimes works, but it actually screws up reasoning sometimes

danielhanchen··on Qwen3.8-2.4T
We will investigate Ling!
danielhanchen··on Qwen3.8-2.4T
Thanks for the support and to the community!
danielhanchen··on Qwen3.8-2.4T
Hey :)
danielhanchen··on Qwen3.8-2.4T
Thanks haha
danielhanchen··on Qwen3.8-2.4T
Hey yes - if you could describe what the issues are - we will gladly fix them!
danielhanchen··on Qwen3.8-2.4T
Hey sorry what are the problems that you're experiencing - we're more than happy to help fix them!
danielhanchen··on Unsloth Desktop: Open-Source App for Local Models
Oh thanks for sharing! We launched a Desktop app with fast diffusion, video gen, training support, inference + tool calling, web search, canvas, HTTPS secure remote access via Cloudflared, API / model swapping + more!

It works in [Windows, Linux, Mac, WSL] X [NVIDIA GPUs, AMD GPUs, CPUs]!

If there are any issues / suggestions - we'll be quick to fix / implement them!

danielhanchen··on Inkling: Our Open-Weights Model
Oh thanks for sharing! The llama.cpp PRs should generally be fine for now - I'm fixing a few small edge cases as well!
danielhanchen··on Good results fine tuning a local LLM like Qwen 3:0.6B to categorize questions
Very cool write-up and GitHub repo!
danielhanchen··on Unsloth Joins PyTorch Ecosystem
Thank you appreciate the support! It's all thanks to you guys and the community!
danielhanchen··on Making LLM Training Faster with Unsloth and NVIDIA
Update - Just got rid of the spiced up intro
danielhanchen··on Making LLM Training Faster with Unsloth and NVIDIA
Thank you!
danielhanchen··on Making LLM Training Faster with Unsloth and NVIDIA
Oh thanks :) We're also going to add MTP support soon for Qwen3.6!

95% of it is fully human done - the maths, algos, code snippets, screenshots & benchmarks are done / conducted by us and NVIDIA :)

We did use AI to fix spelling errors + made some nice plots using Chat (ours would look horrible lol)

Update - Just got rid of the spiced up intro

danielhanchen··on Qwen3.6-27B: Flagship-Level Coding in a 27B Dense Model
Sorry on the delay - so it installs https://github.com/Blaizzy/mlx-vlm and other components and sets up the commands - you don't need to use it but we thought it might be easier for folks
danielhanchen··on Qwen3.6-27B: Flagship-Level Coding in a 27B Dense Model
Sorry on the delay - oh haha that would be cool :) We did release 2bit dynamic ones, but unsure if they'll be helpful
danielhanchen··on Qwen3.6-27B: Flagship-Level Coding in a 27B Dense Model
Yes we do! Sorry on the delay
danielhanchen··on Qwen3.6-27B: Flagship-Level Coding in a 27B Dense Model
We use Duck Duck Go - sorry on the delayed response as well
danielhanchen··on Qwen3.6-27B: Flagship-Level Coding in a 27B Dense Model
Thank you and appreciate it! Sorry on the delayed reply as well
danielhanchen··on Qwen3.6-27B: Flagship-Level Coding in a 27B Dense Model
Oh yes LM Link is cool!
danielhanchen··on Qwen3.6-27B: Flagship-Level Coding in a 27B Dense Model
Hey sorry on the delay - we just added API support, so you can access a remote server - it includes optional python, tool call, bash and web search support if you enable them.

For SSH - we haven't yet done that - for now we have a SHA256 encryption approach, but it's not SSH yet. HTTPS will also sadly have to be the end user's setup process as well - we plan to make it better soon!

danielhanchen··on Qwen3.6-27B: Flagship-Level Coding in a 27B Dense Model
Hey! Sorry for not replying sooner - yes we'll keep publishing more KLD - sadly some are saying we are "optimizing" for KLD now since we posted so many haha - but the whole purpose of quantization is to match the BF16 logits as much as possible whilst reducing disk space (ie reduce KLD).

In general so this is funny and a quirk of quantization - sometimes 8bit, 4bit models do BETTER on downstream benchmarks (SWE Bench for eg), since sometimes rounding can actually somehow act as a "regularization" method (this is just my hunch).

So KLD isn't that expensive, since we leverage the trick of causal attention - since causal attention is lower triangular, we can do 1 forward pass on the enter text (say 2048 tokens), and you attain logits for the prediction for every token's position - so this is O(N^2).

However coding benchmarking require actual inference, and cannot use the causal attention trick, and it's best to run them 10 times since temperature = 1.0 is not deterministic - and take an average. We plan to maybe do something like https://marginlab.ai/trackers/claude-code/, which takes a random sample and does it over time.

danielhanchen··on Qwen3.6-27B: Flagship-Level Coding in a 27B Dense Model
Hey so sorry didn't reply sooner - yes the docker used to be I think 4-8GB ish since CUDA sadly itself is 4GB I think, and PyTorch takes the rest. So unfortunately the Unsloth Docker image has ballooned due to this. We tried reducing it as much as possible, but it's hard :( https://hub.docker.com/r/vllm/vllm-openai/tags for eg is around 11GB ish, ad we're 13.6GB ish.

We'll try our best to compress it more, but it's tough

danielhanchen··on Qwen3.6-27B: Flagship-Level Coding in a 27B Dense Model
Apologies as well didn't reply sooner - Studio supports AMD out of the box now! We worked with AMD to make it work! One thing that is still missing is pre-compiled AMD ROCM binaries, which we're trying to see if we can integrate that.

Interesting on diskpart - let me check and get back to you [EDIT] - visual studio build tools, python 3.13, git, cmake, node.js are all msi-based installers - so these are likely the culprits on using diskpart - essentially MSI installers check if there's enough disk space before installing items

danielhanchen··on Qwen3.6-27B: Flagship-Level Coding in a 27B Dense Model
Oh my apologies I didn't respond - if only HN had a notifier haha

Oh yes we added a custom folder button which can pull .gguf files for now from any folder - it supports LM Studio and Ollama ones - but afreed it's still a mess.

One of the goals is to somehow quick search for .gguf folders, and add recommended folders - we currently have folders for Ollama and LM Studio for eg

Page 1 of 20Next →