Qwen2.5-Coder-32B is an LLM that can code well that runs on my Mac
simonwillison.net
simonwillison.net
Best thing about it is that it's an OSS model that can be hosted by anyone, resulting in an open competitive market bringing hosting costs down, currently sitting at $0.18/$0.18 M tok/s [1] making it 50x cheaper than Sonnet 3.5 and ~17x cheaper than Haiku 3.5.
https://llm.extractum.io/model/Qwen%2FQwen2.5-Coder-32B-Inst...
I'm getting a very usable ~18 tok/s running it on 2x NVIDIA A4000 (32GB VRAM).
Both GPUs cost less than USD $1,400 on eBay.
qwen2.5-coder:32b is 19GB on ollama [1]
Prompt: 49 tokens, 95.691 tokens-per-sec
Generation: 723 tokens, 10.016 tokens-per-secSo far I am using zed editor and can't switch to somethints else untill zed editor gets the update which support FIM directive via custom LLMs
The config for chat, you can do:
"language_models": {
"openai": {
"version": "1",
"api_url": "http://localhost:8080",
"low_speed_timeout_in_seconds": 120,
"available_models": [
{
"provider": "openai",
"name": "Qwen2.5-Coder-7B-Instruct-Q8_0.gguf",
"display_name": "llama.cpp",
"max_tokens": 131072
}
]
}
},
While for code completion, you have two choices atm: supermaven and copilot: https://zed.dev/docs/completions.I'd personally rank the Qwen2.5 32B model only a little behind GPT4o at worst, and preferable to gemini 1.5 pro 002 (at code only, Gemini is a model that's surprisingly bad at code considering its top class STEM reasoning).
This makes Qwen2.5-coder-32B astounding all considered. It's really quite capable and is finally an accessible model that's useful for real work. I tested it on some linear algebra, discussed pros and cons of a belief propagation based approach to SAT solving, had it implement a fast simple approximate nearest neighbor based on the near orthogonality of random vectors in high dimensions (in OCaml, not perfect with but close enough to useful/easily correctable), simulate execution of a very simple recursive program (also Ocaml) and write a basic post processing shader for Unity. It did really well on each of those tasks.
In most of my test o1-preview performed way better than Claude and Qwen was not that bad either.
I'd be shocked if this model held up in the comprehensive private evals.
As long as Jensen Huang keeps shitting out nvidia cards, progress is just a function of cash to burn on paying humans to dump their knowledge into train data... and hoping this silly transformer architecture keeps holding up
There is an interesing discussion about it here:https://news.ycombinator.com/item?id=42104964.
I don't know where this myth had originated, and perhaps it was true at least at some point, but you just have to consider that all the recent major advances in datasets had to do with _unsupervised_ reward models, synthetic, generational datasets, and new advanced alignment methods. The big labs _are_ hiring serious PhD level researchers, and most of these are physicists, Bayesians of many kind and breed, not "domain experts." However, perception matters a lot these days; some labs, I won't point, but OpenAI is probably the biggest offender, simply cannot control themselves. The fact of the matter is they LOVE including the public evals in their finetuning, as it makes them appear stronger in the "benchmarks."
Hardly any PhD has the patience or skill for that matter to code robust solutions from scratch. Just look at PhD code in the wild.
I would exactly want to see that, or "make a little interpreter for a basic subset of C, or Scheme or <X>".
More people should host https://github.com/lm-sys/FastChat
From the announcement:
> we selected the latest 4 months of LiveCodeBench (2024.07 - 2024.11) questions as the evaluation, which are the latest published questions that could not have leaked into the training set, reflecting the model’s OOD capabilities.
I have not performed comprehensive evals of my own here - clearly - but I did just enough to confirm that the buzz I was seeing around this model appeared to hold up. That's enough for me to want to write about it.
Food for thought.
Here's Qwen 2.5 Coder 32B for "Generate an SVG of a pelican riding a bicycle" https://gist.github.com/simonw/56217af454695a90be2c8e09c7031...
Hypothetically let's say the benchmark contains "test divisibility of this integer by n" for all n of the form 3x+1. An extremely overfit llm won't be able to code divisibility for all n not of the form 3x+1, and your benchmark will never tell.
But in modern usage it is often rephrased to: "When a measure becomes a target, it ceases to be a good measure"
https://en.m.wikipedia.org/wiki/Training,_validation,_and_te...
On a desktop too, I wonder if it's worth the additional stress and heat on my GPU as opposed to one somewhere in a datacenter which will cost me a few dollars per month, or a few cents per hour if I spin up the infra myself on demand.
Super useful for confidential / secret work though
This could be a way to satisfy both sides, although it only solves the issue of sending internal data to companies like OpenAI, it doesn't solve the "we might accidentally end up with somebody else's copyrighted code in our code base" issue.
Prompt: 49 tokens, 95.691 tokens-per-sec
Generation: 723 tokens, 10.016 tokens-per-sec
Peak memory: 32.685 GBIt fails spectacularly - the code wont even compile and there are many things missing to get a working solution.
What is the end goal here? Reduction of the workforce by 90% while the remaining 10% click buttons to produce buggy, insecure and bloated code?