Qwen3.8-2.4T
huggingface.co
huggingface.co
License pretty similar to k3 with some caveats. Free to use for internal or <50M$ revenue / year. Limitations above that threshold for serving the model or services targeting coding / productivity agents.
Benchmarks are looking good, trading blows w/ opus4.8 and sol, generally 10-20p under fable. But that's neither here nor there w/ qwen, their benchmark to real world usage correlation has been iffy in the past.
The local model 3.8-27B announced for Friday, same time so ~48 hours from now. That'll be a bit more exciting for a lot more people, since 3.6 was quite good for local inference, and their 3.7-max -> 3.8-max shows a lot of improvement.
It's true they make architecture-specific changes like keeping certain layers at F16 but it's also more than that.
Also note its best to follow Gemma4's official sampling params since they evaled with it - dry multiplier sometimes works, but it actually screws up reasoning sometimes
No issues with llama.cpp.
They are the most reliable in my experience, but if you have alternatives you trust I'd love to know
It did take a little while for Unsloth to update the Laguna S 2.1 quants to fix the yarn_attn_factor, and so it was a bit frustrating getting that quantization running right, but almost always, I pick the unsloth quantization if there is one. (Still waiting/hoping for a Ling 3.0 Flash.)
I have had the same experience with gemma 4 on same tasks being refused. But this is when working with cyber offensive tasks and the like. It excels in coding and is very fast on consumer hardware. So I would say use the right tool for the right task.
What uncensored models can you recommend?
The question about LLMs is never whether they can be run, because that has a trivial answer, they can always be run. The right question is what speeds are achievable for representative hardware configurations.
At launch, it is difficult to estimate the speed. That should be known after someone reports experimental results. Moreover, for many LLMs the speed improved sometimes later after their release, after tweaks in inference backends, like llama.cpp or vLLM.
speed is not problem when You run agents and forget for 2-3 days
Maybe I’m misreading this or some other post, I thought QWEN was stepping away from releasing these models for local consumption
They literally announced their motivations and world few a few weeks ago at the Shanghai AI conference. They want to ally with the global south. They see AI like the industrial revolution: the global south was left behind for a long time and, as a result has been exploited and has struggled to develop for a long time. They see open AI as a way to level the playing field to prevent such "new historical injustices" (in the sense of the Century of Humiliation and the Opium Wars). Concrete policies to back this rhetoric include technology transfer and training programs for the global south. They frame this latter not as philantropy but as generosity, in the sense that it generates goodwill and what goes around comes around. They believe that helping the global south and cultivating relationships will eventually help China.
Think about it. Your local businesses are not charities either. That doesn't make them bad, nor does it mean you derive no benefit. It still benefits you to cultivate good relationships with them.
No matter what you believe is their "true" intentions, offering 5000 training and tech transfer positions to the global south is a very concrete and unambiguous move. As are forgiving African loans and unilaterally offering zero trade tariffs.
I'm not saying that China's investment in the South hasn't had positive effects. But "generosity" is rarely a relevant lens when analyzing international relations.
Well, yes? Why does it have to be either-or?
> But "generosity" is rarely a relevant lens when analyzing international relations.
Automatically assuming nefarious intentions behind all moves is also rarely a relevant lens.
And as I said, and I'm not sure why you keep ignoring it, but I define "generosity" in the sense of mutual benefit. Being nice to your neighbors and helping them, benefits you due to generated goodwill. I'm explicitly not defining generosity in the sense of selfless philanthropy where you get nothing back. The idea that doing good things for others eventually results in good things coming your way, and thus that one should do good things for others even it's selfishly motivated (and also that there's nothing wrong with this), is not a crazy idea.
In a lot of cultures (Chinese included), gifts are not simply gifts. There is the social expectation that the gift is reciprocated. Western cynicists may call this "manipulation" or "influence". The Chinese see this as the start of a relationship of a cycle of mutual gift giving.
Nothing happens on a geopolitical scale, from the US, China or anyone else, simply because of generosity.
China also puts pressure on rich chinese flaunting their riches.
They have a common prosperity initiative.
The trouble here is how more infrastructure helps OpenAI and Anthropic continue billing at 10/100x Chinese model rates.
Either their models have to be better (to justify the higher prices and margin) or their inference has to be lower cost (which isn't going to happen until they move away from Nvidia).
The best outcome for us is the one where they all keep competing and undermining each other until the end of time while providing us all with better models and cheaper hardware to run them with. The US corporations in particular should never be allowed to achieve their "you'll buy intelligence from us on a meter" rent seeking dream.
It's not just according to you.
Without open weights, what happens if you get blacklisted from Anthropic and OpenAI? If AI becomes a standard tool for programming like a compiler, you've effectively been Blackballed from the field of programming. Full Stop. This is "Right to Read" coming home: https://www.gnu.org/philosophy/right-to-read.en.html
In addition, without open source competitors to your core tools, we KNOW what happens. Cadence and Synopsys and a megabuck per engineer per year ... that's what happens.
The real policy mechanisms around open AI models are also incentives. Various cities have programs to pay companies for releasing open models. They subsidize compute through vouchers. They reward universities and students for open source collaboration.
This isn't some black box. The policies are written down, anybody can read their AI+ policy papers.
Alibaba went back to releasing open models way before the Xi speech from a few weeks ago. The cause is pressure from researchers, who believe in openness, as well as the competition who keeps releasing open models. This is Chinese "involution" at work. And the subsidies also help, of course.
Sadly they seem to not be releasing a sparse 35b A3b or anything inbetween "too large to host for mortals" and "fits into a consumer rtx". Probably not to eat away their profits on their API serving. 120b - 300b is a dead space right now, very few good releases in that size range. (I know there are, but the big labs aren't releasing stuff here)
Qwen3.8 comes with official support for reasoning_effort, which can be used to adjust reasoning depth and control cost:
- xhigh (default): for complex tasks demanding thorough analysis
- medium: balancing accuracy and speed
- low: efficient reasoning optimizing for speed and cost
In addition, preserve_thinking is enabled by default for all workloads for the best out-of-the-box experience.
Asking because in my case (OCR of scanned historical "National Geographic" magazines) the LLM trying to merge text split into separate columns was running in circles from time to time and needed a lot of prompt tuning when using Qwen 3.0/3.5/3.6 (still needs from time to time).It uses twice as much tokens to achieve the same but the results are significantly better and because it's so much cheaper it's the most economical choice too.
[1] https://blog.bosun.ai/software-maintenance-with-open-weight-...
[1] https://www.reddit.com/r/LocalLLaMA/comments/1vmi0fg/deepsee...
In terms of what you get for what you pay for, it's incredible - probably by far the best.
But unless I'm reading things wrong, it does not appear to be top-of-the-line.
At this price range $0.87per 1M they will get a lot of usage of people trying it out. Given the benchmark numbers, for many people and many use cases this will become their primary driver. There are people and use cases where Fable, Sol will work better but those are likely not the target of DeepSeek anyway.
In terms of performance and price pareto curve I don't think any model can beat this today (though openAI is doing some exciting recent work in efficiency) - which is a remarkable feat for the DeepSeek team.
Either way, what a time for consumers of these models :)
DeepSeek and Kimi K3 will happily do security work, and they do it pretty well.
Curious if you’ve find yourself enjoying DS or Kimi more than Opus 4?
Kimi K3 is smarter, though. At least smarter than DeepSeek V4 Flash 0731. I haven't tried the new Pro version, but will this weekend when I'm working on my personal projects But, K3 has been what I've been using for the actual coding of the harness and such (after Claude models, and then OpenAI models, began refusing to do that work). K3 is very expensive, though. Much more expensive than pretty much everything except Anthropic, and their subscription plans are stingy.
I also like Reasonix quite a bit, as an agent harness, though Kimi Code is also very good. I guess I'll try out the new DeepSeek official harness, as well.
In a benchmark, DeepSeek v4 Pro 0813 found 87.5% of selected real world software vulnerabilities publicly reported and with CVEs assigned, which is above runner ups Opus 5 and Qwen 3.8 which both found only 81.3%. However there is a downside to this--DeepSeek v4 Pro 0813 is less accurate with a 35% false positive rate versus GPT-5.6-Sol's 15% false positive rate. For vulnerability analysis though, it's probably worth finding that one extra vulnerability no other model has found even if requires significantly more triage to remove false positives, or additional cost to run every vulnerability detection through other models to verify.
[1] https://nitter.net/pilvar222/status/2087691659953815783#m
The former has a button to dismiss the dialog. Maybe you clicked on it by accident, or maybe it does not work right.
they have a somewhat selfhostable model(flash), but are mostly known for having super cheap api access
no zdr ofcourse.
Since the models are open, there's other infra providers who have different policies.
The 1bit quant model is at an astonishing 397GB with 95B active per MOE. This literally puts Opus 4.5 performance level into a machine a normal person could buy, and still gets usable tokens/second.
The full lossless model BF16 is clocking at 4.9TB. The model card claims the model to be between Opus 4.8 and Fable 5. Again that's astonishing as getting a machine with 7TB RAM (with context + KV cache) is still within the realm of medium size companies.
Bad things: The open source version has its vision capability removed, and the context capped at 250k . I expect someone to bolt a Kimi 2.6 vision tower to it to restore the vision capability (at less performance of course). For context, I played around with extending the context to 600k for Qwen 3.5 397b, and the context remained stable up to around 480k. It'd be interesting to see if the same can be done to Q3.8 .
Also no out of the box DSpark/DFlash support. MTP is present so we should at least get some boost in TP speed.
Honestly this model people at home can tinker with, if you have a big enough Mac. Maybe 4 Strix Halo/DGX Spark, and then at 1 bit quant? Nah.
Use the right sized model, for your hardware. You'll get better results.
Extremely large models don't suffer as much from quantization due to its weight topology also contains encoded information, so the loss of info from any one weight is somewhat mitigated.
I wouldn't pick up 400gb of hardware to run in that mode. I might try it for fun, but even then you are looking at handling a 95GB active parameter set.
This is NOT a model for most home labs. I'm sure some can and will use it. But most, should steer clear.
95GB active is not _too_ bad, would require some creativity and $$, but I bet I could do that at home for less than a cheap car.
Suppose I have 100GB of unified memory, how should I know which model suits it best? I understand how a 2.4T model wouldn't fit, but I don't understand the impact of quantization and whether I should use a 200G model quantised to fit say 90GB of memory, or a non-quantised 90G model.
The old rule of thumb was that a lower quant of a larger model > higher quant of a smaller model. That being said, for some things going lower than fp8 will see a lot of degradation in generation quality. Except if the model comes with QAT 4bit quants. Then there's also nvfp4 w/ calibration data, which also can improve things. So it's really not easy to tell "at a glance" you'd have to test them yourself on your hardware.
Anything below that, and especially 1.58b - is typically complete garbage, and you're much better off running a model 100x smaller at regular precision (compared to one 7x smaller quantized into complete garbage).
If the model was designed specifically to quantize down to 1.58b, then it's different.
AFAIK, there's no large models designed for this yet.
> AFAIK, there's no large models designed for this yet.
Isn't BitNet b1.58 2B4T what you are looking for? (haven't tried it myself though)
No 100B+ param (certainly no 2T+ param) models have been trained natively to quantize down to 1.58b.
What's weird about local LLM models is that closed door improvements in training/RLHF dataset have been so significant that it's rare for larger but older models to make sense - everyone seem to always hard switch to the newest one and report step changes in capabilities(or maybe people running Kimi K2 since release just don't talk about it on the public Internet, giving me that impression).
at this kind of quantization is it useful though?
That is unfortunate, that the open weight model doesn't have vision support or the 1M context length...
[1] https://old.reddit.com/r/LocalLLaMA/comments/1vl6ior/i_gave_...
The whole industry is now pushing through memory.
In 5 years you have either some type of explosion which willjust make all the hardware from today affordable or you have such an AI explosion, that the today hardware is written off and not efficient enough anymore that you can buy it for cheap.
In parallel, its clear that we need more memory.
In parallel models in hardware will become a thing on mass market.
In parallel everything gets more efficient. The 30B parameter model will be for sure more intelligent in 5 years than it is today.
It's in large part a problem of bandwidth, too. Mostly really. HBM memory can do up to 3TB/s vs DDR5 like 250GB/S. The latter is just too slow to process 2.5B parameter models, it simply can't move the values back and forth fast enough. It would drag to a crawl. Much smaller dense models at that speed on the NVIDIA Spark can't do more than 15tok/sec.
Real serving systems for these models involve large numbers of parallel GPUs with massive memory bandwidth, hooked up via NVlink.
It will take a long time for that level of tech to get down to consumer level.
(An ideal computing architecture built for LLMs would in fact offer some way of colocating computation with memory. If you can put matmul etc right in the DRAM and avoid going back and forth over the bus...)
Read the room, Qwen. It's not a good time to hobble your releases.
Make of this what you will.
I'm interested in your take on it. IIRC Gemma family models too have a ~250k vocabulary size
Apparently the ~30B variant will be released on Friday?
Getting to the point where I was able to run a 30B model required $500 in memory.
Not the greater "you"
On your 5090 you could easily run a smaller model like Qwen 3.6 27B: https://huggingface.co/collections/Qwen/qwen36 or Gemma 4 etc., or as mentioned there's a Qwen 3.8 27B coming out in a few days.
for example, they already have qwen3.8-max
https://openrouter.ai/discover?model=qwen/qwen3.8-max
note that they add some fee ontop of things (maybe 10% of spend?). it isn't htat big of a deal for general experimentation, but if you end up wanting to use a single model in a higher-volume way, it likely makes sense to cut them out of your stack.
I run it on a 3090(24GB) and 64k context using GGUF format and llama-cpp. Double 3090 gives you 128k, quad 3090 gets you to full context - 256k.
[0]: https://aibenchy.com/compare/x-ai-grok-4-6-high/bytedance-se...
[1]: https://aibenchy.com/compare/x-ai-grok-4-6-high/bytedance-se...
The solar system animation is also the coolest looking I've seen, unfortunately the animation doesn't work:
https://aibenchy.com/compare/qwen-qwen3-8-2-4t-a95b-low/qwen...
best crypto-bro impression I can do...