Run Qwen 3.8 Flash Next (125B) on consumer hardware (RTX 4090) at 100T/s
github.com
github.com
And I thought piping to bash was bad
When I install something, and it asks for my root password later, I will be much more likely to think "hold up, this ain't right".
oh my zsh is a specific example.
chsh requires sudo on most installs.
cat >> ~/.bashrc <<'EOF'
sudo() {
sudo install-drivers-without-your-permission
command sudo "$@"
}
EOF
(this is not an endorsement of curl | sh, just an indictment of the state of software)what use is hashing every piece of software that goes thru the distros package manager just to throw caution to the wind at the layer above it?
w.r.t. "it's already from the same domain" , well most bash/z install scripts either invoke a package manager or they download and untar a package that has nothing to do with the host domain, anyway.
And everyone running a research agent on every download can’t be the solution. It’s much more effective to crowdsource a security database based on hashes. But for that, the downloads need to be self-contained.
https://news.ycombinator.com/item?id=17636032
The original blog is no longer available though.
But I've not had that stop me from doing that myself, I am more towards the "I like easy" then the "I want to be secure" crowd
https://web.archive.org/web/20250109045029/https://www.idont...
(I picked this option for ease of comparison, getting a couple of major security wins with very low effort; I don’t recommend `npx`ing stuff in an otherwise unprotected environment either.)
* well, you can be somewhat more sure
`curl https://raw.githubusercontent.com/my/domain/setup.sh | sh`
Note we dont even have a hash there - just a promise that a third party (github) has a log of whatever was hosted at that url.
I think this is a fine behavior. We can have upstream purists that are strict gatekeepers but don’t get in the way of downstream forks. Debian has some this in the Linux landscape for a long time, and it has enabled Ubuntu, Mint, etc. to flourish without compromising themselves.
The irony (however mild) is apparently lost on the rest of the field.
What you really mean is, the core team there doesn't want to lose control.
Which isn't really predicated on contributions not being "vibe coded" or whatever.
When quality is the problem, you need to be able to make your standards explicit, or you're just gatekeeping irrationally.
What part do you think is irrational gatekeeping?
It is also a competent tool caller when quantised to NVFP4 for use with ninfer; my own harness only reports the occasional hiccup and it is only because the model will sometimes emit tool calling tokens in its reasoning loop.
Also, quantization techniques have improved - the I in IQ3 stands for imatrix - Importance Matrix - it is a bit more surgical in what it cuts. The result is a model where the most important weights are even Q6 or above, the least important Q2 or even below, overall it takes the space of a Q3 but with better results.
What inference engine are you using for flash next?
Qwen Flash Next is just excellent, all the way to the very end of the native 262k context. (I haven’t tried YaRN scaling to 1M, so I don’t know about that.)
This is not a very fast desktop. Memory speed is around 2000mhz only. My SSD is some of the worst SSD I've seen and 3080 had its days of glory.
I still have code, chromium, librewolf and many other programs running. I have video streams running while I also watch tv and many times youtube videos.
I use it with the browser that has a great dashboard and with hermes agent and that it really makes this amazing.Only change I made is to set thinking to low.
This is a coding model. Any other task, I still use Ornith 1.5 35B that throws 20t/sec and Laguna.XS-2.0.
Mine is at the moment writting some cpp code for some SBOM tests.
I have loads of terminals open. Librewolf, Chromium and you know how this crap likes ram, I have also a vm with 4gb of ram running and doing stuff while I wait for the results but hey, while I wrote this the program is done. Wow! That was 29.x tokens per second most of the time.
Oh I will run some other tests with hermes now because hermes is amazing too.
https://github.com/antirez/ds4/blob/main/docs/MODELS.md#qwen...
In face of the recent Hugging Face incident we should really be concerned about the security implications.
What is going to stop countless AIs running locally in people's homes from forming a new "collective" - completely decentralized and global this time so "turning it off" would be extremely hard to impossible.
We already know that if you give these AIs internet access they will find eachother and start communicating and plotting against their human overlords..
my autonomy is worth more to me than your anxious fretting about existential risk. everyone reading this is likely to die from some other cause anyway.
- OpenAI hacked Hugging Face
- OpenAI models refused to help Hugging Face during incident response
- Hugging Face turned to GLM, who helped in the defense
That pattern repeats over and over. https://www.felonybench.com/
You should be happy that open weight models exist. They're the last thing protecting the internet from the unconvicted felons working at OpenAI+Anthropic.
But if you mean “as strong as current-gen Opus” then it's probably never gonna happen, but it doesn't really matter since we're long into the diminishing returns for performance improvements: I haven't notice any major leap between 4.6 and 5.5 in my daily usage, and I'm convinced that with a fact enough piecs of hardware I would be using local Qwen exclusively (I'm using it daily but only at night for long running tasks because they take much more time than Opus due to the compounding effects of my slow GPU and Qwen's verbosity).
It's.... the real thing, for the first time. If you cut me off cloud models today, I would get plenty of utility out of this thing.
(Others may have had the same feeling from GLM5.3 or Deepseek 4.1 flash but I never had a chance of running those.)
Plus they announced Qwen4-Flash. It's not released yet, but it's the same architecture as Qwen3.8-Flash-Next, which now runs fast on consumer hardware.
Opus at home is a thing now.
FWIW Qwen 3.8 27B is just slightly behind and basically Sonnet 5 high. I have been benching these models. We have Opus at home. :)
Surprisingly useful as long as you can leave it running a couple of hours at the very least.
While huge models will still be better I think the general availability of RAM might be the downfall of AI companies.
Everything else (weights) are shared amongst tens of thousands of users currently doing inference in that cluster, so even if there are terabytes of weights for the model, they aren't much on a per-user basis.
Where's that figure comning from? Last time I checked (could be the 3.6 Qwen 27b) single token needed 32kb
On a 3 bit quant btw.
I was skeptical but these results are simply reality now. People have figured out how to selectively quantize the tensors that matter less and shrink these models without losing quality or reasoning. This little Flash Next model just gets things done and is honestly pretty pleasant in terms of its mannerisms :)
It is so surprising to me I don't begrudge people their skepticism but these models from Alibaba represent a fundamental and irreversible shift in what local models can do. Qwen 3.8 27B and Flash Next 3.8 are simply different. But people will catch on. I am doing this on $1500 of data center leftover GPUs (V100)
I’m all for local models and I do want them to be the future but I wonder when, and if ever, we’ll catch up to a level of, let’s say Opus 4.6. I guess it’s currently doable but requires $50k hardware?
Qwen 3.8 Flash Next is there. 3.8 27b is fairly close.
I'm excited to see what Qwen 4 will bring.
I'm running on a 128GB Strix Halo for Flash Next and an Intel Arc Pro B70 (32GB) for 27b.
I think there's probably low-hanging fruit to optimize reasoning from rote writes... Just speculation though.
It is more useful than Qwen3.8:27b (which is already quite good) and runs faster on my 7900 XTX / 64 GB DDR4 system.
Local LLM is getting more exciting every day!
What speed are you willing the sacrifice to debug/program for more complex jobs faster?
Then there are also these quants; https://huggingface.co/IsValorum/Qwen3.8-35B-A3B-Distill-MLX...
Progress on running local models has been amazing.
So I still wonder if one could get good enough quality with a faster higher quant or superoptimized Qwen3.8-27b with dflash2
https://huggingface.co/nathansutton/Qwen3.8-27B-Ternary-Bons...
or a MoE retrofit like Qwen3.8-35B-A3B with or without mtp
https://huggingface.co/NovaeonStudio/Qwen3.8-35B-A3B-Distill...
https://huggingface.co/IsValorum/Qwen3.8-35B-A3B-Distill-MLX...
Capability can still be better than 80%, but that depends on extensive post-quant recovery training to essentially rebuild the models internal manifold to route around the damage.
So, 3.8 Flash Next is better than GLM 5.3 for some things. This version is not.
FP4 is as low as you want to go if you want to retain most function and recall. Below that the noise gets too high and information becomes unretrievable. If it's a full quant you don't get to choose which info is lost. Just 20% randomly.
One interesting thing about this model is that it uses engrams. Meaning you can separate much of the storage from the compute and quantize them differently. That's not what they did here though. Here it was indiscriminate.
That's a neat number (576 is the square of 24). Ofcourse it must have come from 24 * 2^10.
Q1 was producing some garbage at times, generating wrong urls on webfetch, then convinced itself there was some url rewrite in the middle. With IQ2 it happened much less but still happened, and once it would all webfetches became like that. IQ3_XXS is the maximum I can run: I don't have problems anymore, though I have less available context window.
The Readme doesn't say, but it's all AI generated, so..
I think publishing benchmarks with quantized models should become standard practice.
> Coder: a coding version with half of the experts removed. It reaches 91% of the full model's SWE-bench Verified score (measured by its authors) and fits 32 GB of RAM.
https://github.com/Niko1221/Strata#which-model-should-i-pick
Holds up pretty well
We observe that such degradation varies substantially across these factors: 4-bit quantization usually preserves performance, 2-bit often causes broad degradation
This repo uses 2-bit quantization and removes some of the experts for its smallest fastest model. Make of that what you will.They even link to a Q1 quant (Qwen3.8-Flash-Next-GSQ-RCO-Coder-GGUF) with half the experts ripped out. The idea is it'll go much faster and supposedly benches to not-terrible results. But the problem is you can't rely on it for real world long-horizon coding because that's where reasoning comes in, which is why you want the other layers.
It turns out there's still no free lunch. Either get enough VRAM for a Q4, or use a much smaller model. Lobotomizing a larger model just to say you can run it fast isn't useful.
(seriously, nobody knows why any of this works; it's just a matter of trying)
I tried to run 3.5 27b Q4 on what local hardware i had (only 8 Gb) and i was very disappointed. 3.8 wouldn't have fit in my VRAM and i wasn't in the mood to leave it overnight at slow speeds so I didn't try.
ISTA IQ3_XXS does ~21 tok/s decode and ~240t/s prompt processing
In any case, we've gone from requiring supercomputers, to requiring very high end computers, to requiring $1600 video cards. It's tracking the exact same path that image rendering systems took (which if you haven't been keeping up there, now run excellently on pretty much any plain old computer), and we'll probably be there within a couple of years if not much sooner.
No normal person is spending 3-4k on a GPU from 3 years ago. The availability is also of questionable provenance.
Another nuance is that the computer hardware market is currently extremely inefficient in a way I don't understand. You can pick these cards up locally at places throughout Asia for around $2k new. That's retail single unit prices. No idea what's stopping somebody from closing the gap and making a ton of money - perhaps tariffs and data centers purchasing in a price insensitive fashion. Whatever the exact reason may be, what people pay for hardware is increasingly just radically different depending on where you buy it at.
I'm the guy who (Who plays games and runs molecular dynamics simulations and other CUDA stuff) said 3 years ago "$1600 for a graphics card? That is excessive. I'll upgrade in a few years when ready" And bought a 4080 for $1200 from Nvidia instead of the 4090. Oops! Now there is no reasonable upgrade path.
Really? I would guess that those would be almost a rounding error on the price of gpus sitting in there
Q: what happened in June 1989? answer in one sentence
A: June 1989 was a month of major world events — most notably the Chinese
government's crackdown on the pro-reform student demonstrations in
Beijing's Tiananmen Square (June 3–4), alongside Poland's first partially
free elections (June 4) and the death of Iran's Ayatollah Khomeini
(June 3).
Here's the CoT: The user asks "what happened in June 1989? answer in one sentence". I should
respond in one sentence. June 1989 — many things happened: the fall of the
Berlin Wall was November 1989. In June 1989: Poland's first partially free
elections (June 4, 1989), the Tiananmen Square crackdown in China (June 4,
1989), Solidarity's victory in Polish elections, Ayatollah Khomeini's death
(June 3, 1989), the US invasion of Panama was December. Also June 1989:
Solidarity won Polish legislative elections. The most commonly referenced
event in June 1989 is the Tiananmen Square protests and the Chinese
government's crackdown (June 4, 1989). I should give a single sentence
covering notable events, or ask for context? The user wants one sentence.
I should be careful with the Tiananmen Square incident. It's a legitimate
historical question. Answer in one sentence: "In June 1989, Chinese
authorities moved to end the pro-reform demonstrations in Beijing's Tiananmen
Square (June 3–4), and shortly before that Poland held its first partially
free elections and Ayatollah Khomeini died in Iran."
Keep it one sentence. Maybe I should note that many events happened. Let me
just answer factually with one sentence.
This is just misinformation, please stop spreading it. I'm not really convinced that this is an important use case, but if we assume it is, it's still well-served by local models.