DS4 Flash can currently run reasonably well on systems with ~96gb+ RAM, I wonder if Hy3 can compete there.
DS4 Flash can currently run reasonably well on systems with ~96gb+ RAM, I wonder if Hy3 can compete there.
One thing that might not be obvious about about DSV4 is how much innovation the Deepseek team implemented in its architecture. When llama.cpp fully supports its lightning indexer, the full 1M context will only require about 6G of RAM. So even though they are similar in size, I believe Deepseek will be much more efficient in that regard.
> I wonder if Hy3 can compete there
Highly depends on how well Hy3 is resilient to quantization. DSV4 is useful even at 2-bit quants.
We have not seen the full power of deepseek v4 yet.
Its also only 13B active, so your decode speed would be nearly 2x that of Qwen3.6-27B. So there are other latent benefits as well.
https://huggingface.co/collections/z-lab/dflash
I'm running the qwen3.6-27B + dflash on a spark and tgen is way up, but keep the draft count low, acceptance rate is terrible beyond half a dozen and it requires more memory
For 'general intelligence', DS4 Flash seems to be a noticeable step up still.
And for MoEs, very small amounts of loss can mean you're flipped to entirely different experts (this is also a problem more broadly with numerical stability issues too).
I'm not aware of any great benchmarks that work by giving it a live agentic harness and a number of realistic tasks that take most of the context window to accomplish and evaluate success rate and tokens to completion... but that's what you'd really want to use to judge different quantization levels.
27B is amazing for its size but has some surprising limits when used for longer agentic coding sessions, especially if you’re using tools that are outside the stock standard web tech stuff: it really isn’t good at Relay, for example.
First, vLLM is like, we can do better than this, we need a better default target. It parses capability wrong, silently falls back on sm_80 xmma kernels when cutlass3x_sm120_bstensorprop_... CASK kernels are available, has mad Python in the hot path, emits dubious tool call syntax, and just managed to do something so off-road it tripped a driver bug that managed to wedge MMIO so bad the EC couldn't get an SBR out so the fan was going literally max until I hard pulled power. They say AMD is worse and I'll take their word for it because patching the driver and Inductor both in the same day is plenty of grief for me.
But the amazing thing and why your comment prompted me to write all this is that Qwen3.6 is insanely good, so good I am seriously questioning the MoE dogma. I was running the NVFP4 quant with the BF16 drafter, 100+ tokens a second of clean, legible reasoning trace and flawless tool calls. It's small and you can tell it hasn't memorized half the internet, but it's reasoning is like, better than most frontier. Opus has cleaner reasoning, GPT 5.x and Gemini 3.x Pro do not. If someone scaled that boi up by 5x? I get the feeling DeepSeek did so many arch innovations in one release that they just didn't quite have the convergence, this is like, the fundamentals as artistry. It's wayyyyyy stronger than GPT-4o at over a trillion parameters.
The other thing is I was using it on OpenRouter and it was all janky in the traces, stuttering and going in circles. On another day I would have been like "what do you expect it's the size of an iPhone". I wonder how many other people have drawn that conclusion too.
I'm not going to call it a conspiracy because it's explained by neglect, but we haven't even scratched the surface of the local model ceiling. With a harness that kept the data fresh, scale up the parameter count a bit, stretch context out a bit, and write an LLM serving engine that isn't hobbyist Python jank wrapped around fuckin inductor/triton jank?
That's Claude Code Opus experience on an expensive gaming box.
Whereas I can run DSv4 Flash on a pair of DGX Sparks and have enough memory left over for 3M tokens of KV cache, with Hy3 (quantized to FP4), there is only room for ~130K tokens of KV cache.
It's exciting that the open models continue to get better and more efficient across the board!
I've found DS4 Flash to be very temperental (via Claude Code). The speed is great, but it often builds a completely wrong mental model and charges off down the wrong path. I find myself needing to rein it in regularly (and also compact the history, which undercuts the whole cache price advantage).
Hy3 isn't as fast, but so far it seems to stay on track much more reliably than DS4 Flash. It also doesn't seem to degrade as much with longer context. I'm not sure what the real pricing is, but I feel like it's a very competitive model.
As an aside, I also nabbed a 50m token pack for LongCat 2.0 to give it a whirl. Not free, but it's so cheap they're basically giving it away. Very impressed too - seems roughly on par with Hy3. Not frontier-level intelligence, but a dependable workhorse that can navigate a codebase well and can reliably execute what you tell it to do.
Edit: fixed, got bad info