Hy3
hy.tencent.com
hy.tencent.com
I tried the preview model 41 days ago and got a pelican with a "change pelican color" button: https://static.simonwillison.net/static/2026/hy3-preview-pel...
tencent/Hy3. New Apache 2.0 licensed model from Tencent in China
Is there a Tencent AI lab elsewhere (MiniMax have some association with Tencent, for example)?Also they have large European / South African shareholders.
As of today, it has fallen to 8/9th on the rankings. I don't see a reason where you would use this model over competitors. However, price economics are bit confusing, as currently the effective input price of Hy3 via OpenRouter is now the same as DeepSeek-hosted DeepSeek Flash V4.
I mean it's still a small model, but at least the benchmark scores (incl. on DeepSWE) went up significantly.
It costs as much as Flash, but the benchmarks are on par with Pro (or above in some cases).
Of course, benchmarks are mostly meaningless -- the only real benchmark is the actual work you give it :)
Oh and very good world knowledge for the size: better than than DS4 Flash
The simplest way I'd put it is, teaching a model to write coherently (follow rules, patterns, etc.) is easy enough: just use teacher forcing. Teaching a model to write creatively is easy enough: just use RL and punish it for not being creative.
Teaching a model to write well and creatively takes learning two partially opposing objectives that spike the learning requirements in ways that smaller models really struggle with.
How are you scoring creativity in an unsupervised manner? That seems anything but easy.
Once creativity is being measured in isolation, getting multiple responses from the model is enough to measure creativity a ton of different ways: wordfreq to identify overused phrases, getting multiple responses for the same prompt and promoting the least similar as preferred for policy optimization, etc.
But that's of limited use for stuff like getting diverse names and such. You want creativity and coherency, and if you just punish the model for using an overused phrase, the first thing it does is strongly learn a new overused phrase (or gibberish).
(Also I don't think you mean unsupervised. You probably mean without humans [since LLMs struggle to judge creativity], but that's not what unsupervised means.)
> enough to measure creativity a ton of different ways ...
The things you listed seem more like temperature than creativity to me. At this point it occurs to me that this is likely yet another case of highly misleading technical jargon. Suffice to say that truly creative writing requires something entirely different than unusual sentence structure - in fact it doesn't require unusual phrasing at all.
Re unsupervised, it seems the misunderstanding here follows naturally from the previous difference in word meaning. Hopefully you see the difficulty of scoring long form answers for the creativity of the underlying ideas, as well as the impossibility of using a labeled dataset to train on such a criteria.
And even in domains that lean heavily on "usual phrasing", like technical writing, human writing has notably higher perplexity compared to another LLM's outputs: https://www.sciencedirect.com/science/article/abs/pii/S10766...
With such a low baseline for what's unusual, you do need to get the LLM writing unusual phrases relative to its baseline. Otherwise you get things like repeated n-grams and overused constructs ("it's not X it's Y"), and suddenly the output is predictably not perceived as creative by humans even if you were to insert some otherwise creative or novel premise.
Getting the model to break out of that baseline without disrupting the model's ability to follow technical rules, maintain logic and reasoning, etc. is the difficult part.
-
Also you're again saying unsupervised then following up with descriptions that sure sound like you're referring to RL and supervised learning respectively this time. (supervised learning can improve creativity by the way
Sure, that is also somewhat challenging and is necessary to get human sounding prose. However doing so is not sufficient to produce "creative" literature by any reasonable metric.
> you're again saying unsupervised then following up with descriptions that sure sound like you're referring to RL and supervised learning respectively this time.
Are you sure it isn't you who is confused about the usage of those terms? I merely suggested that both preparing and making use of labeled data (ie supervised learning) seemed like it would prove quite difficult here. Quoting from wikipedia (https://en.wikipedia.org/wiki/Unsupervised_learning):
> Unsupervised learning is a framework in machine learning where, in contrast to supervised learning, algorithms learn patterns exclusively from unlabeled data.
I take it you enjoy works of literature with inconsistent world building?
Or do you mean professional as opposed to creative writing? Because the bar is even higher for that.
Writing isn't so easy after all.
Virtually all logic or reasoning is, in one way or another, part of the support for writing. It’s what separates actual writing from generating nonsense that happens to fit grammar rules.
The specific details depend on the domain, of course, but I can’t see how anyone familiar with the output of writing can think that there is little logic or reasoning in doing it well.
1: I don't think you're right in this instance, but that's beside the point.
DS4 Flash can currently run reasonably well on systems with ~96gb+ RAM, I wonder if Hy3 can compete there.
Edit: fixed, got bad info
Whereas I can run DSv4 Flash on a pair of DGX Sparks and have enough memory left over for 3M tokens of KV cache, with Hy3 (quantized to FP4), there is only room for ~130K tokens of KV cache.
It's exciting that the open models continue to get better and more efficient across the board!
Its also only 13B active, so your decode speed would be nearly 2x that of Qwen3.6-27B. So there are other latent benefits as well.
https://huggingface.co/collections/z-lab/dflash
I'm running the qwen3.6-27B + dflash on a spark and tgen is way up, but keep the draft count low, acceptance rate is terrible beyond half a dozen and it requires more memory
For 'general intelligence', DS4 Flash seems to be a noticeable step up still.
And for MoEs, very small amounts of loss can mean you're flipped to entirely different experts (this is also a problem more broadly with numerical stability issues too).
I'm not aware of any great benchmarks that work by giving it a live agentic harness and a number of realistic tasks that take most of the context window to accomplish and evaluate success rate and tokens to completion... but that's what you'd really want to use to judge different quantization levels.
27B is amazing for its size but has some surprising limits when used for longer agentic coding sessions, especially if you’re using tools that are outside the stock standard web tech stuff: it really isn’t good at Relay, for example.
First, vLLM is like, we can do better than this, we need a better default target. It parses capability wrong, silently falls back on sm_80 xmma kernels when cutlass3x_sm120_bstensorprop_... CASK kernels are available, has mad Python in the hot path, emits dubious tool call syntax, and just managed to do something so off-road it tripped a driver bug that managed to wedge MMIO so bad the EC couldn't get an SBR out so the fan was going literally max until I hard pulled power. They say AMD is worse and I'll take their word for it because patching the driver and Inductor both in the same day is plenty of grief for me.
But the amazing thing and why your comment prompted me to write all this is that Qwen3.6 is insanely good, so good I am seriously questioning the MoE dogma. I was running the NVFP4 quant with the BF16 drafter, 100+ tokens a second of clean, legible reasoning trace and flawless tool calls. It's small and you can tell it hasn't memorized half the internet, but it's reasoning is like, better than most frontier. Opus has cleaner reasoning, GPT 5.x and Gemini 3.x Pro do not. If someone scaled that boi up by 5x? I get the feeling DeepSeek did so many arch innovations in one release that they just didn't quite have the convergence, this is like, the fundamentals as artistry. It's wayyyyyy stronger than GPT-4o at over a trillion parameters.
The other thing is I was using it on OpenRouter and it was all janky in the traces, stuttering and going in circles. On another day I would have been like "what do you expect it's the size of an iPhone". I wonder how many other people have drawn that conclusion too.
I'm not going to call it a conspiracy because it's explained by neglect, but we haven't even scratched the surface of the local model ceiling. With a harness that kept the data fresh, scale up the parameter count a bit, stretch context out a bit, and write an LLM serving engine that isn't hobbyist Python jank wrapped around fuckin inductor/triton jank?
That's Claude Code Opus experience on an expensive gaming box.
One thing that might not be obvious about about DSV4 is how much innovation the Deepseek team implemented in its architecture. When llama.cpp fully supports its lightning indexer, the full 1M context will only require about 6G of RAM. So even though they are similar in size, I believe Deepseek will be much more efficient in that regard.
> I wonder if Hy3 can compete there
Highly depends on how well Hy3 is resilient to quantization. DSV4 is useful even at 2-bit quants.
We have not seen the full power of deepseek v4 yet.
I've found DS4 Flash to be very temperental (via Claude Code). The speed is great, but it often builds a completely wrong mental model and charges off down the wrong path. I find myself needing to rein it in regularly (and also compact the history, which undercuts the whole cache price advantage).
Hy3 isn't as fast, but so far it seems to stay on track much more reliably than DS4 Flash. It also doesn't seem to degrade as much with longer context. I'm not sure what the real pricing is, but I feel like it's a very competitive model.
As an aside, I also nabbed a 50m token pack for LongCat 2.0 to give it a whirl. Not free, but it's so cheap they're basically giving it away. Very impressed too - seems roughly on par with Hy3. Not frontier-level intelligence, but a dependable workhorse that can navigate a codebase well and can reliably execute what you tell it to do.
I looked into DeepSeek's architecture a little bit and the main focus was how can we save as much money as possible. They did a lot of cost cutting with the attention mechanisms. This allowed them to offer an insanely cheap price even on massive contexts, but seems to have come at the cost of performance?
At least, that's my guess, when I see smaller models costing more and outperforming, I think, "they must have denser attention?"
I would.
I'll try it again now that it's out of preview and has been updated with more post-training. It presumably can't be worse, so maybe it's better enough to compete with a 31b model.
I haven't tested it yet so I cannot comment on the quality (nor the comparison with 10x smaller (!) models)
Of course the bigger model embeds more knowledge, but when neither model has the knowledge necessary to perform the task, hy3 makes idiotic decisions all the time whereas gemma 31b has a decent hit rate.
hy3 feels like someone who's read a lot of books and says the right words but has nothing of substance between their ears, gemma feels like a reasonably intelligent person who doesn't understand the domain, the latter is muuuch easier to work with than the former.
I've only used the Hy3 preview, so I don't want to judge too harshly, yet. But, I wasn't very impressed with it a couple of months ago.
perhaps at-home LLMs will bring me back to that. fun days of hacking and thermodynamics.
Not really at gpt 5.5 tier though, and probably below glm 5.2...
But most of all it just works for me for most things I tried and it's exceedingly cheap so there is no reason not to use it, if you need a foss model.
Edited: gpt-5.4-mini not the base gpt-5.4
GPT5.4 xhigh DeepSWE - 52%
A lot of contaminated benchmarks in the blog post about Hy3, needs real testing though I have a distinct feeling it's benchmaxxed like a lot of Chinese models.
-Fable + Gpt 5.6 Sol
-Opus + Gpt 5.6 Terra + Grok 4.5 + Muse Spark 1.1
-Open Chinese models: GLM + et family
The economics is on the Fable tier people are willing to spend a lot on it and on the Open tier you have to give it away to drive usage. The bottom tiers are also getting more and more competitive.
But, it performs very well for its size. I just looked it up, and it's much smaller than I thought it was when I was testing it. 310B A15B is tiny for how well it performs. I guess that explains why it's so cheap.