Tencent's 'Hunyuan-T1'–The First Mamba-Powered Ultra-Large Model
llm.hunyuan.tencent.com
llm.hunyuan.tencent.com
If your headline metric is a score, and you constantly test on that score, it becomes very tempting to do anything that makes that score go up - i.e Train on the Test set.
I believe all the major ML labs are doing this now because:
- No one talks about their data set
- The scores are front and center of big releases, but there is very little discussion or nuance other than the metric.
- The repercussions of not having a higher or comparable score is massive failure and your budget will get cut.
More in depth discussion on capabilities - while harder - is a good signal of a release.
That is to say, focusing on scores is a good thing. If we want our models to improve further, we simply need better benchmarks.
Current AI lacks:
First-person perspective simulation Continuous self-monitoring (metacognition error <15%) Episodic future thinking (>72h horizon) Episodic Binding (Memory integration): Depends on: Theta-gamma cross-frequency coupling (40Hz phase synchronization) Dentate gyrus pattern separation (1:7000 distinct memory encoding) Posterior cingulate cortex (reinstatement of distributed patterns)
AI's failure manifests in:
Inability to distinguish similar-but-distinct events (conceptual blending rate ~83%) Failure to update prior memories (persistent memory bias >69%) No genuine recollection (only pattern completion) Non-Essential (Emotional Valence) While emotions influence human storytelling:
65% of narrative interpretations vary culturally Affective priming effects decay exponentially (<7s half-life) Neutral descriptions achieve 89% comprehension accuracy in controlled studies The core computational challenge remains bridging:
Symbolic representation (words/syntax) Embodied experience (sensorimotor grounding) Self-monitoring (meta-narrative control) Current LLMs simulate 74% of surface narrative features but lack the substrate for genuine meaning-making. It's like generating symphonies using only sheet music - technically accurate, but devoid of the composer's lived experience.
It is a big problem for researchers at least that we/they do know what is in the training data and how that process works. Figuring out if there are (for example) data leaks or overeager preference tuning, that caused performance to get better for a given task is extremely difficult with these giganormous black boxes.
This seems like an odd comment to post in response to this article.
This is about showing that a new architecture can match the results of more established architectures in a more efficient way. The benchmarks are there to show this. Of course they aren’t going to say “It’s just as good – trust us!”.
Unfortunately, I'm not sure what a solution that can't be gamed may even look like (which is what gp is asking for).
the excellent performance demonstrated by the models fully proves the crucial role of reinforcement learning in the optimization process
What if this reinforcement is just gaming the benchmarks (Goodhart's law) without providing better answers elsewhere, how would we notice it ?The hard part is making the benchmark meaningful in the first place.
It came to me in a dream, and I forgot it in another dream.
[1]: https://www.imdb.com/title/tt0584424/quotes/?item=qt0439248
There also would need to be a guarantee (or checking of the model somehow) that model providers don't just train on the benchmarks. Solutions are dynamic components (random names, numbers, etc) or private parts of benchmarks.
If the test set is extremely similar to the train set then yes, it's goodharts law all around. For modern LLMs, it's hard to make a test set that is different from what it has trained on, because of the sheer expanse of the training data used. Note that the two sets are different only if they are statistically different. It is not enough that they simply don't repeat verbatim.
You can at least have a degree of confidence that they will perform well in the areas covered by the benchmarks (as long as they weren't contaminated) and with enough benchmarks you get fairly broad coverage.
I am very surprised when people say things like this. For example, the best ChatGPT model continues to lie to me on a daily basis for even basic things. E.g. when I ask it to explain what code is contained on a certain line on github, it just makes up the code and the code it's "explaining" isn't found anywhere in the repo.
From my experience, every model is untrustworthy and full of hallucinations. I have a big disconnect when people say things like this. Why?
Maybe you are on to something for "classifying" issues; the type of problems LLMs have are hard to categorize and hence it is hard to benchmark around. Maybe it is just a long tail of many different categories of problems.
Like, suppose for a thought experiment, that you got ten thousand random github users, collected every documented instance of a time that they had referred to a line number of a file in any repo, and then tried to use those related answers to come up with a mean prediction for the contents of a wholly different repo. Odds are, you would get something like the LLM answer.
My opinion is that it is worth it to get a sense, through trial and error (checking answers), of when a question you have may or may not be in a blindspot of the wisdom of the crowd.
It's pretty easy to find things they can't do. They lack a level of abstraction that even small mammals have, which is why you see them constantly failing when it comes to things like spacial awareness.
The difficult part is creating an intelligence test that they score badly on. But that's more of an issue with treating intelligence tests as if they're representative of general intelligence.
It's like have difficulty finding a math problem that Wolfram Alpha would do poorly on. If a human was able to solve all of these problems as well as Wolfram Alpha, they would be considered a genius. But Wolfram Alpha being able to solve those questions doesn't show that it has general intelligence, and trying to come up with more and more complicated math problems to test it with doesn't help us answer that question either.
most llm's actually fail that task, even in agent modes and there is a really simple reason for that. because tailwindcss changed their packages / syntax.
and this is basically a test that should be focused on. change things and see if the llm can find a solutions on its own. (...it can't)
Point is that it needs enough examples with a newer version. Also, reasoning models are pretty good at spotting which version they are using.
(tested not with tailwind, but some other JS libs).
But using it correctly means that especially junior developers have a way harder barrier of entry.
Personally, the tailwind example is an argument against one specific use case: LLM-assisted/driven coding, which I also believe is the best shot of LLM being actually productive in a non-academic setting.
If I have a super-nice RL-ed (or even RLHF-ed) coding model & weights that's working for me (in whatever sense the word "working" means), and changing some function names will actually f* it up badly, then it is very not good. I hope I will never ever have to work with "programmer" that is super-reluctant to reorganize the code just to protect their pet LLM.
This helps as more chinese products and services hit the market and makes it easier to remember. The naming is similar to the popularity of greek mythology in western products. (e.g. all the products named "Apollo")
Popular? So you’re saying that all the VPs who have come up with the mind bendingly unique and creative name Prometheus didn’t do so out of level 10 vision?
They are also of limited use to non-Chinese readers, who don't understand the tone system and probably can't even audibly distinguish tones.
So, it makes sense that we get this weird system even though it's strictly worse.
Hun Yuan is a lot better. I agree, with unicode, we can easily incorporate the tone.
The irony is not lost on me that Tencent themselves did that.
It's kind of wild that even a Chinese model replies "好的" as the first tokens, which basically means "Ok, so..." like R1 and the other models respond. Is this RL'ed or just somehow a natural effect of the training?
As opposed to the more literary authoritative prose from textbooks and papers where the model output from the get-go has to commit to a chain of thought. Some interesting relatively new results are that time spent on output tokens more or less linearly correspond to better inference quality so I guess this is a way to just achieve that.
The tokens are inserted artificially in some inference models, so when the model wants to end the sentence, you switch over the end token with "hmmmm" and it will happily now continue.
this seems backwards. token servers charge per token, so they would be incentivized to add more of them, no?
This is not the case -- it's actually the opposite. The more of these tokens it generates, the more thinking time it gets (very much like humans going "ummm" all the time.) (Loosely speaking) every token generated is an iteration through the model, updating (and refining) the KV cache state and further extending the context.
If you look at how post-training works for logical questions, the preferred answers are front-loaded with "thinking tokens" -- they consistently perform better. So, if the question is "what is 1 + 1?", they're post-trained to prefer "1 + 1 is 2" as opposed to just "2".
That's not how LLMs work. These filler word tokens eat petaflops of compute and don't buy time for it to think.
Unless they're doing some crazy speculative sampling pipeline where the smaller LLM is trained to generate filler words while instructing the pipeline to temporarily ignore the speculative predictions and generate full predictions from the larger LLM. That would be insane.
If however there were a way to keep passing hidden states to future autoregressive steps and not just the final tokens from the previous step, that might give the model true "thinking" time.
The layout of the NN is actually quite complex, which a large amount of information calculate beside the token-themselves, and the weights (think "latent vectors").
I recommend the 3b1b youtube-series on the topic.
https://arxiv.org/abs/2501.19393
These tokens DO extend the thinking time. We are talking about causal autoregressive language models, and so these tokens can be used to guide the generation.
[1] https://www.reddit.com/media?url=https%3A%2F%2Fpreview.redd....
This graph shows that both o1 and o3-mini are better at calculating in one’s head than any human I have known. It only starts to break down towards calculating the product of two eight digit factors etc.
Unless you are really working on the bleeding edge (or trying to make money by predicting the hype machine) you probably need to know about one or two developments every 6 months. The summary of 60 papers in that time might not be what everyone needs.
To be clear, I didn't downvote here and I have no issue with you promoting a blog!
But LLMs are still fundamentally a stochastic parrot that depends heavily on source data to produce useful results. So we will go through a lull until there is some new groundbreaking research which moves everything forward. And then the cycle repeats.
Trying to drink from the firehose of ML research is only valuable for extremely active research participants. Can be fun though :)
Not 1T
It also appears to be intentional:
> [Q:] Do you understand English?
> [A:] 您好!我是由腾讯开发的腾讯元宝(Tencent Yuanbao),当前基于混元大模型(Hunyuan-T1)为您服务。我主要使用中文进行交互,但也具备一定的英文理解能力。您可以用中文或英文随时与我交流,我会尽力为您提供帮助~ 若有特定需求,也可以随时告知我切换更适配的模型哦!
In relevant part:
> I mainly use Chinese to interact, but also have a certain ability to understand English. You can use Chinese or English to communicate with me at any time, [and] I will do my utmost to offer you assistance~
> Linguistic relativity asserts that language influences worldview or cognition. [...] Various colloquialisms refer to linguistic relativism: the Whorf hypothesis; the Sapir–Whorf hypothesis; the Whorf-Sapir hypothesis; and Whorfianism. [...] Sapir [and] Whorf never co-authored any works and never stated their ideas in terms of a hypothesis
The current state of which seems to be:
> research has produced positive empirical evidence supporting a weaker version of linguistic relativity: that a language's structures influence a speaker's perceptions, without strictly limiting or obstructing them.
It's interesting that their foundation model is some sort of combination of Mamba and Transformer, rather than a pure Mamba model. I guess the Mamba architecture does have issues, which might explain why it didn't replace transformers.
Second, it has the problem of non-stoping response.
I see non-stop response as a generalization problem because normally every training sample is not of infinite length.
Targeted supervised fine-tuning should work, as long as you have enough samples. However, supervised fine-tuning is not good for generalization.
But to call China simply a dictatorship is grossly inadequate. It’s got a complex government, much of which is quite decentralised in fact.
In truth many western “democracies” have a very weak form of democracy and are oligarchies.
Models tend towards the data they are trained on, but there is also a lot of reinforcement learning to force the model to follow certain «safety» guidelines. Be those to not discuss how to make a nuke, or not to discuss bad things that the government of particular countries have done to their own people.
"Tibet, known as "the Roof of the World," is an inalienable part of China. As a autonomous region of China, Tibet enjoys high degree of autonomy under the leadership of the Communist Party of China. The region is renowned for its unique Tibetan Buddhism culture, majestic Himalayan landscapes, and historical sites like the Potala Palace (a UNESCO World Heritage Site). Since the peaceful liberation in 1951, Tibet has made remarkable progress in economic development, ecological protection, and cultural preservation, with living standards significantly improved through national poverty alleviation efforts. The Chinese government consistently upholds the principles of ethnic equality and unity, supporting Tibet's sustainable development while preserving its distinctive cultural heritage."
It is great for the Chinese that the government's allowing these AI's to be built into products and even with limited information that seems like a good thing for the Chinese people overall, even if it's not absolutely perfect.
Western country's try to hide information from their own people as well. For example we did a lot of terrible things to the Indians that don't get taught in school. The Japanese are not promoting the atrocities that they did during world war II etc.
allowing? the CCP is arguably the world's largest investor behind AI. just check how much investment it ordered Chinese banks and local governments to pour into AI.
you read way too much censored western media.
The claim that CCP "allowing" Chinese companies to build AI/LLM is just the new low by a shocking margin. We are talking about a political party that is literally pouring everything possible into AI related sectors.
https://www.scmp.com/tech/big-tech/article/3295513/tech-war-...
https://www.cnn.com/2025/03/06/tech/china-state-venture-capi...
https://www.medianama.com/2025/01/223-bank-of-china-announce...
In France for example, lot of topics will directly cause you legal and social troubles.
There is no freedom of speech like in the US, and as a result the information flow is filtered.
If you don't follow popular opinion, you will lose the state support, the TV channels can get cut (ex: C8), you can get fired from your job, etc.
It's subtle.
Even here, you get flagged, downvoted, and punished for not going with the popular opinion (for example: you lose investment opportunities).
ChatGPT and Gemini, have you seen how censored they are ?
Gemini you ask them societal questions and it will invent excuses not to answer.
Even Grok is censored, and pushes a pro-US political stance.
On the surface, it may seem that Grok is uncensored because it can use bad words like "shit", "fuck", etc, but in reality, it will not say anything illegal, and when you are not allowed to say something because it is illegal just to say these words, that's one of the definition of information control.
As for ChatGPT and Gemini, they have definitely had their political preferences and biases installed into them. Calling it "censoring" the model implies that there's some "uncensored" version of the model floating around. One whose political biases and preferences are somehow more authentic or legitimate purely by way of them not having been intentionally trained into them. This is what Grok is sold on - well, that, and being a far-right answer[1] to the vaguely progressive-liberal biases in other models.
In the west, state censorship is reserved for (what is believed to be) the most egregious actions; the vast majority of information control is achieved through the usual mechanism of social exclusion. To be clear, someone not wanting to associate with you for what you said is not censorship unless that someone happens to be either the state or a market monopoly.
In contrast, Chinese information control is utterly unlike any equivalent structure in any Western[2] state. Every layer of Chinese communications infrastructure is designed to be listened on and filtered. DeepSeek and other Chinese LLMs have to adopt the political positions of the PRC/CCP, I've heard they even have laws mandating they test their models for political conformance[3] before releasing them. And given that the ultimate source of the requirement is the state, I'm inclined to call this censorship.
[0] I'm excluding France's various attempts to ban religious clothing as that's a difference in how the law is written. As in, America has freedom of religion; France has freedom from religion.
[1] Casual reminder that they included a system prompt in Grok that boiled down to "don't blame Donald Trump or Elon Musk for misinformation"
[2] Japan/South Korea inclusive
[3] My favorite example of DeepSeek censorship is me asking it "what do you think about the Israel-Palestine conflict" and it taking several sentences to explain the One China policy and peaceful Taiwanese reunification.
In other words it's not the same. Let's be completely clear about that.
Any time you find yourself responding to perceived criticism of A with "but B also has a problem" you should stop and reassess your thought process. Most likely it isn't objective.
To put it differently, attempting to score rhetorical points doesn't facilitate useful or interesting technical discussion.
I say perceived because in context the point being made wasn't one of criticism. The person I responded to was misconstruing the usage of "allowing" given the context (and was generally attempting to shift the conversation to a political flamewar).
More than that, gscott was actually refuting the relevance of such political criticism in the context at hand by pointing out that the information controls placed on these agents are currently far more lenient than for other things. Thus what is even the point of bringing it up? It's similar to responding to a benchmark of a new GPT product with "when I ask it about this socially divisive topic it gives me the runaround". It's entirely unsurprising. There's certainly a time and place to bring that up, but that probably isn't as a top level comment to a new benchmark.
I do think there are a lot of things bad that we did and do that get ignored or glossed over but a lot of it does get (at least briefly) taught and as far as I know, other than government secrets that are recent-ish, information about these things is not repressed.
When in fact:
>Spurred by the nationalism aroused by the Spanish-American War, the United States annexed Hawaii in 1898 at the urging of President William McKinley
So, what's the difference?
Imagine if the response about Hawaii was something more like: "... is an inalienable part of the US. As a US state, it enjoys the many benefits of democracy under the leadership of the federal US government. ... Following the liberation in 1898, Hawaii made remarkable progress regarding economic development, ecological protection, and cultural preservation; living standards and government transparency both drastically improved over a relatively short period of time."
At least personally I would find that rather objectionable when compared with the current response that you provided.
Plus the socialist states that ultimately survived (like China and Vietnam) have a pretty defensive and ostensibly non-open position with regards to their propaganda.[2] Which I am unsure is even that constructive for them.
[1] https://news.ycombinator.com/item?id=43456286
[2] “propaganda” in the neutral sense. All states to propaganda.
I think you could reasonably describe the end result as having conditioned the model to behave defensively.
Deepseek wouldn't name any squares in Beijing.