Yi: Open Foundation Models by 01.AI
arxiv.org
arxiv.org
> Yi-34B-Chat model landed in second place (following GPT-4 Turbo), outperforming other LLMs (such as GPT-4, Mixtral, Claude) on the AlpacaEval Leaderboard (based on data available up to January 2024).
> Yi-34B model ranked first among all existing open-source models (such as Falcon-180B, Llama-70B, Claude) in both English and Chinese on various benchmarks, including Hugging Face Open LLM Leaderboard (pre-trained) and C-Eval (based on data available up to November 2023).
It’s hard to guess how long before any flavor of an “open” model will consensus match what was released in 2023 let alone potentially exceed it.
A big part of the race seems like it will depend on how high gpt-5 can raise the bar. If it’s only incrementally things may converge quickly.
I'm not sure why there is such a big gap between the release of the models and the publication of the paper.
EDIT: Okay, this appears to be a new set of models with the same name, based on the same models from November but now with multimodal capabilities.
- GPT-4-Turbo: 50.00%
- Snorkel (current 2nd, Mistral 7B fine-tune): 34.86%
- Yi 34B Chat (current 6th): 29.66%
- GPT-4: 23.58%
Thoughts:
- Just saying that it came 2nd is quite misleading, the difference in score is significant.
- Not sure what's up with this benchmark, I've never seen GPT-4-Turbo vs GPT-4 performing so differently.
- The Snorkel model is impressive with just 7B parameters. The Yi authors claim that their success is based on good training data cleaning. This seems to be key at least for this benchmark. Snorkel has also always been all about that, using programmatic methods to generate lots of quality training data.
The benchmark is bad.
https://huggingface.co/spaces/lmsys/chatbot-arena-leaderboar...
Model creators train models including open-source benchmarks in the data, either intentionally to achieve better scores or inadvertently through leaks from various sources.
The best tests of these models are people who want to use AI to solve real problems attempting to do that with various models. If they work, report that they worked. Also, publish the work and result pairs permissively when possible to evaluate that and use it for fine-tuning, too.
Just in case anybody else is excited then misled by their tagline "Building the Next Generation of Open-Source and Bilingual LLMs".
https://github.com/01-ai/Yi/blob/main/MODEL_LICENSE_AGREEMEN...
1) Your use of the Yi Series Models must comply with the Laws and Regulations as well as applicable legal requirements of other countries/regions, and respect social ethics and moral standards, including but not limited to, not using the Yi Series Models for purposes prohibited by Laws and Regulations as well as applicable legal requirements of other countries/regions, such as harming national security, promoting terrorism, extremism, inciting ethnic or racial hatred, discrimination, violence, or pornography, and spreading false harmful information.
2) You shall not, for military or unlawful purposes or in ways not allowed by Laws and Regulations as well as applicable legal requirements of other countries/regions, a) use, copy or Distribute the Yi Series Models, or b) create complete or partial Derivatives of the Yi Series Models.
“Laws and Regulations” refers to the laws and administrative regulations of the mainland of the People's Republic of China (for the purposes of this Agreement only, excluding Hong Kong, Macau, and Taiwan).
Is there a benchmark for bias in model outputs? No doubt China has one, somewhere, except it's not skewed towards prevention.
I'm hopeful that this is where copyright law lands. It seems like this might be the disposition of the regulators, but we'll have to wait and see.
In the meantime, maybe you should build your product in this way anyway and fight for the law when you succeed. I don't think a Chinese tech company is going to find success in battling a US startup in court. (I would also treat domestic companies with model licenses the same way, though the outcome could be more of a toss up.)
"Break the rules."
"Fake it until you make it."
Both idioms seem highly applicable here.
[1] I think this should be a viral condition. Finetuning on a foundational model that incorporates vast copyrighted data should mean downstream training also becomes public domain.
https://www.gnod.com/search/ai#q=Two%20cars%20have%20a%20100...
I gave it 3 tries and each time, Yi picked one of the cars as the winner.
I've been watching for many months now, how LLMs got better and better at solving it. Many still struggle with it, but the top ones nowadays mostly get it right.
On the other hand, it does feel fun that the top ones appear to solve it, and I understand why it feels cool to have a computer that appears to be capable of solving these puzzles. But really, I think this is just specificity in training. There is no theoretical or empirical basis for LLMs having any reasoning capability. The only reason it can solve it is because side the creators of these top models specifically trained the models on problems like this to give the appearance of intelligence.
The LLMs lay out how to go about figuring out the answer, do a series of calculation steps and then come up with an answer.
If you add "Please answer in just one short sentence." to the prompt, even the top ones get it wrong.
Pause tokens (thinking tokens) are also an interesting method to achieve that and seems to have a positive effect on performance:
Yes there is. Learning to predict the next token implies a lot of things, among which is also logical reasoning. The chain-of-thought approach shows that when you stimulate this behavior, you get higher accuracies.
Deep learning models are specifically designed for automatic pattern recognition. That includes patterns of reasoning and problem solving.
> The only reason it can solve it is because side the creators of these top models specifically trained the models on problems like this to give the appearance of intelligence.
That's not how deep learning works, and not how machine learning works in general. The models can automatically recognize patterns of reasoning then apply those methods to problems it has never seen before.
> The only way it can do so is not through reasoning, but by having been trained on a structurally similar puzzle.
This is a fundamental misunderstanding of how it works. The large deep learning models have 100+ layers, modelling extremely abstract features of the data, which include abstract patterns of problem solving and reasoning. They are not simply regurgitating training examples.
Geoffrey Hinton - Mapping Part-Whole Hierarchies into Connectionist Networks (1990)
https://www.cs.toronto.edu/~hinton/absps/AIJmapping.pdf
"The paper, titled "Mapping Part-Whole Hierarchies into Connectionist Networks" (1990), demonstrated how neural networks can learn to represent conceptual hierarchies and reason about relations like family trees.
Specifically, Hinton showed that by training a neural network on examples of family relationships (parent-child, grandparent-grandchild, etc.), the network was able to accurately model the inherent logical patterns and reason about new family tree instances it had not encountered during training.
This pioneering work highlighted that instead of just memorizing specific training examples, neural networks can extract the underlying logical rules and reasoning patterns governing the data. The learned representations captured abstract concepts like "parent" that enabled generalizing to reason about entirely new family tree configurations."
His answer was "Neither win" and it took him 1 minute and 24 sec using no pre-defined algorithm or heuristic.
He said his process of thoughts was:
"I figured it would take 10 hours for car A to finish 100 miles and it would take twice that long for car B. Since Car B is already halfway there when car A starts, then they would arrive together"
I as 40 year old man, approached it intentionally naively (eg. I did not go looking for an optimal solver first) by making a drawing and attempting to derive the algorithm. It took me ~3 minutes to come to the same conclusion but at the end I had a series of equations, but no algebraic proofs.[1]
So now you have a human child reference metric if you want it.
[1]https://twitter.com/AndrewKemendo/status/1766872572300235022
GPT-4 correctly solved the problem when it was reworded to: "There is a 100 mile race with two participants: car A and car B. Car A travels at 10 miles per hour but does not begin driving immediately. Car B travels at 5 miles per hour and is given a 10 hour head-start. After 10 hours, car A begins to move as well. Who wins the race?"
I think humanity waged war on 01
Very slow, but this is unquantized and there's probably a lot of demand.
Data work is rarely sexy, but (almost) always useful.
Did they release the corpus?
For example, Bytedance has already been caught using the OA API to generate data for their models because they are having such a hard time catching up to OA - and evading bans for doing that, and also instructing employees on how to lie & cover it up: https://www.theverge.com/2023/12/15/24003151/bytedance-china...
Do you think that a small Chinese startup like 01.AI, which by their own admission had to "bet the farm" to buy enough GPUs to train the Yi models at all https://www.bloomberg.com/news/articles/2023-11-05/kai-fu-le... , and which were completely silent about cloning the American LLaMA architecture until people analyzed the released checkpoints and noticed it looked awfully familiar, is going to be above such tactics...? In this economic/geopolitical context? Especially when everyone seems to be doing it, not just Bytedance?* (01.AI claims that, the architecture aside, they didn't simply further train LLaMA models but trained from scratch. You can decide for yourself how much you are willing to believe this.) I wouldn't bet a lot of money on it, and that's why I don't expect to see any large comprehensive data releases from 01.AI for the Yi models.
* This is one of my theories for why so many disparate models by so many different groups all seem to weirdly converge on the same failure modes like 'write a non-rhyming poem', and why GPT-3.5, and then GPT-4, seemed to be oddly difficult to surpass, as if there were some magnetic force which made reaching near 3.5/4 quality easy for 'independent' models, but then surpassing somehow difficult. Everyone is lying or mistaken about 3.5/4 data getting into their corpus, and the sugar-rush of imitation learning fools you into thinking you're making a lot of progress, even when your overall approach sucks. (As Andrej Karpathy notes, neural nets want to work, and so even if you have serious bugs in your code, they will still work pretty well - and simply permanently fall short of their true potential. Cautionary recent example: https://twitter.com/karpathy/status/1765473722985771335 )
You can't hide this. The latent space remains mostly fixed after pre-training. It all depends on the seed for the initial random init. Further pre-training won't move it enough. Because of this property, you can even average two fine-tunings from the same parent model, but never on models trained from different seeds.
For Andoird, they’ll just allow it and your batter will last 30min after a few questions
If it maxes out all cores and memory for 30 minutes then it won’t really work for anything
It is processor and therefore battery intensive but it already won't kill your battery inside of 30 minutes. Obviously it will be worse for resource usage than an app if it's always kept running by some OS level process and set as the processing layer for every trivial thing but it seems like cheaper input handling could decide to promote some input up to being evaluated by an LLM or not.
I've also tried a few Android LLM apps, all running more than 30min.
Current LLM models are not running constantly on the phones to drain your battery. They just run when responding to a prompt. It by no means consumes more battery than a heavy game.
I frantically tried anything available on Groq to improve performance of my GPT-4 based chatbot - it's incomparably bad - and the more of them I see, the more I believe OpenAI has fundamentally no competition at all at the moment.
No exception with the above, also pretty bad (IMHO worse than GPT-3.5).
> We have developed a data cleaning pipeline with great care to effectively clean and filter low-quality data and eliminate harmful information from text data.
I'd imagine it's likely similar for Yi.
China censors those events. They pre-trained with a specific focus on Chinese text, and integrated more native Chinese text than most models do.
Doesn't require any additional filtering on their behalf to have the model reflect that, and if anything the fact that they're mentioned in english implies the opposite of your hypothesis.
If they were going to filter Tiananmen Square, the lift to filter it in English would not be any higher.
As others have reported, English and Chinese queries return different replies on topics that are not kosher in China.
What’s the risk that such models could be used for nefarious purposes by providing propaganda/biased/incorrect/… responses that on a cursory glance seem factual.
I have the feeling you're asking something more specific, something more of a direct interference coming from politics and not just the natural "point of view" about various topics that are present in the chinese training corpora that is understandably different from western corpora.
Do you have anything specific in mind about something that you expect the Chinese government to feed as propaganda that is not already widely being sculpted into the chinese text corpora available on the internet?
I don't have anything specific, and it doesn't have to be different from "chinese text corpora available on the internet", it's just that these models can become yet another channel of distribution for the "chinese text corpora available on the internet" especially if they are unknowingly/naively picked up and used as the foundation by others to build their offerings.
How's the economy of that different for the Russian campaigns? Do they have a larger pool of English fluency to draw from or is the urgency of the operation higher in their case?
(I genuinely would like to learn more about this topic)
In terms of state directed actions, RU/USSR has longer history/experience with foreign subversive influenc operations. They're not taking out full page ads to post editorials on western news paper like PRC, which is just clunky. Bulk of PRC work is focused on now largely defunct United Front presence in west, or on PRC social media platforms in Chinese to target diasphora in Chinese etc. It's running joke in PRC that PRC foreign propaganda department is staffed by incompetent old guards who can't even out influence anti PRC EpocheTimes.
Noam Chomsky and Edward Herman wrote extensively about propaganda in democratic societies in their 1988 book Manufacturing Consent. A nice introductory excerpt is here, and the first two or three paragraphs are enough to begin to see the argument:
https://chomsky.info/consent01/
Put as briefly as possible: propaganda in totalitarian societies is simpler. They just use force to remove people who say the wrong things, and state media to broadcast the “right things”. In democratic societies, institutional power still wants to protect itself, and this is achieved through more complex means, but it is nonetheless still rather effective.
Yes, but in this particular case I'm coming from a viewpoint where I view China as a hostile power. So, at the moment, my worry is about that.
In future, if US slips into authoritarianism, which TBH it might depending on the outcome of the next election, what you note would become a very real problem.
So, putting it differently and more neutrally, is there any research being done on evaluating a political, and other, bias in a model or is it just being all put in the bucket of hallucination?
The point of Chomsky’s work in this case is to show that authoritarianism does not make propaganda more or less likely, it just changes the means by which propaganda is created and reinforced. Chinese propaganda is easier to identify as a foreigner, but the propaganda of your home country has a much more significant effect on your life. The nature of living with pervasive propaganda is that it is hard to see or consider how your life would be different without the propaganda, and that’s what makes it so dangerous.
> is there any research being done on evaluating a political, and other, bias in a model or is it just being all put in the bucket of hallucination?
It’s a better question, and again one that we should ask regardless of the model’s origins.
And yes it makes sense that you don’t see the propaganda in the news you watch. My dad is a very smart man who has always liked Fox News and he genuinely can’t see the propaganda in it.
There are also ways that both Fox and CNN share the same views, and this narrow window of thought on the specific subjects they share in common is a key aspect of Chomsky’s Propaganda Model. In areas that get the left and the right emotionally charged at each other, there can be a pretty wide range of opinions expressed. And then in areas that affect certain ways that state and corporate power influence our lives, there is often complete agreement and zero discussion of dissenting ideas. These ideas are represented as base assumptions about the fabric of society that are considered so obviously true as to melt in to the back drop of the discussion, and be invisible to regular viewers.
Those are the ideas that are so vital for us to understand. Nothing will ever change as long as we keep arguing about trans people in bathrooms, gay marriage, and Taylor Swift’s political opinions. That lack of change is what the institutions in power want. What we need to understand to truly change things is ideas of radical democracy, worker power, equitable distribution of resources and universal rights for people and nature.
Those subjects will never be seriously discussed on Fox News or CNN.