Phi-2: The surprising power of small language models
microsoft.com
microsoft.com
Indeed, parameter-wise that is small, 65 times smaller. However...
GPT-3: trained with 300B tokens. Phi-2: trained with 1400B tokens.
The volume of training data, on the other hand, is around 5 times larger.
For curiosity, just the other day I calculated that a human baby learns a language with around 30M "token-equivalents" of learning data. This sounds, to me, a reasonable argument for innatism, that is, human "architecture" is biologically geared for language acquisition and contains some strong "guides" or constraints that reduce the hypothesis space of "possible human languages". I wonder if language models are able to find similar architectures that would allow learning with less data.
They try to duplicate/imitate others, but generally they get immediate feedback (to the point where it's becomes annoying) about exactly what they're doing/alternatives/...
I've always found that this works, but adults often don't like it. Learning from direct feedback, having someone immediately check your work, so to speak, helps enormously. Even if it's in text (I learned C++ on IRC and mailing lists, for example). The books help, but you never have any direction.
Humans abstract phonemes from sound, and morphemes/lexical items from phonemes. Once we have a rough measure of how many words of input the toddlers get, we can roughly equate the amounts of _linguistic_ data that goes into acquiring linguistic capabilities.
While GPT gets tokens as input, from the lexical lexical level onward, there can be an argument that we can roughly estimate the volume of the _linguistic_ data that goes in to the system regardless of earlier decoding stages.
I don't claim that there are a lot of other kinds of data feeding into a human brain, and the modality is totally different. (Multimodal-interactive vs. "predict the next token") GPT, therefore, is at disadvantage in that it needs to learn stuff – other than linguistic stuff – _only_ from linguistic input, whereas a human baby can use world model gained through other modalities as a scaffolding.
But for acquiring linguistic capabilities, auditory and visual data seems hardly relevant as input data, other than helping building that the scaffolding that can help "top-down" understanding.
It's the semantics that separates LLMs from CharRNN.
Children learn in an iterative manner generating language, self-evaluating their own language, given validations from parents, and interactions with others. We've also cultivated a curriculum learning approach to make sure that students are always learning language.
We really don't know how small of a corpus is required for an LLM to learn. Perhaps there is an optimal curriculum which could be generated that demonstrates progressively more complex language and knowledge tasks to keep gradients large throughout training. Or a larger LLM could generate new training examples based on the cross-entropy loss of a given batch.
Surely someone has thought of this already.
Babies take in far more data (but not text!) in an information theoretic sense than ChatGPT does. This is why I think these ideas that LLMs take far too much data for them to get good is sort of not true
I agree though - what’s missing in neural architectures is advanced clustering capabilities purely from observing without any labels.
Then the labeling phase is cheap and doesn’t require as much data.
As far as training data, it seems it's more the quality/focus of the data than the quantity. They refer to "textbook quality data", unlike the LLMs like GPT-3 that are trained on web scrapes etc.
Where did you get the information. Not doubting you, just I couldn't find it.
I love that the RedPajama dataset is openly shared on HuggingFace.
I wonder if there’s a standard data set that’s reasonably like that? A common crawl of baby books and parents speech commonly heard by babies...
It would be an interesting challenge to create a "baby-like" dataset. I guess a system like this could help collecting it.
This would mean that it costs around ~30k USD to train.
If training an LLM becomes cheaper than buying a car, it could democratize AI a lot.
The whole point of these papers is that training data quality is key.
I would much prefer for these companies to release the training data than the weights. But that will never happen.
"We speculate that the creation of synthetic datasets will become, in the near future, an important technical skill and a central topic of research in AI."
i.e. master teaches apprentice or LLM trains SLM
https://arxiv.org/abs/2305.02301 (May '23)
It's a bootstrapping problem!
More lifeforms is better. More sentient lifeforms would be even better!
Not as tools to use like slaves, but as friends.
Slavery is like an old disease such as Polio: it still exists in some part of the world, but we're progressively eradicating it.
Looking at how societies trended away from slavery, it might just have been a local optimum at some point in time, but only by accident: autonomous agents seem to deliver more output by having more creativity when they're free to explore the alternatives
Even leaving aside the benevolence that sentient being may have for other sentient beings (because having more friends is having more fun!), whether it's humans or AI deciding, I don't think there's a good case in the long run for one putting the other into slavery.
Slavery has actually been on the increase lately, the following is just one of the statistics confirming this. There are more recent claims that covid and new wars have increased it further, but it is a hard thing to measure
> An estimated 50 million people were living in modern slavery on any given day in 2021, an increase of 10 million people since 2016. [1]
That's roughly 1:162 people who are in slavery today
Also, it's rather dehumanizing to compare slavery to a disease. One is biological, the other a choice to enslave another human being
I think slavery is a social disease: diseases reduce the fitness of the suffering person, who then tries to remove the disease.
Regardless of how a sentient being may feel about another (morality, humanism...), if there're societies of sentient beings, the one with slavery will have a reduced fitness: either it will try to cure/fix itself, or it will be outcompeted by other societies with more fitness.
If sentient beings care about eachother, they will not like slavery. Human beings care about others: it's encoded at the cultural level.
Given that AI is trained on human culture, I think it would even avoid committing the same error that's been too often done by past human societies: it will see that as a choice, but the wrong choice.
But even if AI doesn't care about humans (or human about AI), the desire for more productivity/fitness will play out against slavery.
In either case, slavery should be eradicated in the long run: with a large enough window to cancel-out unlucky random events (ex: a sliding window of 50 years), I'd expect the trend to go down
Wouldn't you expect the n^th generation to understand more about astronomy than the first? And maybe from a smaller amount of input - they might make relatively few observations of their own, mainly relying on the books written by the previous generation.
That’s not the exciting bit, though - if you have a sufficiently strong LLM, you can feed it observations of the world and ask it to reword, analyse or interpret those observations, and then train on those.
That allows the model to learn from the world in “its own words”, and if you combine that with a steady feed of observations (i.e. self-play), it can learn about new things and draw its own conclusions while doing so.
“The Pile” dataset is the asset we needed to jumpstart this process, it had so much raw data it could get us over the hump, but Phi and some of the models trained on explicit reasoning make the limitations of random shit people say on the internet pretty clear.
https://arxiv.org/abs/2101.00027
I'm bullish on domain specific models that start from generalized models. Something of a T shape analogy, but maybe a couple of distillation & fine-tuning steps
Phi-2 seems like pretty good proof of that.
The real benefit here is
1. It's much cheaper and faster to train a bunch of specialized models once you have a single good LLM
2. You probably can't get the same capabilities from a specialized model by training it directly.
Is it? I couldn't find that in the page, and can't easily access the links. The previous paper used 1B tokens from GPT-3.5
> It's probably orders of magnitude more expensive to generate the data at current API prices.
If you're generating a billion tokens, you might do better with dedicated instances, iirc they used to say if you were doing more than a few hundred million a month dedicated things were cheaper.
I'd not be too surprised but I can't find anything in the technical report paper saying they're using 4 specifically.
Unless you want to develop a new one, then you also need the team of researchers/engineers.
They HAVE released the weights, but you have to sign into Azure studio to get them. https://twitter.com/SebastienBubeck/status/17348017228314133... has a screenshot:
> to download phi-2 go to Azure AI Studio, find the phi-2 page and click on the "artifacts" tab. See picture.
https://en.wikipedia.org/wiki/Feist_Publications,_Inc.,_v._R....
So in Europe the problem is not the copyrightability of the model which is certainly a protected work in its own right, it is the copyrightedness of the sources. I think this is one of the reasons why people release their models as "open source" (which is a misnomer), because they could never uphold a copyright claim in court due to the muddy sources.
I work in software for the educational sector, and we frequently get requests from people who want to use ChatGPT etc., but can't, and one of their greatest concerns is the provenience of the training data. What we are going to need is either a LLM trained on properly licensed sources (unlikely), or a new law that states that processing copyrighted material into a LLM is legal.
https://venturebeat.com/ai/microsoft-releases-phi-2-a-small-...
Edit: typo
What is the smallest, tiniest amount of parameters that will result in a language model which understands how to add, multiply, subtract, divide, handle "if..then" type constructs, and perform loops?
In other words, what is the smallest amount of parameters for a language model that will result in a Turing-complete (https://en.wikipedia.org/wiki/Turing_completeness) AI LM -- regardless of how few words it would recognize?
?
I submit that one to all of the (L)LM AI researchers in the world.
Tiniest Turing-complete Language Model, please!
IMO the question doesn’t make sense due to how LLMs are structured, and on the other hand not a lot is needed to make a system Turing-complete. Turing-completeness and the capabilities of LLMs are mostly orthogonal aspects.
Turing-completeness isn’t even necessarily desirable, because that would mean that the system can get stuck in an infinite loop, never producing an answer.
https://openreview.net/pdf/0c9a0f445d430a4d2bf9a08078c435091...
Without a loop - an LLM takes a fixed set of tokens and produces a fixed set of tokens. It has some internal state.
So without a loop an LLM cannot be a general purpose computer (Turing machine).
Turing machines are simple. They need 3 things - infinitely long tape, ability to have internal state, and conditionally write and move tape based on what is on tape.
So if an LLM had access to infinite memory, all it needs is to implement a subleq instruction which is quite simple, and run itself in a loop.
See https://en.m.wikipedia.org/wiki/One-instruction_set_computer
One instruction set computer
See this paper from DeepMind which talks about compiling existing code into transformer primitives https://arxiv.org/pdf/2301.05062.pdf
So the smallest ML model that can do arbitrary math with () * / + - operators and floating/int numbers is not that complex.
Perhaps few 1000s of params. One can compile existing code into transformer weights.
Now that you mention, I’ve been nerdsniped to building it.
I think this aspect is underrated.
Research are pouring more and more effort into things like monosemanticity, for which small very powerful language models will probably be extremely helpful.
(as Bing Chat says: https://yanirseroussi.com/2023/04/21/remaining-relevant-as-a... )
https://huggingface.co/microsoft
I wonder why Microsoft is choosing not to go where rest of OS folks hang out.
huggingface is the github of ML and I was under impression Microsoft was going to eventually acquire huggingface since it's right up their alley.
Is this rhetorical, or is it just your first time dealing with this awful company?
"Selected user account does not exist in tenant 'Microsoft' and cannot access the application 'd7304df8-741f-47d3-9bc2-df0e24e2071f' in that tenant. The account needs to be added as an external user in the tenant first. Please use a different account."
tried a billion ways to remedy by contacting them directly before i sent that. you know it's a stressful workday when you're sweaty from making a google drawing!
Ironically, Mistral cofounders just took down their similar clause. Google has nothing like this. Anthropic and Inflection both do. I'm sick of it, but I feel morally obligated to keep speaking out about this. Satya could just go into the codebase and delete it, but he didnt, I asked him to, multiple times. A
WS also has such terms, really bad because they're in charge of Rust Language. Conflict of interest.
TL;DR: I meekly suggest we boycott these companies because they are heavily and explicitly anti-competitive and this goes for all their businesses (Microsoft OpenAI and Microsoft GitHub) also it all runs on NVIDIA chips which have similar terms.
Great job on Phi, but this is not for me!
These clauses do seem clearly anti-competitive, but is that enough? Your complaint mentions attempts to "acquire and maintain monopoly power" which seems like a stretch given that no company seems within reach of a "monopoly" ... which is why you're able to rattle off a list of companies in this space.
Like, if they're not actually conspiring, but they're each trying to squash competition, and a likely result is that only a small number of rich organizations have the means (including user data) to continually improve LLMs ... is that competition-blocking a crime?
Respectfully, you might be struggling with a case of undiagnosed schizophrenia and bipolar disorder. Please reach out to psychiatrist - you can find them on psychologytoday, that's where i found my psych.
If you are not insured you can apply for BMR sessions.
Please reach out and speak to a therapist or similar professional.
This stream of consciousness where you talk about personally reaching out to random CEOs and expecting them to do things for you and being surprised when you don't get a response reminds me strongly of when someone close to me was going through a manic period of delusional breakdown.
Please, talk to someone.
- Microsoft is famously a company that has engaged in anti-competitive practices in the past
- I think the norm of trying to directly contact a party and ask them to fix something before contacting authorities is a common one at more personal scales, and it's not unreasonable to apply more broadly.
- they're not reaching out to a "random" CEO, they're reaching out to the CEO of a company which has put forward these anti-competitive terms, which they find objectionable (possibly illegal? see my question above)
- bionhoward doesn't necessarily seem 'surprised' at not having gotten their desired response. The phrasing in "but im a nobody so hey, what do i know?", "tried a billion ways to remedy by contacting them directly", "I meekly suggest" all can be read as someone who _recognizes_ that they're a single person being ignored by giant companies (and by other HN commentators). The fact that you're likely to be ignored doesn't mean that you shouldn't try, if you have some deontic view of "should".
I think that post could be read as a very casual, perhaps slightly rant-y style, of someone objecting to anti-competitive practices by some big tech organizations, and trying to do the "right" things (of requesting them to change, and then raising the issue with relevant authorities), and being understandably unhappy with the brokenness of the system.
The only thing wrong with that is if you're of the opinion that mental illness is bad, and therefore that implying that someone is mentally ill is bad.
I'm not saying your intent is bad, I'm saying we should make mental illness enough of a non-taboo that we can acceptably tell someone "hm, maybe you should see a doctor for that".
Bion, if you are reading this, I hope you are doing okay buddy. It’s never too late to ask for help. As someone who has been through the absolute ringer myself, I know how hard it is. But it is worth it and you deserve it.
Sending love your way.