Tiny Language Models Come of Age
quantamagazine.org
quantamagazine.org
I have an article on building micromodels that discusses some of this: https://neuml.hashnode.dev/train-a-language-model-from-scrat...
Modern computers don’t have less ram than the old fashioned room size computers. Amount of RAM seems much more analogous here than the size of the hardware running the software. Of course the latter will go down, but the former? I don’t see why.
So in this way you've already lost your bet, as we already know that it's possible to create vastly smaller models with similar performance.
Nevertheless, I have not actually seen a model exhibiting high-level LM capabilities (of the GPT-4 level) that was distilled.
Just because a model can match in perplexity with fewer parameters is not the same as matching the performance as subjectively measured by humanity interacting with the model.
Regardless, this isn't the trend we've seen with software RAM usage - information density is only decreasing as hardware becomes more abundant and comparatively less engineering effort goes into fitting everything into as small a portion as possible.
Could we have libraries of small models that are all experts on niche subjects that basically collaborate?
The advantage is that it’s easier to scale and distribute lots of small models on commodity hardware than it is to run giant models that require heaps of contiguous ram.
I’d love to have some time to play with versions of this.
If I do experiments, I’m making models specialized from foundational to programming to specific language to specific tasks. Each pass will have training data so it picks up general concepts. Then, it will be trained to focus more on what it needs. Eventually, what’s slushing around in its neurons should be just the right information. Or it won’t work.
Anyway, my concepts that I published are at this link:
https://gethisword.com/tech/exploringai/index.html
(See alternative models section.)
> But having more fine-tuned control over model parameters is a pattern that will emerge
This assumes parameters do something interpretable. IMO, it’s more that parameters themselves are meaningless/dumb and patterns evolve via interaction of many such units (like in ants)
But all these smaller models are based on those 100b+ models! Like OP on TinyStories relies on larger models twice, to generate and then evaluate. Or the recent wave of small models which are all based on extracting data from GPT-3/4 to borrow the capbilities+RLHFing for free.
You can talk about how interesting it is and how much of an overhang neural nets have in terms of being overparameterized (which is an important AI safety/capabilities issue - the first AGI will be the largest, slowest, and worst one, by a long shot), but the one thing it doesn't tell you is that smaller models can replace big models entirely. Because they are parasitic on said big models, so you still need to train the big models to begin with.
(The smaller models also seem to lose a lot of the qualitative capabilities of larger ones, like meta-learning. Their TinyStories model does only stories. I doubt you could get it to 'dax a blick'.)
I would, but they don't say what their dataset is that I can find anywhere, and the only thing they say about their instruction-tuned is that it's trained on 'publicly available' datasets. You know, the ones where a lot of them turn out under the hood to be drawing from the OA API or other pretrained models in some way or another...
> Especially not with pre-filtered/classified pre-training data.
Indeed not! But what exactly is prefiltering or classifying all that data...?
Maybe for niche use cases like handling customer support requests small models will do well, but GPT4 is not going to be condensed to a few billion anytime soon imo.
Since it's not, it isn't.
Now that I've said it, I wonder if there's any evidence human brains work that way.
Currently I can only really use smaller models for micro-tasks like sentiment analysis or classification, but any type of problem solving has to be left to GPT-4.
Hear the other timeline from Darius Amodei. https://a16z.com/improving-ai/ We may be entering a regime of multi-billion training costs on custom hardware. Open-source AI will be like open-source CPUs. Cool but the real thing even at the smallest scale comes out of monopsonized (is that a word?) inaccessible infrastructure
That is a very interesting result! This may point to the idea that the notion of time itself (as history/historical knowledge of the state of the system at an earlier point in its evaluation/evolution/computations) being more "present" (or pronounced, "relative to results", "apparent in results") in neural networks which are deep (many layers) rather than wide (more neurons per layer).
Which, if true and widely confirmed to be so by other researchers -- would be an interesting discovery indeed!
Also, if true (and it's a big 'if'!), this may help researchers in other areas better understand the relationship between time and information better (i.e., Leonard Susskind, ER=EPR, etc.: https://en.wikipedia.org/wiki/Leonard_Susskind , https://en.wikipedia.org/wiki/ER_%3D_EPR , https://www.youtube.com/results?search_query=leonard+susskin... )
As an analogy, it would be strange if psychology was a part of physics. Physicists would (I think) not be amused if their work literature became flooded with papers about the human psyche. Even if the underlying "hardware" is physics.
Doesn't that support the argument even more? CS and math share a lot more then CS and AI, yet CS and math are different disciplines.
Neural nets came from biology and rely on math.
The art of application
Of course, computing is closer to pure math than applied maths. Computing is applied pure maths, perhaps.
The majority of what computer science practitioners do is general purpose coding, which doesn't have nearly as much to do with the underlying math.
Or, in other words, how proficient would an applied math person be at building a front end? And how many of their skills would they be able to leverage?
That's the overlap, or lack thereof.
The newcomers with their strong opinions seem to have become confused because programmers hit APIs and think they're doing "AI".
Software architecture and neural network architecture are not synonyms.
IOW, an intellectually-linked discipline that drives so much revenue (and thus has so much work to be done in it) that it comprises its own field (about 75% of which is unique "applications-of-thing" problems).
I don't know that this particular aspect of the "AI revolution" needs some disruptive paradigm shift.
I'm hesitant to appropriate "engineer" into what we do. There are certainly people who work with code who earn that title. There are also many who don't.
I do think it would be healthy for people who work with ML to separate more decisively from people who work with general purpose code. There are enough unique problems and solutions in ML that a clear community would better serve the field's maturation.
As opposed to getting an endless summer of "Why don't you just" software developers fouling things up, because it's "similar".
Within a few years the fraction of papers in AI about new architectures, training, hyperparamter optimization, etc will be dwarfed by papers about things like controlnets, LoRA combinators, multi-model dispatch networks and few-shot embedding methods.
Popularity and status quo doesn't change the definition of the underlying theory. I will push back forever on some demotion of the importance of ML (i.e. theory) as distinct from some hype-driven notion of AI.
Then the focus should be to identify something like a fundamental unit of intelligence in order to formalize a higher-order science out of the foundations. An analogy can be drawn between physics and chemistry: we needed to properly identify "the atom" and its component parts to get anywhere with the science of chemistry. But it took a whole lot of physics to get to that point. It seems similar with the ML-cum-AI transition where we'll still need to dig very deep into statistics and information theory before being able to abstract them away in favor of higher-order concepts.
To me it seems we're really far from anything like that yet. Like at least a couple decades if not more. Friston's got some cool ideas that make me think he may have his name on some stuff later on but again the theory on learning systems is barely getting started.
It feels like a good direction would be to "grow" these models from a linguistic kernel, then supply them with a more organized / reproducible "fact bank" / memory store.
We try to mimic this right now with prompt engineering, but in the end it's always just word soup. Having layers to it, like a language synthesis layer, logic layer, memory layer -- where the logic is a constraint on the language, and the memory provides the working data -- it could lead to a model that hallucinates far less.
However, the returns over increased capabilities vastly overwhelm the costs of running larger models.
It doesn’t matter that a model requires $100,000 worth of hardware if it can automate the work of several information worker paid $200k annually.
> Eldan and Li presented the same challenge to OpenAI’s GPT-2, a 1.5-billion-parameter model released in 2019. It fared far worse — before the story’s abrupt ending, the man threatens to take the girl to court, jail, the hospital, the morgue and finally the crematorium.
We already see hints of this with MoE, but something entirely new wouldn't surprise me.
Is it not? Do we just ignore it because it's not fair or they can't prove it?
How can we get from needing a million stories to being able to use just 10 stories?
You likely recall the reversal curse - if an LLM trains "A is B" it doesn't automatically deduce "B is A". You need to explicitly append the second part to the training set to make it complete. Fragmented or incomplete information in the training set persists fragmented in the trained model. Training is blind, only inference is intelligent.
Similarly, many training examples contain apparent information that conceals implicit deductions, for instance a math problem - the statement is evident and apparent but the chain of thought is implicit, concealed.
What I am getting at is a general principle - LLMs need to "study" the original training data and make it comprehensive, elicit those implicit deductions, and enable LLMs to traverse the conceptual space in all directions.
Study, then train.
In other words, to study is to digest the raw data, to reconnect fragmented information, to generate insight, to execute the instructions and see the result, generally to unfold what is hidden.
It's a matter of utilizing LLMs in generative mode to prepare the dataset before training. Microsoft has generated a 150B token synthetic dataset for Phi-1.5 and it demonstrated 5x efficiency gains. They will likely increase 100x. The promise of faster, cheaper inference and fewer errors is very alluring.
> The success of the TinyStories models also suggests a broader lesson. The standard approach to compiling training data sets involves vacuuming up text from across the internet and then filtering out the garbage. Synthetic text generated by large models could offer an alternative way to assemble high-quality data sets that wouldn’t have to be so large.
And they completely ignore the fact that those models are building their "synthetic stories" based on the knowledge from all that vacuuming text from across the internet, as if now we had already solved the problem of sourcing human data without the need for further vacuuming.
What is the status of synthetic data generated from a model trained on copyright infringing content?
How about the text generated from an open source model, trained on open or licensed data, but having copyrighted material in the prompt for reference.
Does going through an AI model wash copyrights away? Does any hint of copyrighted data in the training corpus or prompt invalidate the right to publish the results? Only when they are similar enough to the copyrighted content? How about when the content itself is pretty common and not unique at all, like a solution to bubblesort?