I have an article on building micromodels that discusses some of this: https://neuml.hashnode.dev/train-a-language-model-from-scrat...
I have an article on building micromodels that discusses some of this: https://neuml.hashnode.dev/train-a-language-model-from-scrat...
Modern computers don’t have less ram than the old fashioned room size computers. Amount of RAM seems much more analogous here than the size of the hardware running the software. Of course the latter will go down, but the former? I don’t see why.
So in this way you've already lost your bet, as we already know that it's possible to create vastly smaller models with similar performance.
Nevertheless, I have not actually seen a model exhibiting high-level LM capabilities (of the GPT-4 level) that was distilled.
Just because a model can match in perplexity with fewer parameters is not the same as matching the performance as subjectively measured by humanity interacting with the model.
Regardless, this isn't the trend we've seen with software RAM usage - information density is only decreasing as hardware becomes more abundant and comparatively less engineering effort goes into fitting everything into as small a portion as possible.
But all these smaller models are based on those 100b+ models! Like OP on TinyStories relies on larger models twice, to generate and then evaluate. Or the recent wave of small models which are all based on extracting data from GPT-3/4 to borrow the capbilities+RLHFing for free.
You can talk about how interesting it is and how much of an overhang neural nets have in terms of being overparameterized (which is an important AI safety/capabilities issue - the first AGI will be the largest, slowest, and worst one, by a long shot), but the one thing it doesn't tell you is that smaller models can replace big models entirely. Because they are parasitic on said big models, so you still need to train the big models to begin with.
(The smaller models also seem to lose a lot of the qualitative capabilities of larger ones, like meta-learning. Their TinyStories model does only stories. I doubt you could get it to 'dax a blick'.)
I would, but they don't say what their dataset is that I can find anywhere, and the only thing they say about their instruction-tuned is that it's trained on 'publicly available' datasets. You know, the ones where a lot of them turn out under the hood to be drawing from the OA API or other pretrained models in some way or another...
> Especially not with pre-filtered/classified pre-training data.
Indeed not! But what exactly is prefiltering or classifying all that data...?
Currently I can only really use smaller models for micro-tasks like sentiment analysis or classification, but any type of problem solving has to be left to GPT-4.
> But having more fine-tuned control over model parameters is a pattern that will emerge
This assumes parameters do something interpretable. IMO, it’s more that parameters themselves are meaningless/dumb and patterns evolve via interaction of many such units (like in ants)
Maybe for niche use cases like handling customer support requests small models will do well, but GPT4 is not going to be condensed to a few billion anytime soon imo.
Hear the other timeline from Darius Amodei. https://a16z.com/improving-ai/ We may be entering a regime of multi-billion training costs on custom hardware. Open-source AI will be like open-source CPUs. Cool but the real thing even at the smallest scale comes out of monopsonized (is that a word?) inaccessible infrastructure
Could we have libraries of small models that are all experts on niche subjects that basically collaborate?
The advantage is that it’s easier to scale and distribute lots of small models on commodity hardware than it is to run giant models that require heaps of contiguous ram.
I’d love to have some time to play with versions of this.
If I do experiments, I’m making models specialized from foundational to programming to specific language to specific tasks. Each pass will have training data so it picks up general concepts. Then, it will be trained to focus more on what it needs. Eventually, what’s slushing around in its neurons should be just the right information. Or it won’t work.
Anyway, my concepts that I published are at this link:
https://gethisword.com/tech/exploringai/index.html
(See alternative models section.)
Since it's not, it isn't.
Now that I've said it, I wonder if there's any evidence human brains work that way.