Whether or not that will pan out with "fine tuned" versus "generalized" versions of the same data-eating algorithms remains to be seen, but I suspect the bitter lesson might still apply.
1: http://www.incompleteideas.net/IncIdeas/BitterLesson.html
Whether or not that will pan out with "fine tuned" versus "generalized" versions of the same data-eating algorithms remains to be seen, but I suspect the bitter lesson might still apply.
1: http://www.incompleteideas.net/IncIdeas/BitterLesson.html
AI will be used to select the net to automatically load. Nets will be cached, branch predicted etc.
The future of AI software and hardware doesn't yet support the scale we need for this type of generalized AI processor (think CPU but call it an AIPU.)
And no, GPUs aren't an AIPU, we can't even fit whole some of the largest models on these things without running them in pieces. They don't have a higher level language yet, like C, which would compile down to more specific actions after optimizations are borne (not PTX/LLVM/Cuda/OpenCL.)
In 5 years, the specialized model would still beat the generalized ones. They would just be different from the ones today.
Also, it's not only the prompt engineering. Training ChatCPT was a decades long process, with billions of dollars invested, and it takes a small cluster to run the thing. How would the other models compare if given similar resources?
Besides that, on the context of this thread, the bitter lesson has absolutely no relation to the specialized vs. general purpose model dichotomy. It's about the underlining algorithms, that for this article are all very similar. (And it's also not a known truth, and looking less and less as an absolute truth as deep learning advances, so take it with a huge grain of salt.)
It would also be less robust in the face of exceptional situations that should make it doubt the reliability and relevance of its own training.
For instance, an autonomous driving system that doesn't know the first thing about the zombie apocalypse could put its passengers' lives at risk by refusing to run over "pedestrians".
A specialised diagnostic system might not notice signs of domestic violence that a GP would see. Being able to connect the dots beyond the confines of some specialised field is extremely useful in many situations.
I think that argument works better the other way. I'm outspoken in arguing many of the edge cases in driving are general intelligence problems we're not about to solve by optimising software over a few billion more miles in a simulator, but I don't want that intelligence so generalised that my car starts running down elderly pedestrians because there's been a lot of zombie literature on the Internet lately. I'm pretty confident there's more risk of people dying from that than a zombie apocalypse.
I'd prefer my self-driving car not to learn about the zombie apocalypse, lest it starts running over pedestrians.
Anyway, even if generalist models are truly better at everything, they’ll never be faster or cheaper than their specialist counterparts, and that matters.
The big flaw in this paper, to me, is that no one knows what GPT-4 has been trained on. It could include all the same training sources as the specialized models for all we know. If that's the case, it's a bit curious that it requires prompting to get the accuracy, but again, we'll never know.
I'd go so far as to say this paper isn't even giving any useful information here, aside from the fact that GPT-4 can be made to work well in these specific domains with prompting.
It’s a philosophical position as much as a technical one