Can generalist foundation models beat special-purpose tuning?
arxiv.org
arxiv.org
One could argue universal general intelligence is simply the ability to optimally specialize as necessary.
I think one aspect that is overlooked when people are involved is that we create specialized or fine-tuned models precisely for the reason that a general approach didn’t work well enough. If it had, we would have stopped there. So there’s a selection bias in that almost all fine-tuned models are initially better than general models, at least until the general models catch up.
There’s little doubt that GPT4 is going to be most capable at most tasks either OotB or with prompt engineering as here. But that doesn’t mean it’s the right approach to use now.
same will happen for temporal information that stratifies and changes, ie, asking what a 1950s scientist understands about aquantim physics.
It would simply have two different submodels inside its ubermodel.
Interesting philosophic topic how to ensure objectives aren't affected too much. You can make this mind experiment even as a human. There are parts of our goal system we might want to adjust, and parts we'd be horrified to, being very reluctant to accidentally risk affecting it.
Whether or not that will pan out with "fine tuned" versus "generalized" versions of the same data-eating algorithms remains to be seen, but I suspect the bitter lesson might still apply.
1: http://www.incompleteideas.net/IncIdeas/BitterLesson.html
AI will be used to select the net to automatically load. Nets will be cached, branch predicted etc.
The future of AI software and hardware doesn't yet support the scale we need for this type of generalized AI processor (think CPU but call it an AIPU.)
And no, GPUs aren't an AIPU, we can't even fit whole some of the largest models on these things without running them in pieces. They don't have a higher level language yet, like C, which would compile down to more specific actions after optimizations are borne (not PTX/LLVM/Cuda/OpenCL.)
In 5 years, the specialized model would still beat the generalized ones. They would just be different from the ones today.
Also, it's not only the prompt engineering. Training ChatCPT was a decades long process, with billions of dollars invested, and it takes a small cluster to run the thing. How would the other models compare if given similar resources?
Besides that, on the context of this thread, the bitter lesson has absolutely no relation to the specialized vs. general purpose model dichotomy. It's about the underlining algorithms, that for this article are all very similar. (And it's also not a known truth, and looking less and less as an absolute truth as deep learning advances, so take it with a huge grain of salt.)
It would also be less robust in the face of exceptional situations that should make it doubt the reliability and relevance of its own training.
For instance, an autonomous driving system that doesn't know the first thing about the zombie apocalypse could put its passengers' lives at risk by refusing to run over "pedestrians".
A specialised diagnostic system might not notice signs of domestic violence that a GP would see. Being able to connect the dots beyond the confines of some specialised field is extremely useful in many situations.
I think that argument works better the other way. I'm outspoken in arguing many of the edge cases in driving are general intelligence problems we're not about to solve by optimising software over a few billion more miles in a simulator, but I don't want that intelligence so generalised that my car starts running down elderly pedestrians because there's been a lot of zombie literature on the Internet lately. I'm pretty confident there's more risk of people dying from that than a zombie apocalypse.
I'd prefer my self-driving car not to learn about the zombie apocalypse, lest it starts running over pedestrians.
Anyway, even if generalist models are truly better at everything, they’ll never be faster or cheaper than their specialist counterparts, and that matters.
It’s a philosophical position as much as a technical one
The big flaw in this paper, to me, is that no one knows what GPT-4 has been trained on. It could include all the same training sources as the specialized models for all we know. If that's the case, it's a bit curious that it requires prompting to get the accuracy, but again, we'll never know.
I'd go so far as to say this paper isn't even giving any useful information here, aside from the fact that GPT-4 can be made to work well in these specific domains with prompting.
I think one of the problems this paper ignores isn't whether you can get a general purpose model to beat a special purpose one, but whether it's really worth it (outside of the academic sense).
Accuracy is obviously important, and anything that sacrifices accuracy in a medical areas is potentially dangerous. So everything I'm about to say assumes that accuracy remains the same or improves for a special purpose model (and in general that is the case- papers such as this talk about shrinking models as a way of improvement without sacrificing accuracy: https://arxiv.org/abs/1803.03635).
All things being equal for accuracy, the model that performs either the fastest or the cheapest is going to win. For both of these cases one of the easiest ways to accomplish the goal is to use a smaller model. Specialized models are just about always smaller. Lower latency on requests and less energy usage per query are both big wins that affect the economics of the system.
There are other benefits as well. It's much easier to experiment and compare specialized models, and there's less area for errors to leak in.
So even if it's possible to get a model like GPT-4 to work as well as a specialized model, if you're actually putting something into production and you have the data it almost always makes sense to consider a specialized model.
For prototyping and for smaller use cases, it makes a lot of sense to use a much more general model. Obviously this doesn't apply to things like medicine, etc. But for much more general things like: check if someone is on the train-tracks, or number of people currently queuing in a certain area, or if there's a fight in a stadium; I think multi-modal models are going to take over. Not because they're efficient, or particularly fast; but because it'll be quick to implement, test, and iterate on.
The cost of building a specialized model, and keeping it up to date, will far exceed the cost of an LVM in most niche use cases.
If you expect your model to be used a lot, and you don't have a way to distribute that pain (for instance, having a mobile app run the model locally on people's phones instead of remotely on your data centers) then it ends up being a cost balancing method. A single DGX machine with 8 GPUs is going to cost you about the same as a single engineer would. If cutting the model size down means you can reduce your number of machines that makes increasing headcount easier. The nice thing about data cleaning is that it's also an investment- you can keep using most data for a long time afterwords, and if you're smart then you're building automated techniques for cleaning that can be applied to new data coming in.
If you were approaching this problem anew today, you'd probably try with GPT-4 and Claude, and then see what you could achieve by finetuning GPT-3.5.
And, yes, for a given level of quality, the finetuned GPT-3.5 will likely be cheaper than the GPT-4 version. But for radiology notes, perhaps you'd be happy to pay 10x per even if it were to give only a tiny improvement?
To put it another way, the researchers at Rad AI consumed every paper that was out there including very cutting edge stuff. This included reimplimenting GPT-2 in house, as well as many other systems. However, we didn't have the same data that was used by OpenAI. We also didn't have their hyperparameters (and since our data was different it's not a guarantee that those would have been the best ones anyways).
So with that in mind it's possible that Rad AI could today be using their own in house GPT-4, but specialized with their radiology data. In other words them using a specialized model, and them using GPT-4, wouldn't be contradictory.
I do want to toss out a disclaimer that I left there in 2021, so I have no insights into their current setup other than what's publicly released. However I have no reason to believe they aren't still doing cutting edge work and building out custom stuff taking advantage of the latest papers and techniques.
They introduce medprompt, which combines chain of thought reasoning, supervision, and retrieval augmented generation to get better results on novel questions. This is a cool way to leverage supervision capacity outside of training/fine tuning! But they compare this strategy, applied to GPT-4, with old prompting strategies applied to MedPalm. This is apples and oranges -- you could easily have taken the same supervision strategy (possibly adapted for smaller attention window etc) to get a closer comparison.
Second, MedPalm is a fine tuned 500b parameter model. GPT-4 is estimated to be 3-4x size. So the comparison confounds emerging capabilities at scale, in addition to prompting impact, with anything you can say generally about the value of model specialization.
source: six years in the book industry in California
Theoretically, all these date could be cataloged in a way where you can type in symptoms x y z, demographic information a b c, medical history d e f, and get a confidence score for a potential diagnosis bases on matching these factors to results from past association studies. A system like that I imagine would be easier to feed in true information and importantly would also offer a confidence estimate of the result. It would also be a lot more computationally simpler to run I’d imagine given the compute required to train a machine learning model vs keyword search algorithms you can often run locally across massive datasets.
Happy to help!
A good analogy is building a webapp - would you prefer to hire a developer with 30+ years experience in various CS domains as well as a PhD or a specialized webdev with 5 years experience at a tenth of the rate?
https://www.medrxiv.org/content/10.1101/2023.07.13.23292418v...
Which is still a significant feat. But I doubt it's doing anything more than the sort of semantic lookup that is already well-known to be within its capabilities.
GPT-4 almost disappeared last week. Aside from the obvious AI bus factor and "continuity of science" concerns, I find it absurd that so many papers and open source ML/LLM frameworks and libraries exist that simply would not work at all if not for GPT-4. Have we simply given up?
I thought this was hacker news.
GPT-4 is still the flagship model "of humanity"; it is still the best model that is publicly available. The intent of the paper is to determine whether the best general model can compete with specialized models - how do you do that without using the best general model?
I have high hopes that models will end up a bit like Linux servers, where everyone is building on open foundations.
'Closed source' science happens in every industry. While you may not think of what DOW is doing at science, it is and huge companies like that progress by their internal workings.
What you're stating is something else, that maybe we should fund sciences, such as AI more, but we really did that in the past and it had it's own fits and starts. Then transformers came out of Google and private industry has been leading the way in AI. If you want to keep up with cutting edge and the 10s of millions needed to train them, you'll need to use these private models.
Hackers hack stuff from private companies all the time. Hackers use proprietary software. There is no purity test we have to pass.