Outperforming larger language models with less training data and smaller models
blog.research.google
blog.research.google
And part of the reason why single-hidden-layer networks aren't enough even in continuous memoryless Euclidean cases is, again, because of how loss functions work; you're unlikely to converge on a good approximation with very few hidden layers.
I would suspect what you have instead is a single model attached to data sources, where the model doesn’t have to have so much compressed fact, and instead can rely on higher level summary.
TLDR, the essay explores how LLM could evolve into front-end routers that connect users with specialized tools, leading to a future where federated models determine the best-suited system to answer specific queries. Not too different from today's federated search approaches.
I ended up with the premise that each of them had their relative strengths and weaknesses, and it would actually be best to use all of them, but only in their own areas of strength. Then have them use something akin to a shared blackboard where they could all read the results from the other systems and write their results as well. That the sum of all the available algorithms working together, each in the areas in which they were best, would result in better outcomes than any one algorithm could achieve on its own.
My professor was not impressed. I only got a C.
Now, the story of how my dad had to work his ass off to finally get the College of Engineering to force the professor to actually give me a grade that he owed me, when all the professor really wanted to do was focus on his new job at one of the big airlines -- well, that's a story for another time.
Also interesting is that this isn't an inconceivably clever and out of the box idea. It shows there's still a lot of low hanging fruit to explore, and the future of LLMs isn't set in stone yet. Could be that the real deal is a mixture of experts trained in this style. It's exciting that it feels the holy grail is close to being achievable if only the right combination of ideas is tried.
Why is that?
As I understand it, rather that training the smaller model to produce the CoT/rationale as a part of the (decoded) response, it actually has two output layers. One output is for the label and the other is for the rationale. The other layers are shared, which is how/why the model is able to have an improved "understanding" of which nuances matter in the labeling task.
There are still many low hanging fruits. I have probably seen dozens of variations of chain-of-thoughts, tree-of-thoughts, graph-of-thoughts, self-ask, self-critique, self-plan, self-reflect, etc.
I would want to ensure that the system gains and retains the fundamental structures and skills that you know it needs to effectively and accurately generalize. While maintaining those things you then feed it lots of diverse data to learn the exceptions and ways the skills can be combined. But somehow you need to ensure those core skills and knowledge throughout. Maybe you could do that just by including outputting those understandings or manipulations in addition to the final answer. Similar to what the paper does.
For example, a code generation model might be required to output a state machine simulation of the requested program.
AlpaGasus: Training A Better Alpaca with Fewer Data https://arxiv.org/abs/2307.08701
Textbooks Are All You Need II: phi-1.5 technical report https://arxiv.org/abs/2309.05463
Skill-it! A Data-Driven Skills Framework for Understanding and Training Language Models https://arxiv.org/abs/2307.14430
I can very well believe that empirical this can work. (I haven't checked the literature.) My point was merely that given my priors, this isn't intuitive.
Or did the authors count the amount of training data for the LLMs to the required training data for the destined/task-specific models?
https://blogger.googleusercontent.com/img/b/R29vZ2xl/AVvXsEj...
Yes.
They are counting the amount of data you need to collect to solve your problem. I can grab a pretrained LLM, and the data I have to collect in that instance is what I need to fine tune it.
The datasets behemoth LLMs are trained on include a lot of noise that derail progress. They also just contain a lot of irrelevant knowledge that the LLM has to learn or memorize so an obscene amount of parameters is required.
When you're not trying to teach a language model the sum total of human knowledge and you provide a high quality curated dataset, the scale barrier is much lower.
But LLMs do gain useful "emergent" properties from training on massive lower quality datasets.
It's just that given what we know about neural networks, it's often easier and simpler and more effective to increase the amount of training data than to change anything else.
If someone knew all the math and science in Wikipedia, for example, I think they'd probably be forgiven for not knowing every regionalism.
And if you want it to be superhuman then you're by definition not capable of knowing what's important, I guess.
But that looks like a 'smeared' out probability distribution on the next token. Not like the text produced by an unsure human.
Nevertheless, human feats can act as an existence proof of what is possible. Including of what might be possible for a neural network.
(I'm not sure whether a large language model necessarily needs to be a neural network in the sense of a bunch of linear transformations interleaved with some simple non-linear activation functions. But for the sake of strengthening your argument, let's assume that we are assuming this restrictive definition of LLM.)
I wonder what ideas there are about how one could estimate the optimal size.
Apple ships the Mac Studio which support up to 144GB of usable GPU memory.
Would be amusing if they were to release a Mac Pro with 300+ GB and dominate the LLM serving space.
- HW design
- UI
- UX
That Macs come with awesome VRAM is only because #3. Apple has no explicit incentive to make LLM-ready machines.
Is there any framework that can batch LLMs on Metal? I don't think GGML or MLC have it yet.
Otherwise that is just another reason they wouldn't be good for LLM hosting at this moment.
Anyway, the real disruptor is Intel. They could theoretically barge in with a 2x48GB Arc card and undercut the market that AMD/Nvidia refuse to dive into because of their Pro card clients.
Its really a shame that Apple (and AMD/Intel or pretty much any other infrence vendor) are not directly contributing to llama.cpp. The feature set is amazing and growing at a stunning pace.
Absolutely not. The LLAMA license [1] is clear that it's not open source. It's for non-commercial, research only, and only by explicit permission from Meta. The weights were leaked, on 4chan [2], illegally, according to the license. Very very few people are using it legally. This interpretation is clear from its wording, and also matches the interpretation of our flock of lawyers.
[1] https://github.com/facebookresearch/llama/issues/266
[2] https://levelup.gitconnected.com/metas-chatgpt-is-now-illega...
But, Llama 2 is not open source [1].
[1] https://blog.opensource.org/metas-llama-2-license-is-not-ope...
Or it'll get dropped within a year.
This is a make it or break it problem for Google and they will get it right or they won’t get it at all.
I use Google for basic bitch things like finding a company website or what time is it in Tokyo.
In that respect it's not replacing all my searches, but it seems to be replacing the "hard ones" where it's hard to compete with Google at a disproportional rate. If that is actually the case, it'd spell bad news for Google whether or not it kills search - if it becomes cheaper/easier to compete by offering a mix of less complete search with an OpenAI integration, it opens the door for far more attempts at competing with them.
Similarly, product searches are second largest category in which they make tonnes of money. This is also done by people as they don't really like to search on amazon or other ecommerce sites. This is also a huge money spinner for them.
Both of these are not going anywhere as both of these are tactical spends.
Now let us come to long tail. These are again big money and are at risk for Google. However, you have to understand that Goog ads are clicked by most tier 2 users. We, techies, do not really click at ads. We go for organic ranking (mostly). We are the base of chatGPT right now. Tier 2 and lower users don't really use chatgpt.
Even if they do, they would not do it for product discovery or site discovery as it has too much friction: go to chat.openai.com, type in your question, it responds in slow, jerky manner vs just type in browser bar what you are thinking.
To top it, Chatgpt also has stale data. Moreover, it is heavily lobotomized to not give any controversial or edgy answers. This curtails usefulness of chatgpt.
It's worth mentioning that the Google you're talking about was way way different than it is today. Google Search was a startup. Google Search + Email was a small company. Google Search + Email + YouTube was a midsize firm. Now they're a humungous megacorp that's slow to make necessary changes when there's a paradigm shift like LLMs.
OAI had a big head start for sure.
Edit: Apple is certainly taking the lion's share of profits in the phone space. Since Google isn't investing as heavily anymore in Android, the development has slowed considerably.
0: https://www.statista.com/statistics/272698/global-market-sha...
Lately you could argue they're being overtaken by their competitors, especially in terms of productization. But they still hold pretty cards IMO.
In AI they were giants in what looks post-ChatGPT like a mediocre field. Their search now is looking very jaded.
The trajectory here isn't remotely like their past performances; it's not a safe bet to assume they'll win through with Bard or anything else.
The agility of OpenAI and the revolutionary impact of gpt3+ has made the former incumbents like Google look like posturing, self-satisfied, giant lumbering has-beens. They aren't getting back on top without massive internal changes.
Acquisition doesn't count
Also, with more recent changes I can get decent summaries much easier (and at no cost to me) for every document in a google drive.
In my day to day, those features can be pretty powerful tool, even if its not at GPT4 level 6 months after GPT4 and Bard were released in March.
To my knowledge, no one is really integrating search at the level that Google is with generative language models.
I would not be surprised if ChatGPT moves some of the plugins to the free tier as well, to stay competitive over the longer term.
Without their papers, there wouldn't be GPTs.
All this AI angst going on right now is ridiculous.
Yeah, every technology gets a hype cycle these days.
But not every technology has research cycles this fast.
In fact, I can't think of anything that's had research cycles this fast.
Huh? My answer as a human would have been "race track", as that is probably "where the people are" (during a race).
Did I fail? Am I a poor language model? Or is the whole thing just tea leaf reading to begin with?
But my point is really that speaking of the correct answer with a question as vague and open to interpretation as this one is absurd.
You're being asked a multiple choice question:
- There's an implication of intended vagueness/indirection and multiple possibly suitable answers.
- Textual entailment is a very common type of multiple choice format, which would lean towards a more plain semantic correlation
In a more natural conversational setting, you'd get a different answer:
https://chat.openai.com/share/00fed9d6-e3de-4319-9c76-ae1800...
https://chat.openai.com/share/dc9a796c-870c-44ee-b421-31c24b...
No offense to you, but I never would have picked "race track". That answer doesn't make much sense since the original prompt doesn't mention anything about racing, and the definition of a populated area fits well with the question being asked.
“Populated areas” does not mean much. Montana is a populated area. It’s not a highly populated area but it’s populated alright.