What?
No, what the article said was:
> At that pace, it doesn’t take long before the cumulative effect of all of these fine-tunings overcomes starting off at a size disadvantage.
>Indeed, in terms of engineer-hours, the pace of improvement from these models vastly outstrips what we can do with our largest variants, and
> the best are already largely indistinguishable from ChatGPT.
^ The author did not note chat gpt is better, the author claims that the 7B koala model is 'largely indistinguishable from ChatGPT'.
and:
> While ChatGPT still holds a slight edge, more than 50% of the time users either prefer Koala or have no preference.
Which is highly misleading.
The koala authors rated their model by passing it to 100 people using the mechanical turk, noting:
> To mitigate possible test-set leakage, we filtered out queries that have a BLEU score greater than 20% with any example from our training set. Additionally, we removed non-English and coding-related prompts, since responses to these queries cannot be reliably reviewed by our pool of raters (crowd workers).
So.
What you have is a model that performs pretty well for some trivial conversational prompting tasks.
What you DO NOT have, is something that is: "largely indistinguishable from ChatGPT".
Anyway, regardless of the creative interpretation of the authors writing, the point that I'm making is that your point:
> So, the winning strategy is whatever strategy allows your model to compound in quality faster and to continue to compound that growth in quality for longer.
Is founded on the assumption from the post that:
> While the individual fine tunings are low rank, their sum need not be, allowing full-rank updates to the model to accumulate over time.
ie. If you fine tune it enough, it'll get better and better in an unlimited fashion.
Which is provably false.
If I have a 10-parameter model, there is no possible way that the accumulation of low rank fine tunings will make it the equivalent of a 7B, 13B of 135B model.
It is simply not complex enough to do some tasks.
Similarly, smaller models like 3B or 7B model, appear to have an upper bound on what is possible to achieve with them regardless of the number of fine tunings applied to them, for the direct and obvious same reason.
There is an upper bound on what is possible, based on the model size.
The 'best' size for a model hasn't really been figured out, but... I'm getting pretty sick of people saying these 7B models are as good as 'ChatGPT'.
They. Are. Not.
People will go to the best models, with the best licenses, but... those models are, it seems, unlikely to be fine tuned smallish models.