The biggest takeaway from the article for me was GPT 4 has 100 trillion parameters, 500x more then GPT 3. that type of exponential scaling is going to hit an upper bound really quickly. So when you say it “moves further with each update” you should realize that the improvements are coming from throwing massively more scale at the problem and not some underlying improvement in the technique, and also balance the amount of improvement with the scale itself.
obviously gpt-4 is nowhere near 500x better than gpt-3. let’s say it’s 20% better (very generous imo). can they realistically 500x the model again? and if so, is that going to be worth an additional 4% gain to the original model quality? numbers are completely made up and math is probably wrong but i think i’m hopefully making my point, that diminishing returns will quickly become a blocker with this type of scaling.