Model performance doesn't seem to be monotonically increasing with size. The 84B parameter model is the best on most tasks.
What happened to "you just train it on more data and performance goes up exponentially"? There were so many charts proving this and we could even project at what level of data/compute world-conquering superintelligence would inevitably emerge.