Wait, so there's a way to make a model as smart as GPT but with less parameters? Isn't that why it's so good?
"We find that current large language models are significantly under-trained, a consequence of the recent focus on scaling language models whilst keeping the amount of training data constant."
It's difficult to evaluate a LLM's performance as it's all qualitative, but Meta's LLaMA has been doing quite well, at even 13B parameters.
I think what we have access to is a fair bit slower.