at the very least, even if that's not the case, inference will be drastically less gpu heavy by then I suspect.
at the very least, even if that's not the case, inference will be drastically less gpu heavy by then I suspect.
"We find that current large language models are significantly under-trained, a consequence of the recent focus on scaling language models whilst keeping the amount of training data constant."
It's difficult to evaluate a LLM's performance as it's all qualitative, but Meta's LLaMA has been doing quite well, at even 13B parameters.
I think what we have access to is a fair bit slower.
If you focus on english only, this can easily reduce the paramters 5fold
LLMs seem to be comfortable with hundreds of programming languages, DSLs and application specific syntaxes so how does supporting a couple more natural languages become so expensive?
I see how more training data would be needed, but I don't understand how that maps to a greater parameter count.