what does he mean when he says 1T or bust? Is he referring to 1 trillion parameters? Are you saying that GTP-3 has 2.7 trillion parameters? Does it mean that to get to GPT-3 level it needs 100x more amount of dataset?
I interpreted it as aspiring to a trillion paramters but I'm not sure.
GPT-3 has 175 billion parameters. So they need to scale by 64x. They already have a comparable amount of data than what was used by OpenAI, so it's about scaling the numbers of GPUs.
I see so that means this GPT Neo is 64 less powerful?
Accuracy and numbers of parameters don't scale linearly together. It varies widely depending on exactly what you are measuring accuracy on as well etc.
But a very approximate rule of thumb would be to say that accuracy scales with the log of the parameter count (for the same architecture).