(it's all blurry)
(it's all blurry)
https://www.cerebras.net/wp-content/uploads/2023/03/Downstre...
https://www.cerebras.net/wp-content/uploads/2023/03/Scaling-...
https://www.cerebras.net/wp-content/uploads/2023/03/Scaling-...
EDIT: Looks like it scores better with less training - up until it matches GPT-J/Pythia/OPT and doesn't appear to have much benefit. It maybe scores slightly better then GPT-J which is pretty "eh", I'm not sure if GPT-J level performance is really useful for anything? NeoX 20B outperforms it in everything if you don't care about the amount of training needed.
Does the better performance for less training matter if that benefit only applies when it's only performing a lot worse then GPT-J? It appears to lose it's scaling benefits before the performance is interesting enough to matter?
edit: scratch that, it seems the AJAX endpoint returns 504 more often that not.
But I'm not sure anymore that it wasn't initially blurry... Perhaps I'm hallucinating, like large language models.
Current image displayed is https://www.cerebras.net/wp-content/uploads/2023/03/Scaling-... , will see if it changes.
Really, really bad mark on whoever is in charge of their web marketing. Images should never look that bad, not even in support, but definitely not in marketing.
edit: so this post is more useful, 4k res using Edge browser