I don't think that's a totally fair comparison.
We don't have overcomplicated distributed training infra. They are pretty much as complicated as needed.
Scaling vertically (having a bigger chip) is very hard. There are tons of tradeoff when making the chip, and it's overall an insanely complex problem. That's why Cerebras is a 8 years old company and yet you would be hard pressed to find anyone using them still.
And even if you give me a Cerebras chip that works perfectly, it will still be much easier for me to buy two of those chips and link them together in distributed training mode, than it will be for Cerebras to build a chip that is 2x the size.
The scale of the current generation of clusters to train models the size of GPT-4 are in the range of 25,000+ GPUs with 80GB of memory each, so no matter your chip size, complicated distributed infra is a necessity. Even assuming everything on Cerebras' marketing page is fully accurate, you would still need to distribute the training over 500+ of those massive chips to replicate a 25k GPU cluster.