Why do they only host small models rather than the 2.4T version? Is the I/O and interconnect between the wafers bad due to the limited beachfront relative to the massive size of the chip?
The CEO was on Gradient Dissent a couple years ago: https://www.youtube.com/watch?v=qNXebAQ6igs
[1] https://www.cerebras.ai/blog/accelerating-gpt-5-6-sol-ultraf...
It would look bad for cerebras if other people are hosting the 27b version and show a higher TPS than cerebras.