Cerebras had less chip perimeter to hook up external memory I/O and is memory capacity limited with just SRAM. SRAM circuit size hasn't been scaling nearly as well as logic on recent nodes, but if scaling there had continued to from when Cerebras started it may have worked out better.
They'll probably still have to do advanced packaging putting HBM on top to save things.
They could maybe enable some cool real time inference stuff like VR SORA, but that doesn't seem like much of a product market for the cost yet.
Maybe something heavy on inference iteration like an o1 style model that trades training time for more inference, used to process earnings reports the fastest or something zero sum latency war like that will be a viable market. A real time use case that may be viable with cerebras first could be with flexible robotics in ad hoc latency sensitive environments, maybe warfare.
If models keep lasting ~year timescales could we ever see people going with ROM chips for the weights instead of memory? Has density and speed kept up there? Lots of stuff uses identical elements to help make the masks more cheaply, so I don't think you could use something like EUV for ROM where every few um^2 of die is distinct.