All or a vast majority of of the cerebras manufacturing capacity was going to a few companies that aren't publicly available inference providers on openrouter, for their own internal use.
or
The asking price of the S-3, no matter how speedy it might be, for small/medium size customers made it economically prohibitive to purchase and use to sell public inference vs. buying more common nvidia b200 or whatever.
imo the issue is that most openrouter demand is inauthentic activity (things that anthropic and openai models will refuse to do like pretend to not be bots when interacting with humans)
So I plugged 288 trillion tokens/month (OpenRouter's current rate), 500 billion MoE model average, and the math comes out to be around 620 B200 GPUs minimum.
So basically, OpenRouter's volume must be absolutely tiny compared to the volume hyperscalers are getting.
[1] https://x.com/ren_stocks/status/2056946641815396718?s=20
A Ferrari can seat up to 4 people[0], which is about the same as the average car. Capacity doesn't change much.
Meanwhile, a bus/subway system is meant to support millions of people. Tokyo's metro has to support up to 37 million people. You can't do that with Ferraris.