I'm not sure I understand why they would route it to other models. It can't be that they don't have enough compute. Maybe worried that the answer from their own models would be bad? Doesn't really make sense, but I could be missing something.
How do you know?
Every cloud provider (AWS, Azure etc) is struggling with meeting LLM demand.
Source: first hand info
Because there was an article written in Bloomberg about how they're looking to sell their extra compute [0]
fair q. call_ ids and gAAAAA blobs arent damning but the rs_ reasoning ids embed a unix timestamp that matches the session to the second, then OpenAI's 819x marker. plus the summary is in OpenAI's summarizer voice.
its just a best guess.