Inference at scale can be complex, but the complexity is manageable. You can do fancy batched inference, or you can make a single pass over the relevant weights for each inference step. With more models using MoE, the latter is more tractable, and the actual tensor/FMA units that do the bulk of the math are simple enough that any respectable silicon vendor can make them.
In any event, one can make some generalizations about the companies involved. Nvidia makes excellent hardware that everyone wants and charges large enough markups that their margins are around 90%. AMD is chasing the big buyers to sell their products. Google spends a lot and is a mature company, and they seem uninterested in selling chips that compete with Nvidia, but they certainly care about revenue and profit. OpenAI, Anthropic, etc and, perhaps oddly, Meta don’t seem to care too much about profit, but they certainly spend enough money that it would help them to get more bang for their buck. Alibaba, etc buy whatever Nvidia gear they can get, but they have a lot of incentive to find a domestic supplier, and Huawei seems quite interested in becoming that supplier. And there are plenty of US startups (Cerebras and others) going after the inference market.
Maybe someone knows which providers are selling access roughly at cost and what their prices are?