betting on a compute bottleneck sounds like a recipe to get thrashed when the bottleneck relieves itself.
At the investment scales being discussed, CUDA/architecture and other advantages do not matter - you could spend 1 billion on building a new chip architecture. The ram/fab inputs have been a commodity market for years. Heck, even the model bottleneck doesn't seem real when it's only 1-4 billion or less to get a state of the art model.
At some point the compute bottleneck will be relieved, you can see NVidia hedging their strategy with both open models and on-device chips targeted for local inference. The 200 dollar a month plan will absolutely be taken over by local hardware in the future.