The key to running LLM services in prod is setting up Gemini in Vertex, Anthropic models on AWS Bedrock and OpenAI models on Azure. It's a completely different world in terms of uptime, latency and output performance.
Unless you convince MS to let you at the "Provisioned Throughput" model. Which also requires being big enough for sales to listen to you.