Microsoft absolutely runs their own models on their own hardware, at scale, and they have done so for years just like every other hyperscaler -- Project Brainwave was first publicly talked about as far back as 2018. The generative LLM craze is a recent phenomenon in comparison. They are absolutely going to go all in on putting AI functionality in Bing, in Excel, in Windows, etc etc. To do that, you need hardware.
None of this is really strange. It also wasn't strange when Google announced H100 systems while also pushing TPUs they developed. Microsoft has Jensen on stage because customers of Microsoft Azure demand Nvidia products. Customers of Google Cloud demand Nvidia products. So, they provide them those products, because not providing them loses those customers. It's that simple. Everyone involved in these deals acknowledges this.