For inference, Nvidia's strengths seem to be not very important.
You can do inference on a single GPU. And AFAIK the software stack is not important for inference either. Because you don't have to experiment with the sofware. You just need to get it to run and then you will run it for a long time unchanged. Groq for example runs LLAMA on their custom hardware, correct?
And I expect hardware for inference to become a bigger market than hardware for training.