They all build GPUs/TPUs that can run LLMs, correct?
They all build GPUs/TPUs that can run LLMs, correct?
For inference, Nvidia's strengths seem to be not very important.
You can do inference on a single GPU. And AFAIK the software stack is not important for inference either. Because you don't have to experiment with the sofware. You just need to get it to run and then you will run it for a long time unchanged. Groq for example runs LLAMA on their custom hardware, correct?
And I expect hardware for inference to become a bigger market than hardware for training.
https://xcancel.com/carrigmat/status/1884244369907278106
The only reason you need all those GPUs is because they only have a fraction of the ram you can cram in a server.
With AMD focusing on ram channels and cores the above rig can do 6-8 tokens per second inference.
The GPUs will be faster, but the point is inference on the top deepseek model is possible for $6k with an AMD server rig. 8 H200's alone would cost $256,000 and gobble up way more power than the 400 watt envelope of that EPYC rig.
You can also train Llama on TPU: https://cloud.google.com/tpu/docs/v5p-training