My take: Today, using AI for search-related problems is still not cost-effective for most use cases. That being said, the landscape is evolving quickly. First, in some areas, an individual search creates more value than in others. An individual consumer doing a Google Search is totally different from a lawyer searching for reference material. Areas where the individual search creates more value can already benefit from AI today. Second, LLMs become exponentially cheaper, driven by more cost-effective computing but also more cost-effective models. Look at the pricing of GPT4o-mini vs GPT4 (the original). The models are comparable in performance for many search-related problems, but the price has decreased by 200x in 1.5 years ($0.15 vs $30 per 1M token). If that price trend continues, more and more search use cases will benefit from AI.
Anyone who wants to use CUDA continues to be forced to buy Nvidia hardware.
Now people say CUDA is better optimised than ROCm. Fewer claims that ROCm never works, more that it leaves performance on the table.
You've just offered that people are using AMDGPU in production for inference. I think some are using it for training.
That software moat is looking rather empty to me. And good riddance, the cuda programming model is nasty.
The MI300A is gorgeous. Massive APU thing. Single block of fast memory that x64 threads and gpu kernels have atomic RMW access to. It is the evolution beyond the separate accelerator concept and the whole language patterns around "offloading" kernels is kind of nonsense on it. The gpu and x64 cores are peers.
Viable alternatives to cuda are coming online at the same time as the hardware moves beyond it. It's a liability, not a moat.