They are but it takes a lot of time.
Most of the big players - Google, Meta, OpenAI, Amazon, and Microsoft all are actively developing TPU/NPUs that would be used instead of the H100/A100's everyone is using for machine learning.
Google (tensorflow/jax), Meta(pytorch), Microsoft(onyx), Openai(triton) and Apple (mlx) each have software stacks for optimizing models for multiple platforms.
It takes a lot of time to develop the silicon and software stack. As a result everyone is using H100's in the interum until the hardware/software catches up. Google has been using their TPUs already.
There's other companies like Groq that are also developing NPU/TPU like devices.
And some PoC work: https://vortex.cc.gatech.edu/publications/hotchips-poster.pd...
Neither are specialized TPU/NPUs but they do fast vector operations.
When you include Nvidia's huge margins, this bet is less risky, but it's still only one that the largest companies can make.
Software alone cannot do it, you need to make a bet on hardware.
It's so challenging, capital intensive, and takes so long to bring these things online (especially if it's not your core competency), that but by the time you got something working in this paradigm, it's possible that a better paradigm/approach will have emerged.
Joking aside, it is quite expensive and I'm pretty sure you'd run into legal issues due to vertical integration, depending how you perform it. I'll put it this way, a few years ago the headlines were about how China was poaching TSMC workers and offering huge salaries[0,1]. I haven't seen Chinese chips become competitive yet, so it looks like >5 years.
Beyond that, it's more than the chip. AMD has caught up to Nvidia in hardware. But no one is rushing to buy AMD cards because they are still not as good. Nvidia's secret sauce is Cuda (MKL is still an advantage to Intel). The naivity of the Tiny Corp was thinking that everything could be resolved in a few weekends of hacking. But Cuda is deep in a lot of projects. Like a project's backend's backend's backend deep. You got decades of engineers using Cuda and the huge momentum around it. While most programmers will never touch Cuda or GPU code, almost everyone touches things that connect with them (and this is every single day for many people. We could also say the same thing about optimization libraries like MKL). It is something that looks simple because we don't talk about it much or haven't had the experience with. But I assure you, GPU programming is a whole other world. Optimized programming is a very different style of programming and you need a different framework of thinking and problem solving. GPUs add another lay of complexity on top of that. It's why you'll hear so many people complain about writing kernels, but damn, the results speak for themselves.
So you gotta match on the hardware. Then you got to develop great software. Then you got to get that software into the other software that everyone else is using. And you gotta convert people along the way, getting them to turn from a thing they already know and have experience with to a completely new thing.
```edit
The disadvantage of being a first mover is you got to invent everything yourself and page the path, letting others follow. But the disadvantage of being a follower is that to get people to use your road or lane you can't just be equal, you have to be *better*. And you usually have to be significantly so. Momentum is a really powerful force and I think it is highly undervalued.
I think Nvidia is safe for the next few years. They aren't slacking and relying on their momentum. They're still pushing very hard, which only makes it harder for competitors. It's hard to displace a sleeping giant. It's even harder when that giant is fighting back.
```
> And how long would that take?
A long fucking time.
[0] https://news.ycombinator.com/item?id=24129861
[1] https://www.reuters.com/technology/taiwan-raids-chinese-firm...
MI300x is excellent and I'm buying them. =)