- Crypto hardware needed SHA256 which is basically tons of bitwise operations. That’s way simpler than the tons of matrix ops transformers need.
- NVidia wasn’t focused on crypto acceleration as a core competency. There are focussed on this, and are already years down the path.
- One of the biggest bottlenecks is memory bandwidth. That is also not cheap or simple to do.
- Say they do have a great design. What process are they going to build it on? There are some big customers out there waiting for TMSC space already.
Maybe they have IP and it’s more of a patent play.
(I mention crypto only as an example of custom hardware competing with a GPU)