Rich webapps hadn't been invented. Smartphones? If you're lucky your flip phone might have a colour screen. If you've got money to burn, you can insert a PCMCIA card into your Compaq iPAQ and try out this new "802.11b" thing. Java was... being Java.
Almost all the software out there - especially if it had a GUI, and a lot of it did - was distributed as binaries that only ran on x86.
A vast amount of code was only intended to compile and run on a single OS and architecture (circa 2000, that was usually x86 Win32; Unix was dying and Wintel had taken over the world). If some code needed to be ported to another platform, it was as good as a from-scratch re-write.
[0] in case you wanted to use the thing in Visual Basic, which you very well might.
"had". That's what helped prop up their monopoly but it didn't last. These days if can't run your software on another architecture, like ARM, you can run at least on AMD. AMD can basically run the same software as Intel. This isn't the situation for NVIDIA vs everyone else, so far.
There were also ISA extensions. Even if Intel had trouble competing on existing code, they would often extend the ISA to gain a temporary advantage over their competitors by enabling developers to write more optimal code paths that would run only on Intel’s most recent CPUs. They have done less of that ever since the AVX-512 disaster, but Intel still is the one defining ISA extensions and it historically gained a short term advantage whenever it did.
Interestingly, the situation is somewhat inverted as of late given Intel’s failure to implement the AVX-512 family of extensions in consumer CPUs in a sane way, when AMD succeeded. Intel now is at a disadvantage to AMD because od its own ISA extension. They recently made AVX-10 to try to fix that, but it adds nothing that was not already in AVX-512, so AMD CPUs after Zen 3 would have equivalent code paths from AVX-512, even without implementing AVX-10.
Thats where Nvidia learned to "optimize" Cuda software path. Single threaded x87 FPU on SSE2 capable CPUs.
https://arstechnica.com/gaming/2010/07/did-nvidia-cripple-it...
https://www.realworldtech.com/physx87/3/ "For Nvidia, decreasing the baseline CPU performance by using x87 instructions and a single thread makes GPUs look better."
They doubled down that approach with 'GameWorks' crippling performance on non Nvidia GPUs, Nvidia paid studios for including GameWorks in their games.
And what does pytorch et al. use under the hood? cuBLAS and cuDNN, proprietary libraries written by NVidia. That is where most of the heavy lifting is done. If you think that replicating the functionality and performance that these libraries provide is easy, feel free to apply for a job at NVidia or their competitors. It is pretty well paid.
I never claimed it was easy. I meant in my opinion it is in the order of 10s of millions dollars of investment, not a trillion dollar CUDA moat that people comment here.
Anyone that wants off the shelf parts at scale is going to turn to Nvidia.
And we just gave them billions in tax dollars. Failing upwards...
Nvidia is helping power the next generation of big brother government programs.