Jensen Huang did. When people think of tech visionaries they think of Jobs or Musk, but Huang is just as great. He's been bang on the money on the future of this industry since he founded Nvidia, which is how they managed to not just consistently stay ahead of their competition (ATI) or put them out of business (3dfx), but leapfrog them (AMD) by branching in several fields (AI/ML, PhysX, computer vision, self driving, compute etc.) He saw early on that GPUs should push into general-compute and not just be for video games, and he executed well on that.
There are interviews on Youtube with Huang at Standford IIRC, where he discusses his vision of the GPU industry from the early days of Nvidia. Check them out, the guy's not your typical CEO suit focused on the share price, but he's basically a tech genius.
So, to answer your other question about why only Nvidia manage to win compute and not the other GPU companies, it's simple. Huang had the vision for the entire ecosystem from GPU chips, to drivers, to APIs and SW libraries, to partnerships and cooperation with the people and the companies who will use them. Building great GPUs for compute is not enough if you're just gonna throw them on the market without the ecosystem and support behind them, and expect it to be a success. That's what Nvidia gets and the rest (AMD/Intel) don't. So while ATI/AMD had tunnel vision and was focused only on building gaming chips, Huang was busy building a complete GPU-compute ecosystem for their gaming chips with the rest of the industry.
No wonder they are always a shadow of what proprietary ones bring in the box.
And it's easy to see that Huang and Nvidia put their money where their mouth was. The first GTC was 2009. That's 3 years before the famous AlexNet paper that's often credited with kicking off the current AI on GPUs trend.
by wasting two years on Quad rendering? Wonder how much VC money got burned in NV1 and cancelled NV2 debacle.
Rendering 3D graphics for games and the supercomputing used by AI/ML/research both need the same thing: Embarrassingly parallel math calculations with little or no branching in the code.
For example, a feed-forward neural network is just a whole lot of multiplication and addition. Transforming a 3D vertex in space to a 2D screen coordinate is matrix multiplication, which is just a whole lot of multiplication and addition. If you've already designed silicon to perform those operations in a single clock cycle, making it do HPC/SC rather than gaming isn't that big of a switch.
Unix was born to play a video game on a different platform. The Curses library for terminals (Windows users: "widgets" for the console") where born for Rogue.
Today, a damn serious OS like OpenBSD gives the BSD games set by default on the base install.
Text adventures for the Z-Machine had top tier grammar recognition parsers.
Then, a lot of simulation games overlapped the serious simulation software features, such as Sim City.