ADDED: And, yes, Pat was an Intel insider but I think he was still a pretty good choice for the job. Not sure who else I would go with.
ADDED: And, yes, Pat was an Intel insider but I think he was still a pretty good choice for the job. Not sure who else I would go with.
You certainly couldn’t know AI was coming. But GPUs have been a thing since the 90s. And they only kept getting better. They were also the only thing other than CPUs that gamers would pay a high margin on.
Nvidia and AMD both made plenty of money on them. Intel had it easy. They had the best fab and could’ve easily bundled things for system integrators, netting lots of sales.
Instead they had no external GPU and the integrated one when they finally added it was a joke for a very long time. Their newest efforts sound like maybe they’re at least reasonable (though they don’t compete at the top end) but that’s like 15 years too late at least.
I don’t know. Seems to me like it would be a pretty obvious money making venture for them, and if they had gotten over the nonsense of Larrabee/making it out of x86 seems like they should’ve been able to have a really good chunk of the market easily.
The big win no one could’ve predicted is that if you had a good modern competitive GPU five years ago, then you were well positioned to pivot some for AI. At a minimum they would have a horse somewhere near the race.
IMO it seems like they had the Google problem: if you can't sell a billion units it has to be canceled.
Nvidia pivoted to AI almost 2 decades ago (late 2000's). And big tech recognized GPU's potential around the same time.
Google brain was 2011 and tensorflow was publicly released in 2015 (much earlier internally).
The "Attention is all you need" paper came out in 2017. That was for many folks, another smoking gun that GPU's are fundamental to AI.
The point is: Intel has had a looong time to get into GPU's. They just chose not too. And it wasn't because "nobody saw ai coming".
Of course it’s not actually the CPU doing the work, it’s just in the same package.
Part of the difference is that audio can only get so good. But graphics don’t have much of a limit yet. We kept adding more colors and higher resolutions and higher refresh rates, more polygons, and even a second display for true 3-D.
It’s the true embarrassingly parallel problem we’re throwing more silicon at it just keeps making it better.
By the time Intel decided to get “serious“ and develop Larrabee, they clearly saw the market was there and would keep growing otherwise they wouldn’t have even bothered.
And honestly, whether it was on the same package as the CPU or an actual part of it the way normal vector units or the branch predictor are, who would benefit more from having a good graphics story to go on their GPU than the king of CPUs. They could use their advantages to just further cement their moat.
This was already solved 7 years ago, but few games actually make use of it: https://valvesoftware.github.io/steam-audio/
Same problem IBM had. They didn’t like people using “their“ PC architecture, so they decided to make the MCA bus and other proprietary stuff. Force everyone to do what they want.
IBM‘s AMD was Compaq and all the clone makers. They all got together and used their power to prevent the market from being screwed up by IBM.
I could really see Intel wanting to make that kind of move but it never would’ve worked in reality.
I still think ia64 was sunk by the let-the-compiler-do-it tarpit. There were ambivalent aspects such as, at the time, questions of power or chip area. But things like memory latency are just not predictable enough to do strictly in-order pipelines. OoO won.
Had AMD not been around with AMD64, the industry would have had to get it working no matter what, HP and Microsoft were already on the Itanium train, and if production of x86 got slowly replaced by Itanium that was it, Windows and HP-UX would drag the ecosystems into it, and eventually the remaining issues would be sorted out.
For what it's worth, I followed Itanium as an IT industry analyst and wrote a short book recently on the topic. https://s3.amazonaws.com/bitmasons.com/docs/Lessons+from+the...
AFAICT, they expected Itanium to own the high end server/HPC market, while 32-bit x86 would continue supporting PC’s for another decade.
That seems remarkably short-sighted in retrospect.
UNIX systems, I also doubt, as we were still in the days each vendor had their own CPUs.
In the end, Nvidia kind of got lucky with their bet on AI. While CUDA has been a thing for a while, it was limited to mostly academic and workstation workloads. It wasn't until LLMs and things like Stable Diffusion when it really became popular.
Unfortunately for non-NV users, most of these AI workloads when run locally assume CUDA is present and it's a considerable effort to make them work on non-NV GPUs. It's fairly okay for common workloads like running basic LLMs and SD models, but as soon as you get into more advanced/obscure stuff the harder it gets without CUDA.
I see flavors of this narrative quite often. Nvidia has been positioning their GPU's for AI since the late 2000's. And their GPU's were being used for AI at least as early as 2011
To me, it doesn't seem fair to discount Nvidia's foresight. A lot of smart people saw this opportunity even before the invention of LLM's. That's a bit like building a light bulb before discovering electricity.
?? People have been buying GPUs since the late 90's, it was always an obvious growth vector.
And forget AI, we were using GPUs to run statistical programs since the late 00's...
It only seems nonsense in retrospect. No one expected nVidia to improve flops/dollar in their GPUs that quickly.
And Larrabee wasn’t the only architecture deprecated by GPGPUs, Cell Broadband Engine by Sony+Toshiba+IBM was another one.
It is one of the reasons why Playstation 3 is the less celebrated one, and the console generation where XBox was able to pull ahead.
Larrabee, I was at a GDCE 2009 session where it was demoed, and while the idea was promising, it didn't seem to actually pull off in practice, plus there was hardly much hardware where we could actually try it out ourselves.
Xeon Phi didn't made it better regarding adoption, nor did AVX.
> It is one of the reasons why Playstation 3 is the less celebrated one, and the console generation where XBox was able to pull ahead.
In retrospect, Cell was a great architecture. In typical Sony fashion it was simply a few years too early. At the time games were single-threaded and many console games still expected CPU and GPU to be in sync.
On the PS2, the difference between using the vector accelerator or not was 1.4x – many developers never bothered. But on the PS3 this was simply not an option, as the performance difference was 6x (!).
The big issue at the time was dev tooling. Most existing engines had no support for dispatching jobs or any form of pipelining. Adding that required redesigning the engine from the ground up. On top of that the dispatcher/manager and the jobs use two different µarchs. Obviously, game and engine developers at the time hated cell with a passion.[1]
But even on the PC and Xbox, frequency scaling ended with the Pentium 4. Multi-core CPUs became mainstream. As CPUs stalled, most of the performance increase on PCs came from GPUs, that had moved from a fixed-function pipeline to shaders.
Slowly, even PC games had to do the bulk of their work in jobs dispatched across multiple cores. Engines that had adapted to Cell got a massive head start.
Today, the very textbook for game development – "Game Engine Architecture", written by Jason Gregory – uses the architecture of the Uncharted games on PS3 as the prime example for how to build a modern engine. As result, nowadays even small indie games are built that way.[2]
________________________
1. Gabe Newell: "Investing in the Cell, investing in the SPE gives you no long-term benefits. There's nothing there that you're going to apply to anything else" https://www.wired.com/2007/10/valves-gabe-new/
2. e.g., Tiny Glade: https://www.youtube.com/watch?v=jusWW2pPnA0.
Although I think it would still be better to use specific shading languages as the hardware isn't the same architecture, thus there is only so much one can do with traditional languages without extensions, or GPGPU specific algorithms/data structures, but that is me as outsider with interest in the field.
The experts seem to be enjoying exploring CUDA, SYCL, MSL, and similar for such purposes.
https://home.otoy.com/render/octane-render/
If you prefer actual code,
https://advances.realtimerendering.com/s2015/aaltonenhaar_si...
and
"Mesh Shaders - The Future of Rendering"
https://www.youtube.com/watch?v=3EMdMD1PsgY
"GPU driven Rendering with Mesh Shaders in Alan Wake 2"
Intel sucks at software and should have leaned in to their strengths - documenting the hell out of the low level workings of Larrabee. Workloads tuned to it would have been a great moat.
Yeah, but writing efficient CUDA kernels, or D3D compute shaders, is not particularly straightforward either. At least GPU cores have direct access to global memory but still, for good throughput that memory access needs to be coalesced. Then there’re manually managed groupshared memory, atomics, wave/warp intrinsics in D3D12/CUDA, and now these special low-precision matrix multiplication instructions for AI.
People accepted all that complexity because performance is too good to ignore, both flops/dollar and flops/watt. Contemporary mainstream GPUs were just better than SPEs on Cell, or AVX-512 on Xeon Phis.
Specifically, PS3 (2006) delivered up to 200 GFlops FP32, first-generation Xeon Phi (incredibly expensive devices released in 2010) 750 GFlops, but GeForce 460 (released in 2010 for the price of $200) delivered up to 900 GFlops.
But in retrospect, of course it was a missed market. It's hard to disentagle that choice from others in the same era (including the refusal to jump onto EUV or chiplets).