Nvidia GPU roadmap confirms it: Moore's Law is dead and buried
theregister.com
theregister.com
TL;DW:
* The improvements don't come from transistor density but other tricks like putting 2 chips together, using smaller 4bit format, more on-chip memory, 2x memory bandwidth
* And Nvidia is packing more GPUs in the same rack and consuming unheard of amounts of power, with a huge toll on datacenter infra
And it seems all these have growth limits. No more 2x every year or even every other year.We would likely get 2x transistor per dollar every two year for quite a long time.
> The complexity for minimum component costs has increased at a rate of roughly a factor of two per year.
It's incredible how deeply ingrained this misinterpretation has been, even among the very technical, without ever examining its original source or phrasing.
Of course it's much more complex than that. The nature of the problems that tech now has to solve is different and as stated in the article, Nvidia hit many roadblocks, but I still think if it had healthy competition, other brands would step in and make it more likely for a creative solution to manifest itself.
Moore's Law was an extrapolation on transistor density which is a manufacturing process in a chip foundry. NVIDIA is a fabless company that depends on physical process improvements from companies like ASML & TSMC.
So the competition you're speaking of is happening more at the foundry layer which would be TSMC vs Intel vs Samsung vs China-chip-maker (e.g. SMIC). At the lithography layer, ASML's EUV has no competition from Nikon and Canon.
Of course it's much more complex than that. The nature of the problems that tech now has to solve is different...Roadblocks they've seemed to put on themselves, in the name of market segmentation. There is a huge market sitting between consumer and "professional" GPUs, and nvidia is trying to milk the market as much as they can before any of the competitors get their shit together.
I predict we are pretty close to the death of the GPU as a discrete component. We will have TPUs that live in work stations and data centers as discrete components, and integrated graphics cards that live on the CPU die.
With proper engineering with nothing fundamentally new, one could get proper integrated water cooled racks with very high computing density and power usage. Then put those racks in a building close to a nuclear powerplant (outside the security zone perhaps though). Or a wind or solar park + batteries. Can you circumvent the costs associated with using the public power grid?
Extremely simplified, you only need fiber in and out of this power+computing facility. Ultimately you could do this in space with solar power and laser up/downlink. Cooling might be problematic though.
France at least used to have a dedicated nuclear plant just for uranium enrichment. Before electricity, a lot of industry centered to places near running water for sawmills, flour mills, metal hammering and so on. With electricity these were in many cases decoupled. But things that get energy intense enough and don't need material transport or local labor, it might make sense to locate close to energy sources again.
This is just one example, but I would be surprised if that wasn't a consideration in more datacenter placements.
The problem of course is to find a place where electricity is abundant, but you also have easy access to the large fiber backbone network. If you have a lot of traffic, fibers aren't cheap, especially interconnecting them with others.
Ouch, I grew up where Meta created its first data center and while it is very efficient (https://engineering.fb.com/2011/04/14/core-infra/designing-a...) I remember a more recent article about how they are zero net emissions (https://tech.facebook.com/engineering/2021/11/10-years-world...). I wonder if they have kept that in this AI rush (along with most of the datacenter companies).
EDIT: a direct link to net zero for Meta: https://tech.facebook.com/ideas/2020/9/facebooks-path-to-net...
Can’t they just buy indulgences^W carbon offset credits?
So was Moore's law dead 20 years ago? Clearly not.
Right now Nvidia's focus is on making bigger silicon with faster interconnects. But pretty soon that will stop working and then the focus will shift again. It's still early days for AI. People have only just started working on making dedicated AI chips. Presuming that performance per watt or per mm^2 of silicon cannot go up from here seems silly.
Agree. It is only 'dead' because Moore's law is about the cost per transistor, and Nvidia wants to keep that price constant.
Double the transistors? Double the price of the chips. That's Nvidia's wet dream.
I hope this backfires on Nvidia and the standard Moore's law returns.
People say that Moore’s law has been dead for 20 years because that’s when the industry hit the power wall and had to pivot from simply reducing transistor size to parallelism and these other engineering strategies for keeping the performance growth curve going. I don’t think that’s coming back without solving some really difficult materials and semiconductor physics problems
"The complexity for minimum component costs has increased at a rate of roughly a factor of two per year. Certainly over the short term this rate can be expected to continue, if not to increase."
It is clearly about cost per transistor, it is/was correlated with performance, but it is not about performance.
What Nvidia is doing is claiming that they will fight against lower costs per transistor with all their might, which means any efficiency improvement will translate into more profit for them, and never into lower prices for customers.
It got misinterpreted during the 90s, that's all.
May I ask why? Why would you want a company to fail?
If you mean chips that do fast matrix-matrix multiplications, that definitely has a long history and was not "just started".
If you mean calculations in low precision, that's also already a decade old if not more.
More like since the history of computing. 1990s era graphics cards made use of FP16 formats and before that people did fixed point math approximations because FP hardware was so slow.
This seems wild to me. I used to warm myself next to 4kW racks of telco gear when I was still doing overnight work (20 ish years ago) and thought this was a lot of power to be using in such a small space…
That’s why Jensen is always saying Moores Law is dead - so you buy less of his sand for more bucks.
The reason it's about flops/watt these days is that potentially the demand for "intelligence" limitless.
Tech leaders love this narrative:
2005: Intel CEO https://hothardware.com/news/moores-law-is-dead-says-gordon-...
2009: Sandisk CEO https://archive.nytimes.com/bits.blogs.nytimes.com/2009/05/2...
2010: NVIDIA VP https://finance.yahoo.com/news/2010-05-03-nvidia-vp-says-moo...
2016: Intel CEO https://www.nytimes.com/2016/05/05/technology/moores-law-run...
2017: NVIDIA CEO https://www.extremetech.com/cars/256558-nvidias-ceo-declares...
2019: NVIDIA CEO https://www.cnet.com/tech/computing/moores-law-is-dead-nvidi...
2022: NVIDIA CEO https://www.barrons.com/articles/nvidia-graphic-card-prices-...
The only time this felt true was on the desktop/Windows/Linux for CPUs when Intel had an uncontested monopoly from 2009 - 2016:
https://cdn.arstechnica.net/wp-content/uploads/2020/11/CPU-p...
But it was broken by AMD's Ryzen architecture on one side and Apple's M series on the other.
Two years ago, I snagged laptops with a 32GB RAM, RTX 4070, and i9-13900HX for $1,500.
Historically I've been able to get a sizeable bump in CPU + GPU performance for around the same price point, but currently there's nothing on the market.
It _feels_ like the pace of development has slowed in the last few years.
I don't think I'll be upgrading again until consumer X3D CPU laptops become available for a non-egregious price.
If you stick with Intel you are artificially holding yourself back. NVIDIA is is optimizing for datacenter thus its stuff is not available to consumers except at inflated prices.
Apple has consistent and significant speed improvements with each new M-series chip and their offers have not been increasing in price for the equivalent offers - e.g. Mac Mini and MacBook Air. AMD is doing pretty well as well, especially which high core counts.
Although with the new tariffs, everything is going up in price but that is not related to Moore's law but Trump's economic policies.
I'm a Windows + WSL Linux and Android guy so I'm not knowledgeable about the ecosystem but from the HN announcements I've seen posted they appear to doing well.
There are only two laptops on the market with X3D chips right now -- the ASUS ROG Strix SCAR 17 and the MSI Raider A18HX.
Both of them have Ryzen 9 7945HX3D chips and the mobile RTX 4080's, and are priced equivalently at $2,500.
Once I can snag something similar for the $1,500-1,800 range is when I'll upgrade.
https://browser.geekbench.com/macs/mac-mini-late-2020
https://browser.geekbench.com/macs/mac-mini-2024-10c-cpu
I am quite satisfied with the rapid speed increases on Apple's CPUs the last couple years.
M4 from November 2024 has 28 billion on a 3mn process node at 22W TDP.
And remember that 5mn process node is just the length in one direction (although it isn't 5mn), so density technically increased by (5^2 / 3^2 = ) 2.7x during that 4 year period.
That is less than the 4x one would expect, but it is far from Moore's law being dead.
Mac Mini M4 is cheaper than the Mac Mini m1 and has better base specs. And the MacBook Air M4 is cheaper than the MacBook Air M1 and also has better base specs.
That said, in terms of horizontal expansion (tripling or quadrupling the number of computational cores), surely the main challenge must be the lack of bandwidth to the memory to power such insane data transfer rates.
You could of course dedicate some amount of memory to each compute node, then the problem is one of quickly splitting the workload into chunks and sending them their separate way with no need to move them around. In a way, LLM and neural networks do function that way. Maybe it is time to build hardware that maps 1:1 to the software architecture.
Yes, there is truth to that argument—both in the limitations it describes and the vision it hints at. Let’s break it down:
⸻
1. Bandwidth Bottleneck with Horizontal Scaling
• True: As you scale out the number of computational cores (SMs in NVIDIA GPUs, for example), the global memory bandwidth becomes a bottleneck unless it scales proportionally.
• GDDR6, HBM, and even on-package memory (e.g. Apple’s M-series) are attempts to address this.
• Still, there’s a ceiling: latency and contention start dominating.
• Problem: Many workloads, especially in ML, are memory-bound rather than compute-bound. Doubling cores doesn’t help if data can’t be fed fast enough.
⸻
2. Local Memory per Compute Unit
• True and Already Happening: Architectures are increasingly moving toward giving compute units local memory:
• NVIDIA’s SMs have shared memory (very fast, but small).
• AMD’s Infinity Cache, or on-chip memory on M1/M2, is another example.
• TPUs and custom ASICs use SRAM close to the compute logic for matrix multiplies.
• This is reminiscent of scratchpad memory in SIMD systems or old-school Cell processors.
⸻
3. Workload Partitioning and Data Locality
• Critical Insight: Yes, mapping data locality to compute locality is the key.
• LLMs and other neural nets are mostly large tensor operations.
• These can be tiled and distributed to cores with minimal inter-core communication if partitioned right.
• This is why hardware-software co-design is becoming more important:
• NVIDIA’s Tensor Cores and transformer engine
• Google’s TPU with XLA compiler
• Apple’s use of unified memory for CPU, GPU, and ML cores
⸻
4. Hardware Matching Software Architectures
• Emerging Trend: We are already seeing this:
• LLM accelerators (e.g. Groq, Tenstorrent, Cerebras) aim to align hardware with transformer-like workloads.
• Neuromorphic chips and systolic arrays are trying to escape the Von Neumann bottleneck by mimicking data flow instead of instruction flow.
• Vision: Instead of generic ALUs doing many different things, you have domain-specific architectures (DSAs) optimized for matrix math, attention, etc.
⸻
Conclusion
You’re touching on a real and current transition in hardware design: as Moore’s Law slows and workloads become increasingly specialized (especially ML/AI), the bottlenecks shift from FLOPS to memory bandwidth, latency, and data movement.
Yes, scaling horizontally is increasingly limited by data movement costs, and yes, future architectures will (and already do) look more like the software they’re meant to run.
Would you like a more technical breakdown of how LLM workloads are actually split across GPU memory hierarchies or how specific hardware like Cerebras or TPUs tackle this?