Are chiplets enough to save Moore’s Law?
eetimes.com
eetimes.com
In 1965, he predicted a doubling of components every year for the next ten years. In 1975, he changed it to every 2 years. And he was probably only looking out to the next 10 to 20 years.
Because, you know, that's reasonable. Even Moore himself knew that it was impossible to continue indefinitely.
Yes, let's get better. Yes, let's try to find more ways to efficiently compute. But let's not be slaves to words spoken before we were born because they were easy to achieve back then.
Basically AMD has caught up (arguable, but let’s say they have) to Nvidia with respect to driver quality and hardware performance/power consumption, just in time for the world to become more concerned with how a GPU handles AI stuff than how it handles video games.
I've not experienced, at any point, from the Radeon 9550 to the Radeon HD 5750 to the RX 480/580 driver issues that prevented me from playing a game or anything like that. Also, NVidia on Linux has pretty much always been a nightmare. They're still working out Nvidia-specific issues with Wayland over there in Linux world. If you are on Linux, AMD is and has always been the way to go for your GPU if you wanted the best UX.
Why do you say in that regard but then not bring up any examples of that regard? I said AMD is terrible at emulation, VR, and AI. As far as I know everything else might be great. AMD has problems on Xenia, problems on RPCS3, glitches and degraded performance on CEMU, Yuzu, and Ryujinx. The 7000-series has so many problems running regular VR games there's a performance disclaimer on AMD's own site.
I say this as someone who really, really would like to buy a 7900XTX but can't because everything I'd actually want to use it for over my current Nvidia card is hamstrung.
I will say I recently threw away an RX 580 because it couldn’t stay under 100C. I repasted it, and tried setting the fans to 100%, which was quite loud. It was an XFX model too… usually they’re on top of QC. It’s worth like $75 if I found the screws for the fan shroud (took off to try to get better temps via water cooling) but it’s easier to just throw it in the trash can. eBay takes around 12% and then I have to pack and ship the thing. Not worth it.
I just bought an RTX 3090 with a broken power connector notch because it was the only GPU I could find for under $1000 that had more than 20GB of VRAM (for future proofing/model training). It’s just a shit time to need a GPU.
Hardware Unboxed covers this pretty well here: https://www.youtube.com/watch?v=aQklDR8nv8U&t=1365s
You can see that compared to the 3090Ti or the 6950XT* it's sipping power when it's performance limited. I think what many gamers see though, is that the performance difference is about all you'd expect going to 4nm from 8nm, based on performance gains we've seen in the past from other chips.
Anyway, the rough idea is: If I built a 3090 on a 4nm node it would perform the same. Plus, NVidia has a track record of how they behave when the competition isn't up to snuff.
Now, this is silly because using a more leading edge manufacturing process takes a ton of work and is an improvement overall.
[*]: The 6950XT is manufactured on a 7nm node from TSMC. Note that there is weirdness in how each company measures their process node, which I won't go into here but just be aware that a TSMC N7 process might not be 7nm sized parts for all parts, and a lot of it has to do with how well they can layer things.
With the high end chips I think they're simply of the opinion that performance is more important than efficiency (and I don't really pay attention to the lower range too much these days), and thus like most other high end desktop parts, they're all clocked as high as the chip and cooling can handle, even if dialing it back a little would drastically improve efficiency.
I guess chiplets will help with yield issues, which could make some nodes economical (which is of course what Moore’s law is all about). But of course they won’t allow them to get down to transistor sizes that they just couldn’t hit at all.
Lego-ing out IP blocks could be really cool. It would be neat if this gave OEMs room to compete in a more technical field (selling Xeons in Dell cases vs Xeons in Lenovo cases is not very cool, if Dell and Lenovo could pick the parts in their package, that could be interesting).
For instance, let's say you bought a new PC in 2017 with a high end graphics card. Your CPU was probably manufactured on 14nm processes, your GPU on 16nm, your RAM from 16 to 20nm, and your motherboard chipset on 22nm. It's been a similar story throughout computing.
Taking that idea and scaling it down to processor designs themselves is cool and no doubt faces hard technical challenges.
AMD has used the same IO die across multiple generations of Zen hardware. This cuts the dev and validation costs for using new processes.
Fungibility helps reduce the cost of making SKUs. Want more cores? Add more compute dies. Your customers want more big chips then expected? It's easier/faster to adapt production since only the final steps differ. The component pieces are the same.
But hardware still has plenty of runway. The third dimension is barely beginning. Also customizing the hardware to the workload has lot of untapped potential. RISC-V has a committee discussing accelerating dynamic languages for example. Memory tagging for automatic array bounds checking, pointer masking for helping GCs, custom instructions caches, etc. And you can go deeper; why not have a Node Webserver chip?
The time is always paid by someone. If developers decide to write the best code, they pay the price once in terms of development time. If the code is written with a developer velocity first perspective, every user pays the price every time they use the code.
Also, developer velocity vs. program performance is not a correct perspective to look at it. I have used at least half a dozen libraries with great developer ergonomics and have world leading performance (e.g. Eigen, Catch2, easylogging++ from top of my head).
The sluggishness, absurd loading times, and RAM usage for even the simplest applications on blisteringly fast hardware is astonishing.
A numerical Gaussian integration code I have written in 2017 or so can run at 1.7 million iterations per second, per core, on a stock, 3rd generation i7 3770K.
It's implemented in C++, and only written carefully. No fancy optimizations are done.
If programmers paid half the attention their code deserves, I bet most of the software we use would be 2x faster, on average.
I mean as a user everything seems fine to me. Things are getting better, not worse.
On that department, I was able to store two 3000x3000 double precision dense matrices and some big double precision vectors, plus the 3D model representing the object I'm working on in ~250MB IIRC. The code is a scientific materials simulator, BTW.
Eclipse (the IDE) uses ~1GB (after warming up) with my fat C++ config, and that thing is written in Java, and is a full blown IDE with an always on indexer + static analyzer.
VSCode uses 1GB for nothing. Atom used 1GB for nothing. Evernote uses ~900MB on start on macOS. That thing takes notes!
We have the resources and wizardry to keep bad programming in check, but we can do much better if we wanted to.
The tricky bit is finding the spot between “find a library” and “be lazy about it,” if you want to find tasks worth doing well.
It is about running a rebuild if Minecraft on late 90s early 2000s Macs and showing the huge performance differences.
Yes Minecraft is not a great example but it is visually striking.
I'm so sick hearing about "lazy devs." Because I was one! I left a programming career because JS is fun for me and I didn't want it to stop just being a fun hobby. (The money and prestige tricked me into thinking it was my life's calling.)
But I worked with many people who were way smarter and harder working. And we still had to ship shitty code sometimes. Money paid the bills, perfect code was a luxury.
There has to be an economic incentive for this to change. There's hundreds of millions of programmers but only a few dozen relevant tech CEOs.
However, there are also examples of very fast and nice software which are also closed source and sold for good money (OmniGraffle, BBEdit, Pixelmator, iA Writer, MindNode, etc.), and they work really fast and efficient for what they do, because while they're for profit, apparently their management understands that good code brings good money.
Let's not jump to conclusions from ambiguous sentences. I may have clarified it a bit, I guess, but it's very late here.
Here's to hoping that crystal/nim/golang/rust type languages start to get more popular with the kinds of people who would normally use electron.
Note: I don't believe there ever will be a year of the desktop. But for Libre/Open source to thrive, it cannot just be free/open. It has to be good.
Both SW and HW have been leaning on Moore's law a long time.
At least with the current AI Hype we should have enough momentum to sustain the cost of innovation down to 2030. Around 1nm or 10A in TSMC Node's terms.
We will have to wait to see if Apple could further push the limit of CPU performance. Right now they have the highest performance per watt design, with AMD catching up in Zen 5 / Zen 6. May be going all the way to 16 Wide Decode? Surely somewhere along the line we hit the laws of diminishing return. We already went from 4 in x86 design to 8.
I guess the next playstation? Even though I quite dislike the digital jail which are video game consoles, going to be hard to resist to FFXX.
But if you don't consider a chiplet to be an independent package, then yes, maybe they can save Moore's Law.
Personally, once chiplets enter the picture, I think the question is becoming one of semantics, and therefore less interesting.
Well there were no multi-chiplet IC's back in his days
It could be possible to stack less power-hungry chips like cache chips, but probably not the main CPU/GPU core. It will just be another way of extending Moore’s Law.
I'm a bit naive of the actual... physical bits. I'm not clear what makes a chip, and if the cache would be that - or something a bit simpler
For now, AMD is using the technique to add more cache to their CPUs rather than moving all of it to another layer.
but maybe redundancy and binning has made this moot.
This is necessary because packaging chiplets is usually nonreversible. So if you put ten chiplets in a package and one of them is bad, you have to toss the entire package with all ten chiplets and all your packaging costs.