AMD Tackles Coming “Chiplet” Revolution With New Chip Network Scheme
spectrum.ieee.org
spectrum.ieee.org
Smaller chips means that you get more yield in the presence of errors. Intel builds relatively large 600mm^2 chips (like the XCC aka the 28-core Xeon), but AMD thinks the future is to build networks of ~200mm^2 chips, like what they've done with Zen / Threadripper / EPYC.
The advantage for AMD is that they've built a single design: the Zeppelin die. RyZen is simply one Zeppelin. Threadripper is two Zeppelins. And EPYC is four Zeppelins.
That's it. One singular chip design, mass produced over and over again, to handle AMD's entire consumer and high-end line. Keep this one design small to help yields and maybe AMD can get a process advantage over Intel's larger designs.
AMD's "mobile" or "APU" line is Raven Ridge (a 2nd design at 193mm^2) that doesn't use this.
----------------
The above is the current status quo. This "active interposer" that AMD is developing in the article would go above and beyond in terms of integration.
Note that HBM2 (next-generation high-bandwidth RAM) requires the interposer. PCB is not good enough for HBM's protocol. Ditto with Hybrid-Memory Cube (a competing standard). So it seems like the future of computer parts will be the interposer.
The interposer isn't necessary for AMD's CPU strategy however. So the roadmap for this network won't come till 2020+ or later (CPUs). Unless AMD might be building this network for their GPU line?? (But that roadmap is also past 2020). I bet this is all research-and-development, and may never come out as a commercial product.
>AMD's "mobile" or "APU" line is Raven Ridge (a 2nd design at 193mm^2) that doesn't use this.
My guess is that future APU will be chiplet based design as well.
Intel's EMIB is a specific design that is cheaper than a full silicon interposer.
As such, EMIB is a methodology which may enable chiplet designs in the future. Well, I guess Intel did marry that Xeon+FPGA chip earlier this year, so EMIB is deployed today. That's probably as "chiplet" as you can get, since the FPGA is actually on the CPU's cache-coherence network.
Alas: Intel's UPI (ultrapath interconnect) is designed for PCB boards. Its how 2socket or 8-socket chips communicate together. So its not really doing anything "special" with EMIB, its just shrinking down a PCB in some respects.
> My guess is that future APU will be chiplet based design as well.
It really depends on what the "sweetspot" is. Below a certain size, you don't get any yield benefits. While above a certain size, it becomes expensive... and then impractical... to build chips.
AMD's APUs are generally seen as low-cost cheap chips. At under 200mm^2, its likely that APUs will simply be more efficiently manufactured as single pieces of silicon.
Now, if AMD decided to make a "high end" APU, kinda like their Intel+AMD collaboration project Hades Canyon, then maybe that would use chiplets of some kind. From my understanding, the Hades Canyon collaboration project is just running PCIe however, so its not really a chiplet yet.
Does it mean it would be easier for companies to cooperate, so in one chip you'll get best parts from the leaders ? or does IP already solves that ?
The more common approach is via SOCs though.
So in many regards, these "chiplets" aren't new, at least conceptually. What needs to happen is for the protocols to be redefined and respec'd for today's technology.
In particular: the silicon interposer allows for many more connections than previous technologies, as well as much lower power consumption. So protocols need to be designed with those power-requirements and huge "pin counts" so to speak.
Consider HBM2: its a 1024-bit bus. But there's four of them per Vega64, so that's 4096-wires on the interposer connecting the GPU to the RAM.
This is a level of integration never before seen, even in the MCM world. The future is with larger pin counts at far lower energy costs than before.
So its not so much that these problems are purely research. They're closer to engineering. There already exist protocols for cache coherency or fast communications at these levels, they just need to be tweaked for the new scale of things.
And the 80s https://en.wikipedia.org/wiki/Transputer
And, aw heck. We knew it was coming :) https://en.wikipedia.org/wiki/History_of_supercomputing
While a popular meme, this is not actually true. Epyc is actually a totally different die, stepping B2 vs the B1 die used in Ryzen+TR.
https://en.wikichip.org/wiki/amd/ryzen_7/1800x
https://en.wikichip.org/wiki/amd/epyc/7601
The 2700X is actually on a different die as well, and Raven Ridge on another. There will probably be another die for Banded Kestrel, if AMD ever gets around to releasing that. Presumably, their embedded SOC products are their own die as well.
So, about 5 dies per generation, across AMD's lineup (Epyc, Ryzen, APU, Atom, and embedded SOC). They're using about half as many dies as Intel is - still a significant difference, but far from the "one die for the whole lineup!" meme.
The big difference is that they're serving the whole server market with one small die, vs the three that Intel uses.
Of course, a small die isn't all roses either - both mfrs limit you to 8 dies per system, so right now with an 8-core die AMD systems are limited to 64-core systems (dual-socket Epyc) versus the 224-core systems that you can do with octo-socket 28-core Xeons. But, not everybody needs a million-dollar octo-socket system either.
Once AMD takes a node advantage that advantage will be diminished somewhat, but Intel's 10nm woes are a whole different story ;)
I thought steppings were for fixing errata, and once a new revision is qualified, the old one is no longer manufactured.
At the time, there was some speculation that B2 might be the "mirror image" die that Epyc uses 2 of.
http://www.usb.org/kcompliance/view/catalog_search/results_b...
http://www.usb.org/kcompliance/view/view_item?item_key=88142...
However, they now report Pinnacle Ridge (Ryzen 2000) as being on the B2 stepping.
Ryzen 2000 is on a slightly updated 12nm process but does not incorporate any library changes. I'm not sure if you could just drop the existing die onto the new process (given that 12nm is really a 14+), or if you could un-flip the die to the proper pinout using the substrate, but it seems like that might be where they moved to the B2 stepping.
But yeah, AMD produced the B1 stepping for an uncommonly long time. At least through the end of 2017, and the first B2 steppings showed up like June 2017.
Now physics is harder to overcome, the cost of development at the bleeding edge of technology is higher than ever, and the continued desire for larger and larger systems caused the SoC to break apart again. It’ll be interesting to see if this is what the future looks like for silicon-based chips, or if this is a temporary shortcut.
Consider that EPYC is basically a miniturized multi-socket design. Infinity Fabric is really AMD's new protocol built on top of HyperTransport (multi-socket protocol from the past). Before, AMD used to support 8-sockets. But today, AMD stitches 4-chips together and only supports 2-sockets.
From a software perspective (ie: NUMA), EPYC x2 sockets looks like an 8-socket chip of old. In effect, AMD has miniturized the 4x-socket setup in the form of EPYC. And it has also miniturized the 2x socket setup in the form of Threadripper.
----------------------
These Threadripper / EPYC chips have the same downsides as all old 2x, 4x, and 8x NUMA designs of the past. High latency and poor communications between cores.
The thing is: the modern environment is a highly virtualized, highly independent set of systems. Running 8x NUMA efficiently today is as simple as spinning up 8x VMs, one for each NUMA node.
IIRC, people are finding that Intel's 28-core design is far more effective in say... unified Database performance. Intel's design has a true L3 cache which can be used by all 28-cores, while AMD's L3 cache is split between each die. 4x 8MB caches cannot function as a singular cache in a large-scale database application.
But there's enough situations (ie: VMs, multitasking, render farms) where AMD's NUMA + Infinity Fabric is good enough. And with prices anywhere from 1/4th to 1/2 the cost of Intel, AMD's chips these days are certainly worth considering.
Then I played the game "TIS-100" and found out that my intuition was very likely correct.
Assuming 10G Ethernet is 8b/10b like a lot of other protocols, that's 1GB/s over 10G Ethernet.
Here's a $130 SSD with 500GB of storage: https://www.amazon.com/Mushkin-PILOT-500GB-Internal-MKNSSDPL...
That's 2600 MB/s read speeds. Or more than double your 10G Ethernet.
------------
RAID0 8 of them together with ASRock's M2. Quad Ultra, and you've got $1040 of SSDs + $200 for the 2x Quad Ultra cards, or just $1200 for 20GB/s read/write speeds. More than enough to saturate any network I'm aware of.
In fact, someone has already done this: http://www.guru3d.com/news-story/eight-nvme-m2-ssds-in-raid-...
They used a higher-end NVM.e SSD and measured 28GB/s (that's capital B, gigaBYTES) on the Threadripper + x399 motherboard.
Modern protocols tend to be 64b/66b or better. So that's why I listed "assuming 8b/10b", its hard to memorize which protocols are which.
Apparently I'm wrong. 10G Ethernet seems to be a more modern 64b/66b in any case.
I'd be very skeptical of that. See: Cell Broadband Engine.
Either way, the research is one of the many small steps forward to better chips.
Using multiple small dies and tying them together has several advantages. Small dies yield better, so sometimes several small dies are cheaper than one large one. There's also versatility because you can mix and match components.
It's really not that complicated an idea- modular packages are more flexible! What's new is making it work within a compelling power, price, and performance envelope.
With HBM2 being used in AMD's (and NVidia's) high-end products, it seems like the DRAM Channels are going to require an expensive interposer.
But "what else" can benefit from an interposer? If your RAM requires it, are there cheaper or more efficient designs that are built out of a network of chiplets on an interposer, as opposed to building out huge chips all the time?
AMD is already forced to build an interposer for Vega64. Might as well research other uses of it.
You are careful not to say "blockchain" :)
Yes, they are a dev in the cryptocurrency/blockchain space but still, who knows?
https://octavosystems.com/app_notes/osd335x-design-tutorial/
I would like to know how they solved that problem. Is there any public paper or patent explaining that?
That’s obviously not a hardware development, but I feel like the motivation may be similar: make components more modular; stabilize, standardize and align their interfaces.
By making these components “plug and play” the distance between a logical flow chart and the actual implementation is somewhat reduced, making the development of custom components more efficient and agile.