Goodbye, Motherboard. Hello, Silicon-Interconnect Fabric
spectrum.ieee.org
spectrum.ieee.org
I definitely think we will see more chiplets and more standardization on interfaces between chiplets. The focus will be on how to minimize energy per bit transferred (a big topic in Subu's talks) and how to minimize the die area used for inter-chiplet communication. In monolithic silicon, you don't have to think about die area, since your parallel wires between sections might just need a register or two along the way. With chiplets, you typically can't run wires at that density yet, so you still have some serialization/deserialization hardware. But, since it's not crossing multiple high inductance solder balls and PCB traces, you can get away with less. Hopefully also you can get away without area-intensive resynchronization, PLLS, etc.
I think it will definitely be awhile before this kind of integration is used outside of niche cases though. The costs are just insane. You have to pre-test all manufactured chiplets before integration, and that test engineering is nothing to sneeze at. If you don't then you have all kinds of commercials issues about who is liable for the $500k prototype one bad chip broke.
On the bright side, I see the chiplet approach benefitting other integration technologies. For example, wafer level and panel level embedded packaging technologies can be used for 1-2um interconnects now. You won't get a wafer sized system out of it with any kind of yield, but it's probably the direction mobile chips and wearables will go.
Anyway, disorganized info-dump over.
But the title is a bit, well, overpromising or broad. I don't think we'll replace traditional motherboards anytime soon (except maybe in smartphones?). Rather, it will be an incremental progress.
- first, SoC's will be replaced with chiplets
- then we'll start seeing more and more stuff being integrated on this wafer.
- say, instead of a server motherboard with multiple sockets, have all the CPU chiplets on the same wafer and enjoy much better bandwidth than you get with a PCB
- integrate DRAM on the wafer. This will be painful as we're used to being able to simply add DIMM's, but the upside is massively higher bandwidth.
The motherboard pcb per se will live for a long time still, if nothing else then as the place to mount all the external connectors (network, display, pcie, usb, power, whatnot).
With that said, some of these technologies can have a layer of surface mount pads on top. So you have a substrate of epoxy with all your chips and interconnects embedded in it, and then surface mount parts on top. For example, passives, connectors, etc. It would look almost like a motherboard, but with all the chips inside. Of course, for cost and yield reasons, this will be for mobile devices only at first.
Edit: the links below show solder balls. Today this technology is used for packaging, and has been used on chips in phones for years now. In the near future, we should be able to embed or surface mount passives and mechanical components, so maybe we don’t need the PCB.
https://www.semanticscholar.org/paper/3D-eWLB-%28embedded-wa...
https://www.semanticscholar.org/paper/Latest-material-techno...
- The interconnect pitch is huge, 0.3mm-0.4mm. HBM memories have 1000s of I/Os
- The inductance of the solder balls and the impedance discontinuities in the path mean the logic below still has to have big energy-hungry I/O drivers
- If you want to stack more than one die, you need something expensive like through silicon vias (TSV)
Hot parts next to other hot parts increase thermal power density, more heat to remove from a small area.
Colder parts next to hot parts can overheat because of the hot neighbors.
I suspect water cooling may become a must, air just cannot take away enough heat.
And if you put more dies below a heat spreader, you get more surface area, i.e. better heat flow overall (compared to a single die with the same power consumption) from the dies to the heat spreader and from the heat spreader to the cooling plate.
That's also the reason why bigger air coolers don't really do as much as you'd think they should in terms of cooling performance or overclocking, the difference between an NH-U14S and an NH-D15 is really quite small. If the problem is heat dissipation through the fins all you have to do is make the cooler bigger.
Wouldn't that be easier if everything's on a big wafer? It's already been done for a normal motherboard.
Sure, it'll take some more upfront engineering to design a system/rack/datacenter for water cooling than just immersing a server in a tank of inert liquid (flourinert or whatever they use these days), but I'm quite sure that at some point water cooling will be the standard solution in data centers.
One way I imagine this working out is that, instead of just replacing the plastic motherboard with a silicon motherboard, you eventually do away with a single monolithic motherboard entirely. Instead, you have "compute blocks" (comprised of chiplets bonded to a silicon chip, or conventional chips on a conventional circuitboard) that connect with each other via copper or fiber optic point-to-point communication cables, and you can just wire them together arbitrarily to build a complete computer. Like, you might have a couple blocks that house CPUs, one or two that have memory controllers and DRAM, and maybe one with a PCI bus so you can connect peripherals, and you can connect them all in a ring bus. You could house these blocks in a case and call it a server, or connect a lot more blocks and call it a cluster.
The main advantage of such a setup is that you don't have a single component (the motherboard) that determines how much memory, how many processors, or what sort of peripherals you can have.
Also if all these components are reasonably smart and interconnected, it could become more common for the CPU to merely coordinate communication in many cases, so larger chunks of data could easily be handed around different components and the processor only telling them what range of bytes to send where.
* PCBs are cheaper to manufacturer than silicon wafers.
* PCBs can be arbitrarily created and adjusted with little overhead cost (time and money).
* PCBs can be re-worked if a small hardware fault(s) is found.
* PCBs can carry large amount of power.
* PCBs can help absorb heat away from some components.
* PCBs have a small amount of flexibility, allowing them to absorb shock much easier.
* PCBs can be cut in such a way as to allow for mounting holes or be in relatively arbitrary shapes.
* PCBs can be designed to protect some components from static damage.
What I can see on the other hand is some packages end up being dropped into the PCB and soldered at the sides. Sometimes this is done with large through-hole capacitors, where the legs are bent and the capacitor sits in the middle of the PCB (inside a cut hole). Other than ball packages, you could could probably drop the majority into the PCB itself.
The other obvious option for manufacturers will be to put more tech on a single die, but then other problems are also raised. For example, some parts are binned based on their tested results.
That's been done for at least 30 years:
https://www.keesvandersanden.nl/calculators/hp32sii_repair.p...
* PCBs can act as integrated antennas.
* PCBs can easily mount connectors.
Uh, this seems like a pretty serious downside.
> Another drawback of SoCs is their high one-time design and manufacturing costs, such as the US $2 million or more for the photolithography masks
> ... 6 paragraphs later ...
> Patterning wafer-scale Si-IF may require innovations in “maskless” lithography.
No matter how efficient is your power supply, you will be losing electricity very very rapidly within single centimetres.
That's why there is no way to work around the need to move the voltage conversion on chip.
In the future we may even increase IC voltage to reduce the copper losses for very low power, but huge devices.
Reading this reminded me of a remark from Bunnie Huang’s teardown of a dirt cheap ‘gongkai’ cellphone (https://www.bunniestudios.com/blog/?p=4297): “To our surprise, this $3 chip didn’t contain a single IC, but rather, it’s a set of at least 4 chips, possibly 5, integrated into a single multi-chip module (MCM) containing hundreds of wire bonds. I remember back when the Pentium Pro’s dual-die package came out. That sparked arguments over yielded costs of MCMs versus using a single bigger die [...] I also remember at the time, Krste Asanović, then a professor at the MIT AI Lab now at Berkeley, told me that the future wouldn’t be system on a chip, but rather “system mostly on a chip”. The root of his claim is that the economics of adding in mask layers to merge DRAM, FLASH, Analog, RF, and Digital into a single process wasn’t favorable, and instead it would be cheaper and easier to bond multiple die together into a single package. It’s a race between the yield and cost impact (both per-unit and NRE) of adding more process steps in the semiconductor fab, vs. the yield impact (and relative reworkability and lower NRE cost) of assembling modules. Single-chip SoCs was the zeitgeist at the time (and still kind of is), so it’s interesting to see a significant datapoint validating Krste’s insight.”
I wonder if there’s any advantages to si-if from an ewaste (aka reverse logistics, cradle-to-cradle) perspective
https://en.wikipedia.org/wiki/File:Al-Elko-bad-caps-Wiki-07-...
If that doesn't work, yea you pretty much start testing things with a multimeter starting from the power source. But if it isn't a capacitor failure, it's probably ESD or power-surge related damage and not worth trying to fix.
Louis Rossman on YouTube has a lot of stuff on circuit repair that I find good, https://m.youtube.com/watch?v=_at9Jy3dfeY.
But welding grates is more of a side gig of that company, so they only do batches from time to time. One day a new, rather large batch is due, but robot doesn't start up at all, and programmer stays dark. It has power but is dead. What to do? Disassembling the controller/programmer of course. Something super special, running only one "App", written in something esoteric, running on a CPU which was designed to only run that esoteric stuff and nothing else.
And probably a mouse somehow crawled into the case and shat and pissed onto the PCB, and the CPU. Which is corrosive and dissolves the pins of the CPU und the the copper traces of the PCB, turning them into some sort of gel. But not much area at all, so easy to bridge with wires if it weren't for the dissolved pins of the CPU. So i removed that part, cleaned it with compressed air, benzine and alcohol and then very slowly and carfully drilled open the edge of the CPU until i could see the bonding wires from die to pin. Again very carefully soldered wires onto the ones missing pins, bridged that over the broken copper traces on the PCB, hot glued that crazy work, and reassembled it.
Against my expectations it worked! At 10Mhz! For years afterwards. How the mouse made it into the case wasn't obvious because the largest openings had only the diameter of a pencil, and i can't imagine a mouse fitting through that. But it somehow did. Anyways, what i wanted to say is that sometimes you can see what's wrong without knowing electronics at all. Same with the capacitor problem other commenters mentioned. They have to have a flat top, any bulging is wrong, especially when the top cracked open and some gooey stuff leaked out. Or from below.
There is an extremely detailed article about it here: https://en.wikipedia.org/wiki/Capacitor_plague
They are mostly "refurbished" (=repaired and cleaned), and then sold again. If possible, even as new. Or, of course, sent back fixed.
This one's not on his channel but is a good example: https://www.youtube.com/watch?v=g0S1ku9xvDI
I also spoke to someone from Russia who worked as a electronics repair person fixing things we would normally throw away because the price of new equipment is very high compared to the price of an expert's time to fix it.
"We also need to consider system reliability. If a dielet is found to be faulty after bonding or fails during operation, it will be very difficult to replace."
Their proposed solution (which is not repair):
"Therefore, SoIFs, especially large ones, need to have fault tolerance built in. Fault tolerance could be implemented at the network level or at the dielet level. At the network level, interdielet routing will need to be able to bypass faulty dielets. At the dielet level, we can consider physical redundancy tricks like using multiple copper pillars for each I/O port."
Your point is still valid, just wanted to call out their their thoughts on the issue.
Hacked PS3s could reenable the binned off 8th SPE for instance.
But about customizability and upgrades!
I like to choose how much RAM and storage and which ports I want with how many generic, vector, FPGA and neural cores, thank you very much.
And I like to change them later, to upgrade gradually. Even buses.
Like the authors, when I saw AMD's chiplet pictures for the Zen2 chips I felt that we would see this expanded. Intel has also done some interesting optical chip to chip interconnects that would facilitate assembling these newer multichip modules into chassis that route signals to and from the outside world.
The next (and perhaps last) element to fall into place is a way to efficiently cool these systems. One of the problems that large data centers face is not that they want "smaller" boxes but that they need to pull enough heat out of a rack of servers in order for them to reliably function.
But, I don't think we can escape the need for packages and PCBs entirely. At some point, you're going to need to interface with something that either A) you don't or can't control, or B) something at appropriate scale for interfacing with the humans whom the fancy system is ultimately supposed to serve. In either case, here come standard connections that are much bigger than the chiplet dies or the interconnect fabric between them, and thus, the need for PCBs to connect the Systems-on-a-wafer to the outside world.
As such, I think it will be a while before I can pick components out on Mouser's website, and have all those component chiplets fused to an interconnect wafer and delivered to my home or employer's shipping dock (though that would be damned awesome).
I agree. People have been talking about that future since at least the 1990s (back then SOC didn't mean a somewhat standard chip with a lot of peripherals, but literally a custom die with your own hardware on it). "Hardware/software co-design" was part of the jargon of the day. I'm not holding my breath.
But I do imagine seeing a lot of consumer products go this way. Not a $3 IoT light switch/malware vector, but anything that gets rid of connectors is a win on both the BOM and reliability standpoint, but anything in the $50-$500 price range with volume over a million units is probably worth it.
Like I can realize a PCB prototype in my apartment (or rather, my parking lot without telling my landlord), or rent a CNC machine at a makerspace/public library to do it. If I need the thing fabbed with a nice solder mask and silk screen, I can have it made for < $20 domestically with under two weeks lead time. Component sourcing is even easier.
But where do I go to have the chiplets I need for the circuit? Organize the logistics to ship them to the clean-room where they can be packaged on this fabric? How many widgets do I need to ship for this to be viable, or for the contract fab to not laugh at me? How do I prototype? How long is the lead time?
It just seems like there's a lot in the way of this being viable for run of the mill projects.
AFAICT this technology will have no impact on the low end hobbyist end of the market, but rather (if all works out) enables those $zillion behemoths to produce ever faster systems since they won't be as bottlenecked by off-chip bandwidth as they are with today's PCB's.
As the article states, you will not make those at home any time soon, unless ”maskless“ fabbing becomes a thing.
I'm going to guess that when PCBs were introduced, you couldn't get them for a couple of dollars from China either.
These kind of things tend to become cheaper over time, as manufacturing, competition (newer or competing components) and availability grows.
So I assume it's at least possible that at some point, a few decades from now, they're cheap :)
I see no reason why infinity fabric could not be routed through a silicon interconnect, but it would be wasteful. The whole point of silicon interconnects that the hardware protocols like pci-e are no longer necessary and can use much lower power/cheaper ways to transfer data.
I guess Apple will hire them pretty soon.
Congratulations on the insurance money from your building burning to the ground.