Intel enters a new era of chiplets
servethehome.com
servethehome.com
There are situations when you really need that, but they mostly involve imagers. The Advanced Scientific Concepts flash LIDAR had a two-chip stack, with the light detectors made with InGaAs technology. The counters and timers were ordinary CMOS. This also shows up in some IR sensors.
New nodes are fine for logic, but it takes time for analog IPs to move to new nodes (and some may not). So what AMD did, using an advanced node for compute/logic and a less advanced one for I/Os should be typical. You can also see this in broadband cellular modems, where the baseband part is on an advanced node and the RF on an older one.
When you look at processes offering, you often have variants optimized either for peak performance (frequency) or maximum efficiency. The peak performance would be the natural choice for (big) CPUs, and an efficiency node better suited for a GPU or any massively parallel accelerator where efficiency is more relevant than peak frequency (on this, I think Intel planned to use TSMC for their HPC GPU, could be related: they can focus on high perf for their CPUs).
Well, one advantage is that it can be a cost/efficiency savings on a few levels.
For example, in the case of Zen2/3, the main CPU chiplet is at 7nm, but the I/O die is at 12nm. If I had to guess, generally speaking it is useful in cases where parts of the final module would benefit from higher transistor density versus others; Memory controller tech generally gets fewer updates than the CPU/APU itself, so it allows faster design cycles; you already know the existing I/O controller works, one less portion to re-qualify.
The same on I/O where there are speciality node much better suited for I/O, and their rate of progress is quite slow compared to other part of the system.
But most of these are only good for Desktop and Server. And the world is now largely Smartphone, or Laptop Energy efficiency chips. Where integrated SoC still provides the best results in efficiency.
As a consequence, if your design has CPU dies and I/O dies, using a smaller process only for the CPU dies is likely to be a good trade-off.
https://www.anandtech.com/show/1656/2
https://pcper.com/2006/11/intel-core-2-extreme-qx6700-proces...
Also, let's be honest. First-gen Epyc was glued-together nonsense. It was very much NUMA, it performed weird, the cause was the weird chiplet design (no IO die at that time). It was fair to criticize it on that basis. It had far less cache and much higher latency than the designs that followed, and it felt far more NUMA as a result.
Hyperscalers agreed and turned their noses up at it after some test installations, and decided to wait for Rome - which was much much better in those areas. But Naples was bad and that shouldn’t be sugarcoated.
I don’t really get why people are always so miffed by Intel saying that, it was a criticism that had been leveled against them, and they were right in applying that criticism to AMD, it was a quirky server design glued together out of consumer dies, it suffered all the same downsides as Intel’s own “glued-together” designs.
Didn’t they really just… call a spade a spade?
Their whole sour grapes culture is just bizarre to me.
Had a chat with an old timer whl was doing horizontally scaling compute with thr IBM A400.
25 years later it's the starry eyed wonder of the 2010s.
However, they appear to have managed Altera into the ground, and time will tell if AMD makes the same mistakes with Xilinx.
I remember when you had to buy an extra processor to get floating point.
* https://en.wikipedia.org/wiki/X87
At one point there were video game(s) with 'extra' functionality that was only available with this hardware 'upgrade':
* https://en.wikipedia.org/wiki/Falcon_3.0
(Get off my lawn.)
Back in the day, as another comment mentioned, the PPro had a 'chiplet' style configuration where The CPU and Cache were on the same chip but separate dies. The problem with this was the CPU and Cache had to be bonded to the chip first, then tested, and if either was bad, game over. Additionally, at the time die size was at more of a premium, in the case of a PPro, 256Kb of cache was close-ish to 2/3 the size of the CPU die. [0]
The P2, Katmai P3, and the Athlon 'Classic' (Pluto/Orion) used offboard cache; This was far better from a yield standpoint (I'm assuming the cache chips could either be tested before install, or were easier to rework) but limited their speed.
It's crazy to think that the Katmai P3 itself had around 9.5 Million Transistors, but the 512Kb of cache was another 25 Million on it's own!
[0] - https://en.wikipedia.org/wiki/Pentium_Pro#/media/File:Pentiu...
Very cool
Like when they used to sell mainframes with excess processor capacity then you would pay to unlock the processor that was already there if you need it. If not it was simply manufactured to sit unused in a mainframe its entire life, then be thrown in the trash.
I didn't specifically see anything that said this in the article but there is a TON to digest in there though mdular hardware always has that built to waste vibe. Even if they claim the opposite.
Chips always have to be binned. But previously a chip would have to be binned down to it's worst component I guess -- if they had a chip with great CPUs but the GPUs were a little wonky, and they didn't have an appropriate processor line for that combo, they'd have to bin the whole thing down to low-tier. Now they can instead match up the good CPUs and the good CPUs.
Plus they'll be able to satisfy some of their lust for SKUs by mixing and matching tiles, rather than making a bazillion slightly bins.
That's so Intel. All 37 variants of the Intel Core i9 CPU: [1]
[1] https://www.intel.com/content/www/us/en/products/details/pro...
What you want to look at is how many SKUs for an architecture, like say Alder Lake for Desktop[1]. Do we really need 5 to 7 SKUs at 16, 12, 6, or 4 cores, but only 2 SKUs at 10 cores, and 4 SKUs at 2 cores?
[1] https://ark.intel.com/content/www/us/en/ark/products/codenam...
Only 2 SKUs at 10 cores seems a little weird, I wonder if they are 12 core parts with some cores disabled or something like that.
Edit: Note it is just the i5 [...]K's from Q4 '21 that have 10 cores. It isn't that surprising that the early enthusiast parts are a little weird, right?
Whether you look at just the desktop segment (29 SKUs from two dies) or the whole family (94 SKUs from four dies), it's an incredibly overcomplicated product stack.
If demand for that SKU is higher than for the 12 core SKU and margin justifies it, then yes. It happens with every vendor, specially when node been out for some time - yields are much better and there aren't this many down-binned chips to handle the demand.
Intel has a hard-on for segmenting their product lines. To the point they "launch" new products like the i9, which is just what an i7 used to be. They also deliberately cripple products, like selling SKUs with VT-x disabled. Not to mention them keeping ECC memory out of the entire desktop market for basically all of history.
Ha, that has been a long time, no? I remember buying a computer with a Q8300 in 2008 or so and there I had to assure a specific revision.
You're licensing the processor capacity. You're not paying for the actual piece of silicon, you're paying your fair share of the R&D and fab investment that went into it. You want to use more, you pay more. The same as software.
The amount of silicon in the chip is, what, a small fraction of the amount of silicon in a single grain of sand? A whole processor is tens of grams of material total. The cardboard boxes a standalone chip comes in probably weigh more. I wouldn't get worried about "waste" here.
And obviously current chip shortages have nothing to do with it, nobody's including extra unlockable chips in a shortage for that chip.
If a manufacturer calculates it's more profitable to include unlockable chips, the very fact that it's more profitable means it's a more efficient allocation of resources, therefore not wasting anything. We're not running out of silicon generally speaking, and this isn't a case where there are big environmental externalities not being accounted for.
People seem fine with software unlocks for software but often confuse the price of hardware with the cost to manufacture the same.
IBM still does this.
Enterprise actually likes it, because it allows zero-downtime upgrades if needed down the road - you don't have to stop your database for a half hour while the tech upgrades your CPU, you just drop in a new license file that says hey, they paid up, here's a key signature to enable the other cores! and suddenly you have a faster CPU, or more memory, or whatever. This is actually a highly-requested feature on AMD Epyc processors from enterprise customers as well, afaik, so, maybe coming soon.
(and it's really not a good thing even in the consumer space... like the "pay to turn on hyperthreading" processor that people flipped out about, that segmentation never went away, AMD and Intel still artificially disable cores for market segmentation, and sometimes still artificially disable SMT (eg 4700U). But now you aren't allowed to pay to turn them on... so if you change your mind and realize you need SMT, now you have to buy a whole new laptop. It's hugely wasteful and expensive. I bet a lot of people who bought a 6600K or 6700K wish they could have paid another $100 to turn on hyperthreading but nope, their only option was paying $350 on ebay for a used 7700K!)
Consumers don't like it cause they're jealous that there are parts of the chip they can't use yet.
Enterprise LOVES this cause they don't want to pay for parts of the chip they can't use yet, but also forsee wanting to use it in the future. Comes down to a CapEx vs OpEx optimization and having that ability to fine tune that balance is a No 1 requested feature.
As I go though life I keep getting more aware of as silly as it sounds how many things are designed to strip "joy" from our lives. Even if you only wanted 5/10 of the cores or whatever in this thing and sure you paid less. There will always be that nagging feeling in the back of your head that there is a part of what you paid for being intentionally kept away from you. Effectively stripping your ability to have joy/happiness for your purchase. You don't deserve full use of your product because you didn't hustle/work/strive/stress/struggle hard enough to deserve it. Its just anti-human as all things are becoming now day.
Of course if you want to rent 2 units, the landlord would be happy to give you the keys if you pay the rent for the second unit!
I think the people who get sad at using 5/10 cores (and paying less for it) are being unreasonable.
It's buffet syndrome, wanting it all and being too greedy.
Now if the manufacturer made you pay all the cost of 10 cores and only letting you use 5, that's a different issue.
Cores+™ - experience the performance of multiple cores
GHz+™ - awesome speed for the latest games and office spreadsheets
Memory GiB+™ - make room for multiple applications
Note that monthly subscriptions to this Intel Premium Experiences™ will continue regardless of whether re-verification is successfully carried out.
The new Intel chips would arrive in 2023.
5 years before that Intel shipped Coffee Lake.
5 years before that Intel shipped Haswell.
5 years before that Intel shipped Yorkfield Core 2 Quads.
And to comment on topic, this mix of processes, each optimised to the task at hand and all on the same package sounds perfect for Intel. But I wonder how accessible it is for regular TSMC customers. We've had chiplet presentations, I don't recall anything like this being offered, despite requiring extremely high bandwidth for our application.
I wouldn't be surprised that after a period of evolution and leap frogging it ends up in pretty similar spots. They have similar constraints and tools after all.
But chiplets are actually first for cost optimization, and second for process optimization IMHO.
The first cost optimization is to leverage the better yield for a small die compared to a monolithic die. The density of defects for a given process means that N small chiplets will be cheaper than one monolithic die with same number of cores: one fault will kill the whole monolithic die where it will kill only one chiplet. A SoC only uses good chiplets ("known good dies" or KGD is the term of the art). That's what has driven AMD too chiplets in the first place.
The process optimization can also be for cost saving more than performance: if an older node is acceptable, it's cheaper.
Then there is the design saving, and this shows up in Intel chiplets presentation: by developing M CPU and N GPU chiplet variants for example, one develop M+N tiles but by mixing and matching can offer MxN SoC. One has to add the chiplets interconnection complexity (extra work vs a monolithic design), but this may still save on design development. And design development costs increase a lot with each new node.
Chiplets may also become a way to extend the life of a design. We may see at some point chips embedding a mix of "old" and "new" chiplets. This could extend the useful life of a chiplet design, and give more time to amortize the increasing design costs.
So I guess a good way to see chiplets is as a cost management tool first, in the face of ever increasing design costs on advanced nodes. With the nice side effect that they also open the possibility to optimize the process per tile, if needed.
With a chiplet standard and the possibility to treat chiplets as silicon IPs today, a chiplet target market may be extended too. For example Intel as a fab may sell its chiplets (or part of their catalog) to their fab customers. This is yet another way to absorb increasing development costs, but we're not there yet.
And in case people are not familiar, Intel's version is called Foveros, so it isn't available to TSMC customer. TSMC has a few similar packaging tech on offer, but AFAIK they are not as advanced as Intel, although Intel's Faveros is also a little more costly.
It is interesting we basically cycle back to old computer model with north bridge, south bridge, CPU, GPU etc all linked together.
TSMC is part of the UCIe (Universal Chiplet Interconnect) consortium, so I'd assume they have some capability. But the other members are ARM, Intel, AMD, Qualcomm, and Samsung... so I'm not sure if it's a matter of you have to be 'working in that club' or if TSMC can provide help on custom solutions.
Also, Apple's 7nm chips outperformed AMD/Intel 7nm chips.
[1] https://www.cpubenchmark.net/cpu.php?cpu=AMD+Ryzen+7+5800U&i...
[2] https://www.cpubenchmark.net/cpu.php?cpu=Apple+M2+8+Core+350...
This Reddit post summarizes the efficiency and speed advantages of the M1: https://www.reddit.com/r/hardware/comments/nii37s/comment/gz...
I care about openness and performance (watts are irrelevant if reasonable like they are now).
Your list, to me, sounds like marketing. Move the goal posts to these arbitrary points, declare victory.
Intel and AMD chips go to OEMs
Once you have critical mass being able to have enough experienced engineers up and down the whole stack, and you’ve figured out an organizational structure that allows them to work together efficiently, I don’t see how you can beat the vertical integration in a world where the fabs Apple can use are actually better than Intel’s own fab.
I know nothing about chip design, but like most things, I feel like one exceptionally brilliant tech lead surrounded by a group of say ~100 merely very smart collaborators can make a world class ARM based processor-design. A mere $100mm annually in hiring and overhead?
The rest of Apple’s advantage comes from being able to actually hire the best, and being their own customer at scale, which means being able to buy first place in line at the fab with billions of dollars of cash.
This is perhaps less true now that the chip has so many specialized areas on the die? Like, does the neural engine get allocated a certain mm^2 and certain number of bus lanes, and then a fully separate team of 100 designs it? I suspect the neural engine part of the chip is actually super simple to design, it’s the tight coupling with the OS and getting apps to properly leverage it which is tricky.
Seeing your comment at the top makes me not want to ever open any silicon topic here again as I'm sure it will be full of these kind of low effort comments vomiting marketing garble on how AS is the best and everything else is doomed to failure, while not bringing any useful info or arguments on-topic.
But in any case, there's plenty of things to be said about this article. About one year ago (random link with relevant quotes : https://www.pcgamer.com/intels-3d-chip-tech-is-perfect-so-it... ), Intel was mocking AMD for using a chiplet approach, before announcing today that it was - clickbaity title aside - going to change everything.
The sad truth is, both Intel and AMD are in the exact situation. AMD went chiplet in order to make their performance cores at TSMC, and their less critical cores ("IO") at GloFo.
And Intel will be doing the same thing tomorrow (again, random link on the topic: https://www.tomshardware.com/news/intel-ceo-visits-tsmc-agai... ) by producing their performance cores at TSMC and their less critical ones on their own processes.
In both cases, this is just a question of using a very limited resource (TSMC's best in class process) the more effectively that you can (by throwing extra engineering at making a chiplet design that works).
And it's supremely relevant to the discussion to talk about how Apple, by throwing capital at a company (TSMC) that was, for the couple of decades I used to cover this, at best 2 years behind the best in class, today where they are (far far in front).
We could definitely have a long discussion about the hubris that led 2015 Intel where they are today (completely stuck with a 7 year old "+ paint coatings" aging 14nm "performance" process), or how Gelsinger is trying to make the best out of the situation (I personally think he's immensely qualified and Intel's best hope, though that may not be enough to bring Intel back to where it was), but at the end of the day, Apple threw a wrench in what seemed like an unshakable performance lead from Intel by spewing a bit of money left and right (they didn't only bet on TSMC early on, they threw money at GF for example, and it wasn't that massive early on from my understanding), and the silicon world hasn't been the same since.
All of what? This topic is about Intel chilpets, which many of those will end up in datacenters where most Intel chips go, and that's not where Apple sells chips for.
Not every chip made and sold in the world revolves around laptops, tablets, smartphones or the apple ecosystem.
So can we please talk about Intel's chiplets impact on the industry and less about Apple silicone which has nothing to do with this?
This is all about semi manufacturing, and the position that TSMC now has in the fab space, thanks to Apple (yes, really, that's what I went on about in the previous comment). No part of my comment referred to arm, architectures, anything of the sort, just that Apple's money, applied broadly at first in the semi manufacturing space, then in a very very targeted way, took TSMC from tier 2 manufaturer to the best in class.
If you look back a few years, only x86 chips were having volume at the bleeding edge of manufacturing process. Gpus, Smartphone, everything else was one node back at the very least. Apple threw money and orders with a massive volume (iPhone + iPad is pretty close in units to the x86 cpu market, above 300M roughly off the top of my heard) at TSMC and that early + continous investment helped them fast forward their processes while Intel is still stalled in 2015.
Apple is using TSMC today (the best bits), AMD is using TSMC today (the second best bits) for the performance part of their chiplets and so will Intel tomorrow for the exact same reason. This is the relevant bit that I was pointing at.
IMO GloFo's spin-off worked out very poorly for AMD in the short term, but long term it let them partially-leapfrog Intel much as they had done 20 years prior with the K6/K7.
There's two main things that IMO give the M1/M2 their 'magic';
- Dram on die (helping their PPW, especially single threaded PPW) - Tight integration between OS and CPU.
This is, perhaps, where the x86 consortium has fallen into a challenge in the face of tight integration; The majority of that group would likely shriek at the idea of a DRAM on CPU, "here you go that's all you get" idea. I saw it a lot when I slung PC hardware; Folks who would insist on having as much upgrade-ability as possible, but never actually bought the upgrades between PC purchases. Even still, RAM is the main thing I personally find myself still upgrading on either purchased or older PCs.
That being said, It would be interesting to see if they try doing DRAM chiplets for these; I'm sure some 'ideal state' would be where DRAM chiplets + slotted RAM cause the chiplets to be dedicated to integrated GPU resources, or act as a form of L4 cache for one or more banks of DRAM.
Memory people usually either buy as much as they need, or buy some and then add more. I've done it just about every PC build of mine, sometimes completely swap memory.
GPU generally gets upgraded, unless you're one of those who buy the best every year.
Storage for sure gets upgraded.
Currently sending this from my skylake i7 that seen 3 different GPUs, 2 different RAM kits, I lost count how many times I've upgraded storage.
I would probably get mad if I couldn't upgrade memory down the line. It maybe makes sense for laptops, but I don't see why do that on desktop. While M1 has memory right there, its latency is higher than intel and amd.
Which has very little to do with Apple PR, but everything with how CPUs/GPUs are overwhelmingly made from silicon.
[1] https://trends.google.com/trends/explore?date=today%205-y&ge...
Not a single second of thought is ever spent by the architects/designers on optimizing “absolute performance”. We only care about perf/area and perf/watt. It is the marketing teams that try to hype up gamer performance. Overclocking/high voltage performance requires the engineering knowledge of a freshman intern: go raise the voltage/freq, run the test program, make a SKU.
Source: worked on CPU/GPU arch/design for 20 years, including at Intel.
The takeaway from the story: connecting a bunch of shitty chiplets together makes a shitty SoC. Except this time, Intel paid TSMC a buttload of money to make their shitty design.
As such it's not something Apple PR invented, but rather hijacked.
List of places I tried taking conversation over that time, but it was ignored or read as complaints about Apple. (see Disclaimers in footer if you read these and think 'Wow, he just wanted to talk about why Apple was bad')
- M1 was the first processor on a particular node; so there was a short term opportunity to do an apples to apples comparison by taking down M1 numbers and waiting for upcoming launches
- it wasn't as trivial as "manufacturing on the improved node" for AMD, but it was for Qualcomm
- performance of ARM vs. X86 could be teased out by tracking tuple of node x manufactor of chips and being patient; projections and tracking of Qualcomm & AMD chips performance
- the initial M1 was beaten by Tiger Lake in desktop & sustained performance cases, which was two(!) nodes behind
- performance and noise issues from Apple optimizing for absolute fan silence always, leading to them only kicking on at extremely high speeds far into the performance workload, that had already been throttled
== DISCLAIMERS ==
1. I am a happy M2 MBP owner and think its the best chip.
2. My more nuanced view, summarized is that it is almost restrained in that the hardware got bigger, somehow, and in software there are growing pains as drivers adopt from iPhone use case to mostly-plugged-in use case. To wit, throttling seems optimized for ad copy around fan noise at low workloads than the user.
3. If you feel these discussions were focused on denigrating Apple, please recommend curious thoughts to have about chips that aren't denigrating Apple
Apple Silicon is the most exiting thing to happen in the field in decades. Apple handles platform transitions really well so it may seem less of a tectonic shift than it actually is.
It’s normal for people to be enthusiastic amongst such facts.
One silicone part has an impact over the entire computing sphere, while the other is relevant only within the Apple ecosystem. So bringing up AS in all threads about generic computing chips that anyone can buy directly is pointless as Apple does not cater to that market.
That's like me bringing up how fast my Ferrari is in threads about utilitarian vans and pickup trucks. Sure, they're both cars, but they don't compete in the same segments, so Ferrari making an even faster car has no impact or threat on the market of vans and pickups.
I also seems undeniable that Apple's move to custom silicon shook up the entire market, challenged preconceptions. It was a pivotal moment for the industry as a whole, not Macs only.
When Apple will sell its chips to consumers, PC OEMs, datacenters, console manufacturers, AI and car companies, with full documentation and Linux support, in order to compete on the same turf with Intel and AMD, then we can talk about an industry wide revolution. But until then, this "revolution" will not extend beyond the Apple/Mac ecosystem where Apple has no competitors, and who's market share in the general PC space is still tiny.
For now, Apple Silicon is just sorta a blip because most people buy them for coding environments and editing pictures, and apple decides to dedicate all of those performance per watts on having Apple Music boot up as fast as possible on boot.
For me, I get the popcorn out for all the ongoing GPU drama between AMD, Nvidia and now Intel. All of that seems to be the sequel to Pirates of Silicon Valley.
1st paragraph: SOC are not upgradable, that’s the whole point of the trade of and most of what makes this tradition exciting. What will Apple do with the Mac Pro is the most important question in this industry for the last couple of years.
2nd: what? Apple Music what?
3rd: what does that movie, which was about Apple and Microsoft, had to do with any of the current chip industry and its many players?
EDIT: F this downvoting, people seem to not get the obvious fruit pun.
There is no such thing as a "best chip on the market". Best chip for what? You're confusing the word SoC/CPU with "chip" which is a very generic word.
The best "chip" is the one that best suits your individual application or business needs, but there is no such thing as a best chip on the market. That's why Apple is only a tiny fraction of the computing market share and so many other chip vendors are still in business, because every application requires different chips.
M2 doesn't solve every needs neither as a CPU (since you can't buy it outside the Apple ecosystem), neither a a generic "chip". Why can't you accept that?
For example, have a look at https://openbenchmarking.org/vs/Processor/AMD%20Ryzen%20Thre... (user benchmarks of M1 and Threadripper). Compiling Linux on the 2990WX appears to be about 4 times faster than on the M2. (There are lots of other examples of one of the two CPUs being faster than the other but compiling Linux is the most time-expensive task I regularly do on my 2990WX. The energy usage in this task on the 2990WX is almost certainly a lot higher of course; this will be true for most tasks. However, the 2990WX is also 4 years older of course, manufactured in a different node, not very optimized for power saving and not operated in a very power saving mode.)
Between what they're doing to silicon, the outrageous App Store behavior, and how they flaunt that they're a quazi-government entity, I think this could find broad support in Congress and the DOJ.
Nobody can compete with Apple, and that's a bad thing for everyone.
They need a slap, hard.
We’d still be in the Stone Age if everybody had that mentality
They made a chip that's vastly better than anything out there at a given power consumption level. They are not attempting to use this advantage to corner the silicon market; indeed, they are neither licensing the design, nor selling the chips outside their own end-user hardware.
How does any of that say "antitrust" to you?
It would be a healthier ecosystem for startups and competitors and make for faster total sector growth.
That said, fanboyism makes people blind and uncritical, and (at least to me) its apparent company like Apple needs some good old criticism, rather than blind worship. Otherwise they will fall (if not fallen) into "we know whats best for you and you have no say in it" like with cough cough "child porn" filters or battery-gate.
I truly honestly don't trust their "we are more secure" marketing pitch, especially as non-US person.
At the end, its just another corporation driven by huge army of managers with main focus on salaries and bonuses. The idea that they are somehow morally better than everybody else when they keep hiring from companies like Facebook is pretty dangerous and goes back to beginning of my post.
Respect is earned, while fanboyism is rarely a good thing as by definition it's something rather biased.