Apple’s A14 Packs 134M Transistors/mm²
semianalysis.com
semianalysis.com
I also don't think the author understood the TMSC presentation. TMSC clearly said that is used a "model" of a typical SOC of 60% logic, 30% SRAM, and 10% analog. Then they said that for each category of thing, you could expect 1.8x, 1.35x, and 1.2x of shrink. If you do the math, that means an overall shrink for a 'typical' SOC that conforms to their model would be 1.57x (approximately).
That Apple achieved 1.49x would suggest they got 95% of the process shrink effectiveness.
Then there is the cost per die and thus cost per transistor discussion. It is important to remember that this is likely the most expensive these wafers will ever be. The reasoning for that statement is that during a process node life-cycle the cost per wafer is set initially to capture "early movers" (who value density over cost). Much like any product where competition will emerge "later" there is a window early on to recapture value which can pay back your R&D and capital equipment investments. As a result the vendor sets the price as high as possible to make that repayment happen as quickly as possible. Once paid back, the price provides profit as long as it can be supported in the presence of other competitors (in this case, I would guess that role is played by Samsung). The GSA uses to publish price surveys of wafers on various nodes over time but they don't seem to do that any more. Anyway, so the cost per transistor will go down from this point but how much depends on how much margin is in the current wafer price.
So I agree that the cost per transistor is not going down as quickly as it has in the past, and its possible that this node may not get to be as low as the previous node. I'm curious how it compares when you look at 7nm introduction price per transistor vs todays price per transistor. And if you get the same ramp with the 5nm node what that would mean.
Which has Samsung in production of their 7nm node this year as well.
Consequently, almost all of Nvidia's current troubles stem from being unable to do all of their fabrication at TSMC, and the yields for Samsung 8nm being very poor.
Nvidia recently canceled the 2x RAM variants of their 3070 and 3080 cards (and not because of insufficient GDDR6x, only the 3090 takes that), have stopped shipments of GPUs for new cards (they are only continuing to ship to fill existing orders), pushed the 3070 release date until after the RDNA2 announcement; most of this stems from Samsung's poor yields no matter how inexpensively Samsung is selling those wafers for.
This also isn't the first time Nvidia used the Samsung footgun.
TSMC does NOT say they have a 1.8x shrink for N5, they say for LOGIC you can get that, but for SRAM and Analog the results are 1.35x and 1.2x. Had they summed that together for a "typical SOC", which they also discuss (and one presumes that Apple makes typical SOCs) then the "theoretical" shrink is 1.57x for SoCs.
The challenge is that whereas at one time the node size was that of a "gate" (which could be 4 transistors), in a marketing race for smaller numbers fabs started emphasizing "feature" size.
Because of this change, the "theoretical shrink" is a function of what kinds of circuits you're putting down. Pure logic? You get one number, two gates connected together for a flip-flop, you get another number, a voltage regulator, or ADC filter, you get another number.
So doing the analysis the author claims to do, can ONLY be done if you know what percentage of the part you are making on the new process is what. I am under the impression that they missed that.
1) In the way back times, (think Intel 8080A) the complexity of the chips was advertised in "logic gates". More gates = more impressive chip.
2) But logic gates weren't equivalent from one process to another, and so it switched from "logic gates" to "transistors." More transistors => more impressive chip. (this is when I left Intel for Sun Microsystems)
3) But not all transistors are created equal, and there were things (like copper metal layers) that made chips better even it it meant you couldn't fit as many transistors so "line size" was what was important. Smaller line size => more impressive chip.
4) But now people had redesigned transistors so that they could be packed more densely and the limiting factor was how much silicon you needed for the gate (NMOS/CMOS) and since that wasn't a whole transistor, it was just a "feature" of the transistor, "feature size" became the new marketing term. Feature size was measured in nanometers and so the smaller nanometers implied more features per unit area.
It has all evolved over time so that it is harder and harder for any sort of comparative analysis between processes seems to make any sense at all these days.
These days, much like the TSMC presentation that is excerpted in the original article, semiconductor fabs rely on comparative measures like "same stuff would be size <x> on this process vs size <y> on the previous process." All the really interesting parameters to me are things like how that effects leakage (thus idle power) and voltage thresholds (thus idle power and maximum frequencies).
I'd love it if there was some sort of SI unit you could demand which would give you a better comparison metric but I don't think we'll see that. Everybody wants to be "the best" and that is most easily achieved when you can dynamically define the metric for "best."
https://www.youtube.com/watch?v=1kQUXpZpLXI (from 13:50, but the rest of the video is just amazing)
>I'm curious how it compares when you look at 7nm introduction price per transistor vs todays price per transistor. And if you get the same ramp with the 5nm node what that would mean.
First Gen N7 being ~10K+ per wafer while N5 being around ~13K with higher yield compare to N7 in the same stage. So N5 should still be cheaper.
The meaningful achievement is how many discrete electrical components are composed into a given area. Not some arbitrary dimension of some cherry picked subset of these components.
I disagree. The meaningful achievement is how power-efficient, fast, and cheap you can make a given chip. (Secondarily, how small and how durable wrt cosmic rays; but for most purposes these are not super important.)
If that follows as a result of many discrete electrical components being packed into a small area, great; but the latter isn't intrinsically interesting.
Efficiency, performance and cost are strongly related to density. Price less so; that's a function of supply and demand.
They are related, of course, but the important stuff can be measured more directly by looking at how well programs work on a given computer. I think it's just a little odd to praise one cherry-picked, arbitrary metric for being less game-able than another cherry-picked, arbitrary metric. Especially when we have metrics to hand that are a lot closer to what real people care about in a computer. I certainly have never shopped for a CPU based on the number of transistors, but I have made purchasing decisions based on things like cinebench and passmark scores, which try to get at what ultimately matters to me (i.e. how many FPS a CPU will drive in the games I play).
Manufacturers use “x nm”, yield (a metric correlated with price, but also completely uninteresting for consumers), etc. because they tell chip designers what they need to know.
They avoid benchmark scores because they bring the CPU design in the picture as a variable.
From the article:
“Even if SRAM scaling kept up, the cost per transistor would still have remained flat from N7 to N5.”
Is average over an area the meaningful achievement or is the meaningful achievement the smallest individual gate length? Neither are super useful without additional context.
I hesitate to say there are none, because I'm sure some task that is highly sensitive to latency might potentially benefit, but if you build the whole chip that way, it would actually be a regression.
What I fear is that we've hit the point where this is no longer a safe assumption. That we will be having people chase feature size numbers that don't actually result in a proportional increase in chip density.
Transistors per square millimeter is closer to a measure we actually care about (speed of light and clock speed) instead of a bench number that doesn't measure anything except perhaps instructions per watt.
I don't think anyone actually cares about transistors/mm^2 at all, I think what we care about is perf or perf/watt for our specific workloads. I don't care if the chip is built with vacuum tubes if it is fast, efficient (per dollar and watt), and physically fits in the device I want it in.
Second, one of us is reading that number wrong. They said 1.49x, not 49%.
In any other conversation, 2x is reducing feature size by 50%. 3x is 1/3 of the original, or reduced by 2/3. That means 1.8 is 45% smaller, and 1.49 is 32.9%.
Similarly, if you cut the pitch of a circuit in half you should see 4x as many transistors. If you could keep shrinking the space between transistors while shrinking the transistor, then going from "7" to "5" node should have been a 1.96x factor for areal density, not the 1.8x they claim, or the 1.49x Apple achieved.
I'm not saying they screwed up. As soon as nodes stopped measuring literal transistor size, it wouldn't take long for the names to be aspirational instead of descriptive. It's something they can name the project early on when the set of potential tech has been selected and some estimates have been made. For building a team it's fine. But I'm not on that team, I'm a customer (and current or former shareholder).
I think the consumer cares about the transistors per mm^2 (after the voltage and the instructions per second), not the node number. Especially when each foundry uses the same number to describe different densities. I shouldn't have to keep remembering that TSMC-7 = INTC-10. Numbers that are actual numbers, please.
It's the silicon equivalent of measuring one's BMI.
There is a little freedom in the Z axis, but not very much. But if speed of light matters to performance, then a chip design that increases the z axis decreases the euclidean distance between any two gates, which should (or at least could) matter to performance, right?
It barely matters. Gate delays and thermal limits outweigh distance by a huge factor. If you need to go further distances then you can wait one cycle and cover a relatively huge length.
At the cost of increasing pipeline depth, right? We've been wrestling with that forever.
We've been wrestling with the number of pipeline stages vs. the number of transistors in the critical path forever. Not so much physical distance.
As long as the measure is in mm^2, not mm^3, I think (hope?) that would avoid any perverse incentives against breakthroughs that allow you to add more layers to a chip and still maintain yields.
https://www.extremetech.com/computing/315186-apple-books-tsm...
so it seems apple paid to get to the front of the line
It seems to me apple's marketing and product strategy has more to do with bringing in that money than actual engineering and design, contradicting the original assertion.
Marketing might get you there for a short time, but maintaining that growth and long term high customer satisfaction doesn't come without great engineering and design.
Apple's involvement with TSMC also provides TSMC with an opportunity to learn from Apple. Nobody books an entire pure-play foundry without taking the time to figure out what they want to do with that manufacturing capacity in the first place.
Delivering a CPU on a new node process is not a trivial achievement, along with the significant design changes required to achieve the potential gains of the node. Neither is Apple’s business ability to lock in x months of exclusive use of that node. Apple have done an amazing job here, hand in hand with TSMC.
What makes you think coordinating something like this is an easy task for anyone? As if you could ever distil this down to the efforts of a single organisation.
Even that number is not so straightforward.
Tr/mm^2 = 0.6 * (NAND2 Tr)/mm^2 + 0.4 * (Scan Flip-fop Tr)/mm2
How about bit count? (at 4 bits per transistor)
I don’t think that’s all that Apple specific as there are increased ISPs and security chips on other phones as well.
Also, one can play games on an iPad, and games will soak up arbitrary amounts of CPU and GPU power.
So Genshin Impact can look and run beautifully XD
- Less heat produced > allows for smaller heatsinks and, as a result, smaller devices or bigger components
- Performance bottlenecks become less likely
- Features like 120hz display refresh rate become possible without the user noticing degradation
- Faster OS boot and app starts
- Ability to do more work locally vs using a server (see the many ML features)
- Tech debt in system code is less of an issue - code can be shipped early and optimized later
You may well be benefitting from running a task longer, at a slower speed, and using less transistors.
RAW processing is actually a really interesting application, and rendering time depends heavily on what you do with the bits from the sensor.
If you use non-linear interpolation, 2-50x the time to build.
If you do highlight recovery, 2-5x the time to build.
If you only need half the native resolution, build time may be reduced 2-4x.
Basically "viewing a RAW image" can take 100 milliseconds or 15 seconds (on the same hardware!), depending on how you're interpreting the sensor data.
(Source: playing with dcraw and libraw to rasterize raw images for PhotoStructure)
There's a lot of corner cases there where certain consumers value it a lot.
Also performance per watt is a huge deal. The expanded power envelope that improved PPW brings allows 120hz displays which are battery hogs.
I'm already missing the glorious few months that PUBG could be played on Mac via Stadia.
Changing the chip isn't going to change the code base of the operating system.
I told the customer that using a caching layer such as a CDN would help paper over the worst of the inefficiencies in their application and the network stack.
That was true! The download times halved.
However, benchmarks with F12 developer tools showed that while downloads reduced from 200ms to 100ms, the overall load times were still seconds, of which about 50% was actual CPU time.
(This is typical of large, complex sites built by non-experts or organisations where performance is not a primary concern.)
Web sites like these perform very noticeably better if you have more CPU cores or faster CPU cores.
Essentially, CDNs are easy to deploy and 5G provides nearly gigabit download speeds, so the bottleneck has shifted back to the CPU for a lot of the web.
I guess your point is that sloppy web developers / web companies make poorly performing code and shift the burden of handling it out to visitors’ machines, so those machines have to keep improving to not be left behind?
That might be true, but is pretty depressing.
Faster computers enable easier programming strategies, which improve net productivity. People are expensive, so giving them better "power tools" improves efficiency.
We don't complain that coal miners just need to shovel more efficiently. We give them enormous trucks and hydraulic mining shovels.
I think the cynicism here is that if only these frontend developers would just learn to optimize we would all finally be better off. I think there are two factors that push against this though. First, if you learn to optimize then you can charge more for your labor, and you will likely get a job somewhere else. Second, companies exist for profit, and optimizing is not often the most profitable next step.
That's true and reasonable. And after this 5nm node, TSMC has 3nm and IIRC 2nm. We are at the point where throwing more hardware at slow software is becoming a non solution.
The good news is that in many cases there is a LOT of performance on the table on the software side.
That's not to say that there isn't an enormous amount of improvement available in software. Projects like simdjson prove that there's a 2x or more improvement available for many high-performance use cases by migrating to vector instructions.
I'm currently agonizing over an application that has an 80-120ms response time for 5 hops (including db writes). Yet I know many people who code in my language would find that number amazing for the work being done in those hops.
No SIMD. Nothing too crazy. Just well thought out processing flow and proper separation of the stupid.
The thing is - most web software isn't comparison shopped. It's developed on contract. I work at a company that does a lot of this, and, hypothetically, maybe we'd get hired to do an internal app for some retailer (imagine Best Buy or Walmart). When you build something like that, there's a gun to your head about time-to-deliver. They don't care if it's good — it's just got to meet a "reasonably decent" metric of quality. The only thing they really care about is getting it released on a certain timetable.
And all of this is because of contract law - timetables are provable breach-of-contract. You promised you'd make something by February, but you weren't done until April - that's something that's a provable failure-to-deliver. That's got financial penalties. But quality? Quality's gotta be astoundingly bad to be "provable". "Lag" or "the site's kinda slow" doesn't hold up in a lawsuit. (These things rarely lead to lawsuits, but do lead to 'punishment' by a refusal to renew a contract, knowing that the contractor can't sue the company because the contractor provably failed to deliver what was agreed on.) So they optimize for what's punishable according to the terms of the contract.
That's why all of this corporate enterprise software sucks - because nobody gets punished if it's slow or crappy; they just get punished if it's late, or if it's provably broken, because that's the stuff that's easy to pull out of a contract for.
And it's true that the people paying for it could go elsewhere, but page performance probably weighs pretty low on that list.
I would also use my iPad as a synthesiser (eg. Magellan) but then record the audio to a "real" computer, ie the Mac, via a recording interface.
Getting iSH on the App Store did take some work on our side; we were in contact with Apple to get it approved and in compliance with the guidelines. You may have seen that the version of iSH on the App Store had package management functionality removed. So far we've been able to keep the ability to run Linux programs, as well as import files, unchanged.
One interesting thing to see from Apple is that they are not just increasing the performance of their chips, but they are investing heavily in making certain operations faster. For example, in the time since tbodt started working iSH Apple has reduced the cost of uncontended atomics to almost zero overhead (this is especially good for Apple because it immediately makes reference counting code faster, which both Swift and Objective-C spend a lot of time in). iSH doesn't currently do anything fancy there, it just uses full locks everywhere, so this is something that we would be interested in pursuing at some point.
So while it would make sense that a new process would have new design rule that didn't map well to older EDA tooling, it's harder to see that argument holding for SRAM. If SRAM isn't scaling, why is general logic?
0.021um per cell [1] for TSMC N5 vs 0.027um per cell [0] for TSMC N7.
Even if Apple did everything right, they'd still be behind due to the limits of the process.
[0]https://fuse.wikichip.org/news/2408/tsmc-7nm-hd-and-hp-cells...
[1]https://semiwiki.com/semiconductor-manufacturers/tsmc/283487...
They're often literally made out of different metals that have different resistance characteristics at different wire gauges, for example.
Interconnects was a very famous case, but many aspects of electrical engineering change when you get that small, because hey if you shrink the diameter of a wire by half guess what happens to the volume.
Semiconductor engineering is the field of dealing with a thousand of these tiny problems, and an improvement in the photolithography wavelengths is actually an abstraction of solving the thousands of tiny problems involved with all the different materials.
Back in the old days we used to get a doubling of performance every year or two. Then Intel got into trouble and has been producing the same chips for what 4-5 years? Essentially just minor changes.
It’s not really Apple’s fault that the rest of the industry is lagging so badly on performance. My Nokia was fast enough, of course no one “needs” a faster phone. But it only appears so much faster because the rest of the industry has stagnated.
Had Intel kept their tick-tock cadence then these new "faster” phones would hardly be considered fast.
Also the current gen high end laptop CPUs from Intel are still 14nm. Which they introduced in 2015 with broadwell [1]. Others have shown that’s it’s very possible to scale beyond 14nm, so there is good reason to blame Intel here!
1. https://ark.intel.com/content/www/us/en/ark/products/87719/i...
I have a work-provided 16" MBP that does all these things effortlessly.
I'd suggest popping off the bottom and cleaning out the dust, I do this about once a year.
If one has support for VP9 and one doesn't, then most of that is not because the cores changed in performance, it's because the video decoder got rearranged.
https://en.wikipedia.org/wiki/Transistor_count
Has a big table listing a number of historical CPUs.
I think we just have to accept that silicon scaling is slowing down.
Isn't eDRAM an IBM-only technology?
Like if you have a box of dimensions 10x10x10cm the theoretical maximum number of dice of dimension 1x1x1cm you can fit in the box is 1000 dices. Chances are you won’t get 1000 dices in to the box as the world is more complex than the theoretical model we used for the scenario.
I've got a side-gig/hobby making stuff, and the the 2x tolerance works pretty well as a rule of thumb for me, but the project type, material, and application probably play a big part in those considerations. Any industrial/mechanical/materials engineer care to weigh in?
Apparently, R&D spend went up 50% from 7nm to 5nm. I'm curious to see how many flavors of the A14 Apple cooks up given the comparatively high die cost.
See e.g. https://www.cultofmac.com/658650/iphone-11-pro-max-productio...
> Apple A13 processor that (is) supposedly is priced at $64.
Worth it? Apple thinks so.
https://www.extremetech.com/computing/315186-apple-books-tsm...
[1] https://www.techpowerup.com/272139/samsung-foundry-to-become...
>The model is based on an imaginary 5nm chip the size of Nvidia's P100 GPU (610 mm2, 90.7 billion transistors at 148.2 MTr/mm2).
[1] https://cset.georgetown.edu/wp-content/uploads/AI-Chips%E2%8...
Is there a way to read the original report that this post was based upon?
Can't their other clients reach the same density?
One consideration is that while we think of each transistor as individual and isolated, in reality they are simply wells of n and p type silicon with some oxides and metal thrown in. Because of this we can get unwanted parasitic components and latches. A BJT transistor for example is just a pattern of three silicon types, so it is somewhat inevitable that these will occur, whether they cause issues, and to what extent depends on the design.
Packing everything as tight as possible gives noise concerns, thermal concerns, signal integrity issues ect.
This is not to devalue the achievement of TSMC, but to state that design work is still a work as well.
If you are just making a relatively low-speed L3, getting rid of the heat produced in the SRAM itself will never be a problem. In general, the problem in SRAM layered on chip is considered to be that the stacked dies are a very good insulator, so the hot active die on the bottom will have heat dissipation problems.
https://finance.yahoo.com/news/intel-now-ordering-chips-tsmc...
I sincerely doubt Intel will ever disclose it.
At 4 bits per transistor and 3 pins, how do you count the samsung bits? :)