> it's TSMC's flagship process that enables most of its competitive advantage.
Eh. Right now Apple is matching a (desktop) 5950X at laptop level TDPs - the iGPU is way lower than a desktop 3080 but so is a mobile 3080 at 115W. If you doubt that number, a Pugetbench result just showed up. TBH the specifics of the"which processor is 5% better than the other here" slapfight doesn't actually matter here, compared to the fact that Apple is doing it in what, 60W? And the AMD+NVIDIA system is using a 300-350W GPU and a 142W PPT CPU. Yeah, that's high-clocked kit, but with results like that there's zero reason to doubt Apple's claim of it beating a (mobile) 5800H (probably sustained 62W PPT) with a 115W mobile 3080.
M1 Max: https://www.pugetsystems.com/benchmarks/view.php?id=60176
Desktop 5950X/3080: https://www.pugetsystems.com/benchmarks/view.php?id=60613
pugetbench premiere pro definition: https://www.pugetsystems.com/labs/articles/PugetBench-for-Pr...?
https://i.imgur.com/dpBHGxN.png
Apple has never fibbed or bent the truth on their marketing benchmarks - and they have sandbagged before, the numbers have sometimes been higher than presented. They have nothing to be ashamed of here, this is an extravagant chip, it's twice the transistors of a 3090 to make your desktop 3060 ti / mobile 3080 run at 45W, and a near-fastest-in-market product in 15W for the CPU, enough to saturate the GPU for most things you'll want to do. In order to do that, they put the equivalent of an 8-channel DDR4 bus on it (it's 16-channel on DDR5 and DDR5 is significantly faster) and then used tiny LPDDR5 packages they could stack. Oh, and you can allocate any of the system RAM as VRAM, so you can do large-memory VRAM tasks like machine learning (albeit slowly of course - it's desktop 3060 Ti performance). For a portable workstation it actually pretty well owns.
Also the regular ol M1 had as much cache as a 9900K. For its 4 performance cores + 4 efficiency cores. The whole A15 design is just an extravagant display of transistors, Apple spared no expense, it's just a design exercise in "what if we built it ourselves and ran up the scoreboard and called that our 'external equivalent' expense". Apple will spend $300 to build the meme 5nm (nearly) icelake-SP-sized mobile workstation chip because a 5800H or 12Gwhatever would cost them $200 anyway and this thing dunks on those processors. Vertical integration matters at their scale, just like Google and Amazon and others.
With Zen4 on DDR5 and TSMC 5nm (actually N5P - it's a better node than what Apple is using for A15/M1) next year we'll see how true that actually is about it "all being node advantage".
I bet AMD can catch up to laptop-power-budget A16 (the next-gen Apple core that Zen4 will compete against - also on N5P) in performance with 142W PPT processors ("105W TDP") like their 5950X equivalent, I'm kinda doubtful they can close a factor-of-3 gap in perf/watt at laptop-level PPTs.
Yeah, Apple is clocking a lot lower and that's where some of the efficiency comes from... but look at the IPC difference here. They're matching processors that are running roughly twice as fast, at peak efficiency clocks. It's an insanely wide architecture, and it's hard to scale x86 like that - x86 instruction decoding has bad transistor-order-complexity for higher widths, due to variable instruction lengths and other problems. You can make all the execution units you want in the core, but it's hard to keep them fed if you can't decode enough instructions. x86 has already done a lot of the obvious tricks like instruction caching/etc.