Oh, so, I got lost in the appendix explaining network topology, but to circle back on this: yes, yield, and cost. It would be a very large total area too even if it yielded well. Just too expensive in general.
Alder Lake 8P8E: 215.25mm2
Alder Lake 6P0E: 162.75mm2
Raptor Lake 8P16E: 257mm2
For comparison:
8700K: 149.6mm2
9900K: 174mm2
10900K: 206.1mm2
Zen2 CCD: 74mm2
Zen3 CCD: 83.74mm2
That's actually pretty big for a consumer processor already. And it's all in monolithic 5nm(-tier node), which isn't cheap even if it yields fine. So them having a uarch that's at a pretty bad area disadvantage isn't good, and tbh they obviously aren't delivering on any kind of efficiency promise.
Physics is getting hard and wafer costs are spiraling pretty bad, which is why AMD is exploring advanced packaging/etc. Doesn't always work though - like RDNA3. Data movement still seems to be very expensive, although 2.5d and 3d stacking (and direct-bonding) will mitigate this somewhat. But advanced packaging means moving a lot more data, and you have to be careful of what lives on what side of what links. Cache being on the other side of the infinity links (not infinity fabric!) in RDNA3 seems like potentially a specific problem with the design, since you pay the cost for the data movement to the cache and not just the data movement for the memory.
I think you're right they probably could do it if they wanted etc, maybe sell it as a pseudo-HEDT (especially if you can glue together a pair of dies directly to 2x the normal core count - and if you can glue together 2x16C all-P-core designs that's fine for HEDT for a lot of things imo!). But the price would probably be fairly high (16C would be like, probably $700-900) and the power would still be quite high (intel does not win at any power bracket right now even with limits, it's just less bad if you limit it to 150W), etc. Maybe some of the power stuff would go away if you got rid of the split-brain big/little clusters on a ring thing, but, even if you went with 16 P-cores on a ring, the latency would still go up a lot, and you'd notice it because the stuff you want it for is gaming/etc. The latency would hurt gaming IPC a decent chunk imo, or you'd have to go to a double-ring like broadwell.
It's a mess and this is the point where the ringbus scaling craps out, is my point with the latency discussion. It seems hard to have more than about 8C or 12C per "tier". Even Bergamo (AMD's new e-core variant of Epyc) is 16C of Zen4C per CCD, but it's 2 CCXs of 8. Broadwell dual-ringbus is 2 tiers of 12 cores each. The subsequent Intel chips moved to the mesh. Alder/Raptor do 8P+4 e-core clusters (12 nodes). Etc. You can add more tiers of 8-12, but about 8-12 nodes per tier seems to be the limit that scales well due to interconnect bandwidth/etc, just historically imo. Interesting convergence.
(plus a couple nodes for pcie agents and iGPU and memory controller and shit I'm not counting here, not stops just just cores)
https://www.anandtech.com/show/10158/the-intel-xeon-e5-v4-re...
I think strategically they want and need to keep selling the e-cores though, it's not what's right for you, it's what's right for them and their migration path. Some of these pieces it's hard to see how you do everything in a single go - it's tough to go from "everything is the same" to "lol CMT with 3 slow/1 fast thread controlled by this thread director that wants to talk to your OS scheduler". But the theoretical end-state of "big.little within a core cluster" or "within a CMT core" is pretty neat at least, that would mitigate the latency problems of dedicated "little core clusters". And this is one of jim keller's ideas apparently, while he was at intel (briefly, lol)
Now again, to say something nice here: the p-cores are pure out-of-order monsters. Very wide decode units, lots of execution resources, etc. It is the same as the "zen4 vs zen4-X3D" split, stuff that prefers zen4 over x3d also really prefers raptor cove, it's an execution monster and it does it all on just 8 p-cores. it just also uses more energy to do it, and cache can handle some other situations where the working set helps cover some useful working set. And the e-cores do give you a ton of the performance equivalence of having a wider AMD processor in MT workloads (if they're not latency-sensitive, ie cinebench and video encoding), just not at particularly great power compared to AMD's 16 p-cores clocked much lower.
I like my Atom processors a lot (and I've looked seriously at denverton etc). They're not bad cores at all, but Alder/Raptor just put them in the worst possible place, with too much voltage (no DLVR!), meaning you might as well goose the whole thing and go for power etc, and the latency isn't flattering.
I would have loved the idea of Sierra Forest-HEDT with like 192 / 256 cores (or whatever specifics) enabled or whatever, if they could get that to a relevant price for enthusiasts it'd be amazing, the HEDT market is super dead and the e-cores are good enough nowadays. That would be a super high-value place to put a product, if it's not moving adequately in the server market etc. Give me 5820K level value for something the client platform cannot do, and Intel benefits from actually getting a foothold on a market that is willing to tinker and build stuff.
But everything intel does has come so late that it almost doesn't matter, it's a sidegrade at best, at much worse efficiency etc. Just buy a 7700X or 7800X3D or 7950X/X3D. Let alone the slaughter in the server market etc. Sapphire Rapids is not great either, lots o weird power shit, and maybe it needs another stepping to fix it? ok, or, if you're a hyperscaler, amd will ship you zen5 samples probably within the next 3 months, and they can actually execute and deliver.
Hard to see that changing anytime soon either, I don't even see green shoots, I think they're in a death spiral tbh. I see the 2.5gbe nic team failing over and over (I225/I226), I see sapphire rapids taking far too many steppings and base die changes, I see DLVR still not working, I see no plan for AVX-512, I am guessing they have continual problems with integration/packaging, I see every product being a one-off with no reusability, etc. They're too important to let go under but they're in deep shit and they have to do a turnaround on an understaffed underpaid employee force etc, while battling deep internal-culture rot and middle-management warfare etc. It's gonna be a while before they're relevant imo.
I, too, have looked at big.little in 12/13th gen laptops and just ehhh is the world ready for that yet? I thought 11th gen was kinda attractive for that reason, the last all-p-core uarch (oh and you get avx512 too, etc). Orrrr you just buy something with a 7940HS/HX/whatever and get 8 big old Zen4 cores... with AVX-512.... and that's every product segment and it's going to get worse, pretty much. Intel still coasts on massive availability but at a technical level AMD is clowning them in pretty much most enthusiast or enterprise use-cases.
(one exception, usually, is system stability. AMD's AM4 USB dropout glitch isn't really fixed despite lots of effort, neither is the AM4/5 fTPM stutter bug, and a physical TPM header is something to look for on an AMD board lol, because fTPM is broken and causes random stutter (TPM operations getting blocked by some single-threaded UEFI process in vendors' UEFI implementation, is the internet speculation). Intel mostly is better about not having that shit, for now. In the past I have heard of a lot of problems with AMD chips in servers too, just weird linux problems etc, (and of course segfault affected a lot of early Zen1/1000-series chips in a lot of things, the scope was downplayed p. bad), but that's scandalous hearsay and tbh today I think Asrock X570D4U-2T or ROMED8-2T or GENOAD8X-2T owners etc are happy, people would report problems etc. Intel does have the support story of being the default... right up until they won't anymore.)
They have some neat ideas but AMD is gonna leap forward again with Zen5 too, everyone has neat ideas that will be coming to fruition in 3-5 years. Zen4 was the easy stuff - a pretty minor port of zen3 to 5nm, with DDR5, with AVX-512, and some tweaks to open up architectural or timing bottlenecks to push clocks etc, while they did the DDR5 switchover stuff. Some cleanup (and they got it pretty much to 6 GHz in peak 1T lol) but all the interesting stuff is coming next year in zen5, it's gonna be much wider etc (similar to intel's own width increase with golden) etc, it's expected to be a pretty significant increase. this is a big rework of the whole thing to clean up and scale higher. so this 14th-gen stuff will be going up against zen5 being probably 20-30% faster in general performance, it'll be a pretty decent uplift. They're in trouble in pretty much every product segment already, it really seems like they struggle to even get the product out the door these days.
https://en.wikichip.org/wiki/intel/microarchitectures/alder_...
https://en.wikichip.org/wiki/intel/microarchitectures/raptor...
https://www.techpowerup.com/297506/intel-raptor-lake-core-i9...
https://en.wikichip.org/wiki/intel/microarchitectures/coffee...
https://www.techpowerup.com/267649/intel-core-i9-10900k-der8...
https://en.wikichip.org/wiki/amd/microarchitectures/zen_2#Di...
https://wccftech.com/amd-ryzen-5000-zen-3-vermeer-undressed-...
https://wccftech.com/amd-epyc-bergamo-cpu-die-detailed-16-ze...
https://en.wikipedia.org/wiki/List_of_Intel_Core_i7_processo...
https://en.wikipedia.org/wiki/Intel_Graphics_Technology
https://en.wikipedia.org/wiki/Tegra#Models
https://en.wikipedia.org/wiki/CUDA#Version_features_and_spec...
(just wanted to call out: wikichip is a lovely site, lots of randomly useful information there. and wikipedia also has a number of useful lists of cpus/gpus/mobile SOCs/etc with characteristics listed, and good sources for CUDA compute capability etc. Don't sleep on wikipedia as a quick reference for what the boost is on random xeon sku xyz, and so on.)