> Delightfully fun write up. I made a pretty superficial jab, but you've really painted a great picture of microprocessor design/tradeoffs as they've happened.
PaulGPT aims to deliver. Not actually an AI, just frequently accused of being one because I read some shit and it triggers Opinions and Tangents. And I don’t stake any position lightly, I have More Opinions why I’m right lol. It leads to Controversy. But I’m perfectly willing to defend my opinions against counterarguments and ultimately I’d rather mald and then admit I’m wrong.
> Definitely have a deep love of the "communication processor" grade gear,
I really wish Denverton had been more available. I can't even bite on surplus enterprise gear because there's barely any out there. Same with xeon-D, sick on paper but way too expensive. Plz make the intel accelerator thingy just onboard everything and also useful for zfs checksumming, that would be a gamechanger for ZFS on NVMe :\
> I did really like Lakefield, which tried to be a ultraportable capable 1P 4E alike system
Yes as I have commented, I really have been a fan of Kabini (Athlon 5350), Airmont (N2808), and Goldmont Plus (J5005) and recently I landed a pair of Skylake NUC7i7 for $125 a pop as well. They still are compelling for certain "microserver" applications given their extreme low price - a $50 CPU+mobo or a $125 booksize changes the expectations. 10 years ago it was $50 for a 5350 and mobo, 8 years ago it was $125 for a 2GB/32GB ECS Liva X, 3 years ago it was $125 for a J5005 NUC, recently $125 for a barebones with thunderbolt support? Yes, I like cheap machines even if their power is limited, $150 for a barebones machine that offers a faster capability at low TDP or some other unique capability is fine with me.
I am looking to use a RPi4 for a local NTP stratum-1 server with GPS and maybe use some of the nucs or other minipcs for freeIPA or a wireguard bastion or similar. With sufficient RAM a J5005 NUC actually made a really nice thin client during COVID WFH - swapping completely tanks performance and 16GB+ makes it perfectly fine even with lots of tabs/etc.
I have my eye on the Atlas Canyon NUCs too, which are finally available in quantity. The only things I don't like are the reduction to single-channel (Which seriously impacts performance vs the expected scaling, especially in iGPU) and the continued lack of Thunderbolt/USB4 - I like that it finally has M.2 NVMe and some other niceties but it really needs USB4 so you can plug it into more powerful stuff if desired. External expansion is going to be very baseline once USB4 reaches saturation and it’s not going to be that many years.
> rumors, that the Intel Meteor Lake SOC die might have it's own integrated E-core island? That seems insanely bright; just turn off the core complex,
Yeah that would be a cool workaround to the data-movement power penalties of chiplets/tiles. I mentioned elsewhere but having to have IF links powered up just to have cores idling along is an obviously dumb thing and yeah that’s a good solution for it, turn off the IF links and just run shit on the IO die.
Data movement and general idle-power is obviously the penalty of MCM and the more data you move across the more (smaller) chiplets the higher it is. It’s all just some new asymptotic limit of scalability (which direct bonding like cu-cu will change again, at the cost of thermals).
Intel claims EMIB is less than an interposer… I’m not really clear how the fiberglass-style interposer vs silicon interposer vs EMIB all stack up in practice. EMIB is 900% crucial to Intel’s future though, remember that TSMC offers advanced packaging solutions and (just like sapphire rapids getting their shit together so Intel has a viable cell library+process to sell to custom foundry) getting their shit together on EMIB is going to be a mandatory requirement for custom foundry’s success. The idea that a couple major intel chiplet/tile based products are seeing a lot of delays is generally concerning.
> I think that task of actually understanding when real P-cores really should be brought up is an interest challenge facing consumer computing today
I think phones have pretty good solutions for this but consumer and server both may have a different optimum for poweriness and boostiness. I think an optimum solution probably is not computable for the same reasons most “sufficiently advanced compiler-magic” doesn’t work - you’ll only really know at runtime. Still I am highly in favor of whatever hinting schemes we can come up with - good dev behavior can make a lot of difference.
Apple’s cores are really interesting in this area. Blizzard is extremely fast and small (like 1/3 of a Gracemont core transistors for similar-ish performance? I’ve never seen exact numbers but broadly that’s how it stacks up) and honestly Avalanche is sick, especially for any sort of JVM task. It’s just generally good at JIT, it’s not just x86 it really just crushes JIT compared to other architectures and JVM falls into that too. Cinebench is underselling the perf/w at load (because of less frontend load on x86) and I think the idle power stuff is undersold too. When it’s really on Linux I think the numbers are going to be impressive.
What’s a real easy answer to “when should I run p-cores”? Just have a real fucking fast e-core and if the e-core complex is getting overrun on a prolonged basis, pull out the really big guns.
> Not as relevant to the discussion so far, but just gonna mention: rumor-mill this week is that Genoa, the large Zen4 epyc chips, which were expected real-soon-now, are alleged to be facing some significant delays.
I did not know this, welp. Intel catches a bit of a lucky break. I think their mindshare is damaged though, even if Genoa were 6 months after SPR-SP nobody would really care tbh, intel’s been a lot later and AMD shows signs of better execution these days.
Sapphire Rapids is a good chip though. It’s not “lol throw your AMD shit in the dumpster” tier good, but, the game is back on, it’s good enough Intel can sell it, especially if they continue to be willing to cut deals. They need cash, they’re desperate, and this is obvious even now, if you’re paying more than half list price for intel you’re a complete fucking chump and quarter or less is more typical. Same for the 10-series during pandemic and the 12/13th gen prices, Intel is willing to deal to keep the fabs busy and keep the cash coming.
https://www.youtube.com/watch?v=_2yjjHzifL8
> Calling out Knights Landing as an SMT4 chip with a big vector unit is extremely on target, extremely interesting.
That’s never occurred to me either but the framing of “lol what if bulldozer but with AVX-512” rubbed both those nerves at the same time. No fuck you what if that already exists and everyone hates it!?!? ;)
(And then the transformer model takes over, beep boop paulGPT online, that's an interesting one for the following reasons ;)
> That question of how bad the scheduling complexity really is is a compelling question
Yes I agree, that is the money question. How much area did AMD save by having some fixed portions of the pipeline that don’t have to be scheduled between threads? How much does Sun save on Niagara? Or Intel on Core-SMT2 or Phi-SMT4? This would be an extremely interesting three-way from chips+cheese or similar, how well did all those sets of tradeoffs work and why were respective decisions made for those designs/use-cases? They all made different decisions around their frontend, I should take a peek at agner fog’s microarchitecture on those uarchs sometime.
> I suspect it probably is not really that huge a barrier.
That’s my guess as well. Alternating threads on the frontend may be the worst of it. If you have fixed decode/fetch units per-thread ala HyperThreading (or the option of a split of 5/0, 3/2, 2/3, etc) maybe that’s most of the squeeze. IDK though.
> Still, we seem lament to explore much beyond SMT2, for the time being, but maybe that's fine.
No… there is another ;)
https://en.wikipedia.org/wiki/POWER9
TBH I really really want a TALOS II setup, that is going to be one of those things that I snipe in 10-15 years when they’re cheap, it’s a neat piece of hardware. I have an AMD Quadfather and a KNL pcie coprocessor (we have a discord, folks!) and some other assorted nostalgia-tech too.
https://discord.gg/2qJXMTmE (expiring link to control spam, feel free to ping me on another tech thread later if anyone needs)