Testing AMD's Bergamo: Zen 4c
chipsandcheese.com
chipsandcheese.com
As I said before, I do believe that this is the future of CPUs core. [1] With RAM latency not really having kept pace with CPUs have more performant cores really seems like a waste. In a Cloud setting where you always have some work to do it seems like simpler cores but more of them is really the answer. It's in the environment that the weight of x86's legacy will catch up with us and we'll need to get rid of all the waste transistors decoding cruft.
I suspect the future will be something between these extremes (tons of dumb cores or ever-more-complicated cores to try and squeeze out IPC), though.
In practice, this means that decode increases for branchy code, but non-branchy code will be limited to just 3 decoders. In contrast, an ARM X4 or Apple M4 can decode 10 instructions under all conditions.
This also play into ideas like Larabee/Knights processors where you basically want the tiniest core possible attached to a massive SIMD engine. x86 decode eats up a lot of real estate. Even worse, x86 decode adds a bunch of extra stages which in turn increase the size of the branch predictor.
That's not the biggest issue though. Dealing with all of this and all the x86 footguns threaded throughout the pipeline slows down development. It takes more designers and more QAs more time to make and test everything. ARM can develop a similar CPU design for a fraction of the cost compared to AMD/Intel and in less time too because there's simply fewer edge cases they have to work with. This ultimately means the ARM chips can be sold for significantly less money or higher margins.
How this affects decode clusters is left as an exercise to the reader.
Not if there are fewer than 10 instructions between branches...
…but when we start swarming simple cores, that cost starts to rise. Each core needs to be able to decode everything. Now when you can a 100 cores, even if the cruft is just 4%, that means you can have 4 more cores. This is for free if you are willing to recompile your code.
Now, it may turn out that we need more decoding complexity than something like RISC-V currently has (Qualcomm has been working in it), but these will be deliberate, intentionally chose instead of accrued, that meet the needs of today and current trade offs, and not of the eart 80’s.
One big one is probably around consistency model[1] and such which affects atomic operations and synchronizing multi-threaded code. Usually not directly though, I typically use libraries or OS primitives.
Are there any non-obvious (to me anyway) ways us "typical devs" rely on x86/x64?
I get the sense that a lot of software is one recompile away from running on some other ISA, but perhaps I'm overly naive.
Generally the answer is "we bought this product 12 years ago and it doesn't have an ARM version". Or variants like "We can't retire this set of systems which is still running the binary we blessed in this other contract".
It's true that no one writing "fairly standard software" is freaking out over the inability to work on a platform without AVX-VNNI, or deals with lockless algorithms that can't be made feasibly correct with only acquire/release memory ordering semantics. But that's not really where the friction is.
Otherwise, via translation techniques, getting 80% of native performance isn't unheard of. Which would be very fast relative to any such Ivy Bridge server.
Transitioning away from x86 definitely is feasible, as very successfully demonstrated by Apple.
For us the biggest obstacle is that our compiler doesn't support anything but x86/x64. But we're moving to .Net so that'll solve itself.
... is slowly dissapearing. Even on Windows 10 is very hard to run Win32 programs from Win95, Win98 era.
It is not. Or at least not the future, singular. Many applications still favor strong single-core performance, which means in, say, a 64-core CPU, ~56 (if not more) of them will be twiddling their thumbs.
> It's in the environment that the weight of x86's legacy will catch up with us and we'll need to get rid of all the waste transistors decoding cruft.
This very same site has a well-known article named “ISA doesn’t matter”. As noted though, with many-core, having to budget decoder silicon/power might start to matter enough.
That seems backwards to me. Narrower, simpler cores with fewer execution engines have a much easier time decoding. It's the combinatorics of x86's variable length instructions and prefix coding that makes wide decoders superlinearly expensive.
master jart@studio:~/blink$ ls -hal o/tiny/blink/x86.o
-rw-r--r-- 1 jart staff 23K Jun 22 19:03 o/tiny/blink/x86.o
Modern microprocessors have 100,000,000,000+ transistors, so how much die space could 184,000 bits for x86 decoding really need? What proof is there that this isn't just some holy war over the user-facing design. The stuff that actually matters is probably just memory speed and other chip internals, and companies like Intel, AMD, NVIDIA, and ARM aren't sharing that with us. So if you think you understand the challenges and tradeoffs they're facing, then I'm willing to bet it's just false confidence and peanut gallery consensus, since we don't know what we don't know.The problem is that superscalar CPU needs to decode multiple x86 instructions per cycle. I think latest Intel big core pipeline can do (IIRC) 6 instructions per cycle, so to keep the pipeline full the decode MUST be able to decode 6 per cycle too.
If it's ARM, it's easy to do multiple decode. M1 do (IIRC) 8 per cycle easily, because the instruction length is fixed. So the first decoder starts at PC, the second starts at PC+4, etc. But x86 instructions are variable length, so after the first decoder decodes instruction at IP, where does the second decoder start decoding at?
But in the specific case of tracking variable length instruction boundaries, that happens in the L1i cache. uOP caches make decode bandwidth less critical, but it is still important enough to optimize.
x86 decoders take a tiny but still significant silicon and power budget, usually somewhere between 3-7%. Not a terrible cost to pay, but if legacy is your only reason, why keep doing so? It’s extra watts and silicon you could dedicate to something else.
[0] https://chipsandcheese.com/2021/07/13/arm-or-x86-isa-doesnt-...
Like I said, you its only a small amount of extra silicon you’re paying the x86 tax with, but with the world mostly becoming ARM-compatible, there’s no more reason to pay it.
Also worth noting that a high-end GPU already has 16K cores, although the definition of a "core" in that context isn't as clear-cut as with normal CPUs.
These server CPUs are still being made with 5nm or 4nm technology. Sure, that's just a marketing number, not a physical size, but the point is that there are already firm plans from both Intel and TSMC to at least double the density compared to these current-gen nodes. Another doubling is definitely physically possible, but might not be cost effective for a long time.
Still, I wouldn't be surprised to see 4K vCPUs in a single box available in about a decade in the cloud.
After that? Maybe 3D stacking or volumetric manufacturing techniques will let us side-step the scale limits imposed by the atomic nature of matter. We won't be able to shrink any further, but we'll eventually figure out how to make more complex structures more efficiently.
It'll be a combination of nascent technologies on the cusp of viability.
First, something like this will have to be manufactured with a future process node about 2-3 generations past what is currently planned. Intel has plans in place for "18A", so we're talking something like "5A" here, with upwards of 1 trillion transistors per chip. We're already over 200 billion, so this is not unreasonable.
Power draw will have to be reduced by switching materials to something like silicon-carbide instead of pure silicon.
Then this will have to use 3D stacking of some sort and packaging more like DIMMs instead of a single central CPU. So dozens of sockets per box with much weaker memory coherency guarantees. We can do 8-16 sockets with coherent memory now, and we're already moving towards multiple chiplets on a board that is itself a lot like a large chip with very high bandwidth interconnect. This can be extended further, with memory and compute interleaved.
Some further simplifications might be needed. This might end up looking like a hybrid between GPUs and CPUs. An example might be the IBM POWER server CPUs, some of which have 8 hyper-threads per physical core. Unlike POWER, getting to hundreds of kilocores or one megacore with general-purpose compute might require full-featured but simple cores.
Imagine 1024 compute chiplets, each with 64 GB of local memory layered on top. Each chiplet has 32 simple cores, and each core has 8 hyper threads. This would be a server with 64 TB of memory and 256K vCPUs. A single big motherboard with 32 DIMM-style sockets each holding a compute board with 32 chiplets on it would be about the same size as a normal server.
M1 really is a counterexample to all theory that Jim is saying etc. The real proof would be if same results were also reproduced on M1 instead of Zen
Apples M series looks very impressive because typically, at launch, they are node ahead of the competition, so early access deals with TSMC is the secret weapon this buys them about 6 months. They also are primarily laptop chips, AMD has competitive technology but always launches the low power chips after the desktop & server parts.
Aren't Apple typically 2 years ahead? M1 came out 2020, other CPUs from the same node level (5 nm TSMC) came out 2022. If you mean apple launches their 6 months ahead of the rest of the industry gets on the previous node, sure, but not the current node.
What you are thinking about is maybe that AMD 7nm is comparable to Apple 5nm, but really what you should compare is todays AMD cpus with the Apple cpu from 2022, since they are on the same architecture.
But yeah, all the impressive bits about Apple performance disapears once you take architecture into account.
There only seems to be comparisons between laptop CPUs which are quiet limited.
Apple is just that good.
I have doubts about that. I-cache word lines are much larger than instructions anyway, and it was the reduction in memory fetch operations that made THUMB more energy-efficient on ARM (and even there, there's lots of discussion on whether that claim holds up). And if you're going for fixed-width instructions then many instructions will use more space than they use now, reducing the overall effectiveness of the I-cache.
So even if you can prove that a fixed-size decoder uses less power, you will still need to prove that that gain in decoder power efficiency is greater than the increased power usage due to reduced instruction density and accompanying loss in cache efficiency.
That's not too bad for the first instruction in a line but the second instruction is dependant on how the first instruction decides, and the third dependent on the second. Etc. So it's not only a big multiplexer tree, but a content dependent multiplexer tree. If you're trying to unpack multiple instructions each clock (or course you are. You've got six schedulers to feed) then that's a big pile of logic.
Even RISC-V has this problem, but there they've limited it to two sizes of instruction (2 and 4 bytes), and the size is in the first 2 bits of each instruction (so no fancy decode needed)
The UltraSPARC T1 was designed from scratch as a multi-threaded, special-purpose processor, and thus introduced a whole new architecture for obtaining performance. Rather than try to make each core as intelligent and optimized as they can, Sun's goal was to run as many concurrent threads as possible, and maximize utilization of each core's pipeline.
In my experience, people just don't know how to build multi-threaded software and programming languages haven't done all that much to support the paradigm.
Multi threading is still the domain of gnarly bugs, and specialists writing specialist software.
The only kind of forward looking thing I've seen in this area is the Bend language that has been making strides a couple months ago.
And besides all that, Amdahl's law still exists - if 10% of the program cannot be parallelized, you're going to get a 10x speedup at most. Every consumer grade chip tends to have that many cores already.
Go? Rust? Any functional language?
Just fire up the Windows Process Explorer and look at the CPU graphs.
Rust in my opinion, is the biggest admission of failure of modern multithreaded thinking, with having classes like 'X but single threaded' and 'X but slower but works with multiple threads', requiring a complex rewrite for the code to even compile. It's moving all the mental burden of threading to the programmer.
CPUs have the ability to automatically execute non-dependent instructions in parallel, at the finest granularity. Yet if we take a look at a higher level, on the level of functions, and operational blocks of a program, there's almost nothing production grade that can say: Hey, you sort array A and B and then compare them, so lets run these 2 sorts in parallel.
It is not even there. Windows (7,10) has difficulties splitting jobs between I/O and processor. Simulations take hours because of this and because Windows like to migrate tasks from one core to the others.
In Linux, I/O threads are real, with true asynchronous I/O only being recently introduced with io_uring.
1) In the cloud there are always more requests to serve. Each request can still be serial. 2) Stuff like reactive streams allow for parallelisation. The independent threads acquiring locks will forever be difficult, but there are other models that are easier and getting adopted.
It should be noted that the successor of Bergamo, Turin dense, which is expected by the end of the year, will have 12 compute chiplets, for a total of 192 Zen 5c cores, bringing thus both more cores and faster cores.
With the right preprocessing, it should be possible to essentially stream from NIC straight to CPU, process some stuff, and output it straight to the NIC again without ever touching RAM - although I doubt current DMA hardware allows this. It'd require quite a bit of re-engineering of the on-wire protocol to remove the need for any nontrivially-sized buffers, but I reckon it isn't impossible.
this already exists [0]
[0] https://www.intel.com/content/dam/www/public/us/en/documents...
It's perfectly possible to load the same OpenWRT release on a modern x86 and run in the same amount of ram with much higher performance. That's how I do my routing on a 1gbit/s synchronous fiber connection and I can do line-speed wireguard just fine.
Distributions like OpenWRT require so little RAM because they select very conservative options when compiling everything, are very selective about what they include, build against minimal C libraries like musl libc, make use of highly integrated system utilities like busybox, and strip things like debugging symbols from the resulting binaries.
Otoh if you just mean try to write software to minimize the need to hit memory, that's totally reasonable -- and what you have to do if you want the best performance.
Until someone figures out how to do 3D stackable SRAM similar to how SSDs work, SRAM will always consume most of the area on your chip.
There is some outstanding uncertainty about cache coherency vs performance as N goes up which shows up in numa cliff fashion. My pet theory is that'll be what ultimately kills x64 - the concurrency semantics are skewed really hard towards convenience and thus away from scalability.
Are you saying that the x86 memory model means RAM latency is more impactful than on some other architectures?
Is this (tangentially) related to the memory mode that Apple reportedly added to the M1 to emulate x86 memory model to make emulation faster? - presumably to account for assumptions that compilers make about the state of the CPU after certain operations?
It makes life easier for programmers of multithreaded software, but at the cost of high synchronization overhead.
Contrast to e.g. Arm, where programmers can avoid a lot of that synchronization, but in exchange they have to be more careful.
* IIRC from some recent reading.
[memory (shared)]
⇕
[L3 cache (shared)]
⇕
[core 0] ⟺ [L1 cache (private)] ⟺ [L2 cache (shared)] ⟺ [L1 cache (private)] ⟺ [core N]
⇕
[core 1] ⟺ [L1 cache (private)]
Each CPU (or CPU core) has its own private L1 cache that other CPU's/CPU cores do not have access to. Now,
code running on the CPU 0 has modified 32 bytes at the address 0x1234 but the modification does not occur directly in the main memory, it takes places within a cache and changes to the data now have to be written back into the main memory. Depending on the complexity of the system design, the change has to be back propagates through a hierarchy of L2/L3/L4 (POWER CPU's have a L4 cache) caches until the main memory that is shared across all CPU's is updated.It is easy and simple if no other CPU is trying to access the address 0x1234 at the same time – the change is simply written back and the job is done.
But when another CPU is trying to access the same 0x1234 address at the same time whilst the change has not made it back into the main memory, there is a problem as stale data reads are typically not allowed, and another CPU / CPU cores has to wait for the CPU 0 to complete the write back. Since multiple cache level are involved in modern system design, the problem is known as the cache coherency problem, and it is a very complex problem to solve in SMP designs.
It is a grossly oversimplified description of the problem, but it should be able to illustrate what the parent was referring to.
I think the sibling comment explains it - x86 makes memory consistency promises that are increasingly expensive to keep, suggesting that x86’s future success might be limited by how much it can scale in a single package.
(for 95% of usage - somen niche things can utilize resources better, those not coincidentally often being the focus of benchmarks)
https://www.amd.com/en/products/processors/technologies/3d-v...
I believe the Knight-series chips were killed because they ran into a bunch of issues and politics with Knight's Mill then killed it in favor of their upcoming GPU architecture only for that to have issues and be delayed too.
I suspect that a RISC-V version of Larabee would work better as the ratio of SIMD to the rest of the core would be a lot better.
Cutting all that stuff and saving power means you can squeeze in a couple more cores and improve performance that more. 32 registers a significant advantage too.
RVV also allows wider vectors than AVX. Without reordering, you can get stronger performance guarantees from a single, wide SIMD than from two narrow SIMD (it’s also possible to look at the next couple instructions and switch between executing one at full width or two at half width each without massive investment).
And we haven’t even reached peak advantage yet. RISCV has the option for 48/64 bit instructions with loads of room to add VLIW instructions. While most VLIW implementations require everything to be VLIW, this would be optional in RISCV. Your core could be dual issuing two wide VLIW allowing a single thread to theoretically use four vector units. If you add in 4-8 wide SMT, the result is something very GPU-like while still using standard RISCV.
Maybe x86 can do this too, but I don’t think it would be anywhere near as efficient and it would certainly eat up a lot of the precious remaining opcode space.
The first cores using the final extension are hitting the market this year, so addition is happening really quickly.
Epyc Naples was also 32c and also 2017. 4- and 8-way SMT wasn't something AMD or Intel ever bothered with, no, but IBM also did with the POWER line. It seems to help in some workloads, but probably isn't generally useful enough for volume solutions.
Would be interesting to see a compute bound perfectly scaling workload and compare it in terms of absolute performance and performance per watt between Bergamo and Genoa.
Zen4C is said to be a little better performance per watt than Zen4, but it's not clear how much.
It's got a lower clock ceiling. Not much lower than acheivable clocks in really dense zen4 Epyc, but a lot less than an 8 core Ryzen.