AMD EPYC 9754 Benchmarks for the 128-Core Bergamo
phoronix.com
phoronix.com
I'd also love to see software like Firefox trying to leverage these abundant puny cores, because the average corporate laptop now does 8 threads or more and that number is only going to get higher, with mixed beefy/puny cores.
Zen 5c, although the ~25% increase in per Core improvement doesn't seems that interesting in these kind of workload. Until AMD could fit 256 Core in a Single Socket. Definitely achievable with 2nm, although we are looking at 2026+.
Also worth noting that AMD could drop in ARM chiplets as well as x86, so there is some flexibility there as well if there's enough demand. As it stands, this will bring a lot of compute density for a lot less power draw than previous generations of servers.
Latest GPUs have 10,000+ CUDA Cores, so, if you can paralyze the work, evidently pretty damned far.
Of course we'll need ever faster versions of PCI-E to come out to feed these beasts with enough data they don't sit idle most of the time.
More comparable would be the ~130 streaming Multiprocessors of a H100.
https://stackoverflow.com/questions/58071834/why-does-each-t...
NVIDIA has never given a good explanation about what they mean by the "instruction pointer" that belongs to each NVIDIA "thread". It certainly does not mean what in means normally, i.e. a special register that contains the address from where the next instruction will be fetched for execution. I believe that this "instruction pointer" refers to a register where the actual instruction pointer is saved when a "thread" is stalled because it has diverged into two branches after a condition test and only one of the branches continues to be executed, while the other branch must be executed later, with the complementary predicate.
These saved instruction pointers are presumably used for scheduling the "threads" to be executed by the SIMD lanes provided by the hardware, in such a way as to satisfy the cross-lane dependencies.
It really depends on the cache and memory architecture. You need to be able to feed data to the CPU efficiently.
There's a point where you won't be able to pack more transistors on a chip, but if Cerebras' work is an indication, chips themselves can get quite large.
But remember these chips are not designed with desktop users in mind. These chips are more similar to the ones inside very large computers that are built to host hundreds of simultaneous virtual machines on behalf of hundreds of users.
Sadly, we won't see desktops based on these unless we custom order them.
And STH recently posted an article about how they're used: https://www.servethehome.com/100m-usd-cerebras-ai-cluster-ma...