Everything I've learned building the fastest Arm desktop
jeffgeerling.com
jeffgeerling.com
Never thought I'd live in a world where drivers are published for Linux first. This is great.
Nvidia sells ARM servers with 8-16 GPUs to my knowledge.
So that makes sense
The original 1.2 teraflop CPU number also smells funny; 128 cores * 2 neon units / core * 2 doubles / neon unit * 2 flops / cycle * 2.8 billion cycles / second = 2.8 TFlops peak, an even halfway decent BLAS will get you to 80% of that, for 2.3 TFlops. Double that number for single precision. Either a completely untuned BLAS or benchmarking a problem that's much too small for the CPU under test.
The CPU benchmark used HPL linpack following a standard-ish Top500-style benchmark, mostly because I think it's fun to see how various single CPU systems compare to historic 'supercomputers' on the official lists.
There are different ways to calculate (and benchmark) flops, the way I'm doing it is with this open source project: https://github.com/geerlingguy/top500-benchmark — which can use OpenBLAS, Blis, or ATLAS, and I've tried all three un-tuned.
I also worked with some Ampere engineers to run their optimized version for Ampere Altra Max: https://github.com/AmpereComputing/HPL-on-Ampere-Altra
On their own server systems with 8 channels of memory they topped out around 1.6 Tflops. There are other tests you can run and get more or less Tflops, but I based my own results on the top500 test.
See more discussion about the results and testing in the following links:
https://github.com/geerlingguy/top500-benchmark/issues/19 https://github.com/geerlingguy/top500-benchmark/issues/17 https://github.com/geerlingguy/top500-benchmark/issues/10
(And see some of the issues linked back to the Ampere repo in those issues.)
Edit: And regarding the video card mention—it was meant more as a generic reference (e.g. running A100 or 4090 since you can go much further there than Altra Max), and not specifically to the 4070 Ti... but I can see how that is not as clear!
I'm also curious whether you think Apple's decisions on memory architecture (despite being non-upgradeable) will have a leg up in the long run. You mentioned that memory bandwidth tops out around 174GB/s. Although you handily beat the Mac Pro in a multicore benchmark thanks to core count, one of the Mac Pro's claims to fame is its memory bandwidth of 800GB/s, as well as its unified memory architecture.
Fast & efficient cpu cores, perhaps slightly memory bandwidth limited, sounds like a great compile/build machine.
Any chance of doing a Linux kernel compile, and time that? Bonus points for 2nd, 3rd etc run with much of the stuff in RAM.
Curious minds would like to know.
But it can only get better from here. I'm glad that PC users are starting to benefit from Apple's "desktop ARM" revolution.
But on "gaming rigs" where people are content to eat nvidia's absolute bed-shitting power consumption, and where replacing ram is a rallying cry? Not going to happen. It'll take until they're dragged along by the other two segments, sometime in 2-3 hardware generations. Maybe 2029-30 at the earliest.
I like your estimate, but I think it's on the conservative side by at least a couple of years. The migration from x86 to ARM will be slow until it isn't.
¹ https://www.statista.com/statistics/1003576/gaming-pc-shipme...
i'm sure people gaming on windows will appreciate that too -- given the success of the steamdeck. if gamers could get 2K or 4K gaming. on machines using 15-30W whether ARM or x86. then similarly us dev's would benefit as well.
I don't see PTC, Cadence, Autodesk etc. releasing any ARM binaries soon.
Windows has been multithreaded for decades across all layers, even classical Win32 written in C is using multiple threads.
Android, ChromeOS and macOS likewise.
In the end, single core performance being that relevant is quite niche case in modern software stacks.
Even in multithreaded desktop applications, it's rare to see them effectively use more than 8 threads.
I would still rather have 128 M1-class cores than 128 Neoverse-N1 cores :)
macOS, ChromeOS, Windows and Android are heavily multithreaded, even when there is one main application thread, the underlying OS APIs are using auxiliary threads.
Easily observable in any system debugger.
The easy way to get around this is to have heterogeneous cores. Ampere could have easily slapped 2 A76 cores for every 32 Ampere cores to prevent this problem.
The harder way to get around this is to multithread the partitioning algorithm, but that will convince more people that multithreading isn't worth it, because their Ryzen runs the same code faster even though it might only have 12 cores.
Hopefully Qualcomm Oryon will spice things up on the client side a bit. Maybe we can get some real HEDT designs after that.
I'm excited for the future of ARM desktops, but that also means they need to significantly increase single core performance. It looks like maybe the Snapdragon X Elite will get us there? Or at least significantly closer.
I would love to see an EPYC-style Arm CPU with M1 or better cores.
The benchmark results are great, with the exception of the single core speeds. We just need non-Apple ARM to catch up in the single core speeds across the board -- they seem particularly slow here compared to everyone else. Even Intel and AMD I think outdo Apple ARM now on single core.
I think it is a bit of a shame that Apple's top of the line M1/M2 go to 24 cores and the rumoured M3 CPUs only go up to 32 cores at most. I think the Ultra should have 64 cores personally. Rendering and video people want that.
Yeah they do currently. But, if the single-core performance jump we see between A15 and M2 (same architecture), about 25%, can be replicated between A17 and M3, we should expect Apple to retake the single-core performance lead, as M2 is only ~8-12% behind. Its not unreasonable to expect a perf jump like this, given its mostly just shoving more watts through the same core architecture, and Apple's tremendous efficiency lead gives them more headroom to do that than Intel's 14th gen over 13th.
Its also a totally valid question whether M3 will be based on A17 or A16. I'm not sure if we have strong rumors to point one way or the other. We'll know more next week.
Because there are a lot of svchost.exe instances fighting for this core. /s
I've seen this blend of argument being made towards multicore CPUs when AMD and Intel started selling them. Critics started claiming these new chips were being designed like Gilette razors, and there was no way desktop applications, not even AAA games, could ever saturate so many cores. It took Supreme Commander to be released to get these critics to finally shut up.
The moral of the story is that today's software was designed targeting yesterday's hardware. Nevertheless, even today you'd be hard-pressed to find a single application that doesn't run multiple threads, and on top of that you'd be hard-pressed to find a single desktop user that only runs a single application when working on a computer. Even smartphones today are packing CPUs with 8 cores.
And finally, HN being a tech-related forum, most of the people here work on software. Your typical run-of-the-mill build process is embarrassingly parallel. This means that all it takes to saturate that many cores is for anyone to hit "build". And I won't even go into ML.
> The moral of the story is that today's software was designed targeting yesterday's hardware.
The problem with CPUs w/ high core counts is that GPU can be more efficient in most cases. The tasks that specifically need CPU are likely complex and dynamic - in terms of codeflow and dataflow - so that they can't benefit from SIMD/SPMD. I don't think "desktop" need any such complexities, especially when the actual complexity of desktop software hardly increased for decades except web browsers.
The Dev Kit is probably what most people here would want, but if you probe around the site you can also find the fully loaded (PSU, case, etc) option too.
I am trying to think of alternative desktop environments for the use of multiple cores. Perhaps in the future we shall all do parallelisable matrix multiplication and translate and scale relationship vectors with our desktops and mash up new applications trivially. So we can use all 128 cores
And RAM goes a lot deeper too. Look at these two sticks of RAM. See how the one on the right has twice the number of memory modules? That allows the individual stick of RAM to pump through data more quickly than the one on the left, even though both of them are rated at DDR 3200 and CL22.
The bus width is defined by the DIMM standard. A 64-bit DIMM (which is what he is showing) always has a 64-bit data interface, regardless of what devices are on the DIMM. The DIMM with fewer devices on it just has 16-bit-wide microchips on it. That doesn't have anything to do with the latency or clock speed of the chips themselves.
And it's strange that he calls each microchip a "module". I would call that a microchip or a device. A "module" is an assembly that contains multiple microchips.
I just get the vibe that he's someone skilled at putting systems together but doesn't have any real engineering background.
https://frankdenneman.nl/2015/02/20/memory-deep-dive/#:~:tex....
You've put off the vibe of someone that assumes they're correct without double checking themselves ;)
Its really only a concern in the realm of ultra large memory applications: science, machine learning, HPC, etc