Cortex X2: ARM aims high
chipsandcheese.com
chipsandcheese.com
There's still no real synergy between CPUs and GPUs, even though they get less different with time. No one seems to have a decent plan for how to balance local memory resources with cache-coherency over CXL. Inter-system networking seems to be frozen in 1980 (just with faster serdes). Does in-chip and inter-chip optical change anything, or does it just mean littering the place with tranceivers?
it seems like everyone settles for ugly, barely usable interfaces like CUDA, OpenCL, and then scabs them over with a higher-level interface like tensorflow.
Now consider that with Phoenix Point APUs you have Zen 4 CPU, memory controllers, IO controller, RDNA3 GPU, and the XDNA "transputer-like" coprocessor, all hanging off common cache coherent transport (Infinity Fabric). All the hardware parts are there.
I think we could argue that tensorflow and/or pytorch has displaced any HSA interest, too. These are programming interfaces that do an OK job of abstracting from the hardware details, and are totally embraced by the most demanding field (AI).
In the wee hours of this morning, I was trying to figure out why my attempts to build LLVM 17.0.3 from source were failing because it couldn't find an HSA-related symbol.
My impression was that the build-system's support for using HSA was a little wonky, but I'm not sure if that's fair. IIRC I worked around the issue by not trying to build LLVM's openmp code.
As caches have gotten bigger I’ve wondered when or if we will treat them as local memory instead of trying to maintain an illusion of cache coherency at a high cost.
But if you never call those, it is possible to access the cache explicitly. Coreboot have a custom in-tree compiler, ROMCC [1] that compiles C to a subset of x86 that never emits memory operations.
They use this to treat cache as RAM [2] for a brief period during early boot before the RAM itself gets initialized.
1. https://github.com/wt/coreboot/tree/master/util/romcc The whole compiler is a single C file
2. https://www.coreboot.org/data/yhlu/cache_as_ram_lb_09142006.... Perform s/LinuxBIOS/Coreboot/ in your head as you read
I have seen a patent application for a technology from {big hardware company} addressing exactly this. Dunno if the approach will work, but there are definitely plans
If your main interest is CPU cores even a relatively old version will be enough. From memory recent versions (5 and 6, 6 being the latest) have mostly added content on GPU and data center computing, but the sections on CPU architecture are rather stable.
Organization and Design has different editions for RISC-V, aarch64 and MIPS. The MIPS one is what I learned computer architecture on and it's super solid.
Cloud? Oracle free-tier and otherwise Hetzner.
In this class, there are many models with RK3588, having e.g. dual 2.5 Gb/s Ethernet ports (with the possibility of attaching more Ethernet NICs on USB 3 or on PCIe M.2 adapters) and supporting PCIe 3 x4 M.2 SSDs and/or eMMC (and the attachment of more SSDs on USB). (The model that I like most is NanoPC-T6, which exploits best all the interfaces of RK3588, but without adding things that should better be added externally, only when they are needed, like the additional USB hub present in many other models.)
A cheaper option, but with much slower peripheral interfaces, is the new Raspberry Pi 5 model.
Nevertheless, a homelab server with an ARM CPU makes sense only for developing Aarch64 applications.
For just doing the job there are many small and cheap fanless computers with Intel N100 (4 E-cores) and 4 to 8 2.5 Gb/s Ethernet ports (typically sold on Amazon as firewall appliances). For only a few dollars extra it is possible to find much faster cheap small computers (from companies like Minisforum or Beelink) with older AMD Zen 3 CPUs, like the 6-core Ryzen 5 5600H.
For a higher price of $500 to $600, there are small computers with AMD Ryzen 9 7940HS, which can support e.g. dual M.2 PCIe 4 x4 SSDs and SATA SSDs, and dual 10 Gb/s Ethernet NICs (on Thunderbolt), dual 2.5 Gb/s ports on the MB + many other 2.5 Gb/s ports on USB, while being faster than big and expensive servers from some years ago.
Unfortunately, I assume that this product has been canceled a few months ago, when Intel sold their NUC business to ASUS.
There are a few small computers that support ECC and which use obsolete Intel Tiger Lake or Tremont-core-based Intel Elkhart Lake CPUs, but those CPUs are a dead end, being slow and supporting instruction sets that are different from the current mainline Intel CPUs, so I would not recommend any of them.
The best remaining choice depends on which is more important, the size and the power consumption or the price of the server.
For very small size and low power consumption I am not aware of any good solution at a reasonable price, because even when some of the Arm CPU SoCs or Intel or AMD mobile CPUs support ECC, I have not seen any such computer board that includes the ECC support. There are some industrial computers with ECC, but those are expensive for what they offer.
If only the cost is the problem, and second-hand servers are avoided because the server to be bought is intended to be used for many years, then a server with desktop Intel or AMD CPUs must be used. The MBs with the Intel W480 chipset are expensive, so the cheapest solution is to use one of the AM5 MBs that specify ECC memory support, e.g. from ASUS or ASRock Rack, together with one of the cheaper Ryzen 7000.
Another option is an older AM4 MB, like the Mini-ITX ASRock Rack X570D4I-2T ($400 due to including dual 10 Gb/s Ethernet ports), which has the advantage of using cheaper older Ryzen 5000 CPUs, with cheaper DDR4 ECC memory, so the total system cost would be reasonable.
The only disadvantage of the desktop Ryzen CPUs when used as servers is that, even if they have excellent energy efficiency when they are actually running programs, they have a relatively high idle power consumption, because only the cores are shut down when doing nothing, while the I/O die has a permanent consumption around 20 W or more. Therefore one must choose between the low idle power consumption of a few watts of the laptop CPUs and the ECC memory support of the desktop CPUs.
Because in my home lab most servers alternate between times when they are used intensively with times when they stay idle for hours or days, except for one server that is connected permanently to the Internet, all the others are used with Wake-on-LAN, so they are shut down when idle, for negligible power consumption.
The recompilation of the Linux kernel may be more difficult, because the right configuration file and modules must be selected before doing it, but it should be possible as most support for RK3588 is included in the mainstream kernel. Also U-Boot (the boot loader that loads the Linux kernel) should be recompilable from sources.
The hardware is more trustworthy than that of Intel or AMD computers, because it comes with the complete schematics and the technical reference for RK3588 is much more complete than for any Intel or AMD CPU.
A hardware backdoor could have been implemented only in the Ethernet interface of RK3588, but that is not used in NanoPC-T6, which uses Realtek Ethernet interfaces on PCIe lanes. Any hardware backdoor in those would have required a close and secret cooperation between a major Taiwanese company and a major mainland Chinese company, which is unbelievable.
This is normally provided in all Arm SoCs by binary blobs. Nevertheless, at least for the Mali GPU there is a reverse-engineered driver in the Linux kernel, which might be usable with RK3588.
In any case, this hardware acceleration is the same in all RK3588 boards, regardless of the vendor, it is not specific to NanoPC-T6.
Software from board manufacturers shouldn't be used anyway, if not because in a few years it is often discontinued and not updated anymore because they're pushing newer models. Thankfully we have Armbian and Dietpi which are the distros of choice for all boards that don't run major PC oriented distros (and a nice alternative for those that do). The number of boards supported by these two distros is astonishing:
https://www.armbian.com/download/
The NanoPC-T6 is already supported by Armbian build system:
I bought orange pi 5+ to replace my x86-based odroid-h2+. Will start the installation in the next few weeks. I have seen enough anecdotes people running NixOS on this chip, so feel pretty optimistic about it.
Also, #nixos-on-arm Matrix channel is amazing.
Why? Cost, performance, and power consumption should be in the same ballpark as x86, but support for x86 is still significantly better. There's nothing magical about running a workload on ARM. The math obviously changes when you have thousands of machines.
Design looks good & beefy. But implementation is the other part of the equation. Or thermal throttling depending on device it's in.
E.g. Wikipedia says X3 has 512-1024kB of L2 cache per core, M1 has 3MB.
[Edit] + Buying up all state of the art production capacities so competition is one node behind.
There is no Apple secret sauce.
As long as the others don't want to go that route - and they seem not to be in need to cut into their margins (AMD shows how X3D helps with performance).
I think what is interesting especially for Intel/AMD is that Xiaomi drops legacy 32 bit ARM and translates apps to 64bit.
Dropping 16/32bits can reduce die size which can be used for larger caches for the same price.
One of Apple's actual secret sauces is they can make their big caches fast. Typically latency increases with cache size so it's a tradeoff. Apple trades off less here. And it's not some "only fast because tsmc" it's just really solid engineering at both the architectural and physical design level.
"The memory is not inside the package"
vs.
"The SoC and RAM chips are mounted together in a system-in-a-package design." [0]
Every mobile SOC does the same? All Intel SOCs do this? Which one? Can you point out the 16Gb of RAM in this Meteor Lake SOC?
https://images.anandtech.com/doci/20046/Meteor_Lake_Hotchips...
The Wikipedia article on Meteor lake doesn't even mention memory at all [1]
The drams on an apple chip are still bog standard lpddr. Most benchmarks find the actual memory middle of the road at best.
Critically they aren't magically on the die or any more inside the package than most other high end mobile chips.
2. "It is in a package like no other vendor, but it's not changing performance"
3. ???
ram is ordinary POP, you got lied to by Apple marketing. If you acted on this marketing and spend money then re-programming will be very difficult with brain actively fighting on every step to prevent cognitive dissonance.
It is cool to live in the future where 243 GB/s is middle of the road.
It is still impressive that Apple pulled it off 2 years ago, IMO.
Anyway the point is, this is not a meaningful performance benefit as it's still just off the shelf LPDDR5. In fact the M SoCs tend to underperform in memory latency tests.
"Anyway the point is, this is not a meaningful performance benefit"
Do you have a benchmark to read? This "Still LPDDR5" is hand waving.
M2 is a 5nm (N5P) chip, AMD laptops already use 4nm.
This is such a funny statement. Do people think Apple is dumping wafers into the ocean? Or buying the capacity and not using it?
The economic reality is that Apple can pay more for cutting edge process because they have higher prices and margins. So, people paying a premium for hardware get more advanced hardware.
How is this in any way surprising? Is the theory that if only Apple wasn’t willing to pay a premium, TSMC would sell the same wafers cheaper to other manufacturers? Wouldn’t that make TSMC 1) dumb, and 2) less profitable and therefore less able to invest in the next process?
So just like every other non-server ARM SoC?
Except for secret sauce like super wide instruction decode and enough registers to keep all their execution units filled[0], sure I guess there's no secret sauce.
Caches are only useful when they're serving execution units and Apple packed their chips with them. That's special sauce. If it wasn't special then every ARM chip would have the same levels of performance. It's not like the M1 was Apple's first chip. The A-series have been kicking the shit out of other ARM chips for almost a decade. If Apple didn't have any special sauce in their chip designs this wouldn't have been the case. It's not like Qualcomm doesn't have good chip designers and hasn't tried to compete with Apple's chips.
Coincidentally, both of those caches in Apple M design are unusually large - 192KB for instruction cache size and 128KB L1 data cache size - per core (!). The same goes for L2 cache size - 3MB per (performance) core.
When compared to bleeding edge _server_ CPUs from AMD and Intel, it's crazy to see that those figures are by several _magnitudes_ larger in the Apple M design. E.g. Zen3 Epyc - 32KB of instruction cache size, 32KB of L1 data cache size and 512KB L2. Intel Xeon Gold - 32KB of instruction cache size, 64KB of L1 data cache size and 1.25MB of L2.
From what I read, before that, there’s “paying billions to get state of the art production capabilities built”
Chances are that capacity wouldn’t be there without Apple’s money, so if Apple didn’t exist, it still wouldn’t be available to others as rapidly as it is now.
> There is no Apple secret sauce.
They didn’t always have loads of money, so, historically, there must have been something else than “they have loads of money and large margins, so can afford to buy the best”.
I think there still is something more than that. For example, it also is about having the courage to decide that milled aluminum is a better way to build laptop chassises, so spending billions on buying/creating the capacity to build millions of such chassises is a good idea, or to decide that, at their size, building your own CPUs is worth doing.
I think part of their secret sauce also is that they have higher standards for what they want to sell. Take for example foldable screens. They must have prototypes with them, but don’t have a product because they don’t deem them good enough.
But the Apple that stayed alive in the 90’s-early 00’s is pretty different from modern Apple. Modern Apple makes some of the best chips out there. Old Apple stayed afloat by selling a Unix clone on commodity x86.
I don't think that's quite true. It's clearly a combination of better microarchitecture (very wide decode, 128 byte cache lines, etc), and also massively bigger area budgets. Maybe more the latter, but it's pretty clear that Apple is right at the top of the "good microarchitecture" leader board.
The M1/M2 is an amazing chip, however you need to consider that these chips are in different price brackets hence the performance discrepancy.
The new Qualcomm Oryon chips are more powerful but also much more expensive. So the X lineup is actually quite reasonable if you want smartphones to be "affordable".
https://www.androidauthority.com/qualcomm-snapdragon-8-gen-4...
Edit: this is the benchmark I’m basing my comment on finding it a bit slower than the 2022 A16.
https://www.notebookcheck.net/Alleged-Snapdragon-8-Gen-3-Gee...
Xiaomi 14 ( using Snarpdragon Gen 3 ) is announced and reviewed. Pre-Order started and shipping in November.
>being almost as fast as one which shipped several years earlier
The A16 has been shipped for a little more than a year. I could have compared it to A17 Pro and the answer would still be the same. Since A17 Pro is only slightly faster with a Clockspeed improvement. And that is comparing in Cortex X4 vs A17 Pro Clock to Clock, X4 would be landing within 10% range.
>is … crazy?
Considering 99.99999999% of the internet said Apple will always be 3-5 years ahead and ARM ( or Snapdragon or any other non Apple CPU design ) will never be able to match Apple, despite being shown how X3 is the proof and X4 finally shown it. Yes. I find it pretty crazy.
> Considering 99.99999999% of the internet said Apple will always be 3-5 years ahead and ARM ( or Snapdragon or any other non Apple CPU design ) will never be able to match Apple
Serious citation needed on that. Most people said Apple did a good job and that Qualcomm needed to step it up considerably to stay in the game. If they have, that’ll be great for many millions of buyers so competition will have worked exactly the way we wanted it to.
Even the successor Cortex X3 is already available phones (Snapdragon 8 Gen 2 or Dimensity 9200). It benches around 1800 on Geekbench, which is comparable to a Zen3+ low power U laptop core, or an lower clocked Intel 12th gen U laptop core.
Fro comparison, the latest Raspberry Pi 5 features A76 cores, which benches around 900 in phone implementations, comparable to 8th gen Intel cores. Apple's A13 scores around 1600, the M1 around 2200.
Ah yes already looking forward to everyone cutting support for armv7 package building on apt, just like they did for v6 when v8 was 'the thing'. This rolling cycle of incompatibility and obsolescence is so goddamn infuriating.
Or maybe everything will just be 64 bit from now onward, idk.
IIRC, ARMv8 has both 32-bit (AArch32) and 64-bit (AArch64); yes, ARM's naming is confusing. What's being dropped is 32-bit (AArch32), similar to what's happening in the x86 world (and AFAIK also the Linux on mainframe world), and for similar reasons.
Right now there are steps in between which means they care more about selling the design and its variations. If there is a PPA miss you could point the finger in a few directions
[1] https://en.wikipedia.org/wiki/Andrew_Grove#Only_the_Paranoid...
BTW Andy was not the CEO in times of this inflection point.
A better target would be Itanium: that was in his era and showed the danger of picking the wrong gamble. They wanted a proprietary platform which couldn’t be copied by AMD, Cyrix, etc. but they staked the whole exercise on a highly speculative CPU design and made a number of horrible tactical errors like trying to screw a few million in compiler licensing out of developers at the time when their multi-billion bet-the-company investment crucially depended on developers porting to a chip which was critically dependent on having the best compiler. Being paranoid about preventing competition lead them to be shockingly cavalier about their plans being robust.
Heh. Nice
In that chapter, Andy Grove recognized the merits of both technologies, they even developed RISC chips. The preference for CISC was maintaining compatibility of the 386 line of extremely successful microprocessors recognizing that following another technology was resource intensive and against the extremely popular and fruitful 386 line of business.
The analysis is clear.
BTW, I think this is a book recommended to any entrepreneur since it addresses atemporal concepts and you return to the book anytime you gain more experience to clarify your experiences.