As a key exhibit, AVX-512 native code destroys Apple Silicon. To be clear, I like and use Apple Silicon but they can’t carry a workload like x86 and they seem disinterested in trying. Super-efficient for scalar code though.
Now, like with everything in life, of course, there's highly-specialised datapaths like AVX-512, but then again these only contribute towards single-threaded performance, & you yourself said that "High-performance workloads are not single-threaded programs." Now, as your compute network grows larger the implementation details of the memory fabric (NUMA, etc.) become more pronounced. Suffice to say the SoC co-packaging of CPU and GPU cores, along with some coprocessors, did wonders for Apple Silicon. Strix Halo exists, but it's not competitive by any stretch of imagination. You could say it's unfair, but then again, AMD MI300A (LGA6096 socket) exists, too! Do we count 20k APU's that only come in eights, bundled up in proprietary Infinity Fabric-based chassis towards "outperforming ARM64 in high-perf workloads"... really? Compute-bound is a far cry from high-performance, where the memory bus, and idiosyncrasies of message-passing are King as number of cores in the compute network continues to grow.
I think their slogan could be "unlimited, coherent, persistent, encrypted high-bandwidth memory is here, and we are the only ones that really have it."
Disclaimer: proud owner of thoroughbred OpenPOWER system from Raptor
Really, 1TB/s of memory bandwidth to and from system memory?
I don't believe it since that's impossible from HW limits PoV - there's no such DRAM that would allow such performance and Apple doesn't design their memory sticks ...
It is also no more special with their 512-, 768- or 1024-bit memory interface since this is also not designed by them nor it is exclusively reserved to Apple. Intel has it. AMD has it as well.
However, regardless of that, and regardless of the way how you're the one skewing the facts, I would be happy to see the benchmark that shows, for example, a sustained load bandwidth of 1TB/s. Do you have one since I couldn't find it?
> You can get somewhat better with Turin
High-end Intel/AMD server-grade CPUs can achieve a system memory bandwidth of 600-700GB/s. So not somewhat better but 3x better.
5x is false, it's more like 4x. Apple doesn't use memory sticks, they use on-SoC dram ICs.
The M3 Ultra has 8 memory channels at 128-bit per channel for a total of 1024-bit memory bus. It uses LPDDR5-6400 so it has 1024-bit * 6400000000 bits = 819.2 gigabytes per second of memory bandwidth.
You're realistically going to reach power/thermal limits before you saturate the memory bandwidth. Otherwise I'd like to hear about a workload that'll make use of the CPU, GPU, NPU, etc. to make use of Apple's marketing point.
https://web.archive.org/web/20240902200818/https://www.anand...
> While 243GB/s is massive, and overshadows any other design in the industry, it’s still quite far from the 409GB/s the chip is capable of.
> That begs the question, why does the M1 Max have such massive bandwidth? The GPU naturally comes to mind, however in my testing, I’ve had extreme trouble to find workloads that would stress the GPU sufficiently to take advantage of the available bandwidth.
> Granted, this is also an issue of lacking workloads, but for actual 3D rendering and benchmarks, I haven’t seen the GPU use more than 90GB/s (measured via system performance counters)
They did not (or, rather, could not) measure the theoretical peak GPU core saturation for the M1 Max SOC because such benchmarks did not exist at the time due to the sheer novelty of such wide hardware.
So, which part of "We are talking about the memory bandwidth available to the CPU cores and not all the co-processors/accelerators present in the SoC" you didn't understand?
Personally, I learn, reflect, and gain a lot from engaging in conversations with other people, as – not infrequently – the others complement my understanding of points I might have previously not considered or missed. I call it knowledge, and knowledge is power.
Now, I'm not sure whether it's genuine to compare Apple Silicon to AMD's Turin architecture, where 600 GB/s is theoretically possible, considering at this point you're talking about 5K euro CPU with a smudge under 600W TDP. This is why I brought up Sienna, specifically, which is giving comparable performance in comparable price bracket and power envelope. Have you seen how much 12 channels of DDR5-6400 would set you back? The "high-end AMD server-grade," to borrow your words, system—would set you back 10K at a minimum, and it would still have zero GPU cores, and you would still have a little less memory bandwidth than a three-year old M2 Ultra.
I own both a Mac studio, and a Sienna-based AMD system.
There are valid reasons to go for x86, mainly it's PCIe lanes, various accelerator cards, MCIO connectivity for NVMe stuff, hardware IOMMU, SR-IOV networking and storage, in fact, anything having to do with hardware virtualisation. This is why people get "high-end" x86 CPU's, and indeed, this is why I used Sienna for the comparison as it's at least comparable in terms of price. And not some abstract idea of redline performance, where x86 CPU's by the way absolutely suck in a single most important general purpose task, i.e. LLM inference. If you were going for the utmost bit of oompf, you would go for a superchip anyway. So your choice is not even whether you're getting a CPU, instead it's how big and wide you wish your APU cluster to be, and what you're using for interconnect, as it's the largest contributing factor to your setup.
Update: I was unfair in my characterisation of NVIDIA DGX Spark as "big nothing," as despite its shortcomings, it's a fascinating platform in terms of connectivity: the first prosumer motherboard to natively support 200G, if I'm not mistaken. Now, you could always use a ConnectX-6 in your normal server's PCIe 5.0 slot, but that would already set you back many thousands of euros for datacenter-grade server specs.
From https://www.ibm.com/support/pages/ibm-aix-power10-performanc...
> The Power10 processor technology introduces the new OMI DIMMs to access main memory. This allows for increased memory bandwidth of 409 GB/s per socket. The 16 available high-speed OMI links are driven by 8 on-chip memory controller units (MCUs), providing a total aggregated bandwidth of up to 409 GBps per SCM. Compared to the Power9 processor-based technology capability, this represents a 78% increase in memory bandwidth.
And that is again a theoretical limit which usually isn't that interesting but rather it's the practical limit the CPU is able to hit.
Also there's to note the substantial overprovisioning of the lanes to handle lane-localized transmission issues without degrading observed performance.
See my reply to adjacent comment; hardware is not marketing, and LLM inference stands to witness.
The opposite case is also possible. You can be compute limited. Or there could be bottlenecks somewhere else. This is definitely the case for Apple Silicon because you will certainly not be able to make use of all of the memory bandwidth from the CPU or GPU. As always, benchmark instead of looking at raw hardware specifications.
All of it, and it is transparent to the code. The correct question is «how much data does the code transfer?»
Whether you are scanning large string ropes for a lone character or multiplying huge matrices, no manual code optimisation is required.
[0]: https://developer.apple.com/documentation/accelerate
[1]: https://ml-explore.github.io/mlx/build/html/usage/quick_star...
Yes.
> Have they really cracked it at homogenous computing […]
Yes.
> have it emit efficient code […]
Yes. I had also written compilers and code generators for a number of platforms (all RISC) decades before Apple Silicon became a thing.
> […] for whatever target in the SoC?
You are mistaking the memory bus width that I was referring to for CPU specific optimisations. You are also disregarding the fact that the M1-4 Apple SoC's have the same internal CPU architecture, differing mostly in the ARM instruction sets they support (ARM64 v8.2 in M1 through to ARM64 v8.6 in M4).
> Or are you merely referring to the full memory capacity/bandwidth being available to CPU in normal operation?
Yes.
Is there truly a need to be confrontantial in what otherwise could have become an insightgul and engaging conversation?
Others also have. The https://lemire.me/blog/ blog has a wealth of insights across multiple architectures, which include all of the current incumbents (Intel, Apple, Qualcomm, etc.)
Do you have any detailed insights? I would be eager to assess them.
They already are open enough to boot and run Linux, the things that Asahi struggles with are end-user peripherals.
> OTOH server aarch64 implementations such Neoverse or Graviton are not as good as x86_64 in terms of absolute performance. Their core design cannot yet compete.
These are manufactured on far older nodes than Apple Silicon or Intel x86, and it's a chicken-egg problem once again - there will be no incentive for ARM chip designers to invest into performance as long as there are no customers, and there are no customers as long as both the non-Apple hardware has serious performance issues and there is no software optimized to run on ARM.
That's for entertainment and for geeks such as ourselves but not realistically for hosting a service in a data center that millions of people would depend on.
> These are manufactured on far older nodes than Apple Silicon
True but I don't think this would be the main bottleneck but perhaps. IMO it's the core design that is lacking.
> there will be no incentive for ARM chip designers to invest into performance as long as there are no customers
Well, AWS is hosting a multitude of their EC2 instances - Graviton4 (Neoverse V2 cores). This implies that there are customers.
Why not? Well form factor is an issue. But you can easily fit a few mac pros in a couple Us. Support is generally better then some HP or Dell servers.
AWS has a bit of a different cost-benefit calculation though. For them, similar to Apple, ARM is a hedge against the AMD/Intel duopoly, and they can run their own services (for which they have ample money for development and testing) for far cheaper because the power efficiency of ARM systems is better than x86 - and like in the early AWS time that started off as Amazon selling off spare compute capacity, they expose to the open market what they don't need.