How to get 1.5 TFlops of FP32 performance on a single M1 CPU core
jott.live
jott.live
In addition to running on the CPU, M1 Max devices have three separate kinds of hardware-accelerated `gemm`: the GPU, the ANE (Apple Neural Engine), and this special matrix coprocessor. Here's a fairly detailed post that benchmarks each:
https://tlkh.dev/benchmarking-the-apple-m1-max
And here's a great post about the justification for having so much special-purpose hardware:
https://medium.com/swlh/apples-m1-secret-coprocessor-6599492...
As for the matrix coprocessor, Apple's built-in BLAS implementation (Accelerate.framework) uses this chip. You can link Numpy against this to benefit in your Python programs, for example. Here are some old instructions: https://gist.github.com/MarkDana/a9481b8134cf38a556cf23e1e81...
All this represents yet another cycle on the Wheel of Reincarnation... (http://catb.org/jargon/html/W/wheel-of-reincarnation.html)
Yep, thermal throttling is a thing, and sometimes all you need is either useless silicon padding or some specialized, most of the time dark silicon to both make it feasible to cool and prevent it from melting.
AMX co-processor 2 TFLOPS FP32
GPU 8 TFLOPS FP32
Neural Engine 5.5 TFLOPS FP16Just power efficiency?
Or do you have to manually split the computation between them?
However, I do have some experience with having the same code run on the GPU and the CPU. In my work, we have tried breaking images (usually frames of video) into various sized chunks and processing on both the CPU and GPU at the same time. Our conclusion is that the overhead of using both outweighs any benefit you’d get. The GPU is so much faster than the CPU, there’s no point in involving the CPU at all. These experiments were done several years ago, so perhaps the landscape has changed since then, but that was what we found.
https://highperformancegraphics.org/slides22/Journey_to_Nani...
https://advances.realtimerendering.com/s2022/index.html#Lume...
They're great presentations with a lot of depth in the notes. I think videos are around somewhere if you prefer that.
Two specifics I'd mention:
It seems a lot of games now use feedback between frames as a way to tolerate the latency of moving data between CPU and GPU. Eg the CPU will use GPU crunched data from the previous frame as a source for CPU crunching that optimizes what data gets passed to the GPU next.
The other is that fixed functionality is moving into shaders. Unreal 5 uses a mix of hardware rasterization and software rasterization in a shader (and path tracing now as well). There the tradeoff between the two is triangle size in pixels.
Isn't this adding new cores directly onto the main chip? That doesn't sound like it fits to me.
And at this point GPUs have been straddling both sides of the divide for decades, depending on the particular device form factor and the necessary power.
The only thing I would actually say has gone through a cycle lately is the crypto accelerator for mac SSDs.
These are coprocessors, which are a very different thing from just another CPU core. For one, they use a different architecture (instruction set, registers/memory, etc.).
The "wheel of reincarnation" refers to features/capabilities on coprocessors eventually being folded into the main CPU. While CPUs have adopted insights from GPU implementations, GPU functionality has never been fully folded into CPUs (software rasterizers don't count).
Well that's why I didn't say "just another CPU core". But fine, I don't want to argue semantics.
> The "wheel of reincarnation" refers to features/capabilities on coprocessors eventually being folded into the main CPU.
Then that's definitely not happening here, and hasn't happened to x86/arm since they gained floating point, right?
Ok but it doesn't actually justify why AMX & ANE both exist. It makes kind of a vague handwavy "well AMX latency is better[1] and that's useful[2]"
1: but not measured, so not actually known, and with a note that it's been called out that the AMX understanding is incorrect, so is this point even still accurate?
2: but not elaborated on in the slightest or any comparison of test workloads
So why do both AMX & ANE exist? CPU team did AMX before the ANE team showed up with something bigger & better? Are they actually used for differing workloads or simultaneously?
For instance, we did benchmarks of spaCy (natural language processing) transformer models across various Apple Silicon SoCs and MPS was 1.9x (M1) to 5.5x faster (M1 Ultra) while providing far more performance per Watt. E.g. using MPS on an M2 MacBook Air used 4W less energy while being 2.7x faster than AMX.
Full benchmarks are at the end of this post:
It always seemed to me like SIMD/AVX/etc would eventually come for the GPU's lunch money... How many more product generations of "SIMD on steroids" before this is practically true?
The latency factor is the biggest thing for me. The GPU is a turtle compared to CPU-bound techniques. I can see emerging applications for this in real-time/streaming where every millisecond counts.
GPUs are an optimization to try to use the excess Moore's law we have to get to the ghost of Dennard's law.
Building something that can reliably output reasonable-quality 3d graphics without relying on specific GPU technologies will give you a much broader realm to operate with.
I believe something along this path is the solution for streaming gaming. I perceive the failure of Stadia, et. al. as being a consequence of trying to bolt streaming onto existing GPU-based, local gaming solutions. Build something from scratch with streaming/latency as a #1 priority, and you can dramatically expand the operational radius of each datacenter (e.g. ~100km per millisecond saved).
https://www.pingdom.com/blog/theoretical-vs-real-world-speed...: “you should probably double the “ideal” response times shown above for a more realistic target to aim at“
So yes, ⅓ of light speed in vacuum seems a decent heuristic.
If you're talking full-screen Angry Birds with, say, a 2x average compositing, you're going to be fine on the CPU; but, energy- and jitter- wise you'll still be happier with the GPU, overall.
NVIDIA's streaming service is doing relatively fine in comparison. They simply share a GPU between several users for anything that isn't demanding enough. They also get around some of the concerns about gaming being turned into another streaming style fragmented mess by not actually selling the games. You simply log into your account on Steam/GOG/whatever and play the games you already own as you might on a local PC.
Additionally, "building something that can reliably output reasonable-quality 3d graphics without relying on specific GPU technologies" doesn't make much sense to me. If it's an accelerator designed to handle relatively modern 3d graphics, due to the programmability of a modern graphics pipeline it's effectively just a GPU. There aren't any underlying technologies that are required to be used as long as they can produce a similar output (mobile GPUs tend to have a different approach to how they implement the graphics pipeline compared to desktop GPUs for instance).
1) 1.5 TFLOPS is already less than the GPUs in most current phones. Like you're not exactly talking a big number here. You're talking 14 year old graphics (when desktop GPUs crossed the 1TFLOP mark)
2) And, this is the bigger issue, this AMX unit only does matrix multiplies. I'd be fascinated to see someone create a 3d renderer with only matrix multiplication. I'm sure it's possible but this 1.5 tflops ain't remotely comparable to a GPU's 1.5 tflops. Like if you tried to do "traditional" pixel shader rendering code on this unit you'd instantly be cutting that 1.5 tflops to 1/16th the performance - you're now at a theoretical peak of a mere 93 gflops of GPU-comparable performance.
You also forgot to mention the usability bit. Branching is bad for GPUs. They either take both branches or in the optimized version, they stall the pipeline until they get a result then attempt to optimize which units need to execute on which branch.
The first is simply inefficient. The last is efficient, but impracticality slow on branchy code.
On the power side, bloated x86 is hardly a good showcase for what such an architecture could do. I also suspect that an automatic thread manager like most modern GPUs use could also add a lot of value.
Branching is bad for wide SIMD workloads of which GPUs are the most common. But just shoving that SIMD under the control of a CPU instead won't fix your branching issues. You're still probably going to choose the path of just executing both branches & masking off the results.
I seem to remember that LRB could test & jmp in 4 cycles, if you were careful. Since it was 4x barrel processed, those jmps were “free”.
I moved to integrated GPUs, later. x86 is bloated, but LRB was not. Also, the decoder — even a big one like x86 — isn’t a major HW problem. I’d say x86’s memory hierarchy is more of an issue.
I'm not sure GPUs will be displaced, looking at the difficulties Larrabe had on the driver side, but I do think we'll see more flexible alternatives becoming popular.
That said I think the “good enough” metric is already there and unless you’re doing hardware ray tracing or extreme details at high resolutions you won’t need or care about a GPU any more.
Latency though isn’t the issue. The times involved for human perception are long and not getting shorter.
And also, as weird as it is, 1.5TFlops isn't actually that ridiculous. We had that performance 14 years ago at 150w with desktop GPUs. 14 years to reduce from 150w to what, 5w?, is cool but also honestly pretty par for the course is it not? Especially for a fixed-function block?
But otherwise that's less than what mobile GPUs can do, although those take more power they're not in entirely different ballparks either
I don't want to give up SO-DIMMs for a few mm thinner laptop, but going from the intel/amd standard 70GB/sec to 400GB/sec is a pretty big incentive.
Regardless, no that's not the drastic shift necessary. It's not even a shift at all. 4, 6, and 8 channel memory architectures have all existed for a while now in data center SKUs. It hasn't changed the desire for GPU compute in the data center
Well the 4, 6, and 8 channel memory architectures have been dramatically slower than the Apple 800GB/sec with dramatically less memory transactions in flight (1 per channel) than Apple's 800GB/sec with 64 channels. The push for Apple like memory systems with HBM from Intel, AMD, Fujitsu and others does show that there's a need for more compute bandwidth and Fujitsu in particular has shown that a healthy CPU with a GPU like memory system can get great performance on real world codes without the hassle of rewriting for CUDA.
With all that said, yes there's a healthy demand for GPUs and codes that run best on GPUs. However I find it refreshing that at last one company is providing significantly improved memory systems in laptops and small desktops.
"the M1 Max isn’t able to fully saturate the SoC bandwidth from just the CPU side; From a single core perspective [..] it’s able to stress the memory fabric to up to 102GB/s. [..] Adding a third thread there’s a bit of an imbalance across the clusters, DRAM bandwidth goes to 204GB/s, but a fourth thread lands us at 224GB/s and this appears to be the limit on the SoC fabric that the CPUs are able to achieve"
https://www.anandtech.com/show/17024/apple-m1-max-performanc...
So peak CPU bandwidth on the M1 Max is 224GB/s. Which is really good for a CPU to be sure, but we're still talking bandwidth numbers far below what you'd need to eliminate the GPU.
> Well the 4, 6, and 8 channel memory architectures have been dramatically slower than the Apple 800GB/sec with dramatically less memory transactions in flight (1 per channel)
Huh? I guarantee you Xeons and Epycs are not limited to 1 transaction in flight, nor is Graviton2.
The AMD APUs in things like the steam deck, PS4, and PS5 (448GB/s fwiw) are also definitely not limited like that, either.
> than Apple's 800GB/sec with 64 channels
M1 ultra is 32 channels. Well, it's really 2x 16 channels since it's a NUMA configuration, like a dual-socket system.
> The push for Apple like memory systems with HBM from Intel, AMD, Fujitsu
lol? Apple's memory system isn't like HBM. It's more like a traditional GPU's but with GDDR swapped out for LPDDR. M1 Max is 512-bit memory bus, just like many GPUs have done over the years. 512-bit bus and it can only hit 400GB/s shows you just how much less bandwidth LPDDR has vs. GDDR, too. AMD & NVidia are hitting 1TB/s with 384-bit buses these days. Which is also what the HBM2e Xeon's are claiming.
Sure, wouldn't want to starve the GPU when the CPU is busy. Not sure if the AMX has it's own connection or shares the CPU complex bandwidth.
> I guarantee you Xeons and Epycs are not limited to 1 transaction in flight, nor is Graviton2
Right, I said 1 per channel, not 1. I've tested this on Xeons and Epycs, but not genoa yet. Generally maximum random throughput tends to be with 16 or so cache misses per socket with 8 channels. From what I can tell about half the latency is cache misses in L1, L2, and L3. Then you queue in the memory controller and wait for whatever memory channel (or two depending on config). By keeping 16 in the queue as soon as one of the 8 channels free you have another request pending for that channel, at least most of the time.
The Genoa looks particularly promising on this front, they upgraded from 8 channels to 12, but the DDR5 dimms actually provide two channels per dimm, so you end up with 24 narrow (32 bit) DDR5 channels. I've already seen results showing that Genoa scales better to high core counts than the zen3 Epyycs, which makes sense with 24 memory requests in flight instead of 8.
> M1 ultra is 32 channels. Well, it's really 2x 16 channels since it's a NUMA configuration, like a dual-socket system.
Heh, well monolithic chips are on the way out. Even the zen2 epycs used multiple pieces of silicon (chiplets) and the new Intel Xeon/Sapphire Rapids do the same. Don't see a particular difference. Even the desktop Ryzens have multiple pieces of silicon (one IOD and one CPU chiplet at the minimum). In any case the Apple M1 Max memory system matches the Genoa and beats the newest Intel Sapphire Rapids with a single piece of silicon, and wins handily doubles that with two pieces of silicon, unless you wait for the unreleased HBM flavors.
Compared to 100% of Intel/AMD laptops and probably 98% of desktops that apple memory system is pretty amazing. Sure there are threadrippers (4 channel), threadripper pro (8 channel), and some similar Intel options, but all are expensive, low volume, and desktop/workstation only.
Ok, then surely you won't take issue with me describing my x86 desktop as having 1.2TB/s of memory bandwidth? It just happens to be in a NUMA configuration and after all, wouldn't want to starve the GPU when the CPU is busy...
> Don't see a particular difference.
You don't see a difference between 1 memory controller and 2 memory controllers? Compare Zen1 TR/epyc vs. zen2 TR/epyc.
Hint: m1 ultra is like Zen1, and probably ain't the future
Oh yeah and Intel did this, too, with the 56 core xeons a couple years ago. And, like AMD, also quickly moved away from it.
Unless it's a rather odd config that's counting the GPU memory bandwidth which is decidedly not NUMA since the GPU mem isn't cache coherent with the main memory.
> You don't see a difference between 1 memory controller and 2 memory controllers?
From what I can tell the multiple memory controllers per chip AMDs performed well on most codes, but certain codes, OSs, and benchmarks didn't handle the different latencies within the socket well. As a result AMD moved to one memory controller per socket in the next generation.
The M1 ultra moved to two chips and 2 memory controllers, but as a result manages to hit 800GB/sec peak which is a pretty big win compared to any normal socket, even the 96 core AMD Genoa or Intel SPR are sitting in the half to 1/3rd range, despite coming our a year or so later.
It makes some GPGPU workloads viable that otherwise wouldn't be, but it doesn't reduce the bandwidth needs of the GPU for traditional GPU workloads so net-net you're worse off overall. You either use DDR and penalize your GPU performance, or you use GDDR (like consoles do) and sacrifice CPU memory latency. Also power efficiency of GDDR is worse, especially compared to Apple Silicon's choice of using LPDDR
HBM2e gets you everything except it's expensive and limited capacity. But it'll be interesting to see how that plays out with the Xeon Max, which can also still supplement the 64GB HBM2e with some absurd channel count DDR5
AMX is an unstable ISA that changes between product generations. That's why it's not publicly documented.
Arm SME is the standardisation of the concept, but is not inmarket yet.
https://community.arm.com/arm-community-blogs/b/architecture...
cblas_sgemm: 36 GFLOP/s
vDSP_mmul: 41 GFLOP/s
That's a pretty big deal if these functions are >30x faster on the M1...!edit: that seems to be verified in the tlkh.dev blog post above. Interestingly, I ran the same code on my bargain-basement 2020 iphone SE, and got 259GFLOP/s! These apple devices are pretty mindblowing.
Yes. Aside from benchmarks, you can easily verify this by profiling an application with Instruments and then inspecting the disassembly.
However, it should be said that AMX does not scale linearly with the number of cores, but with the number of core clusters. So, on the M1 if you use Accelerate in two threads (rather than one), performance will barely improve, because the first thread can keep the AMX unit busy enough.
However, e.g. the M1 Pro and M1 Max have two performance core clusters with AMX units in them. So matrix multiplication doubles roughly two times compared to the M1. Similarly, the M1 Ultra has fours performance core clusters, so matrix multiplication performance is roughly twice that of the M1 Pro/Max and four times that of the M1.
Benchmarks:
You can normally expect to get way more than half the 'outcome' from a neural net with half the ram/compute/time/power budget. So neural nets scale 'down' pretty well.
Apple has done a wonderful job of further locking their user into the golden cage they call a platform.
ML is one of the few applications that benefit from platform-specific optimizations, so if you need every ounce of performance, you have your choice of which walled garden you want to tether your application to. The "lock-in" comes from the specific capabilities of your special-purpose hardware, and for serious applications, you're already thinking hard about whether to design your entire implementation around Apple, NVidia, Google/TPU, or even Android devices. For big models, platform-specific needs influence every aspect of model design, including data/model sharding, quantization, training loops...
For non-scientific applications, it's usual practice to train your model in platform-agnostic ways using PyTorch or Tensorflow or whatever and then deploy it to devices in platform-specific ways, whether that's XLA, CoreML, Edge TPU, Android NNAPI, TensorflowJS, or hell, custom-written GLSL shaders or whatever.
We're just starting to see cross-platform frameworks that abstract model inference: TFLite, PyTorch Mobile, ONNX. To their credit, CoreML can act as a backend for any of these, so you don't even need to worry about your platform.
EDIT: I fell for NVIDIA's marketing. The dense FP16 performance is only half of 284.48, which is 142. Thanks to adgjlsfhk1 for the correction.
It turns out that even without the extra training iterations you often lose surprisingly little in terms of quality of output. In reality you can sparsify a lot more, but 2 out of 4 is so simple and easy to implement in hardware, more complex schemes are much harder to justify.
However, small matmuls (say, <2048 bytes in the K dimension) won't get anywhere near 2x performance.
> An important distinction is that the AMX:CPU ratio is not 1:1; not every core has its own AMX co-processor.
My understanding is there's only 1 of those per regular M1 CPU, maybe 4 on the largest one (Ultra).
So that's actually just about as power-efficient for fp32 as a 3090, which according to wikipedia is 35 Tflops in 350W. Supposedly AMX can do 2x rate for fp16 as opposed to the 3090's 4x rate, so maybe 2x less efficient than a 3090 for fp16.
Interestingly, fp64 hits 370 Gflops at 15W...
A single Google TPUv4 'pod' (entire row of datacenter racks) gives 1,126,400 TFlops.
Thats why your pet ML projects will always be behind those done at big companies.
And googles pods have microsecond latency terabit interconnects, while a fleet of macbooks would have hundreds of milliseconds of latency and low bandwidth...
And Google has many pods...
I'm afraid even a massive team of home users will never beat companies with dedicated hardware.
Well ... SETI@home wasn't trying to create ML boobs for anime waifus! ... I suspect the army of young guys that want to do that would dwarf any other distributed service including bitcoin!
0.5 cycles per instruction max, 3.7GHz clock, that's 7.4e9 instructions per second. If I'm reading it right, that instruction does 16 4-wide dot products, which is ~128 ops. So ~950Gops peak in int8 precision on a server class Xeon assuming no clock throttling.
(edit: flops -> ops)
I understand that this 1.5 TFlops may not be an exact comparison (or maybe it's the same), but if it's even within an order of magnitude, it is beyond mind-blowing, and we've just crossed over into at Exaflops at the supercomputer level.
Is there one such unit per CPU core?
So the title is misleading, even if it is true that you get this performance with a program that uses a single CPU core.
CPU instruction -> AMX instruction -> AMX result -> CPU?
How are these kinds of things usually kept in sync/in a manageable state? Like does the CPU block until the AMX returns?
It's hard not to really appreciate some of the devices we have today. For instance, an RTX 4090 is capable of 660 TFlops of FP8 (MSRP 1600). Would not be surprised if we soon have laptops that can do petaflops of computation soon!
> Let's simplify the problem and implicitly transpose the matrix multiplication. Both A and B (our inputs) will have K (our reduction dimension) as the leading dimension. This doesn't really matter much in practice, but it simplifies our code a lot.
The code is
C[n * 16 + m] += A[k * 16 + m] * B[k * 16 + n];
Which means that actually *m* is the leading dimension of A with stride 16, and for B it is *n* with stride 16.Example: AMX predated the standard ARM matrix multiply instructions. Perhaps Apple will add the ARM versions someday and now can remove AMX without breaking compatibility. Or maybe there will be a non-additive AMXv2.
It might not be a 'trick' per-se, but anyone who intends to use a Mac for work should consider upgraded memory (IMO).
I did feel a little duped when I learned that some M1/M2 machines can only support one external monitor. Now I have to replace my two monitors with a widescreen.
I'm going to keep an eye on ram usage for the next few days. I'm curious what it will look like on a more full workload because if things have been swapping out, I haven't noticed.
You can say that it was my dad's mistake to buy an M1 8GB but I say it was pretty lame of Apple to sell a computer that expensive that can't do basic tasks. And don't even get me started on peripherals like external monitors.
Peripherals have been fine, except for monitors as you mentioned. I do think it's ridiculous that I can only have one external monitor when it's clearly able to support more than that. I can add a third monitor through my iPad or DisplayLink, but both of those methods breaks DRM video.