In any case, I bet the M3 will be able to run rings around the current MacPro.
In any case, I bet the M3 will be able to run rings around the current MacPro.
Yes.
> Compare eg, wind speed (particles) versus the speed of sound (field).
Huh? The speed of particles in the air is even faster than the speed of sound.
Now lets talk about actual processors instead of doing an analogy. 15cm per ns is more than enough to travel everywhere inside your CPU but the vast majority of logic is localized (usually signals stay in the same core). It's only awful if you go off package to DRAM or a second socket but then the budget is often higher than 1ns. Apple probably scores a lot of performance points here because the RAM is so close to the CPU.
• Apple has the benefits of complete vertical integration, both on the hardware and software side.
• Neural engine is essentially Tensor Cores in NVIDIA’s GPU but occupies at least 4x the equivalent die area as tensor cores (no public details on performance yet).
• NVIDIA doesn’t want to make their consumer GPUs too powerful on tensor operations in order to not cannibalise their 1000% markup ML cards.
It’s almost a classic Intel: financial greed and financial engineering, combined with complacency from being long for so long.
Heck, AMD’s new top end card is tied with the 3090 - but $500 cheaper.
NVIDIA is reportedly scrambling to try and get back into TSMC who is going to make an example out of them.
If you want to talk TOPS, an A100 does up to 1.2 exa-ops INT8, and 2.4 exa-tops INT4. That is, an A100 is more than 1000 times more powerful than the M1 at inferencing, while also supporting up to FP32/FP64 weights, and a 3080 is more than 500 times more powerful than the M1 neural engine.
Getting a better result via Rosetta 2 on the single-threaded benchmark is very impressive. I assumed that their benchmark included some vector instructions (which as far as I understand Rosetta 2 does not emulate) so this would mean that these higher numbers are for general-purpose instructions vs vector. That said, I can't find any reference to AVX or SSE in the "Geekbench 5 CPU Workloads" document, although the one for Geekbench 4 does mention them. It would be interesting to see numbers for Geekbench 4 on M1, if it runs via Rosetta at all. I would imagine that the tool is able to detect what is supported to run optimized code for each CPU being evaluated.
[1] https://www.macrumors.com/2020/11/11/m1-macbook-air-first-be...
[2] https://browser.geekbench.com/macs/imac-27-inch-retina-mid-2...
If indeed v5 tests are not particularly about specialized performance but more about real-world use cases (e.g. PDF rendering, SQLite, image compression as others have mentioned) then running Geekbench v4 on M1+Rosetta would give a comparison of emulated x86 without vector instructions versus native x86 with all of its modern capabilities. Now if the M1 wins that…
[1] https://www.geekbench.com/doc/geekbench4-cpu-workloads.pdf
[2] https://www.geekbench.com/doc/geekbench5-cpu-workloads.pdf