• Apple has the benefits of complete vertical integration, both on the hardware and software side.
• Neural engine is essentially Tensor Cores in NVIDIA’s GPU but occupies at least 4x the equivalent die area as tensor cores (no public details on performance yet).
• NVIDIA doesn’t want to make their consumer GPUs too powerful on tensor operations in order to not cannibalise their 1000% markup ML cards.
It’s almost a classic Intel: financial greed and financial engineering, combined with complacency from being long for so long.
Heck, AMD’s new top end card is tied with the 3090 - but $500 cheaper.
NVIDIA is reportedly scrambling to try and get back into TSMC who is going to make an example out of them.
If you want to talk TOPS, an A100 does up to 1.2 exa-ops INT8, and 2.4 exa-tops INT4. That is, an A100 is more than 1000 times more powerful than the M1 at inferencing, while also supporting up to FP32/FP64 weights, and a 3080 is more than 500 times more powerful than the M1 neural engine.
Yes.
> Compare eg, wind speed (particles) versus the speed of sound (field).
Huh? The speed of particles in the air is even faster than the speed of sound.
Now lets talk about actual processors instead of doing an analogy. 15cm per ns is more than enough to travel everywhere inside your CPU but the vast majority of logic is localized (usually signals stay in the same core). It's only awful if you go off package to DRAM or a second socket but then the budget is often higher than 1ns. Apple probably scores a lot of performance points here because the RAM is so close to the CPU.