Is the efficiency really that depend on the architecture (x86 vs arm64)? I always thought Apples lead comes from things like the little big approach and the lower transistor size. Can someone elaborate?
You also save a lot of time and power not having to move data between memory pools by having fast access to the unified memory.
It seems that there’s no one thing that achieves the efficiency numbers we’re seeing. It’s excellent execution across the entire design.