Teardown: Identifying Apple M1's distinct circuit blocks
eetasia.com
eetasia.com
Has anyone done the actual analysis of what the dedicated logic cores on the M1 are? I'm always curious what kinds of work Apple dedicates hardware to.
A determined enough team of people with sulfuric acid and electron microscopes could work out exactly what is going on inside the chip.
(not that anyone would actually do this - maybe Intel? But at least in principle, the information is accessible to non-Apple folks)
Also, what kinds of jobs is that neural engine for? Presumably voice and face recognition to start with, but that doesn't strike me as enough to justify that die space. Good for research of course, but Apple seems to focus on what's practical for everyday users. I've seen stories about CPU designers wanting to use machine learning to manage system cache. I wonder if the NPU will be involved in managing the workload of the M1 chip itself.
Without the neural engine, the processing of tens of thousands of photos can be a real burden on a typical MacBook. Apple has tried to throttle the process so that it doesn’t peg the CPU, but then it takes a lot longer to complete. The neural engine lets the process run full out without consuming a ton of power.
I think the NPU is not as power hungry for inference than the GPU is for the same job.
In addition to being useful for training and inference, the consumer cards use tensor cores for things like mic filtering (RTX Voice) and neural upscaling (DLSS) in games.
General purpose GPU hardware is way more wasteful for matrix math, like maybe >10x waste on power and equally worse performance, than tensor cores.
I didn't realize there was that distinction; I thought GPU's were just optimized for vector arithmetic across the board. What is the difference between general purpose GPU hardware and tensor cores? What does general purpose GPU hardware do that tensor cores do not?
To my knowledge, "tensor cores" and neural accelerators are modeled on something like a coprocessor with a very fast memory bus, and a bajillion of the same small execution unit that can do a single operation in parallel (like a 3x3 matrix multiply) on behalf of the main processor. Like if AVX512 was actually AVX 1,000,000 and only had one instruction, and that instruction did some kind of 3x3 matrix math.
Imagine if you had a very large specialized house made entirely of kitchens (instead of bedrooms, bathrooms, etc), and your roommates were all cooks, you could cook significantly more food at once than in the typical one kitchen small household. You also save power per meal cooked because all of the lights and electricity go to kitchens instead of other rooms.
So on a pipelined CPU, the processor has different pipeline stages that execute in order. An instruction may for example move from fetch, to decode, to load, to execute, to store. The execute step may be executed on a different part of the CPU depending on what kind of instruction it is. Basic arithmetic, floating point, and vector math (such as AVX) can be dispatched through different execution ports and run on different parts of the processor. So a processor may (per core) have a pipeline, an integer math unit, a floating point math unit, and two vector units. Operations running on the execution unit also take some variable amount of time to complete. Having ports to two execution units of the same type available makes it so the processor pipeline can dispatch a second long instruction of the same type before stalling.
I don't actually know how the hardware of a tensor accelerator works, but what I would imagine a "tensor core" to be, is thousands of identical execution units that can only do basic matrix math, and a basic pipeline that is much simpler than a typical CPU.
CPUs and GPUs have highly variable workloads and need a lot of specialized hardware on chip that may not always be in use. This wastes power and means you can't have any one task as densely or efficiently as a dedicated chip. If you're designing a dedicated chip, which has direct access to the main cpu memory (as the apple neural engine has), you can design the chip to directly (streaming) read a large matrix from memory, perform an operation on it, and store the result back into memory.
Normal CPUs and GPUs don't have this capability. They approximate matrix math with lots of individual instructions that just go through a pipeline, stall, cache miss, etc just to do a lot of floating point vector math and store the result back to memory. A dedicated chip can skip all the overhead and just tile efficient matrix math in thousands of execution units. That's why an NVIDIA A100 is 19 teraflops when doing normal floating point vector math, and 150 teraflops when doing fp16 matrix math. It has a section of chip dedicated to efficiently doing the required floating point operations en masse without overhead or extra cycles and cache for fetching instructions.
The hardware requirements for training are much higher than for inference. Roughly:
- Training: Requires massive throughput and compute, so is done on high-end GPUs or TPUs in datacenters (or on researchers' beefy workstations)
- Inference: Compute requirement depends on the size and complexity of the model, but many models can be run on low-end smartphone hardware.
Generally speaking, any time you see "ML hardware" in a consumer device, that hardware is for inference, not training. Apple's Neural Engine is an example, of this. Other current examples include the ARM Ethos-N series and Qualcomm AI Engine, among others. Often this hardware is designed to balance inference performance with power consumption, so that inference can be done efficiently on battery-powered devices.
Training workloads on the M1 use the GPU, not the Neural Engine: https://blog.tensorflow.org/2020/11/accelerating-tensorflow-...
Results: https://marcan.st/transf/latte_stitched.jpg
Edit: sorry, the link was previously the (lower res!) Chipworks shot I'd annotated previously. I got confused. Replaced with mine. It's less clean but higher resolution than what they released.
At the highest zoom I could get I could just about make out individual bits in the eFuse array. This is a Wii U GPU (Latte).
https://marcan.st/transf/latte_otp_slr.jpg
No chemicals needed. I'm probably one of very few people crazy enough to do this mechanically, but it clearly works!
Here's a single shot done with a reversed cheap kit Canon lens of a PS3 GPU (RSX), again with the scraping method.
https://marcan.st/transf/rsx.jpg
I do happen to have access to a SEM too, and obviously that gets a lot more fun. Here's the same eFuse area (again this is still the mechanically deprocessed chip!):
https://marcan.st/transf/sem/latte/20161210_201044.png
I think the mechanical delayering usually gets me to M1 or so, though it's not completely consistent depending on the specific process. Some chips work better than others.
Then again, on older chips you can usually just take a top layer shot and tell apart most of the layers anyway.