Is that really the case? My understanding was that M1 is fast because it's able to keep the chip saturated with instructions due to a large L1 cache and wide instruction decoders. Is anything about that specific to mac workloads?
The memory and instruction architecture may be more 'generic' but it and the neural engine, storage and media controllers, image processor etc will have been shaped and fine tuned by the requirements of the mac.
It is probably the marginal gains of each subsystem being 5-10% better for purpose that gives it the edge.
https://images.anandtech.com/doci/16226/perf-trajectory_575p...
Compare https://slatestarcodex.com/2019/03/13/does-reality-drive-str...
Intel has enough product lines that there are no “unnecessary interfaces”; what were the unnecessary interfaces in the Intel chip used by Apple?
Similarly, so does Qualcomm and the tens of other of ARM licensees - any mass market design can find or customize a core with no meaningful dead weight.
It may be a contributing factor, but I have not seen anything to indicate it’s very important.
I give the “many small things done right” theory much higher likelihood - just as in the case if the iPhone, there wasn’t any specific thing that wasn’t done before - except for a winning combination.
Everything Management Engine related, for one.
So that does not fit the GPs claim that the M1 is fast because Apple does vertical integration and others have to support many systems.
Based on?
Many Apple frameworks tie into those accelerators; including frameworks like Metal that one wouldn't necessarily expect.
Another strength of Apple - as a programmer your workloads can gain automatic acceleration as frameworks are extended to leverage new hardware. Note this isn't anything new; it's been going on for years already.
People need to stop saying stupid stuff like "It's only fast because it's so intensely focused on macOS."
But your point, that there's still a lot of Apple stuff at play there, is a fair one. It will be very interesting to see how (native) Linux runs once that porting effort gets more under way...
In principle whether you are using Python or C++ doesn't matter. It is just an interface. The compiler or interpreter in the back decides the actual performance. Yet it is pretty obvious that the specifications of C++ syntax makes it much easier to create a high performance compiler than the specification for Python.
I have been quite involved with Julia. It is a dynamic language like Python, but specific language syntax choices has made it possible to create a JIT that rivals Fortran and C in performance.
Likewise we have seen from Nvidia slides when the went with e.g. RISC-V over ARM, that the simple and smaller instruction-set of RISC-V allowed them to make much cores consuming much smaller silicon, better fitting their silicon budget.
When you worked as a chip architect didn't the ISA affect in any way how hard or easy it would be for you to make an optimization/improvement in silicon?
I mean if one ISA requires memory writes to happen in order, or have variable length, or left too little space for encoding register targets etc. All that kind of stuff is going to make your job as a chip architect harder isn't it?
Also I don't quite get your argument about modeling the M1 around Mac workloads. We know the M1 is having great performance on Geekbench and other benchmarks which have not been specifically designed for Mac workloads.
Only things I can see with M1 which is specific to Mac is:
1) They do the code needed for automatic reference counting faster. Big deal on Mac since more software is Objective-C or Swift which uses automatic reference counting.
2) They prioritized faster single cores over multiple cores. Hence optimizing more for a desktop type workload than a server workload.
3) A number of coprocessors/accelerators for tasks popular on Macs such as image processing and video encoding. But that is orthogonal to the design of the Firestorm cores.
I don't claim to know this remotely as well as you. I am just trying to reason based on what you said and what I know. Would be interested in hearing your thoughts/response. Thanks.