> The big reason why Apple Silicon is so damn fast is because they were able to shape the design at modeling time exclusively around Mac system level workloads
Is that really the case? My understanding was that M1 is fast because it's able to keep the chip saturated with instructions due to a large L1 cache and wide instruction decoders. Is anything about that specific to mac workloads?