But in the development phase, when you are testing on a smaller corpus of data, to make sure your code works, the on-laptop dedicated chip could expedite the development process.
If ML developers can assume that consumer machines (at least "proper consumer machines, like those made by Apple") will have support to do small-scale ML calculations efficiently, then that enables including various ML-based thingies in random consumer apps.
It can be surprisingly cost-effective to invest a few $k in a hefty machine(s) with some high-end GPU's to train with due to the exceedingly hefty price of cloud GPU compute. The money invested up-front in the machine(s) pays itself off in (approximately) a couple of months.
The "neural" chips in these machines are for accelerating inference. I.e. you already have a trained model, you quantise and shrink it, export it to ONNX or whatever Apple's CoreML requires, ship it to the client, and then it runs extra-fast, with relatively small power draw on the client machine due to the dedicated/specialised hardware.
Unfortunately Apple was very vague when they described the method that yielded the claimed "9x faster ML" performance.
They compared the results using an "Action Classification Model" (size? data types? dataset- and batch size?) between an 8-core i7 and their M1 SoC. It isn't clear whether they're referring to training or inference and if it took place on the CPU or the SoC's iGPU and no GPU was mentioned anywhere either.
So until an independent 3rd party review is available, your question cannot be answered. 9x with dedicated hardware over a thermally- and power constrained CPU is no surprise, though.
Even the notoriously weak previous generation Intel SoCs could deliver up to 7.73x improvement when using the iGPU [1] with certain models. As you can see in the source, some models don't even benefit from GPU acceleration (at least as far as Intel's previous gen SoCs are concerned).
In the end, Apple's hardware isn't magic (even if they will say otherwise;) and more power will translate into higher performance so their SoC will be inferior to high-power GPUs running compute shaders.
[1] https://software.intel.com/content/www/us/en/develop/article...
Now the accelerator in the M1 is only 11 TFLOPs. So it’s definitely not trying to compete as an accelerator for training.