The neural engine is small and inference only. It's also only exposed by a far higher level interface, CoreML.
Where it could still make sense is if you have a small VRAM pool on the dGPU and a big one on the M1, but with the price of a Mac, not sure that makes a lot of sense either in most scenarios compared to paying for a big dGPU.
That's because the M1 has a dedicated matrix math accelerator called AMX [1]. I've used it with both Swift and pure C.
https://medium.com/swlh/apples-m1-secret-coprocessor-6599492...
However, for lower precisions (which is what deep learning uses), you're much better off with a GPU.
30Tflops for a 3080 for vector FP32, but 119Tflops FP16 dense with FP16 accumulate, 59.5 with FP32 accumulate, and if you exploit sparsity then that can go even higher.
FP64 is also supported by AMX, making it quite an impressive region of silicon.
Why is it inference only? At least the operations are the same...just a bunch of linear algebra
Inference also prefers different IO patterns, because you don't need to keep the activations for every layer ready for backpropogation.
It is not really comparable on a step per second level but the power consumption and now GPU memory will make it pretty enticing.