There are a range of ML acceleration possible on existing chips. The basic 4-wide 8 bit integer SIMD extensions in NEON is available on basically all ARM Cortex M4F chips, which is already available 8+ years. It gives 4-5x speedup for neural networks.
The more recent ESP32-S3 has operations with up to 10x speedup, see https://github.com/espressif/esp-nn
Then there are RISCV chips with neural network co processors like Kendryte K210.
ARM has also defined a new set of extensions for NN acceleration, and reference designs for cores being ARM Cortex M85. Chips are becoming available this year.
ST has announced they will have accelerators in several lines.
There are dozens of startups creating accelerator designs and trying to pair them with MCUs.
So we have a bit already, with much more to come in the years to come.