There's a list of MCUs here:
https://docs.edgeimpulse.com/docs/development-platforms/offi...
And some accelerators here:
https://docs.edgeimpulse.com/docs/development-platforms/offi...
This is just stuff that has support in Edge Impulse, but there are many other chips too.
The more recent ESP32-S3 has operations with up to 10x speedup, see https://github.com/espressif/esp-nn
Then there are RISCV chips with neural network co processors like Kendryte K210.
ARM has also defined a new set of extensions for NN acceleration, and reference designs for cores being ARM Cortex M85. Chips are becoming available this year. ST has announced they will have accelerators in several lines. There are dozens of startups creating accelerator designs and trying to pair them with MCUs.
So we have a bit already, with much more to come in the years to come.
On the other hand, esp-nn seems to be code for the xtensa instruction set. I briefly overviewed the instructions. They seem optimized for DSP rather than ML applications. Searching for SIMD returned no arithmetic instructions. Searching for parallel returned instructions for multiply and accumulate. Further, the FPU does not compute any kind of 16-bit floating point numbers.
>ARM has also defined a new set of extensions for NN acceleration Can you provide some more info about this?
There is another chip that is generally available, that has a CNN accelerator/co-processor - the MAX78000 https://www.embedded.com/hardware-conversion-of-convolutiona...
I'm curious, why do you believe int8 will dominate?