PSRAM, Flash, and even SD cards may not have the best bandwidth individually, but they can reach quite impressive rates when you have a shitton of them running all at the same time.
The large scale dedicated hardware systems will still have the edge for performance per watt, but the low entry level and slow incline does make these things quite appealing.
They have a single cycle double multiply per core, and the interpolators give you a heap of ability
The PIO can be awkward, but you can run a bunch of them at once. Going from MCU to MCU you don't even need to involve the CPU cores, PIO to PIO Comms via pins
You are obviously not going to get big TOPS from it because a Trillion is a ridiculous amount anyway. But never underestimate the power of controlling the whole pipeline.
Ultimately none of the other things I'm doing with MCUs are practical, why would this to be any different.
As someone who has done a reasonable amount with PIO, I do not think this is possible. However, that should not stop you. If you get it to work, please ping me.
I can see that being cheaper to bitbang with PIO than to actually compute.
There's certainly some latency stack up, but throughput should be remarkably good.