And here's an FPGA optimized RISC-V RV32IM core for that device (Altera Cyclone IV): https://github.com/VectorBlox/orca
I haven't tried this particular one.
I haven't tried this particular one.
If so, then (naively) could one pack ~10 on that single FPGA? Or does the 'packing overhead' become a big problem? Or does the design use more (say) multiply units pro-rata, so that they become the limiting factor?
You can easily add some logic to let CPUs share pins, too.