Amazing project, I love it.
How about using these Cactus models?
Would it make sense for you to collab with those guys (1) for you to design a cheap but improved, commercialisable version of your $250 chip and (2) for them to tailor their runtime and quantizations to such FPGA hardware?
https://github.com/cactus-compute/cactus
Also have you thought about using a Alveo V80? Still not crazy expensive and could fit bigger models with same approach