* Most of the algorithms that we want to work with in this domain are doing arithmetic operations on ints and floats. This isn't super difficult to do in an RTL, but it's like implementing C++ objects in assembler. You can do it, but you need to think harder than you should.
* FPGAs make you worry about timing. This is a massive shift in thinking for software people. It's also not a value-add; I don't want to care about timing. And it enforces chip-wide dependencies (you can have separate clock domains, but not many of them).
If you simplify the model to "pipelines of arithmetic ops" and then provide an abstraction that eliminates timing (e.g. all ops run in a fixed number of clock cycles and the compiler automatically pipelines them where necessary) then I think you'd have something usable. But this is basically a GPU with a lot of SRAM. Such a constrained problem would run extremely well on any modern GPU or SIMD machine, without the power and cost and obscurity constraints of FPGAs.