The use cases you write about are mostly constrained by design, not software. Configuration of SRAM based FPGAs is rather slow because it requires a scan chain to program each logical element and shift config bits into, and doing it faster requires even more circuitry. You need to multiplex things onto the fabric in practice, you can't "context switch" AKA temporally multiplex very well, you have to spatially multiplex. But FPGAs are already area intensive; a k-LUT needs 2^k SRAM bits for the table, each bit being 6 transistors, on top of the scan chain to program it, and the registers and latches that go with the LUT in a typical logic element, and routing crossbars, and so on. Assuming K=6 then a single LUT would be like ~100x transistor overhead compared to a CMOS NAND gate (not a 1-to-1 comparison, just a ballpark). The SRAM requirements alone are problematic because it scales far, far worse than logic. If you're talking about a modern 7/5/3-nm wafer, area = money, and that's a shitload of money. So, what part of the system architecture do you even put the FPGA on? In the core complex? It can't be too big, then; your users are 99% better off with that area going to more cache and branch prediction. Put it on an older process and stuff it in the package? Packaging isn't free. Maybe just on the PCB? "Bump in the wire" to the NIC, or RAM, or storage? That limits the use, but an out-of-line design means there's less bandwidth available as input/outputs are shared. There are benefits and costs to them all and they all change the use cases and interface a bit. Now keep in mind you might have multiple parallel bitstreams you want to run. All of these choices impact how the software can interface with it and what capabilities it has.
Example: Modern DDR5 has something like 64GB of bandwidth per channel; assuming your design is inline on the bus running at something like 500MHz, you'd need a 128-bit bus, per channel. That clock rate might require deep pipelining, further increasing area requirements, so you can't fit as much other stuff. Otherwise, you need a wider bus and to go slower, but wider buses often scale sub-linearly in terms of area and routing congestion; a 256-bit bus will be more than twice as expensive and difficult to route as a 128 bit one due to limited routing tracks, etc. So maybe you can hit that target, but then you're too routing congested, so you can't fit as many channels as you want in. Ergo, you need bigger/more FPGAs, or serious optimization and redesign. There's no immediate win. You need to explore/napkin math the design space to find the best solution on the pareto frontier, typically. Or just buy a FPGA that's massive overkill, AKA "buy a faster PC", the typical software programmer's solution. But it really isn't plug and play or anything close to that.
It's similar to other niche things, like in-memory GPU databases. They are not held back by CUDA being proprietary. That fact does suck, but it's not really relevant in the grand scheme. They are held back by physical design dictating that parallel accelerators need loads of fast memory to feed the execution units, fast memory is super expensive and takes up a lot of space on the PCB resulting in a physical upper bound on density, and that the working set for such databases typically grows much, much faster than rate at which GPU memory performance/price drops. Past the point of no return (working set > VRAM), their advantages rapidly vanish. Their limitations are in the design, not the software.
FPGAs taught me a lot about hardware/software design. I really like them and want more people to use them. I'm really excited there are fully FOSS flows, even if they have giant limitations. But they are pretty niche and have serious physical design factors to account for; I say that as someone who contributes to, uses, and loves the open-source tools for what they are, and even was lucky enough to play with them for work.