-- -----
So, the answer is you don't need a lot to get big impact.
Let's take some NICs like the Intel XL710-QDA1. These are about $250 used. This is a 40Gbps NIC using just PCIe 3.0 x8. I used DACs to drop power consumption through the floor, and to reduce latency. All of this is presented via ISCSI, and with jumbo frames to eke out a few extra percent throughput. If you're using passive DACs, figure 4W of power per port for the NIC, plus another 1.2W at the switch (assuming four SFP+ fan-out), for one end of that connection. You could also just direct connect between the server and the client.
At this point, you can basically shove older prosumer (say 980 Pro) PCIe 4.0 x4 NVMe sustained transfer over the network. Granted, if you're outside of that STR use case, you'll fundamentally be limited by IOPS. Figure for every 40Gbps, you can throw up to 1,200,000 4K IOPS worth of data over the wire.
If you hit the IOPS cap, increase your link speed. A PCIe 4.0 x16 card can handle two 100GbE ports just fine. Note that as you increase IOPS, you'll eventually hit your operating system's IO scheduler limits somewhere in the 10M-15M IOPS range.
-- -----
The question then is actually keeping the NIC fed if you're going over the network. If you're local, it's at least much easier.
First you likely have data buffered in memory for reads. So RAM will cover you there. For writes, you probably want a FUSE pass-through filesystem on NVMe in front of your real backing store (if it's disk, and you're not pure NVMe), or alternatively a writeback cache. The idea here is that pass-through filesystem is basically a storage tier that sits in front of another volume (or volumes) in an entirely transparent manner, so it still appears as if you're writing to a volume that might just be a RAID with a large amount of disks, but instead it's being written to the (presumably mirrored or striped+mirrored) NVMe first to move it to the final destination later. Alternatively, you could also double it as additional read cache, if you don't need to use is all as a write cache too.
-- -----
That's basically what I've done. I have QSFP+ NICs in the storage server and my HEDT (10GbE, 2.5GbE and 1GbE in the rest of the lab hosts), a mirrored NVMe cache/ingest tier, and a boatload of 16TB HDDs in ZFS across two ZFS volumes behind it. This is further backed by 128GB of ECC memory for ZFS ARC, and all powered by an AMD Ryzen 7 PRO 5750GE that maxes out at just under 39W. The whole system, with the NIC, HBA, two 2TB NVMe drives, eight 16TB HDDs, SATA DOM, 128GB ECC, 8-core 35W TDP CPU, and onboard BMC idles at about 43W, and loads that can primarily hit the cache, sits in the 60-65W range, with an absolute system peak under 140W. These figures can be confirmed by a metered PDU.
This gives me ~80TiB of usable redundant storage, with native prosumer PCIe 4.0 x4 NVMe performance, over the network... with an idle of 42-43W that bumps up to 60-65W for most work.