Taking to the next level - Batching your I/O in software is how you can start saying things like "Transactions per disk I/O", not just "Fewer I/O for those transactions which now fit in fewer blocks due to compression". Batching doesn't have to mean "nightly processing". It can mean "all requests which occurred over the last 100uS". From a user's perspective, this can effectively still be a real-time RPC experience. For systems with very heavy load, this sort of micro-batching can add many orders of magnitude improvement in throughput. Also bear in mind that the more transactions you have available to compress each time means you get better odds when dealing with entropy.
I have personally developed software that can insert 2-5x the stated write IOPS figure with these sorts of tactics. On modern NVMe devices this can mean you start tickling 8 figure transactions per second if the size of each request is very modest.
Sounds like you solved very interesting problems.
https://www.usenix.org/conference/inflow14/workshop-program/...
Using a COW filesystem adds at least some amount of usage, since instead of modifying in place, you'd write a new block and only trim the old block sometime after the new block is committed; but if you don't have snapshots and you have zfs autotrim (and it trims all your old blocks), the commit interval is short (5 seconds by default?), so I wouldn't expect a big difference in effective free space here.
Edit: reading other comments, it looks like their block size is only 64KiB, so this isn't the case, so I don't know the answer. I can only think that perhaps it is an issue because ZFS doesn't deallocate the freed blocks quickly enough and they are making changes fast, meaning a significant amount of disk space is used by blocks that are about to be TRIMed but haven't been yet.
I guess arguably this comes from the lack of availability of ZNS drives, especially ones that aren't outcompeted by a Samsung 980 series drive on workloads that don't exceed the 980's warranty's 600 drive writes over 5 years. And depending on what you're doing, you may even use such drives for the IO-heavy load and retire them after 90% of their TBW to something like fileserver duty.
A zoned-like workload that is append-only within large chunks and uses a Deallocate (Trim) or Write Zeros command to clear chunks should perform well even on standard SSDs. I don't think the exact zone size has much impact, as long as it's comfortably above the erase block size.