The biggest thing I have realized (as have others), is that traditional IO wait strategies don't make sense for NVMe. Even "newer" strategies like async/await do not give you want you truly want for one of these devices (still too slow). The best performance I have been able to extract from these is when I am doing really stupid busy-wait strategies.
Also, single-writer principle and serializing+batching writes before you send them to disk is critical. With any storage medium where a block write costs you device lifetime, you want to put as much effective data per block as possible, and you also want to avoid editing existing blocks. Append-only log structures are what NVMe/flash devices live for.
With all of this in mind, the motivations for building gigantic database clusters should start going away. One NVMe device per application instance is starting to sound a lot more compelling to me. Build clusters with app logic (not database logic).
In testing of these ideas I have been able to push 2 million small writes (~64k) per second to a single Samsung 960 pro for a single business entity. I don't know of any SQL clusters that can achieve these same figures.