Some bits of my plan, especially with respect to how the async runtime in QEMU is designed, were good. However I had vastly underestimated the complexity of one step. The various disk images form a graph that can change on the fly, for example if you make a live snapshot of a VM disk. The method I had thought of to handle changes to the graph would have added a lot of technical debt. Fortunately it was nacked by the maintainer Kevin Wolf and replaced with something better.
For more details on both my plan and what was actually done, see https://kvm-forum.qemu.org/2023/Multiqueue_in_the_block_laye... (slides, also linked from the post) or https://youtu.be/Ubped0PgvZI?si=IsckfZ7uDNYJNp_y (video).
The important thing, at every step, was being committed to getting it done. Some of the intermediate steps were better in terms of bugs fixed, but worse in terms of complexity because you had to juggle both the "good" locks and the legacy global locks.