This feels like the kind of thing that's so obvious it's hard to believe it isn't being pursued... Google could, for example, make a _huge_ splash with the Pixel 10 by presenting this with the option of after-the-fact optical zoom or wide angle shots, or using the multiple lenses for some fancy fusing for additional detail. And to your point, DSLRs have been doing deferred readout in the sense of storing to an on-device cache before writing out to the SD card while waiting for previous frames to complete their write ops... this same sort of concept should be able to apply here.
I don't know much more about the computational photography pipeline, but I imagine there might be some tricky bits around focusing across multiple lenses simultaneously, around managing the slight off-axis offset (though that feels more trivial nowadays), and, as you say, around reading from the sensors into memory, but then also how to practically merge or not-merge the various shots. Google already does this with stacked photos that include, say, a computationally blurred/portrait shot alongside the primary sensor capture before that processing was done, so the bones are there for something similar... but to really take advantage of it would likely require some more work.
But this is all by way of saying, this would be really really cool and would open up a lot of potential opportunities.