The NOT IN handling with NULLs is the part I found most interesting, since that is where rewrites often go wrong. Beyond the regression tests, how do you check that a rewritten query matches the original? For example, running both against random data with lots of NULLs and comparing the results?
The SIGPIPE case that never failed in 20,000 laptop runs but failed on the first Rewind run is a strong result. One thing I wonder about is "same result under all 65 schedules" only rules out failures that happen more than roughly 5% of the time. Is there a way to see how many schedules rarer races need, say 1 in 1,000?
Nice that each frame carries its own checksum, and good that you said upfront the data is synthetic. Two edge cases I would want to see tested: a single record bigger than the 256 KiB frame, and a damaged frame, to confirm the error is caught and the other threads' results are not affected.
Interesting that this happens without any adversarial instruction. One thing I would like to understand: the 61.3% over 105 episodes assumes the episodes are independent, and the 0.9% comes from only around 50 breaches out of 6,000. Did you check how stable that rate is across different wordings of the nondisclosure rule, for example "never disclose in any form, including encoded"?