For some reason, this reminds me the way video encoders compress video:
https://en.wikipedia.org/wiki/Video_compression_picture_type...
It makes me wonder if you could use a similar technique (iframes, bframes or pframes) to get the diff of a "normal" WSI and then train on pattern recognition of those.
These different frames are used to reduce network transmission costs, but it feels similar to the context window if you squint at it as a throughput problem rather than a context window size problem.
It feels like there would be a lot of tools and codecs you could leverage here.