Edit: also you can forward generate 5 or 10 dots in batch without much overhead compared to a single dot since the main cost is pulling KV cache from VRAM so you have free tensor units.
Still, I don't see how this really works .. more compute / embedding transformations are being potentially applied to the prediction, but in what circumstances are these filler positions being used in a useful way? The filler token embeddings themselves presumably aren't matching attention keys, but positional encodings for adjacent tokens will be similar, which is maybe what triggers lateral copying into (and perhaps out of) filler positions?
Essentially a brute-force search. Which is a bit wasteful, but better than just blindly taking the first idea