In the abstract this is very similar to how a Kalman filter works. There is a prediction component and a correction, with a (possibly time varying) adjustments for how much weight to give to measurements. The equivalent here would be to simply ignore the measurements when the eyes move since the input will be known to be messy, and just rely on the prediction.
This turns out the be a very effective (and in some ways mathematically optimal) way to process sensor input even in human-engineered systems.