I don't quite get how diffing frames allows you to find the scores.
TFA mentions comparing a frame with and without - but how do you generate that frame without? If you can already do it, what's useful about doing that?
TFA mentions comparing a frame with and without - but how do you generate that frame without? If you can already do it, what's useful about doing that?
And then he does a good ol' regular crop on the original image to get the UI excerpt to feed the vision model.