I wonder if this can be optimized by letting GPT provide multiple instructions per screenshot instead of just one.
For example in the twitter screenshot, it could use just the one image.
For example in the twitter screenshot, it could use just the one image.
we could try to patch an "interpolation" kinda of thing for change, but also, I'm curious to see if the multi-modal models that are coming out supporting video would be able to actually just "watch the video" in real time, this would be the ultimate solution