Just stream it one frame at a time to the model and eat the latency: https://www.youtube.com/watch?v=IHbJcOex6dk if you need more hand holding.
There's a reason why there's a whole family of models from tiny to huge.
There's a reason why there's a whole family of models from tiny to huge.
I'm on windows.
Ideally I'd like the frames to be dropped, so the inference is done on the last received frame? Is this a standard behaviour?
You really need to have a thread consuming the frames and feeding them to a worker that can run on its own clock.
Under windows, say that I have an RTSP stream (or something similar)
Would you use a single python script with which one of this multithreading solutions?
1 import concurrent.futures
2 import multiprocessing
3 import threading