You Only Look Once: Unified, Real-Time Object Detection
arxiv.org
arxiv.org
I also wonder to what extent merging the detection with underlying P-frame information from the video codecs would help. Knowing that a segment of video just moved to the left would mean the detected object could be moved to the left by the same amount, even if it was passing behind another object. Calculating the movement vectors independently seems silly if you can get that data from the underlying video codec itself.
Edit: to add something more "helpful" to this comment, their paper links to a YouTube channel [1] that shows demos of their method, which I think is great.
The objective function is defined heuristically, and involves about five different sub-objectives (top of page four). Some of the parameters chosen seem to be rough guesses, as does the decision to scale up the images to twice the resolution when moving from classification (the pre-training task) to detection.
It seems miraculous that a process of estimating and refinement, guided by experience, can work on tasks where you have no mathematical guarantee that a good solution can be found. Maybe in time we'll build the theory that explains just why deep learning works so well, but for now I'm just kinda awed and impressed every time one of these stories comes out.
The other possibility would be "maximum a posteriori" which doesn't fit their usage here.