Well if you've already annotated all the bottles in the video then why do you need to train a model to do it?
For instance, they say the one he is holding in one scene isn't counted again in a later scene.
What about the reflection in his sunglasses? Is that 1 or 2 (or 3 if you count the source of the reflection wherever that is).