They need some kind of context input.
-GPS position, intent/goal, domain etc.
I'm at a dog show I would want breed etc.
I'm on the street I just want it come back dog maybe dangerous dog, friendly dog.
Also, would be cool/scary to just get back movable object 1, person 1, living movable object 3 etc. and if I give it multiple scenes from a video it knows person 1 is the same person 1 and if I name (them) Tony it keeps tracking tony.