Plain and boring computer vision would have still required a team who would identify frames to sample before processing, and then annotate, generalize and classify the data.
Not an expert on this but I don't think same kind of manual effort is required anymore.