If you're looking at some generic things that are similar to what CLIP was trained on, this would work. Say you're interested in specific physical security metrics, or monitoring defective parts, or specific things about traffic, etc. CLIP might just say "people walking" or "car in intersection" or "part on conveyer belt" which isn't meaningful enough if all your images are exactly that, but with other small differences.
Another important aspect of this is the amount of frames you need to process. Running CLIP even on 27 million frames (1 day of footage) is super expensive. We've built some infra that makes processing video efficient (forms of parallelization + filtering), without you having to think about it.