You could also stagger the moderation to reduce costs. E.g.
Text analysis: 2 views
Audio analysis: 300 views
Frame analysis: 5,000 views
I would be very surprised if even 20% of content uploaded to YouTube passes 300 views.
You could also stagger the moderation to reduce costs. E.g.
Text analysis: 2 views
Audio analysis: 300 views
Frame analysis: 5,000 views
I would be very surprised if even 20% of content uploaded to YouTube passes 300 views.
I guess it could also be associated with views per time period to optimize better. If the video is interesting, people will share and more views will happen quickly.
There's only so much you can do by guessing the next probably token in a stream. We will probably need something else to achieve what people think that will soon be done with LLMs.
Like Elon Musk probably realizing that computer vision is not enough for full self-driving, I expect we will soon reach the limits of what can be done with LLMs.