Call my a cynic but this feels like a free way for Google to pull more data for training.
Call my a cynic but this feels like a free way for Google to pull more data for training.
I never said it would necessarily be nefarious, but it's the same behaviour of data collection from users of free services to benefit themselves financially. While not always being particularly careful with collected user data.
A slightly related topic is around Google's training on YouTube subtitles. They're able to do this because they host all the content, but they dont allow owners of that content to opt out of that. Again, a free resource that Google get to play with as they feel like.
There is nothing that collects data here, and this is not a Google product that users interact with. Hence, there is no direct financial gain from this work. Google Research has published over 10,000 papers, few of which directly impact the commercial side of Google services.
Your example of automated captioning for videos doesn't seem particularly objectionable. Does it translate, indirectly, to slightly more ad views? Probably, but it's a rounding error in their revenue. I am guessing that few content creators who publish on YouTube find this feature controversial: In addition to free hosting and revenue sharing, they don't have to bear the costs of writing captions while benefiting from accessibility and discoverability.
There are valid complaints against data collection by these tech companies for machine learning. Artists have a point when they condemn generative models trained on their work. And you might reasonably object to the collection of handwriting samples that Google used here, which were scraped from public Imgur posts (Facebook Research's Imgur5K data set).
But there's room for nuance in deciding what uses are fair and acceptable without the knee-jerk reaction of tech company + AI = evil financial motives.