YouTube will show labels on videos that use AI
9to5google.com
9to5google.com
For the past few months, I've been marking videos as "not interested" because they are AI-generated, and I can tell.
But the flip side is that as these tools become more prevalent, it's not immediately clear to me how this line will be defined.
If people are using AI to generate scripts but are still reading them, does that count? Or if they're using AI to generate the images but have written the script, does that count?
It just seems messy, but I'm glad they're taking at least an active approach to it. I also think it will be a sign of how Google as a whole will treat AI generated content over time.
See also: why no consumer backup platform offers unlimited quantities anymore. It only takes like a couple hundred hoarders to bleed you dry, and those guys don't even stand to profit from the activity like the get-rich-quick youtube and kindle schemes are promising.
what if you are animating the content but using some tool with AI components in it.
Not only does it seem messy but in the long run not feasible.
Photoshop using “AI” in its clone tool doesn’t run into those kinds of issues.
> We’ll require creators to disclose when they’ve created altered or synthetic content that is realistic, including using AI tools. When creators upload content, we will have new options for them to select to indicate that it contains realistic altered or synthetic material. For example, this could be an AI-generated video that realistically depicts an event that never happened, or content showing someone saying or doing something they didn’t actually do.
It's for faked realistic videos. Scripts are unrelated.
I just built an AI-generated fakenews app [1] (for fun) and it opened my eyes: we're playing with fire.
The tech is already there: a bit of roop (deepfake) + SadTalker (lipync) + chatGPT, etc and voila! Anyway can create realistic videos / music on the fly! It's both thrilling and terrifying.
AI's involvement in media production isn't just a technical footnote; it's a fundamental shift in the landscape of information and creativity. Just like we scrutinize the origins of our food, we need to dissect the genesis of our media.
What YT is doing here is a first small step. It's time for all tech giants to confront this reality head-on. We're at a crossroads, and the path we choose will redefine our relationship with technology, creativity, and truth.
Thats brilliant. You have to share details on how you built it!
- I think a video with a synthetic actor (even when it's cloning a human) is synthetic and should be labeled as such, whatever the provenance of the script
- I think a video with a human actor, but an AI-written script could also be labeled. The line is blurrier there for me, since some folks have very elaborate prompts which basically amount to "here's my first draft, make it better". But having a straight-up rule is still good. False positives are better than false negatives here.
- And then there's the issue of AI-generated translations (which is what we do at my startup[1] ). I do believe it's fair for viewers to have those tagged as well. And to be able to track provenance to avoid deepfakes.
> YouTube will...require that creators disclose the use of AI in a video
On many of the fake channels, I see comments praising the fake actor for reading the stolen academic material (and ironically many of these comments are likely fake, too!).
If everything that is touched by AI needs to be labeled, then every piece of content will be labeled as AI. If you have a 3 hour recording from Twitch with TTS in the middle, then that would get labeled as AI. Even if they didn't have to, nobody's going to check a 3 hour recording to see that they didn't get a TTS notification.
Put a photo from your phone into a video? Edited by AI.
Put a video from your phone into the video? Edited by AI.
Soon photo editing tools won't even make it obvious that they use AI. Is that "content aware fill" AI? Or is it something else? Will the average person even know?
- between "grammar- and spellchecking", vs "having ChatGPT write the script"
- between "editing your pic with filters", vs "having MidJourney create the pic"
- between "translating your video with TTS", vs "having a synthetic clone of you speak a script from scratch"
The difference is qualitative rather than quantitative. In each case, one is AI-augmented, the second is AI-generated.
And as time goes on this line is going to get even more murky with AI creeping into software like Microsoft Word, video editing suites, Photoshop, et al.
Based on the description it sounds like this has nothing to do with "AI" and is more flag videos that artificially create events that may not have occurred or as some would call it "fake."
More specifically it means: no generative video AI, no AI to write scripts, no AI period. To ** with AI.
The audio path is just as bad, but more fundamentally, how did you opt out of Google’s recommendation AI, and their practice of harvesting view information for ad targeting? Similarly, can you disable their close captioning stuff, and prevent them from using your videos for training?
So, like I said, I avoid using AI myself for generative, creative tasks.
By the way, the sensor is CMOS and not CCD.
Some of that is still stuff you do, but it's stuff you probably do off camera.
It's an extremely fine line.
Yes, it is a fine line, but that does not mean it's better just to bury my head in the sand and ignore it rather than challenge it like some many technologists do.
Is noise cancellation AI? Is my use of GPT to reword one sentence of a speech AI? Is a face-detection autofocus system AI? Is automatically fixing my hair and removing a pimple or two AI?
Paraphrasing, it's using AI to produce visual/audio/environmental deepfakes in a seemingly factual context.
E.g. a comedy sketch doesn't need to be labeled, but if you're putting words in a politician's mouth or adding 12 more rockets to a video that only had 2, then yeah.
In other words it's exclusively to combat misinformation and disinformation.
If I start going on Youtube and watch a bunch of content and my spidey-sense goes off that a lot of these videos are AI-generated (without labels), I'll likely stop going there altogether as well.
And this is not a direct bash on AI-generated content. I think the tech is going to be immensely helpful for all sorts of stuff, including youtube video content creation, but I'm dreading this early adoption period where people pump out low quality junk. I'd rather just avoid it and do something else. I'm sure Youtube is aware of this and is trying to figure out how to control the problem, but the labels just aren't going to help me continue to go to Youtube.
Tough problem, I don't envy the folks at Youtube trying to figure this pickle out.
Let's say it's a subject matter expert (maybe another undiscovered Dr. Huberman) with lack of video production skills to some very custom content that is unique with a subject matter expert.
Being able to explain things very well to create beginners in that case would be greatly benefitted by generative video.
On the other hand, if there is a lack of quality content, and a lack of video editing skills, and it's about ad revenue passively being generated, the new seo optimization... and that kind of fluff can probably be filtered out.
For me, I’ve just discovered we may not hear from enough neurologists in life when specialists can only explain their area, the perspective of the brain is sometimes what I was after all along.
Valuing brain health alone has been beneficial. You may find brain health either by him or someone else is a really beneficial area to spend more time around. I certainly wouldn’t if that information wasn’t reaching me how it currently is.
He has some amazing and very knowledgeable guests on for his fitness videos. The substances and sleep videos also seemed to be very high quality.
Maybe he does participate in some quackery, but everything I've watched seems sound.
I can choose not to watch this stuff, but I think if it starts to trickle too much into the content I want to watch it just won't be worth the effort of trying to weed low quality junk out.
This is all hypothetical vaporware of course. Just saying I wonder if a video encoding solution like this could help people have faith in the video content they see.
When creators upload content, we will have new options for them to select to indicate that it contains realistic altered or synthetic material. For example, this could be an AI-generated video that realistically depicts an event that never happened, or content showing someone saying or doing something they didn’t actually do.
What might remain hard?
Hiring professional voiceover actors, and still doing actual video editing instead of rendering.
The cat and mouse won’t stop, just the floor of what is tolerated will keep rising.
I’m sure there’s people so far ahead in ai video for no reason other than to quietly get ahead and stay ahead.
It won't work: Truth can only be healed with light.
The open source AI community has a lot more to worry about than just trying to maintain a level playing field with corporations.
Also open source licensing should change to prevent them from taking advantage.
Plenty of non ai open source code that’s been taken for free by billionaire corporations that give nothing back, yet more, they order workers around.
Ai can be a tool to replace exactly those that wish to replace everyone else.
Ew, ew, ew, ew, ew. No! I hope that was sarcasm. DRM for everything else so far is already a mistake. Do not put that evil on me, RickyBobby.
There is good prior art happening in the software supply chain, the problem for media content is that you want something like a hardware signature created before software enters the fray
Please, no. We need less IP, not more.
If I wander a few videos too far away I start seeing videos about reptilians controlling the world on channels with 8M subscribers and they show "footage of people shapeshifting caught on camera", completely fabricated alternative history, flat earthers uncovering the grand conspiracy of the globalist etc.
I don't know how AI is worse than this. Also, apparently it's based on the creator disclosing the use of AI in the creation process. I guess the only "authentic" videos will be those of shape shifting reptilians and proof videos that US never landed on the moon. Kind of pointless.
(not a loaded question, and it is possible different companies could emerge that would compete on the basis of their taste/curation/moderation policies. also equally possible it would be too costly/ unprofitable for the market to bear many smaller competing entities).
I think perhaps there's a third option but we just haven't really defined and figured it out yet. Some mixture of crowdsourced taste/moderation plus top down taste/moderation plus unfiltered UGC. Twitter's new user-generated Community Notes might be a good example of a step in a new direction. Social media is still relatively new.
I am on the side of personal responsibility, that is, any content should be associated with its creator, and it should follow them. If someone posts completely ridiculous video, that video should affect their personal lives. If they change mind and apologize, that should be accepted too. That’s basically how real life relations work.
I just find the labeling as AI pointless.
I have not once looked at TV network executives, publishing house editors, museum curators, retail inventory buyers, librarians, gallerists, or magazine editors as any of those things. Why would somebody do so for video curation?
You don’t need some complicated top down bottom up crowdsourced ML blah blah blah. You just need to be able to contextualize content to the curator. Which is what people naturally do when there is an accountable curator.
Perhaps people who grew up recently only know content as endless troves of machine-curated feeds with no accountability or attributability, but that’s actually just a very ahistorical side effect of Section 230.
Every other way of experiencing media has always been through some curated context, with a specific entity you can point at as the responsible curator, and through which you color your experience of the content.
That’s not to say people couldn’t be misled in that model as well, but this whole “curators are the arbiter of truth” thing has no real or historical ground. Explicit curation actually offers the very opposite thing: recognizable, accountable, obvious context.
I think the best answer will be some form of decentralized moderation. All we need is to put curators (users) in a web of trust, and have them cryptographically sign their opinions.
Link?
One thing I'm very cautious about is clicking on rage-bait links from other people. My best friend likes watching dramatic cop-encounters and shares them with me, but I'm always hesitant to click because I'm afraid it will start suggesting them without ability to stop. I can't say if this will work for you, but I aggressively use the "don't show me this" and "tell us why: I don't like the video". I rarely use the "don't recommend channel" (but wouldn't hesitate to use it if I ever got suggested pseudoscience or fake-news bs).
The only workaround I know of, is using incoknito tabs/new profiles/different account. It should be supported out of the box, to watch a video with the option to not include that into your profile.
(my account is messed up beyound hope, for not doing this, I can only start a new one, but I don't use yt that much anyway)
If you click on that kind of shit, of course youtube is going to show you more. I don't know how people think youtube works but it seems like it should be obvious. If you click on nonsense "ironically" just to laugh at it, youtube doesn't know that and is going to show you more. Even if you dislike and "don't recommend" the video/channel, youtube doesn't know why you disliked it. Maybe you disliked it because the video said the aliens are from Venus but you know they're actually from Alpha Centauri. Youtube doesn't know why you disliked it, all they know is that you clicked on it in the first place, so they'll try to find more to show you.
All I click on is videos about ships, airplanes and trains. That's all youtube recommends to me. No wacky shit in my recommendations because I never click on anything even remotely wacky, not even just to laugh at it. Youtube is a mirror that reflects your viewing habits back at you. If it recommends trash it's because you watch trash. Btw I don't even use an account, just a cookie.
Video scripts have been near algorythmic for humans already. Does using chatgpt to make your script count as AI made? What if you gave it an outline, walked through it, and guided it to what you wanted - e.g. grammerly?
If you use machine learning to remove background noise, backgrounds in general?
Generated stock images/props/scenes for video essays?
If you made the entire video by hand but use text2speech?
I think that "the script, images, and voice were all chatGPT" is obviously "synthetic", but that's just the extreme. Humans have been using technology to augment their creation abilities forever.
My fear is a large amount of human made content will be called "synthetic" because some specific part used "AI" (which nowadays refers to literally any procedural/statistical/machine learning I guess?)
Content would then be labeled as "AI generated" if the "human work hours" was less than X hours (like 0.5 hours).
Whereas content would not be labelled as "AI generated" if the "human work hours" was greater than or equal to X hours (like 0.5 hours).
That would likely require sharing with YouTube not just the final work product, but rather all the iterative work product that led to the creation of the final work product (i.e., earlier drafts). When YouTube could calculate the number of "human work hours" based on the incremental creation/edit histories when analyzing all the drafts.
This is a very hard problem and unlikely to actually be solved in this way.
>That would likely require sharing with YouTube not just the final work product, but rather all the iterative work product that led to the creation of the final work product (i.e., earlier drafts).
Not necessarily. YouTube generally trusts the word of the creator. If I upload a video and check the box that says "there is no swearing in this video" then YouTube more or less takes my word for it. They try to do some detection, but it has always been spotty.
I think YouTube would just ask the creator to label the video and that's it. If somebody is found to be in violation too many times then their account gets actioned.
Maybe there is no line to draw, and maybe that's OK.
This all feels so new I'm hesitant to form strong feelings, but I do think transparency while not required is appreciated. With knowledge, I can draw my own conclusion. Without it, I may end up feeling "duped."
Generally I want to continue enjoying human-created art. If something is "more" human-made than not, that's a factor I'll weigh as "better" and more authentic in my mind. But I'm also trending towards preferring indie dramas over CG-filled blockbusters, so I'm not pretending to be any sort of bellwether.