YouTube now requires to label their realistic-looking videos made using AI
blog.google
blog.google
• Makes a real person appear to say or do something they didn't say or do
• Alters footage of a real event or place
• Generates a realistic-looking scene that didn't actually occur
At the very least this will test each of these hypotheses, which we'll learn from and iterate on. I am curious to see the legal arguments that will inevitably kick up from each of these - is color correction altering footage of a real event or place? They explicitly say it isn't in the wider description, but what about beauty filters? If I have 16 video angles, and use photogrammetry / gaussian splatting / AI to generate a 17th, is that a realistic-looking scene that didn't actually occur? Do I need to have actually captured the photons themselves if I can be 99% sure my predictions of them are accurate?
So many flaws, but all early steps have flaws. At least it is a step.
Which is where the real abuse comes in: You post footage of a real event and they say it was AI, and ban you for it etc., because what actually happened is politically inconvenient.
And the only way to prevent that would be a reliable way to detect AI-generated content which, if it existed, would obviate any need to tag anything because then it could be automated.
Which doesn't help you unless non-AI images are all required to be RAW. Moreover, someone who is trying to fabricate something could obviously obtain access to a real camera to emulate.
> You'd have a straightforward case for defamation if your real footage were falsely labeled, and it would be easy to demonstrate in court.
Defamation typically requires you to prove that the person making the claim knew it was false. They'll, of course, claim that they thought it was actually fake. Also, most people don't have the resources to sue YouTube for their screw ups.
Yes, but not to your camera. Sorry for not phrasing it more clearly: individual cameras have measurable noise signatures distinct from otherwise identical models.
On the lawsuit side, you just need to aver that you are the author of the original footage and are willing to prove it. As long as you are in possession of both the device and the footage, you have two pieces of solid evidence vs. someone elses feels/half-assed AI detection algorithm. There will be no shortage of tech-savvy media lawyers willing to take this case on contingency.
But who is the "you" in this case? There can be footage of you that wasn't taken with your camera. The person falsifying it would just claim they used their own camera. Which they would have access to ahead of time in order to incorporate its fingerprint into the video before publishing it.
It’s nice honest users will do that but they’re not really the problem are they.
We do, we ask paid endorsements to be disclaimed.
If you want to publish proof of an event, you should have some pixels on a screen along with some cryptographic signature from a device sensor that would necessitate atleast a big corporation like Nikon / Sony / etc. being "in on it" to fake.
Also since no one likes RAW footage it should probably just be you post your edited version which may have "AI" upscaling / de-noising / motion blur fixing etc, AND you can post a link to your cryptographically signed verifiable RAW footage.
Of course there's still ways around that like your footage could just be a camera being pointed at an 8k screen or something but at least you make some serious hurdles and have a reasonable argument to the video being a result of photons bouncing off real objects hitting your camera sensor.
At which point nobody could verify anything that happened with any existing camera, including all past events as of today and all future events captured with any existing camera.
Then someone will publish a way to extract the key from some new camera model, both allowing anyone to forge anything by extracting a key and using it to sign whatever they want, and calling into question everything actually taken with that camera model/manufacturer.
Meanwhile cheap cameras will continue to be made that don't even support RAW, and people will capture real events with them because they were in hand when the events unexpectedly happened. Which is the most important use case because footage taken by a staff photographer at a large media company with a professional camera can already be authenticated by a big corporation, specifically the large media company.
Of course that says nothing about the issues of corruption of judges in the court system, but that is a "relatively" new issues that DOES absolutely need to be addressed.
(Shoot one could argue that the way certain folks are behaving right now is in itself unconstitutional and those folks should be booted)
Countries all over the world (EVEN IN EUROPE WITH THE GDPR) are a lot less "gracious" with anonymous communication. The UK actually has been trying to outlaw private encryption, for a while now, as an example, but there are worse examples from certain other countries. You can find them by examining their political system, most (all? I did quit a bit of research, but also was not interested in spending a ton of time on this topic) are "conservative leaning"
Note that I'm not talking just about existing policy, but countries that are continually trying to enact new policy.
Just like the US has "guarantees" on free speech, the right to vote, etc. The world needs guaranteed access to freedom of speech, religion, right to vote, healthcare, food, water, shelter, electricity, and medical care. I don't know of a single country in the world, including the US, that does anywhere close to a good of job with that.
I'm actually hoping that Ukraine is given both the motive and opportunity to push the boundaries in that regard. If you've been following some of the policy stuff, it is a step in the right direction. I 100% know they won't even come close to getting the job done, but they are definitely moving in the right direction. I definitely do not support this war, but with all of the death and destruction, at least there is a tiny little pinprick of light...
...Even if a single country in the world got everything right, we still need to find a way to unite everyone.
Our time in this universe is limited and our time on earth more-so. We should have been working together 60 years ago for a viable off-planet colony and related stuff. If the world ended tomorrow, humanity would cease to exist. You need over 100,000 people to sustain the human race in the event a catastrophic event wipes almost everyone out. Even if we had 1,000 people in space, our species would be doomed.
I am really super surprised that basic survival needs are NOT on the table when we are all arguing about religion, abortion, guns, etc. Like really?
We are hundreds of years away from the kind of technology you would need for a viable fully self-sustainable off-world colony that houses 100k or more humans. We couldn't even build something close to one in Antarctica.
This kind of colony would need to span half of Mars to actually have access to all the resources it needs to build all of the high-tech gear they would require to just not die of asphixiation. And they would need top-tier universities to actually have people capable of designing and building those high-tech systems, and media companies, and gigantic farms to make not just food but bioplastics and on and on.
Starting 60 years earlier on a project that would take a millennium is ultimately irrelevant.
Not to mention, nothing we could possibly do on Earth would make it even a tenth as hard to live here than on Mars. Nuclear wars, the worse bio-engineered weapons, super volcanoes - it's much, much easier to create tech that would allow us to survive and thrive after all of these than it is to create tech for humans to survive on a frozen irradiated dusty planet with next to no atmosphere. And Mars is still the most hospitable other celestial body in the solar system.
This is the best argument I've heard for why we should do it. Once you can survive on Mars you've created the technology to survive whatever happens on Earth.
Most people in the world struggle to feed themselves and their families. This is the basic survival need. Do you think they fucking care what happens to humantiy in 100k years? Stop drinking that transhumanism kool-aid, give your windows a good cleaning and look at what's happening in the real world, every day.
They come and ask. You say no? They find cocaine in your home.
You aren't in jail because you refused to hand out data. You are in jail because you were dealing drugs.
Or an APT (AKA advanced persistent teenager) with their parents camera and more time than they know what to do with.
I don't follow. Isn't software backward compatibility a big reason why Android device attestation is so hard? For cameras, why can't the camera sensor output a digital signature of the sensor data along with the actual sensor data?
At least with Android there are too many OEMs and they screw up too often. Bad actors will specifically seek out these devices, even if they're not very technically skilled. The skilled bad actors will 0-day the devices with the weakest security. For political reasons, even if a batch of a million devices are compromised it's hard to quickly ban them because that means those phones can no longer watch Netflix etc.
so you cannot share the original if you intend to black out something from the original that you don't want revealed (e.g., a face or name or something).
The way you specced out how a signed jpeg works means the raw data _must_ remain visible. There's gonna be unintended consequences from such a system.
And it aint even that trustworthy - the signing key could potentially be stolen or coerced out, and fakes made. It's not a rock-solid proof - my benchmark for proof needs to be on par with blockchains'.
You can obviously extend this if you want to add bells and whistles like cropping or whatever. Like signing every NxN sub-block separately, or more fancy stuff if you really care. It should be obvious I'm not going to design in every feature you could possibly dream of in an HN comment...
And regardless, like I said: this whole thing is intended to be opportunistic. You use it when you can. When you can't, well, you explain why, or you don't. Ultimately it's always up to the beholder to decide whether to believe you, with or without proof.
> And it aint even that trustworthy - the signing key could potentially be stolen or coerced out, and fakes made.
I already addressed this: once you determine a particular camera model's signature ain't trustworthy, you publish it for the rest of the world to know.
> It's not a rock-solid proof - my benchmark for proof needs to be on par with blockchains'.
It's rock-solid enough for enough people. I can't guarantee I'll personally satisfy you, but you're going to be sorely disappointed when you realize what benchmarks courts currently use for assessing evidence tampering...
The possibilities are pretty endless here.
Then you give me O, I already know K (you tell me which manufacturer key to use, and I decide if I trust it), and the STARK proof. I validate the proof (including the public inputs K and H_O, which I recalculate from O myself), and if it validates I know that you have access to a signed image I that O is derived from in a well-defined way. You never have to disclose I to me. And with the advent of zkVMs, it isn't even necessarily that hard to do as long as you can tolerate the overhead of running the compression / cropping algorithm on a zkVM instead of real hardware, and don't mind the proof size (which is probably in the tens of megabytes at least).
But did you know large portions of The Mandalorian were produced with the actors acting in front of an enormous, high-resolution LED screen [1] instead of building a set, or using greenscreen?
It turns out pointing a camera at a screen can actually be pretty realistic, if you know what you're doing.
And I suspect the pr agencies interested in flooding the internet with images of Politician A kicking a puppy and Politician B rescuing flood victims do, in fact, know what they're doing.
[1] https://techcrunch.com/2020/02/20/how-the-mandalorian-and-il...
We already understand that with text. We know that to verify words, we have to trace it back to the source, and then we evaluate the credibility of the source.
There have been periods where recording technology ran ahead of faking technology, so we tended to just trust photos, audio, and video (even though they could always be used to paint misleading pictures). But that era is over. New technological tricks may push back the tide a little here and there, but mostly we're going to end up relying on, "Who says this is real, and why should we believe them?"
That idea doesn't work, at all.
Even assuming a perfect technical implementation, all you'd have to do to defeat it is launder your fake image through a camera's image sensor. And there's even a term for doing that: telecine.
With the right jig, a HiDPI display, and typical photo editing (no one shows you raw, full-res images), I don't think such a signature forgery would detectable by a layman or maybe even an expert.
And when they do that, the video is now against Google's policy and can be removed. That's the point of this policy.
I would suspect this helps make moderation models better at estimating confidence levels of ai generated content that isn’t labeled as such (ie for deception).
Surprised we aren’t seeing more of this in labeling datasets for this new world (outside of captchas)
To the extent that they allow Google to exclude AI video from training sets they’re obviously useful to Google.
It's the same as paid reviews: tags and disclaimers exist to make it easier to handle cases where you intentionally didn't put them.
It's not perfect and can be abused in other ways, but at least it's something.
That’s the bar- it’s not going to be infallible but if you don’t find evidence of tampering with the hardware then it’s probably going to be fine.
Everything else - corroboration, interviews, fact checking will remain as they are today and can't be replaced by technology. So I imagine a journalist would reach out to person who recorded thr video, ask them to show their device's fingerprint and ask about their experience when (event) occured, and then corroborate all that information from other sources.
When the news org publishes the video, they may sign it with their own key and/or vouch for the original one so viewers of clips on social media will know that Fox News (TM) is putting their name and reputation behind the video, and it hasn't been altered from the version Fox News chose to share, even though the "ModernMilitiaMan97" account that reshared it seems dubious.
Currently, there's no way to detect alterations or fabrications of both the "citizen-journalist" footage and post-broadcast footage.
Im not saying Youtube shouldn’t have AI labels. Im saying we shouldn’t assume they’re reliable.
No. Having sources of trust is the basis of managing complexity. When you turned the tap water on and bought a piece of meat at the butcher you didn't yourself verify whether its healthy right? You trust the medicine you buy contains exactly what is says on the label and didn't take a chemistry class. That's centralized trust. You rely on it ten thousand times a day implicitly.
There need to be measures to make sure media content is trustworthy, because the smartest person on the earth doesn't have enough resources to critically judge 1% of what they're exposed to every day. It is simply a question of information processing.
It's a mathematical necessity. Information that is collectively processed constantly goes up, individiual bandwith does not, therefore you need more division of labor, efficieny and higher forms of social organisation.
This is a false equivalence that I’ve already addressed.
> When you turned the tap water on and bought a piece of meat at the butcher you didn't yourself verify whether its healthy right?
To a degree, yeah, you do check. Especially when you get it from somewhere with prior problems. And if you see something off you check further and adjust accordingly.
Why resort to anology? Should we blindly trust YouTube to judge whats true or not? I stated that labeling videos is fine but what’s not fine is blindly trusting it.
Additionally, comparing to meat dispenses with all the controversy because food safety is a comparatively objective standard.
Compare, “is this steak safe to eat or not?” To “is this speech safe to hear or not?”
Right now, getting videos which are completely AI/deepfaked to misrepresent, are not subject to the same consequences, simply because either #1 people can't be bothered, #2 are too busy spreading it via social media, or #3 have no idea how to sue the party on the other side.
And therein lies the danger, as with social media, of the lack of consequences (and hence the popularity of swatting, pretexting etc)
There will be no way to say something is true beside seeing it with own eyes.
That said, unless all the AI generators agree on a way to add an unalterable marker that something is generated, at one point it may become undetectable. May.
Faking things is not new, and you've always been right to mistrust what you see on the internet. "AI" technology has made it easy, convenient, accessible and affordable to more people though, beforehand you needed image/video editing skills and software, a good voice mod, be a good (voice) actor, etc.
But these tools make deception easier and cheaper, meaning it will become much more common. Also, it's not just "on the internet". The trust problem this brings up applies to everything.
Look at Russian society for a sneak preview if we don't get this right.
That means the system will be really really awful. So challengers can arise - maybe a challenger that YOU build, and open source!
2. Google search is now no better than DDG, which it also cannot recover from without putting a nail in the coffin of its own search monetization strategy
If you want a randomly accuse me of living in a bubble that's fine. It won't bother me one bit. Me living in a bubble doesn't change the fact that Google is backed into a corner when it comes to search quality and competing with LLMs.
You may be right that they can survive more readily than AOL did.. but I certainly won't help them with anything more than a kick in the pants! ;)
And? Do they have a track record of crafting a false reality for people?
I wonder if this will make all forms of surveillance, video or otherwise, inadmissible in court in the near future. It doesn’t seem like much of a stretch for a lawyer to make an argument for reasonable doubt, with any electronic media now.
This website focuses too much on the technical with little regard for the social a bit too often. Though in general, videos being easily fakable is still scary.
No, it won't. Just as it does now, video evidence (like any other evidence that isn't testimony) will need to be supported by associated evidence (including, ultimately, testimony) as to its provenance.
> It doesn’t seem like much of a stretch for a lawyer to make an argument for reasonable doubt,
“Beyond a reasonable doubt” is only the standard for criminal convictions, and even then is based on the totality of evidence tending support or refute guilt, its not a standard each individual piece of evidence must clear for admissibility.
- Makes a real person appear to say or do something they didn't say or do
- Alters footage of a real event or place
- Generates a realistic-looking scene that didn't actually occur
These are things that have been true of edited video since even before AI was a thing. People can lie about reality with videos, and AI is just one of many tools to do so. So, as you said, there are many flaws with this approach, but I agree that requiring labels is at least a step in the right direction.
I think anonymity is completely dead now. Not due to social networks but AI will definitely kill it. With enough resources it is possible to engage an AI arms race against whatever detector they put.
It is also possible to remove any watermarking. So the only way to prove if the content is made by humans will be requiring extensive proofs to their complete identity.
If you're from a minority or a fringe group, all of your hopes about spreading awareness anonymously on popular social media will be gone.
If someone for example makes a political video and fails to label it, they can delete the video/terminate the account for a breach of service.
Power is power - can be used for good or bad. This labelling is a form of power.
If someone says "I am not a crook" and you edit out the "not", do you need to label it.
What if it is done for parody.
What if the edit is more subtle - where a 1 hour interview is edited into 10 minute excerpts.
Mislabeled videos for propaganda.
Or simply date or place incorrectly stated.
Dramatic recreations as often done in documentaries.
Etc
>> 17th angle - AI Generated vantage point based on existing 16 videos (reference links). <<
Would be • Alters footage of a real event or place
Not everyone works in ML.
> with poorly written legislation
This is a company policy, not legislation.
> designed to garner votes
Represent the will of the people?
> while rival companies take strides
Towards?
> in the technology at rapid rates
YouTube's "Trending" page isn't a research lab.
Even if it was, why would honesty slow it down?
Perhaps AI classification is the mirror opposite to porn, using the test: "I'll know it when I don't see it", ie, if an average user would mistake AI generated content for reality, it should be clearly labeled as AI. But how do we enforce this? Does such enforcement scale? What about malicious actors?
We could conceivably use good AI to spot the bad AI, an endless AI cat and AI mouse game. Without strong AI regulation and norms a large portion of the internet will devolve into AI responding to AI generated content, seems like a gigantic waste of resources and the internet's potential.
I suspect this is Google's actual goal with the tagging system. It's not so much about helping users, rather it's a way to collect labeled data which they can later use to train their own "AI detection" algorithms
Google policy isn't law; there's no court judging legal arguments, it is enforced at Google’s whim and with effectively no recourse, at least not one which is focused on parsing arguments about the details of the policy.
So there won’t be “legal arguments” over what exactly it applies to.
Examples of content creators don’t have to disclose:
* Someone riding a unicorn through a fantastical world
* Green screen used to depict someone floating in space
* Color adjustment or lighting filters
* Special effects filters, like adding background blur or vintage effects
* Production assistance, like using generative AI tools to create or improve a video outline, script, thumbnail, title, or infographic
* Caption creation
* Video sharpening, upscaling or repair and voice or audio repair
* Idea generation
Examples of content creators need to disclose: * Synthetically generating music (including music generated using Creator Music)
* Voice cloning someone else’s voice to use it for voiceover
* Synthetically generating extra footage of a real place, like a video of a surfer in Maui for a promotional travel video
* Synthetically generating a realistic video of a match between two real professional tennis players
* Making it appear as if someone gave advice that they did not actually give
* Digitally altering audio to make it sound as if a popular singer missed a note in their live performance
* Showing a realistic depiction of a tornado or other weather events moving toward a real city that didn’t actually happen
* Making it appear as if hospital workers turned away sick or wounded patients
* Depicting a public figure stealing something they did not steal, or admitting to stealing something when they did not make that admission
* Making it look like a real person has been arrested or imprisonedWhat about music made with a synthesizer?
But since AI can't legally have copyright to their music Google probably wants to know for that reason.
But the person who uses that software certainly can own the copyright to the resulting work.
Hopefully this means that AI generated music gets skipped by Googles DRM checks.
Amount of work is not a basis for copyright. (Kind of work is, though the basis for the “kind” distinction used isn't actually a real objective category, so its ultimately almost entirely arbitary.)
In all of the other cases, it can be deceiving, but what is deceiving in synthetic music? There may be some cases where it is relevant, like when imitating the voice of a famous singer, but other than that, music is not "real", it is work coming from the imagination of its creator. That kind of thing is already dealt with with copyright, and attribution is a common requirement, and one that YouTube already enforces (how it does that is different matter).
> Dream Track in Shorts is an experimental song creation tool that allows creators to create a unique 30-second soundtrack with the voices of opted-in artists. It brings together the expertise of Google DeepMind and YouTube’s most innovative researchers with the expertise of our music industry partners, to open up new ways for creators on Shorts to create and engage with artists.
> Once a soundtrack is published, anyone can use the AI-generated soundtrack as-is to remix it into their own Shorts. These AI-generated soundtracks will have a text label indicating that they were created with Dream Track. We’re starting with a limited set of creators in the United States and opted-in artists. Based on the feedback from these experiments, we hope to expand this.
So my impression is they're talking about labeling music which is derived from a real source (like a singer or a band) and might conceivably be mistaken for coming from that source.
> * Making it appear as if hospital workers turned away sick or wounded patients
> * Depicting a public figure stealing something they did not steal, or admitting to stealing something when they did not make that admission
Considering they own the platform, why not just ban this type of content? It was possible to create this content before "AI".
Valid satire, fair use of the original content: parody is considered transformative. But it should be labeled as AI generated, or it's going to escape onto social media and cause havoc.
It might anyway, obviously. But that isn't a good reason to ban free expression here imho.
It was never historically the case that satire was expected to be labelled, or instantly recognized by anyone who stumbled across it. Satire is rude. It's meant to mock people—it is intended to muddle and provoke confused reactions. That's free expression nonetheless!
How is one to figure out what is real and what is a satire? Times and technologies change. What was once reasonable won’t always be.
Context, source, tone of speech, and reasonability.
- "Times and technologies change."
And so do people! We adapt to times and technology; we don't need to be insulated from them. The only response needed to a new type of artificial medium, is, that people learn to be marginally more skeptical about that medium.
Two recent headlines:
* Biden Urges Americans Not To Let Dangerous Online Rhetoric Humanize Palestinians [1]
* Trump says he would encourage Russia to attack Nato allies who pay too little [2]
Do you really think, if you jumped back a few years, you could have known which was satire and which wasn't?
The fact that we have video evidence of the second is (part) of how we know it's true. Sure, we could also trust the reporters who were there, but that doesn't lend itself to immediate verification by someone who sees the headline on their Facebook feed.
If the first had an accompanying AI video, do you think it would be believed by some people who are willing to believe the worst of Biden? Sure, especially in a timeline where the second headline is true.
1. https://www.theonion.com/biden-urges-americans-not-to-let-da...
2. https://www.theguardian.com/us-news/2024/feb/11/donald-trump...
Yagoddabekidding. That could cover any piece of music created with MIDI sequencing and synthesizers and such.
If AI writes MIDI input for a synthesizer, rather than producing the actual waveform, where does that land?
This is interesting because I was considering cloning my own voice as a way to record things without the inevitable hesitations, ums, errs, and stumbling over my words. By this standard I am allowed to do so.
But then I thought what does it even mean "someone else's" when multiple people can make a video, if my wife and I make a video together can we not then use my recorded voice because to her my voice is someone else.
I suspect all of these rules will have similar edge cases and a wide penumbra where arbitrary rulings will be autocratically applied.
To her, your voice is your voice not someone else's voice.
If you share a Youtube account with your wife, "someone else" means someone other than you or your wife.
The more interesting and troubling point is your use of "synthetic you" to make the real you sound better!
Why?
I presume you wouldn't do that? How about replacing your own voice on a video that you make, allowing your viewers to believe it's your natural voice? Are you comfortable with that deception?
Everyone has evolving opinions about this subject. For me, I don't want to converse with stand-in replacements for living people.
The ethical challenge in my opinion, is that your status as living human narrator on a video is now irrelevant, when you're replaced by voice cloning. Perhaps we'll see a new book by "George Orwell" soon. We don't need the real man, his clone will do.
* Digitally altering audio to make it sound as if a popular singer missed a note in their live performance
Does all the autotuning that singers use in live performance counts?/j
Under the current guidelines, doesn't all music performances that make use of some sort of pitch correction assist are technically "digitally altered"?
A bit funny considering a realistic warning and "live" radar map of an impending, major, natural disaster occurring in your city apparently doesn't violate their ad policy on YouTube. Probably the only time an ad gave me a genuine fright.
> Providers will also have to ensure that AI-generated content is identifiable. Besides, AI-generated text published with the purpose to inform the public on matters of public interest must be labelled as artificially generated. This also applies to audio and video content constituting deep fakes
https://digital-strategy.ec.europa.eu/en/policies/regulatory....
Some discussion here: https://news.ycombinator.com/item?id=39746669
The EU made a rule. YouTube complied. That changes the user experience. They documented it.
I mean you'd expect a pharmaceutical company to mention which rules they comply with at some point, even if not on the actual product (though in the case of medicine, probably also on the actual product).
Also, if every applicable regulation had to be mentioned, it’d be a very long list.
[1] https://epaper.telegraphindia.com/imageview/464914/53928423/...
The solution here is for important institutions to get onboard with the public key infrastructure, and start signing anything they want to certify as authentic.
The culture needs to shift from assuming video and pictures are real, to assuming they are made the easiest way possible. A signature means the signer wants you to know the content is theirs, nothing else.
It doesn't help to train people to live in a pretend world where fake content always has a warning sticker.
We literally have politicians talking about pouring acid on hardware and expect these same bumbleheads to keep their signing keys safe at the same time. The average person is far too technologically illiterate to do that. Next time you go to grandmas house you'll learn she traded her signing key for chocolate chip cookies.
If Apple wanted to sign every photo and document the iPhone they could probably make the whole user experience simple enough for most grandmas.
Some people will certainly give away their keys, just like bank accounts and social security numbers today, but those people probably aren't terribly concerned with proving the ownership of their online documents.
Then your imagination fails you.
If it is automatic/easy, then you have the 'easy key' problem, such as the key is easy to steal or copy. For example is it based on your apple account? Then what occurs with an account is stolen? Is it based on a device, what happens when the device is stolen?
Who's doing the PKI? Is it going to be like https, but for individuals (this has never really worked at this scale and with revocation). Like most social media is posting content taken by randos on the internet.
For prominent people who actually have to worry about being impersonated they could provide their own keys.
The infrastructure could be managed by multiple groups or a singular one like the government. The point isn't to be a perfect system, it's to generate enough trust that what you're looking at is genuine and not a total fraud.
In a world where AI bots are generating fake information about everyone in the world, that kind of system could certainly be built and be useful.
Yeah, that's what I said about PKI. I also said there is confusion between provenance of a statement and its accuracy. Just because it's signed doesn't mean anything about its accuracy, but those that confuse the two will think that just because it is signed, or because it was signed by a certain "trustworthy" party, that indicates accuracy. PKI does not establish the trustworthiness of the other party, it only gives you confidence in the identity of the party who signed something.
George Santos could sign his resume. We know it was signed by George Santos. And yet nothing in the resume could be considered accurate (or even a falsehood) purely because it is signed. That it was proven to be signed by George Santos via PKI is independent of the fact that George Santos is a known liar.
That sounds like a dystopia, but I guess we're going into that direction. I expect that a lot of fringe groups like flat-earthers, lizard people conspiracy, war in Ukraine is fake, will become way more mainstream.
How do you expect people to take the authenticity of this video?
> Besides, AI-generated text published with the purpose to inform the public on matters of public interest must be labelled as artificially generated. This also applies to audio and video content constituting deep fakes
Clearly "AI-generated text" doesn't apply to YouTube videos.
But, it is interesting that if you use an LLM to generate text and present that text to users, you need to inform them it was AI-generated (per the act). But if a real person reads it out, apparently you don't (per the policy)?
This seems like a weird distinction to me. Should the audience be informed if a series of words were LLM-generated or not? If so, why does it matter if they're delivered as text, or if they're read out?
On a local level, I recall how various brands started making a big deal of replacing disposable plastic bags with canvas or paper alternatives "for the environment" just coincidentally a few months before disposable plastic bags were banned in the entire country.
This seems oddly specific to the inverse of what happened recently with Alicia Keys from the recent Superbowl. As Robert Komaniecki pointed out on X [1], Alicia Keys hit a "sour note" which was silently edited by the NFL to fix it.
[1] https://twitter.com/Komaniecki_R/status/1757074365102084464
I will be coming back to this video in several months time to check whether the "Altered or synthetic content" tag has actually been applied to it or not. If not, I will report it to YouTube.
However autotune has existed for decades. Would it have been better if artists were required to label when they used autotune to correct their singing? I say yes but reasonable people can disagree!
I wonder if we are going to settle on an AI regime where it’s OK to use AI to deceptively make someone seem “better” but not to deceptively make someone seem “worse.” We are entering a wild decade.
A lot of people do! Tone correction [1] is a normal fact of life in the music industry, especially in recordings. Using it well takes both some degree of vocal skill and production skill. You'll often find that it's incredibly obvious when done poorly, but nearly unnoticeable when done well.
[1] AutoTune is a specific brand
Still, I find it interesting. If you can't synthetically alter someone's performance to be "worse", is it OK that the NFL synthetically altered Alicia Key's performance to be "better"?
For a more consequential example, imagine Biden's marketing team "cleaning up" his speech after he has mumbled or trailed off a word, misleading the US public during an election year. Should that be disclosed?
seems to comply?
If they don't reject it for that, nothing changes.
This sounds like every thumbnail on youtube these days. It's good that this is not limited to AI, but it also means this will be a nightmare to police.
People using VFX aren't trying to create images in likeness of another existing person to get people to buy crypto or other scams. Comparing the two is disingenuous at best.
The ease and lack of skill required. That brings whole another set of implications.
This is like regulating handguns differently from compound bows. Both are lethal weapons, but the bow requires hours of training to use effectively, and is more difficult to carry discreetly. The combination of ease, convenience, and accessibility necessitates new regulation.
This being said, AI for video is an incredibly promising technology, and I look forward to watching the TV shows and movies generated with AI-powered tooling.
This really depends on what you're doing. There are some great Cinema 4d plugins out there. As the plethora of YouTube tutorials out there clearly demonstrate, multiple professionals, and vast experience, are not required for some of the things they have listed. Tooling and assets costs are 0, in the high seas.
Until Sora is widely available, or the open source models catch up, at this moment it's easier to use something like Cinema 4d than AI.
Do we make all usages of VFX now require a warning, just in case the VFX was generated by AI?
I think this is different to the bow v gun metaphor as I can tell an arrow from a bullet, but I can foresee a future where no human could tell the difference between AI-assisted and non-AI-assisted VFX / art
I believe this is evidenced by the fact that people can go around accusing any art piece of being AI art and the burden of proving them wrong falls on the artist. Essentially I believe we are rapidly approaching the point of it not mattering if someone uses AI in their art because people won't be able to tell anyway
Could a similar argument be applied here? It doesn’t seem like there is much in the way of consequences for lying to Google. But I suppose they have other ways of checking for it, and catching someone lying is a signal that makes the account more suspicious.
The other thing is that they don’t necessarily want to ban all of this content. For example a video demonstrating how AI can be used to create misinformation and showing examples, would be fairly clearly “morally” ok. The policy being that you have to declare it allows for this sort of content to live on the platform, but allows you to filter it out in certain contexts where it may be inappropriate (searches for election coverage?) and allows you to badge it for users (like Covid information tags).
What about the picture you see before clicking on the actual video? This article of course is addressing the content of the videos, but I can't help but look at the comically cartoonish, overly dramatic -- clickbait -- picture preview of the video.
For example, there is a video about a tornado that passed close to a content author and the author posts video captured by their phone. In the preview image, you see the author "literally getting sucked into a tornado". Is that "altered and synthetic content"?
The thumbnail isn't the content itself necessarily.
For context, Pixiv had to deal with a massive wave of AI content being dumped onto the site by wannabe artists basically right as the initial diffusion models became accessible. They responded by making 'AI-generated' a checkbox to go with the options to mark NSFW and adding an option for users to disable AI-generated content from being recommended to them. Then, after an incident of someone using their Patreon style service to pretend to be a popular artist, selling commissions generated by AI to copy the artist's style, they banned AI-generated content from being offered through that service.
My guess is that youtube is going to downrank this content, and may be trying to crowdsource training data in order to do this automatically.
Personally, I've developed a strong aversion to content that is primarily done by AI with very little human effort on top. After how things went with Pixiv I've come to hold the belief that our societies don't help people develop 'cultural maturity'. People want the clout/respect of being a popular artist/creator, without having to go through the journey they all go through which leads to them becoming popular. It's like wanting to use the title of Doctor without putting in the effort to earn a doctorate, the difference just being that we do have a culture of thinking that it's bad to do that.
I think that this policy is not perfect, but it is a step in the right direction.
I’m thinking something as simple as a digital signature that certifies e.g. a photo was made with my phone if I want to prove it, or if someone edits my file there should be a way of keeping track of the chain of trust.
That would allow a much strong verifiability of media. But I'm not sure if that would be possible...
https://blog.google/products/photos/google-photos-features-p... https://blog.google/products/photos/google-photos-magic-edit...
Present day Google is too busy selling AI shovels to quell Wall St's grumbling, to even consider what AI video will to do to the already bad 'needle in a haystack' nature of search.
Note that this will become less useful over time, however, as Google prunes old results from its index.
It's always been possible to write fake news. We've never had to add disclaimers at the top of textual content, e.g. "This text is made to sound like it describes real events, but contains invented and/or inaccurate facts." We feel the need to add this to video because until now, if it looked real, it probably was (of course, "creative" editing has existed for a long time, but that's still comparatively easy to spot).
It's the end of a media era, really.
I mean, yes we have? https://en.wikipedia.org/wiki/All_persons_fictitious_disclai...
> Otherwise please use the original title, unless it is misleading or linkbait; don't editorialize.
I am generally very skeptical of these tags though, I suspect a lot of them are in place to stop an AI consuming its own output rather than any concern for the end user.
God, I wish I could beam this sentence directly into the brain of every single person breathlessly excited about using gen AI to be "a creative".
Once something like Sora is available to the public, its going to be game over. A new bunch of creators will use it to "create" videos and I am sure you will change your mind then.
Typical, people who don't want their feeds filled with AI generated pigswill are to be defeated, not persuaded.
Are they capable of enforcing it? I don't know, but it's clear users understand / don't like the idea of being awash in a sea of AI content at this point.
If they can actually avoid it remains to be seen.
This may spoil the fun in some 3D rendered scenes. For example, I remember there was much discussion on whether a robot throwing a bowling ball was real or not[1].
Part of the problem has to do with all the original tags (e.g. "#rendering3d") being lost when the video spread through various platforms. The same problem will happen with Youtube -- creators may disclose everything, but after a few rounds through reddit and back, whatever disclosure and credit that was in the original video will be lost.
Nothing of use here. As per the usual MO of tech companies they throw the responsibility back on the user. Sounds like yet another bullshit clause that they can invoke when they want to cancel you.
once something sora 1.5 level of ability is there – definitely reverse-sora model which can recognise ai-made videos should be possible to train as well
Pretty much. If Google says "Swiper no swiping" they can point at their policy when lobbying against regulations or pushing back against criticism.
Before surveillance capitalism became the norm, web services told users to not share personal information, and to not trust other users they had not met in real life.
Look at "realistic" photos , it is easy for someone with experience to spot issues, the hangs/fingers are wrong, shadows and light are wrong, hair is weird, eyes have issues. In a video there are much more information so much more places to get things wrong, making it pass this kind of test will be a huge job so many will not put the effort.
Everyone calls this a problem, but it's a predicament because it has NO solution, and I have nothing but contempt for everyone who made it reality.
Creating videos takes quite a bit of time. If AI video generation becomes widely available, pretty soon, there could be more AI content being uploaded to YouTube than human-made stuff.
Presumably, training on AI generated stuff magnifies any artefacts/hallucinations present in the training set, reducing the quality of the model.
I think...
Interestingly it isn't just referring to AI but also "other tools", used to make content that is "altered, synthetic and seems real".
Fair amount of ambiguity in there but I see what they're getting at when it comes to the bigger fish like the president being altered to say something they didn't.
It’s obvious this tech will replace all vfx, lots of elements of a camera pipeline and eventually even video game rendering.
The labelling will become to common it will mean nothing.
Fact checking box like on twitter would be better and if you can't provide it, don't pretend you know anything about the content.
I have a whole lot of shorts content to report..
=)
So this announcement begs the question: is there any way to search for just AI content?
Namely it requires payment under false pretense
Who manufactured the problem here? It isn't just that these systems exist solely for individuals to be selected for by vulnerability, but that the people best poised to protect them are so effectively excluded from any sign of it.
What would it look like if we could see every ad meant for every audience from the same companies? What if we could just look up who was trying to target us? Is responsible, auditable advertisement impossible, or just not able to print money?
Why movies are not labeled but AI video must be labeled? What about comedians impersonating politicians?
If Google or govt is afraid that someone will use AI-generated videos for bad purposes (e.g. to display a candidate saying things that he never said) then they should display a warning above every video to educate people. And popular messengers like Telegram or video players must do the same.
At least add a warning above every political and news videos.
This label is worthless IMO. We're close or already at a point where it's impossible to distinguish between real and AI generated content. Any company offering "AI detection" are scammers. It's just plain not possible.
Forget individual videos for a second and look at youtube-the-experience as a whole. The recommendation stream is the single most important "generative AI" going on ever, using the sense of authenticity, curiosity and salience that comes from the individual videos themselves, but stitching them together in a very particular way. All the while the experience of being recommended videos being almost completely invisible. Of course this is psychologically "satisfying" to the users - in the shortest term - because they keep coming back, to the point of addiction. (Especially as features like shorts creep in).
Allowing the well of "interesting, warm, authentic audio & videos having the secondary gains of working on your psychological needs" being tainted with the question of generated content is a game changer because it breaks the wall of authenticity for the entire app. It brings the whole youtube-the-experience into question, it reduces its psychological stand-in function for human voice & likeness, band-aiding the hyper-individualized lonely person's suffering based content consumption habits. I know this is a bit dramatic, and for sure videos can be genuinely informative, but let's be honest, neither that is the entirety of your stream, nor that is the experience for the vast majority of the users. It will get worse as long as there is a mathematical headroom of making more money out of making it worse, that's what the shareholder duty is about.
When gen-AI came about I was naively happy about the fake "authenticity" wall of the recommended streams breaking down thanks to the garbage of generated sophistry overtaking and grossing out the users. Kind of like super delicious looking cakes turning out to be made of kitchen sponges turning people off of cakes all together. I was wrong to think AI oligopoly would let the opportunity of having a chokehold on the entire "content" business, and here we are. (Also this voluntary tagging will give them the perfect live training set, on top of what they have.)
Once the tech is good enough to generate video streams on the fly, so that all you need is a single livestream, that you won't even have a recommendation engine of videos and instead a team of virtual personas doing everything you could ever desire on screen, it is game over. It might already be game over.
To get out of this the single most important legislative maneuver is being able to accept and enforce the facts that a) recommendation is speech b) recommendation is also gen-AI, and should be subject to same level of regulatory scrutiny. I don't care if it generates pixels or characters at a time, or slaps together the most "interesting" subset of videos/posts/users/reels/shorts out of the vast sea of the collective content-consciousness, they are just one level of abstraction apart but functionally one and the same: look at me; look at my ads; come back to me; keep looking at me.
* Unless you're a powerful state actor then your videos are always 'real'.
The last word on this subject was not written in the 1920s, it's good to revisit old assumptions every century or so, when new forms of media and media manipulation become developed.
The first pass on it is unlikely to be the best, or even the last one.
And just like a prototype that would never end up in production, we'll remain with the first implementation we could think of cough copyright cough
Here's an obvious question that came up (and was resolved differently in different jurisdictions) - can photographs be copyrighted? What about photographs made in public? Of a market street? Of the Eifel tower? Of street art? Can an artist forbid photography of their art? An actor of their performance? A celebrity of their likeness? A private individual of their face? Does the purpose for which the photograph will be used matter?
At what point does a photograph have sufficient creative input to be copyrightable? Is pressing a button on a camera creative input? What about a machine that presses that button? Only humans can create copyrightable works under most jurisdictions. Is arranging the scene to be photographed a creative input? Can I arrange a scene just like yours and take a photo of it? Am I violating your copyright by doing it?
There's tens of thousands of pages of law and legal precedent that answer that question. As a conversation, it went on for decades, with no simple first-version solution sticking.
society in this case = media companies
You have to speak to get heard.
I see YouTube's own guidelines in the article and they seem reasonable. But I think over time the line will move, be unclear and we'll end up like prop 65 anyways.
To be effective, warnings like this have to be MANDATED on the item in question, and FORBIDDEN when not present.
Otherwise you stick a prop 65 "may contain" warning on everything, and it's pointless.
(This post may have been generated by AI; this notice in compliance with AI notification complications.)
> This post may have been generated by AI
I doubt "may" is enough.
A more plausible scenario would be if you aren't sure if all your stock footage is real. Though with youtube creators being one of the biggest groups of customers for stock footage I expect most providers will put very clear labeling in place.
Does blurring part of the image with Photoshop count? What if Photoshop used AI behind the scene for whatever filter you applied? What about some video editor feature that helps with audio/video synchronization or background removal?
As opposed to today, where companies are doing everything they can, stretching the truth, just so they can market their tools as “Using AI.”
If you edit this image by hand you’re good, but if you use a tool that “uses AI” to do it, you need to put the scare label on. Even if pixel-for-pixel both methods output the identical image! Just as a GMO/not GMO has no correlation to harmful compounds being in the food, and artificial flavors are generally more pure than those extracted from some wacky and more expensive means from a “natural” item.
It sounds like the idea is to normalize the use of such an attribution trail in the media industry, so that eventually audiences could start to be suspicious of images lacking attribution.
Adobe in particular seems to be interested in making GenAI-enabled features of its tools automatically apply a Content Credential indicating their use, and in making it easier to keep the content attribution metadata than to strip it out.
At the end of the day it's going to be drenched in contracts and obscure proofs of trust - i.e. some signing cert you can attach to an image if it was generated on an entirely controlled environment that prohibits known AI generation techniques - that technical side is going to be an arms race and I don't know if we can win it (which may just result in small creators being bullied out of the market)... but above the technical level I think we've already got all the tools we need.
1. These two examples are entirely fabricated
I think for it to be effective you'd have to require them to provide an itemized list of WHAT is AI generated. Otherwise what if a content creator has a GenAI logo or feature that's in every video and put a lazy disclaimer.
> (This post may have been generated by AI; this notice in compliance with AI notification complications.)
:D
It's very possible that Prop 65 has motivated some businesses to avoid using toxic chemicals, but it doesn't often help individuals make effective health decisions.
https://www.corporatecomplianceinsights.com/california-warni...
It’s not perfect but it has had a positive effect https://99percentinvisible.org/episode/warning-this-podcast-...
As a non-Californian I’m used to them from the little stickers on seemingly every electronics cable that comes with something I buy.
But from listening to that episode when it came out it sounds like it really has helped a lot, even if it’s also become kind of obnoxious.
seemingly every electronics cable
If it's something you've bought recently the offending ingredient should be listed. Otherwise, my money would be on lead being used as a plasticizer. Either way at least you have the tools to find out now.Like is it one of those things the remove a 1 in a billion chance of cancer, and now have a product that wears out twice as fast leading to a doubling of sales?
First time I was in CA, my then-partner's mother saw a Prop 65 notice and asked why they couldn't just ban the substances.
We were in a restaurant that served alcohol, one of the known substances is… alcoholic beverages.
https://en.wikipedia.org/wiki/California_Proposition_65_list...
Banning that didn't work out so well the last time.
898 Bowdoin St https://maps.app.goo.gl/uHTTd7yYtAibAg1QA
Some of the street view passes the sign is washed out. Click through to different times to see the sign.
[1] https://www.npr.org/sections/health-shots/2023/08/30/1196640...
Lol, this sounds like one of those fabels where an idiot king bans all allergens then a week later everyone is starving to death in the kingdom because it turns out that in a large enough population there will be enough different allergies that everything gets banned.
That already happens for foods.
The solution for suppliers is to intentionally add small quantities of allergens (sesame). [1] By having that as an actual ingredient, manufacturers don't have to worry about whether or not there is cross contamination while processing.
[1] https://www.medpagetoday.com/allergyimmunology/allergy/10652...
Eventually, it may be completely indiscernible, but we aren’t there yet
There may be a few tells still, but those won't last long, and the moment someone can find a new pattern you can make that a negative prompt for new images to avoid repeating the same mistake.
I think we are already there, and it seems like we aren't because many people are using free low-quality models with a low number of steps because its more accessible.
For any physical build there are typ "TITLE 25" such disclosurs that are required for any new-build plans...
Maybe we have TITLE N as designed by AI discolsures that will be needed...
It's marketing-speak and corporate buzzwords to cover for the fact that their LLMs often produced wrong information because they aren't capable of understanding your request, nuance, or the training data it used is wrong, or the model just plain sucks.
Would we tolerate such doublespeak it were anything else? "Well, you ordered a side of fries with your burger but because our wait staff made a mistake...sorry, hallucinated, they brought you a peanut butter sandwich that's growing mold instead."
It gets more concerning when the stakes are raised. When LLMs (inevitably) start getting used in more important contexts, like healthcare. "I know your file says you're allergic to penicillin and you repeated when talking to our ai-doctor but it hallucinated that you weren't."
The reason models hallucinate is because we train them to produce linguistically plausible output, which usually overlaps well with factually correct output (because it wouldn't be plausible to say e.g. "Barack Obama is white"). But when there isn't much data to show that something that is totally made up is implausible then there's no penalty to the model for it.
It's nothing to do with not being able to understand your request, and it's rarely because the training data is wrong.
it translates to "Creates text which contains incorrect or invalid information"
The latter just doesn't sound as good in headlines/articles/tutorials (eg. marketing material).
The software is working as designed, statistics are just imperfect
It doesn't feel right to me either, to use it in the context of generative AI, and I'd support renaming this behaviour in GenAI (text and images both) — though myself I'd call this behaviour "mis-remembering".
Edit: apparently some have suggested "delusion". That also works for me.
Also those two statements are not mutually exclusive.
Errors in statistical models being called hallucinations in the past does not mean that term is not marketing speak for what I said earlier.
https://www.youtube.com/watch?v=wRDfzjxzj3M
> Also those two statements are not mutually exclusive.
> Errors in statistical models being called hallucinations in the past does not mean that term is not marketing speak for what I said earlier.
The implicit claim was that they call this hallucination because it sounds better. In other words that some marketing people thought "what's a nicer word for 'mistakes'?" That is categorically untrue.
I don't think there's any point arguing about whether or not the marketers like the use of the word "hallucinate" because neither of us has any evidence either way. Though I was also say the null hypothesis is that they're just using the standard word for it. So the onus is on you to provide some evidence that marketers came in an said "guys, make sure you say 'hallucinate'". Which I'm 99% sure has never happened.
Webster definition: "a sensory perception (such as a visual image or a sound) that occurs in the absence of an actual external stimulus and usually arises from neurological disturbance (such as that associated with delirium tremens, schizophrenia, Parkinson's disease, or narcolepsy) or in response to drugs (such as LSD or phencyclidine)".
I would fire with prejudice any marketing department that associated our product with "delirium tremens, schizophrenia, [...] LSD or phencyclidine".
Yes: identity theft. My identity wasn't "stolen", what really happened was a company gave a bad loan.
But calling it identity theft shifts the blame. Now it's my job to keep my data "safe", not their job to make sure they're giving the right person the loan.
Calling it "hallucination" implies that there are (other) moments when it is understanding the world correctly -- and that itself is not true. At those moments, it is a word generator that is generating words that DO make sense.
At no point is this a conciousness, and anthropomorphizing it gives the impression that it is one.
There really is no correct word to describe what's happening, because LLMs are effectively philosophical zombies. We have no metaphors for an entity that can appear to hold a coherent conversation, do useful work and respond to commands but not think. All we have is metaphors from human behavior which presume the connection between language and intellect, because that's all we know. Unfortunately we also have nearly a century of pop culture telling us "AI" is like Data from Star Trek, perfectly logical, superintelligent and always correct.
And "hallucination" is good enough. It gets the point across, that these things can't be trusted. "Confabulation" would be better, but fewer people know it, and it's more important to communicate the untrustworthy nature of LLMs to the masses than it is to be technically precise.
If the output is incorrect, that's error. It may not be a bug, but it is still error.
It's a language model, trained on syntactically correct code, with a data set which presumably contains more correct examples of code than not, so it isn't surprising that it can generate syntactically correct code, or even code which correlates to valid solutions.
But if it actually had insight and knowledge about the code it generated, it would never generate random, useless (but syntactically correct) code, nor would it copy code verbatim, including comments and license text.
It's a hell of a trick, but a trick is what it is. The fact that you can adjust the randomness in a query should give it away. It's de rigueur around here to equate everything a human does with everything an LLM does, including mistakes, but human programmers don't make mistakes the way LLMs do, and human programmers don't come with temperature sliders.
The fact that it instead generates syntactically correct code that, more often than not, solves - or at least tries to solve - the problem that is posited, indicates that there is a "there" there, however much one talks about stochastic parrots and such.
As for temperature sliders for humans, that's what drugs are in many ways.
To a degree, people do expect the output to be correct. But in my view, that's orthogonal to the use of the term "error" in this sense.
If an LLM says something that's not true, that's an erroneous statement. Whether or not the LLM is intended or expected to produce accurate output isn't relevant to that at all. It's in error nonetheless, and calling it that rather than "hallucination" is much more accurate.
After all, when people say things that are in error, we don't say they're "hallucinating". We say they're wrong.
> It generates syntactically correct language, and that's all it does.
Yes indeed. I think where we're misunderstanding each other is that I'm not talking about whether or not the LLM is functioning correctly (that's why I wouldn't call it a "bug"), I'm talking about whether or not factual statements it produces are correct.
From there warnings proliferated on so many more products, but getting told that chocolate bars can cause cancer is still a reasonable tradeoff. Especially as nothing is stopping the law from getting tweaked from there.
Comparing it to prop 65 or GDPR makes it look like a probably deeply effective, yes slightly annoying rule...I sure hope that's what we end up with.
A lot of early stable diffusion seemed "realistic" but comparing them to newer stuff makes them stand out at obviously AI generated and unrealistic.
Which would be OK with me, personally. Right now, those cookie banners do serve a valuable function for me -- when I see them, I know to treat the site with caution and skepticism. If AI warnings end up similar, they too will serve a similar purpose. It's all better than nothing.
The ePrivacy directive and GDPR don't literally require cookie banners but the former requires disclosure of specific information and the latter requires consent for most forms of data collection and processing. Even the 2002 directive actually require an option to refuse cookies which many cookie banners still fail to implement properly post-GDPR.
The problem is that most websites want to start collecting, tracking and processing data that requires consent before any interaction takes place that would allow for a contextual opt-in. This means they have to get that consent somehow and the "cookie banner" or consent dialog serves that purpose.
Of course many (especially American) implementations get this hilariously wrong by a) collecting and processing data even before consent is established, b) not making opt-out as trivial as opt-in despite the ePrivacy directive explicitly requiring this (e.g. hiding "refuse" behind a "more info" button or not giving it the same weight as "accept all"), c) not actually specifying the details on what data is collected etc to the level required by the directive, d) not providing any way to revise/change the selections (especially withdrawing consent previously given) and e) trying to trick users with a manual opt-out checkbox per advertiser/service labeled "legitimate interest" which is an alternative to consent and thus is not something you can opt out of because it does not require consent (but of course in these cases the use never actually qualifies as "legitimate interest" to begin with and the opt-out is a poorly constructed CYA).
In a different world, consent dialogs could work entirely like mobile app permissions: if you haven't given consent for something you'll be prompted when it becomes relevant. But apparently most sites bank on users pressing "accept all" to get rid of the annoying banner - although of course legally they probably don't even have data to determine if this gamble works for them because most analytics requires consent (i.e. your analytics will show a near 100% acceptance rate because you only see the data of users who opted into analytics and they likely just pressed "accept all").
WARNING: This video contains content known to the State of Google to be generated by AI algorithms and/or tools.
Ok, beauty face filters are not included. How about character motion animations? How detailed does the after effects plugin need to be before it's considered AI? Can we generate just a background? Just a minor subject in the foreground? Or is it like pornography, where we'll recognize it once we see it?
I fear AI tools will soon become so embedded in normal workflows that it's going to become a question of "how much" not "contains", and "how much" is such a blurry, subjective line that it's going to make any binary disclaimer meaningless.
EDIT: I think these should also include whatever built-in processing is applied to the raw sensor data within the camera itself.
[1] https://helpx.adobe.com/creative-cloud/help/content-credenti...
That future will come, and it will come sooner than anyone's expecting.
Yet all I see is society trying to prevent the inevitable from installing itself (because it's "scary", "dangerous", "undermines the very pillars of society" etc.), instead of preparing itself for when the inevitable occurs.
People seem to have finally accepted we can't put the genie back in the bottle, so now we're at the stage where governments and institutions are all trying to look busy and pass the image of "hey, we're doing something about it, ok? You can feel safe".
Soon we will be forced to accept that all that wasted effort was but a futile attempt at catching a falling knife.
Maybe the next idiom in line will be "crying over spilled milk", because could someone point me to what is being done in terms of "hey, let's start by directly assuming a world in which anyone can produce unrestricted, genuine-looking content will soon come and there's no way around it -- what then?"
All I see is a meteor approaching and everyone trying to divert it, but no one actually preparing for when it does hit. Each day that passes I'm more certain that we will we look at each other like fools, asking ourselves "why didn't we focus on preparing for change, instead of trying to prevent change"?
"Accept it, it hurts less".
I'm not saying it makes the actual situation any better; it obviously doesn't. But anyone can feel the rarefied AI panic in the air growing thicker by the minute, and panic will only make the situation worse both before and after absolute change takes place.
When we don't accept incoming change before it arrives, we surely are forced to accept it after it arrives, at a much higher price.
You asked about preparations: prepare yourself to see governments try (and fail) to regulate what processing power can be acquired by consumers. Prepare yourself for the serious proposal of "truth-checking agencies" with certified signatures that ensure "this content had its chain of custody verified as authentic from its CMOS capture up to its encoded video stream", in which a lot of time and effort will be wasted (there's already people replying about this, saying metadata and/or encryption will come to the rescue via private/public keys. Supposedly no will would ever film a screen!).
The above might seem an exaggeration, but ask yourself: the YouTube guidelines this post is about, the recent EU regulation... do you think those are enough? Of course they're not. They will keep trying to solve the problem from the wrong end until they are (we are) forced to accept there's nothing that can be done about it, and that it is us who need to adapt to live in such a world.
Enjoy the ride, I suppose.
These are all the proper preparation for AI. AI can't generate a private key given a public key. AI can't generate the appropriate text given a hash.
So we build a society upon these things AI can't do.
It has been a good run. We have done things like the tried and true ink stamping to verify documents. We have a labyrinth of bureaucracy for every little activity, mostly because it is the way that has always worked. It has surely been nice for the "administration" to sit around and sip lemonade in their archaic jobs. It has been nice to have incompetent people with no vision being appointed to high places for being born into the right families connected with the right people. That gravy train was surely a joy for those who were a part of it.
Sadly, it won't work anymore. We will need competent people now that actually care.
We need everything to be authenticated now with digital signatures.
It is not even that difficult a problem to solve. The existing systems are far more complex, far more prone to error, far more expensive, and far more difficult to navigate.
AI is giving us an opportunity to evolve. It is a time for celebration. Society will be faster, more efficient, more secure, and much more fun with generative content. AIs will produce official AI-signed content, and unsigned content. Humans will produce official human signed content, and unsigned content. Some AIs will use humans to sign content to subvert systems. But all of this pales in comparison to the fraud, waste, and total abuse of the current system.
These humans would simply generate a public key from the private key, then post it under their human identity. The main threat from AI in the future IMO is not rouge AI, but bad human actors using it for their own nefarious agendas. This is how the first "evil" AI will probably come about.
Personal thoughts:
- we're already a year past the point where it was widely known you can generate whatever you want, and get it to a reasonable "real" threshold with less than a day worth of work.
- the impact is likely to be significantly muted, rather than an exponential increase upon, a 2020 baseline. professionals were capable of accomplishing this with a couple orders of magnitude more manual work for at least a decade.
- in general, we've suffered more societally from histrionics/over-reactions to being bombarded with the same messaging
- it thus should end up being _net good_, in that a skeptic has a 100% accurate argument for requiring more explanation than "wow look at this!"
- I expect that being able to justify / source / explain things will gain significant value relative to scaled up distributors giving out media someone else gave them without any review.
- something I've noticed the last couple years is people __hate__ looking stupid. __Hate__. They learn extremely quickly to refactor knowledge they think they have once confronted in public, even by the outgroup, as long as theyre a non-extremist.
After writing that out, I guess my tl;Dr as of this moment and mood, is there will be negligible negative effects, we already reached a nadir of unquestioned BS sometime between 2010 and 2024, and a baseline be _anyone_ can easily BS will lead to wide acceptance of skeptical reactions, even within ingroups.
God I hope I'm right.
Thank you for the preface you wrote, I completely understand your point of how easy it is to sound like a contrarian online, I'm sure my writing style doesn't help much on that front I'm afraid to admit.
I don't know, we've done a pretty good job at preventing nuclear war so far. We didn't just say "oh well, the genie is out of the bottle now. Everyone will have nuclear weapons soon and there's nothing we can do about it. All wars from now on are going to be nuclear. Might as well start preparing for nuclear winter." We signed treaties and made laws and used force to prevent us all from killing each other.
Should we like deepfakes any better if they're created by a nation state using pre-AI Hollywood production technology, or "by hand" with Photoshop etc.? If 3D printers get better so that anybody can 3D print masks you can wear to convincingly look like someone else and then record yourself on camera, would you expect a different set of rules for that or are we talking about the same kind of problem?
Trying to characterize this as not related to AI just isn't adding to the discussion. Clearly it is a response to the emergence of AI fakes.
[1] And all the other stuff you list
YouTube did this because the EU passed a law about it. The EU passed a law about it because of the moral panic, not because the abstract concept of deception was only recently invented.
It's like having cars already, and speed limits, and then someone invents an electric car that can accelerate from 0 to 200 MPH in 5 seconds, so the government passes a new law with some arbitrary registration requirements to satisfy Something Must Be Done.
Finding out a video is maliciously fabricated is a different problem.
Of course there was. The community guidelines have prohibited impersonation and misinformation for years.
That sounds like a quagmire of subjectivity to enforce. You can argue whether generated content was created to mislead or whether it would cause _egregious_ harm.
Now there's no more arguing. Is it generated or is it real - and is it marked if generated? I still fail to see the downside here.
> Now there's no more arguing. Is it generated or is it real - and is it marked if generated? I still fail to see the downside here.
How is this supposed to lead to less arguing? If there was an easy way to tell if something is AI-generated then you wouldn't need the user to tag it. When there isn't, now you have to argue about whether it is or not -- or if it obviously is, whether it then has to be tagged, because it obviously is and they've given that as an exception.
I'm not getting into the rabbit hole of finding out if it was actually generated - once again, that's a different problem. One that I already mentioned 2 comments back. My point was that there is less subjectivity with this rule. If the content is found to have been generated and isn't marked, then there are clear grounds to remove the video.
What youtube does when there is doubt is not known yet. I don't deal in "well this _could_ lead to this".
They use (with some irony) AI and other algorithms to determine if you're breaking the rules, often leading to arbitrary or nonsensical results, which the review process frequently fails to address.
> I'm not getting into the rabbit hole of finding out if it was actually generated - once again, that's a different problem.
It isn't a different problem, it's the problem. You want people to label things because otherwise you're not sure, but because of that problem exactly, you have no way of reliably or objectively enforcing the labeling requirement. And you specifically have no way to do it in the cases where it most matters because it's hard to tell.
At a mom and pop you could at least talk to a person and figure out what happen.
>I don't deal in "well this _could_ lead to this"
Did you not learn from the entire DMCA thing? Remember the thing where piles tech people warned "Wow, this is going to be used as a weapon to cause problems" and then it was used as a weapon to cause problems.
Well, welcome to the next weapon that is going to be used to cause problems.
There's a handful of cases where yt actually messed up w/ DMCA and considering the sheer volume of videos they process, I'd say it's actually pretty damn good.
So no, DMCA is not a valid reason to assume youtube will handle this improperly.
The alternative to the DMCA is Section 230 of the CDA, which could have just as easily been applied to copyright as it is to anything else if it weren't for the DMCA providing a more abuse-prone alternative.
> And people on the internet don't know what fair-use actually means, so they complain/exaggerate about DMCA takedowns when, surprise, it wasn't actually covered by fair-use.
Abusive takedowns are a huge problem, actually. It's common in cases of businesses sending takedowns for their competitors' websites or videos, for example. They're often completely fraudulent with no merit whatsoever, but the company receiving the takedown has no information on which to base a decision (who created this content? how would they know?), so they just mechanistically execute all of them with no validation.
The ban recourse problem is the opposite.
This is the "keep recourse": "this video is obviously bad, but google doesn't feel like taking it down, and there is nothing I can do about it". Now there is, and it can actually go to a court with a judge in the end, if Google is obstinate.
You didn't have a right to be hosted on Google before, and you don't have now. Of course they can ban you as they like. The thing is, they can't host you as they like, if you're breaking this rule.
Except that the rule can be satisfied just by labeling it, and if there are penalties for not labeling but no penalties for labeling then the obvious incentive is to stick the label on everything just in case, causing it to become meaningless.
To prevent that would require prohibiting the label from being applied to things that aren't AI-generated, which is impracticable because now you need 100% accuracy and there is no way to err on the side of caution, but nobody has 100% accuracy. So then the solution would be to actually make everything AI-generated, e.g. by systematically running it through some subtle AI filter, and then you can get back to labeling everything to avoid liability.
Did the California proposition 65 really result in cancer labels on everything? Or is it just hyperbole? I suppose having a lot of labels is still bad, even if they're not technically on everything.
Who's there
dmcAI takedown notice!
Really, your argument can be generalized to 'why have laws at all, because people will break them and lie about it'.
It isn't a rule against lying, it's a rule requiring lies to be labeled. From which you get nothing useful that you couldn't get from a rule against lying, because you'd need the same proof for either one.
Meanwhile it becomes a trap for the unwary because innocent people who don't understand the complicated labeling rules get stomped by the system without intending any malice.
> Really, your argument can be generalized to 'why have laws at all, because people will break them and lie about it'.
The generalization is that laws against not disclosing crimes are pointless because the penalty for the crime is already at least as severe as the penalty for not disclosing it and you'd need to prove the crime to prove the omission. This is, for example, why it makes sense to have a right against self-incrimination.
This divide is obviously going to play out on two sides.
Proving authenticity may turn out to be as difficult as proving fakeness. People will use this maliciously to flag and censor content they dislike.
RIAA has entered the chat.
"It defrauds us of our hard-earned middle-man cut."Let’s see how long it will take them to collect enough data and train a model to distinguish AI-generated from user-generated videos.
> Generating realistic scenes: Showing a realistic depiction of fictional major events, like a tornado moving toward a real town.
Does the Wizard of Oz tornado scene need a warning now? [0] (Of course not, but it may be hard to draw the line in some places.)
[0] https://www.grunge.com/486387/heres-how-the-tornado-scene-in...
> It can be misleading if viewers think a video is real, when it's actually been meaningfully altered or synthetically generated to seem realistic.
> When content is undisclosed, in some cases YouTube may take action to reduce the risk of harm to viewers by proactively applying a label that creators will not have the option to remove. Additionally, creators who consistently choose not to disclose this information may be subject to penalties from YouTube, including removal of content or suspension from the YouTube Partner Program.
https://support.google.com/youtube/answer/14328491
This is censorship. That's all.