I see YouTube's own guidelines in the article and they seem reasonable. But I think over time the line will move, be unclear and we'll end up like prop 65 anyways.
I see YouTube's own guidelines in the article and they seem reasonable. But I think over time the line will move, be unclear and we'll end up like prop 65 anyways.
To be effective, warnings like this have to be MANDATED on the item in question, and FORBIDDEN when not present.
Otherwise you stick a prop 65 "may contain" warning on everything, and it's pointless.
(This post may have been generated by AI; this notice in compliance with AI notification complications.)
> This post may have been generated by AI
I doubt "may" is enough.
A more plausible scenario would be if you aren't sure if all your stock footage is real. Though with youtube creators being one of the biggest groups of customers for stock footage I expect most providers will put very clear labeling in place.
Does blurring part of the image with Photoshop count? What if Photoshop used AI behind the scene for whatever filter you applied? What about some video editor feature that helps with audio/video synchronization or background removal?
As opposed to today, where companies are doing everything they can, stretching the truth, just so they can market their tools as “Using AI.”
If you edit this image by hand you’re good, but if you use a tool that “uses AI” to do it, you need to put the scare label on. Even if pixel-for-pixel both methods output the identical image! Just as a GMO/not GMO has no correlation to harmful compounds being in the food, and artificial flavors are generally more pure than those extracted from some wacky and more expensive means from a “natural” item.
It sounds like the idea is to normalize the use of such an attribution trail in the media industry, so that eventually audiences could start to be suspicious of images lacking attribution.
Adobe in particular seems to be interested in making GenAI-enabled features of its tools automatically apply a Content Credential indicating their use, and in making it easier to keep the content attribution metadata than to strip it out.
At the end of the day it's going to be drenched in contracts and obscure proofs of trust - i.e. some signing cert you can attach to an image if it was generated on an entirely controlled environment that prohibits known AI generation techniques - that technical side is going to be an arms race and I don't know if we can win it (which may just result in small creators being bullied out of the market)... but above the technical level I think we've already got all the tools we need.
1. These two examples are entirely fabricated
I think for it to be effective you'd have to require them to provide an itemized list of WHAT is AI generated. Otherwise what if a content creator has a GenAI logo or feature that's in every video and put a lazy disclaimer.
> (This post may have been generated by AI; this notice in compliance with AI notification complications.)
:D
It's very possible that Prop 65 has motivated some businesses to avoid using toxic chemicals, but it doesn't often help individuals make effective health decisions.
https://www.corporatecomplianceinsights.com/california-warni...
It’s not perfect but it has had a positive effect https://99percentinvisible.org/episode/warning-this-podcast-...
As a non-Californian I’m used to them from the little stickers on seemingly every electronics cable that comes with something I buy.
But from listening to that episode when it came out it sounds like it really has helped a lot, even if it’s also become kind of obnoxious.
seemingly every electronics cable
If it's something you've bought recently the offending ingredient should be listed. Otherwise, my money would be on lead being used as a plasticizer. Either way at least you have the tools to find out now.Like is it one of those things the remove a 1 in a billion chance of cancer, and now have a product that wears out twice as fast leading to a doubling of sales?
First time I was in CA, my then-partner's mother saw a Prop 65 notice and asked why they couldn't just ban the substances.
We were in a restaurant that served alcohol, one of the known substances is… alcoholic beverages.
https://en.wikipedia.org/wiki/California_Proposition_65_list...
Banning that didn't work out so well the last time.
898 Bowdoin St https://maps.app.goo.gl/uHTTd7yYtAibAg1QA
Some of the street view passes the sign is washed out. Click through to different times to see the sign.
[1] https://www.npr.org/sections/health-shots/2023/08/30/1196640...
Lol, this sounds like one of those fabels where an idiot king bans all allergens then a week later everyone is starving to death in the kingdom because it turns out that in a large enough population there will be enough different allergies that everything gets banned.
That already happens for foods.
The solution for suppliers is to intentionally add small quantities of allergens (sesame). [1] By having that as an actual ingredient, manufacturers don't have to worry about whether or not there is cross contamination while processing.
[1] https://www.medpagetoday.com/allergyimmunology/allergy/10652...
Which would be OK with me, personally. Right now, those cookie banners do serve a valuable function for me -- when I see them, I know to treat the site with caution and skepticism. If AI warnings end up similar, they too will serve a similar purpose. It's all better than nothing.
From there warnings proliferated on so many more products, but getting told that chocolate bars can cause cancer is still a reasonable tradeoff. Especially as nothing is stopping the law from getting tweaked from there.
Comparing it to prop 65 or GDPR makes it look like a probably deeply effective, yes slightly annoying rule...I sure hope that's what we end up with.
The ePrivacy directive and GDPR don't literally require cookie banners but the former requires disclosure of specific information and the latter requires consent for most forms of data collection and processing. Even the 2002 directive actually require an option to refuse cookies which many cookie banners still fail to implement properly post-GDPR.
The problem is that most websites want to start collecting, tracking and processing data that requires consent before any interaction takes place that would allow for a contextual opt-in. This means they have to get that consent somehow and the "cookie banner" or consent dialog serves that purpose.
Of course many (especially American) implementations get this hilariously wrong by a) collecting and processing data even before consent is established, b) not making opt-out as trivial as opt-in despite the ePrivacy directive explicitly requiring this (e.g. hiding "refuse" behind a "more info" button or not giving it the same weight as "accept all"), c) not actually specifying the details on what data is collected etc to the level required by the directive, d) not providing any way to revise/change the selections (especially withdrawing consent previously given) and e) trying to trick users with a manual opt-out checkbox per advertiser/service labeled "legitimate interest" which is an alternative to consent and thus is not something you can opt out of because it does not require consent (but of course in these cases the use never actually qualifies as "legitimate interest" to begin with and the opt-out is a poorly constructed CYA).
In a different world, consent dialogs could work entirely like mobile app permissions: if you haven't given consent for something you'll be prompted when it becomes relevant. But apparently most sites bank on users pressing "accept all" to get rid of the annoying banner - although of course legally they probably don't even have data to determine if this gamble works for them because most analytics requires consent (i.e. your analytics will show a near 100% acceptance rate because you only see the data of users who opted into analytics and they likely just pressed "accept all").
A lot of early stable diffusion seemed "realistic" but comparing them to newer stuff makes them stand out at obviously AI generated and unrealistic.
It's marketing-speak and corporate buzzwords to cover for the fact that their LLMs often produced wrong information because they aren't capable of understanding your request, nuance, or the training data it used is wrong, or the model just plain sucks.
Would we tolerate such doublespeak it were anything else? "Well, you ordered a side of fries with your burger but because our wait staff made a mistake...sorry, hallucinated, they brought you a peanut butter sandwich that's growing mold instead."
It gets more concerning when the stakes are raised. When LLMs (inevitably) start getting used in more important contexts, like healthcare. "I know your file says you're allergic to penicillin and you repeated when talking to our ai-doctor but it hallucinated that you weren't."
The reason models hallucinate is because we train them to produce linguistically plausible output, which usually overlaps well with factually correct output (because it wouldn't be plausible to say e.g. "Barack Obama is white"). But when there isn't much data to show that something that is totally made up is implausible then there's no penalty to the model for it.
It's nothing to do with not being able to understand your request, and it's rarely because the training data is wrong.
it translates to "Creates text which contains incorrect or invalid information"
The latter just doesn't sound as good in headlines/articles/tutorials (eg. marketing material).
The software is working as designed, statistics are just imperfect
It doesn't feel right to me either, to use it in the context of generative AI, and I'd support renaming this behaviour in GenAI (text and images both) — though myself I'd call this behaviour "mis-remembering".
Edit: apparently some have suggested "delusion". That also works for me.
Also those two statements are not mutually exclusive.
Errors in statistical models being called hallucinations in the past does not mean that term is not marketing speak for what I said earlier.
https://www.youtube.com/watch?v=wRDfzjxzj3M
> Also those two statements are not mutually exclusive.
> Errors in statistical models being called hallucinations in the past does not mean that term is not marketing speak for what I said earlier.
The implicit claim was that they call this hallucination because it sounds better. In other words that some marketing people thought "what's a nicer word for 'mistakes'?" That is categorically untrue.
I don't think there's any point arguing about whether or not the marketers like the use of the word "hallucinate" because neither of us has any evidence either way. Though I was also say the null hypothesis is that they're just using the standard word for it. So the onus is on you to provide some evidence that marketers came in an said "guys, make sure you say 'hallucinate'". Which I'm 99% sure has never happened.
Webster definition: "a sensory perception (such as a visual image or a sound) that occurs in the absence of an actual external stimulus and usually arises from neurological disturbance (such as that associated with delirium tremens, schizophrenia, Parkinson's disease, or narcolepsy) or in response to drugs (such as LSD or phencyclidine)".
I would fire with prejudice any marketing department that associated our product with "delirium tremens, schizophrenia, [...] LSD or phencyclidine".
Yes: identity theft. My identity wasn't "stolen", what really happened was a company gave a bad loan.
But calling it identity theft shifts the blame. Now it's my job to keep my data "safe", not their job to make sure they're giving the right person the loan.
Calling it "hallucination" implies that there are (other) moments when it is understanding the world correctly -- and that itself is not true. At those moments, it is a word generator that is generating words that DO make sense.
At no point is this a conciousness, and anthropomorphizing it gives the impression that it is one.
There really is no correct word to describe what's happening, because LLMs are effectively philosophical zombies. We have no metaphors for an entity that can appear to hold a coherent conversation, do useful work and respond to commands but not think. All we have is metaphors from human behavior which presume the connection between language and intellect, because that's all we know. Unfortunately we also have nearly a century of pop culture telling us "AI" is like Data from Star Trek, perfectly logical, superintelligent and always correct.
And "hallucination" is good enough. It gets the point across, that these things can't be trusted. "Confabulation" would be better, but fewer people know it, and it's more important to communicate the untrustworthy nature of LLMs to the masses than it is to be technically precise.
If the output is incorrect, that's error. It may not be a bug, but it is still error.
It's a language model, trained on syntactically correct code, with a data set which presumably contains more correct examples of code than not, so it isn't surprising that it can generate syntactically correct code, or even code which correlates to valid solutions.
But if it actually had insight and knowledge about the code it generated, it would never generate random, useless (but syntactically correct) code, nor would it copy code verbatim, including comments and license text.
It's a hell of a trick, but a trick is what it is. The fact that you can adjust the randomness in a query should give it away. It's de rigueur around here to equate everything a human does with everything an LLM does, including mistakes, but human programmers don't make mistakes the way LLMs do, and human programmers don't come with temperature sliders.
The fact that it instead generates syntactically correct code that, more often than not, solves - or at least tries to solve - the problem that is posited, indicates that there is a "there" there, however much one talks about stochastic parrots and such.
As for temperature sliders for humans, that's what drugs are in many ways.
To a degree, people do expect the output to be correct. But in my view, that's orthogonal to the use of the term "error" in this sense.
If an LLM says something that's not true, that's an erroneous statement. Whether or not the LLM is intended or expected to produce accurate output isn't relevant to that at all. It's in error nonetheless, and calling it that rather than "hallucination" is much more accurate.
After all, when people say things that are in error, we don't say they're "hallucinating". We say they're wrong.
> It generates syntactically correct language, and that's all it does.
Yes indeed. I think where we're misunderstanding each other is that I'm not talking about whether or not the LLM is functioning correctly (that's why I wouldn't call it a "bug"), I'm talking about whether or not factual statements it produces are correct.
Eventually, it may be completely indiscernible, but we aren’t there yet
There may be a few tells still, but those won't last long, and the moment someone can find a new pattern you can make that a negative prompt for new images to avoid repeating the same mistake.
I think we are already there, and it seems like we aren't because many people are using free low-quality models with a low number of steps because its more accessible.
For any physical build there are typ "TITLE 25" such disclosurs that are required for any new-build plans...
Maybe we have TITLE N as designed by AI discolsures that will be needed...