Cali's AG Tells AI Companies Almost Everything They're Doing Might Be Illegal
gizmodo.com
gizmodo.com
-Using AI to “foster or advance deception.”
-Falsely advertising “the accuracy, quality, or utility of AI systems.”
-Create or sell an AI system or product that has “an adverse or disproportionate impact on members of a protected class, or create, reinforce, or perpetuate discrimination or segregation of members of a protected class.“
Do not build an AI system that discovers the best way to target men to get a prostate exam, or the elderly to enroll in a program that may benefit them?
I understand "adverse", but disproportionate implies don't even try to help classes of people. AT ALL!
If you cannot explain why you have made such decisions, then the state will look at disproportionate impact. EG: You hired/lent to/rented to black people 25% less often than you did to white people. "Because the AI said so." doesn't cut it.
I don't like when laws are worded such that we told we shouldn't care what the words say, we all know what we meant, this is a law for getting the bad guys, so don't worry about the actual words we use. You shouldn't need a law degree to know what OR means.
“Adverse or disproportionate impact” is a well litigated phrase. It has a specific meaning in law, which is not immediately obvious from a layman’s definition.
For instance, in Texas motor vehicle code regarding driving under the influence, there is a definition for "Motor Vehicle." While colloquially a person would assume that such a vehicle should have a motor, the definition actually states:
> "Motor vehicle" means a device in, on, or by which a person or property is or may be transported or drawn on a highway, except a device used exclusively on stationary rails or tracks.
When it comes to laws, you have to read the entire law (including definitions) and not just rely on your own understanding of the terms within. Then it gets even more complicated when it comes to so called "Case Law." This is why companies have entire sections of lawyers to inform their managers on compliance.
A large portion of the interview touched statistical inference as it related to ML, specifically how it related to simple neural nets up to deep learning vs classical modeling. The answer I gave aligned with their expectations, which was that models used in lending should not be black box and should be able to quantify which features led to the prediction/output and how much weight they contributed. This was specifically done to address potential discrimination lawsuits.
I have a hard time believing any company that rents, lends, etc. would employ a black box for decisioning. Both private and public lawyers would sue them into oblivion immediately.
https://www.occ.treas.gov/publications-and-resources/publica...
https://www.theguardian.com/us-news/2025/jan/25/health-insur...
See also: "It's not a crime if you do it with an app" [/s] - https://pluralistic.net/2025/01/25/potatotrac/
This space is primarily dominated by AI that replaces mundane work done by humans.
The section of the advisory referencing disproportionate impact is quoting, nearly word for word, a portion of Cal. Code Regs. Tit. 2, § 14027.
So, this section of the advisory essentially amounts to the AG saying "using AI to do something illegal is still illegal".
That does not really answer your question, though.
> What kind of standard is disproportionate impact?...
The kind that does has been codified in CA and other jurisdictions' statutes for a long while now. This means that the standard is extremely well-litigated in the state's courts, and so the answer to "what is disproportionate impact?" is, I think, something like:
"That seems complicated; there's probably a rich case law that provides clarity in some situations but also highlights areas of ambiguity in other situations. If you're in it for profit, and have any questions, get a lawyer who specializes in that area of the law to review your specific circumstance; if you're in it for civics/curiosity, start with the statute then start reading significant case law or law reviews regarding that statute."
It's also the kind of standard that can attract flame wars... hopefully not here, though ;-)
I’m not excusing this design (it could probably be improved), but it may have been intentional.
See below.
> And I am a reasonably competent person on a computer.
Most people are not. It’s generally unwise to design an automated system that assumes computer/tech competence.
> For someone who isn’t maybe this is enough to make them turn away from whatever action they were attempting to do in frustration.
I imagine it’s the other way around. This type of system saves folks with less tech savvy from themselves.
I’m not sure if you’ve designed systems like this before. I have, and I was very surprised at what people thought was reasonable interaction and/or reasonable input.
Confirming choices, perhaps multiple times, before moving forward can save a lot of headache later for everyone involved. The collateral damage is guaranteed “wasted” time for everyone using the chatbot, but the company largely doesn’t care about that — that is, they are more than willing to pass on the cost of their money (e.g., if they hired a human agent) for your time.
- "Falsely advertising the accuracy, quality, or utility of AI systems” - worst case is probably Tesla's Fake Self Driving. Most other LLM systems aren't allowed to make decisions, just blither.
- "Create or sell an AI system or product that has an adverse or disproportionate impact on members of a protected class, or create, reinforce, or perpetuate discrimination or segregation of members of a protected class.“ - hm. Need more instances. Now, using computer systems to create a cartel to violate antitrust laws and push prices up is a thing. There's litigation against landlords for that. But that's not a "protected class" thing, it's an antitrust thing.
Which should be assumed to be legal already, even without the expressly written bill. Copyright maximalism is anti-human.
When I steal a dance you just invented, you're very butthurt about it and run crying to mommy "make him stop copying me!". Then you grow up and bribe Congress to make it illegal. Except for the "growing up" part, that never happened.
-- Cyberpunk Anatole France
____
If I were to steel-man your comment, it would be something like: "Scraping and training must be fair-use because people can be building all sorts of systems with ethical and valuable purposes. What you generate from a trained system can easily infringe, but that's a separate thing."
Also, where does the GNU Public License fall in terms of "anti-human copyright maximalization"? Is it bad because it uses fire, or is it good because it fights fire with fire?
It wouldn't be "fair use". It makes no copies. "Fair use" is the horseshit the courts dreamt up so they could pretend copyright wasn't broken when a copy absolutely needed to be made.
This makes no copies, so it doesn't even need "fair use". Instead, there are people who believe that because they made something long ago that they and their descendants into the far future are entitled to tax everyone who might ever come across that thing let alone actually want copies of the thing.
Your argument must sound intelligent to you, but it starts from a premise of "of course copyright is the only non-lunatic policy people could ever imagine", and goes from there. You can't even think in any other terms.
> Also, where does the GNU Public License fall in terms of "anti-human copyright maximalization"? Is it bad because it uses fire, or is it good because it fights fire with fire?
Stallman is clever to twist the rules a little to get a comparatively sane result from them, but there are others who aren't clever enough to even recognize that that's what he's doing. So, in their minds "what about the gnu license" seems like a gotcha. I won't name those people, but their username starts with Terr and ends with an underscore.
> others who aren't clever enough [...] I won't name those people, but their username starts with Terr and ends with an underscore.
https://news.ycombinator.com/newsguidelines.html
____________
> It wouldn't be "fair use". It makes no copies.
Incorrect, the real-world behavior we're discussing involves unambiguous copies, where LLM companies scrape and retain the data in a huge training corpus, since they want to train a new iteration of the model when they adjust the algorithms.
That accumulation is analogous to photocopying books and magazines that you borrow/buy before returning/selling them again, and arranging your new copies into a clubhouse or company break-room. Such a thing is not usually considered "fair use."
In a hypothetical world where all content is merely streamed into a model, then the question of whether model-weights can be considered a copy with a special form of lossy compression is... separate, and much trickier.
> Your argument [...] starts from a premise of "of course copyright is the only non-lunatic policy people could ever imagine"
Nope, it's just the context of the discussion because it's status-quo we're living with and the one we're faced with incrementally changing. If you're going to rage-post about it, at least stop and direct that rage appropriately.
> Stallman is clever to twist the rules a little to get a comparatively sane result from them, but [you don't] recognize that that's what he's doing.
I already described the GPL as "fighting fire with fire", I don't understand how the idiom didn't make sense to you.
The Attorney General of California is an elected position. He could be recalled but not fired by the Governor.
>Since 1913, there have been 181 recall attempts of state elected officials in California. Eleven recall efforts collected enough signatures to qualify for the ballot and of those, the elected official was recalled in six instances.
https://www.sos.ca.gov/elections/recalls/recall-history-cali...
Like, if the reasoning is "we should recall them because they were mean to the lovely companies :(" then I'd expect the average person to say, broadly, "good" and vote against recall. 'AI' is not particularly popular with the public.
That Scott Wiener? How does he still have a job?
> Likewise, in many contexts it would likely be deceptive to fail to disclose that AI has been used to create a piece of media.
Under this logic a company making pencils is illegal.
Secondly:
The implied basis here is that AI isn't just a product, it's also a service. This isn't you buying a pencil, it's you commissioning the drawing. Most of these products are cloud based SaaS.
And there's also the matter that "it's just a tool" doesn't really apply to foreseeable problems. If a suspicious person shows up out of nowhere buying large quantities of fertilizer, you don't get to go "Well he could be using that fertilizer for anything, not my problem". (This is relevant to AI as pretty much all AI services already have heavy restrictions on their output, this isn't a bunch of researchers publishing a paper and having bad actors implement their own AI based on that. We have companies openly advertising deepfake services.)
Charge the creator not the tool. This seemed immediately obvious to me, but apparently not to most layman (including the California AG, apparently).
Beyond that, I don't see any similarity. In a pencil, all of the "intelligence" is offloaded onto the user. With AI, the company providing the service is playing a more substantial role.
AI is a brand new frontier where they can even pre-program their arbitrary censorship into things you download too! For as much offline censorship as possible too. Wow innovation
We have to do peer to peer, and avoid cloud anywhere and everything like we were asleep for 10 years
But in reality where AI operates the models as well, the analogy is more like "Give me a rough idea and we'll use our pencils to write your letters for you!".
Why it's bullshit: the pencil will not help you create an image. It's literally just a carrier for a medium, graphite. AI will create a high quality image in response to a textual prompt, even if the operator has trouble drawing a triangle. It's good enough to fool some people into thinking that images could be real photographs.
I'm sure you understand this just fine. It's very disrespectful to come in here wasting people's times with such specious arguments.
At that point we just wait for the next Chinese open source model.
Big Tech is the US' golden goose in the race against China. Deepseek shows China is at the doorstep, much closer and more capable than previously assumed. Any thoughts politically about how we can simultaneously crack down on Big Tech, while keeping China in check with sanctions, just went out the window.
"China's not going to respect those laws" is kinda beside the point. If they suddenly decided to cut everyone in the nation's pay in half - or double it - that would have no bearing on what is right for you or I to do.
They literally did exactly that relative to the salaries of the rest of the world, and everyone took them up on it.
In retrospect, keeping China a weak communist nation was so easy. There was even internal dissent in the late 80s. It simply required refusing to make trade deals. US and worldwide wages would have been higher, discontent would have continued fermenting, the party would have remained relatively weak, human rights would not have been so easily sold out to the lowest bidder, the US would probably not have lost 6 million manufacturing jobs in a decade (3x the number of jobs in SV), and we blew it.
This is extremely naive, and fallacious.
As policymaker, you do not know beforehand how countries are going to develop over a 40 year period (not even your own country :P). Thus the only realistic option would've been a catch-all sanction regime against... possible future geopolitical rivals? Non-democratic nations? States with different cultural values? No matter which you pick, sacrificing trade like that would've been extremely expensive and limiting for US growth (might've included India, Africa, Vietnam, Thailand, Japan, Europe, Russia, depending on what criteria you pick).
You might have seen other countries jumping at the opportunity, filling the gap and benefitting immensely in the process, like the EU, or India, Russia, Japan, some pan-African Union... The only certainty in the outcome is that the US in such a scenario would NOT be as wealthy as it is today.
"when you're used to privilege, the loss of it feels like oppression"
It will hurt to fix this, but I don't think it needs to hurt that much I think it would hurt a lot less if we were actually trying to make it happen rather than occasionally being dragged kicking and screaming in that direction.
(i do not have any meaningful ideas how to bring about this kind of change, (maybe a fake giant squid alien in manhattan? :P))
Now we have to pivot and focus on containment in Cold War II.
OpenAI is talking about spending half a trillion US dollars, they have the money to license data.
In music, there is compulsory licensing and companies that use recorded music are able to make the economics work.
It needs to be repeated that these are not simply "tokens", they are the product of millions of individual people that are being appropriated for the financial gain of a very few other people.
Can't. Even if someone has the money (I truly doubt), you can't contact millions of copyright owners (as you report).
I'm in support of them being able to do it, but the right avenue is by working and lobbying hard to change antiquated copyright laws. Being able to disregard copyright only if you have enough billions of dollars on hand is the worst outcome. It's literally laws that only apply to the poor.
I'm sure you already see the folly of that argument.
Anyhow, flowing on, the allegedly totally inefficient governments of this world routinely contact millions and millions of legal entities, and many of them are poorer than Microsoft, Google, or even OpenAI, yet they somehow manage. So it seems to be practical.
Of course, that does not answer the cost thing, we all know governments just print more fiat money...
So we have been told that IP is indeed property and the property owner has a right to compensation for use. Nobody ever told me that I just have to be blatant enough to be scot-free. And I guess Sony, Warner Bros., Atlantic et. al. didn't get the memo either, or why would they sue a single university student for 4.5 million dollars? [1] This seemed and was much too much for a single university student to pay. So "too expensive" is off the table, too. Weird world.
[1] the Tenenbaum case. Tenenbaum was lucky but still broke afterwards.
There is currently no law that states it is illegal to train a model on copyrighted work.
Contacting millions of people is something many businesses on earth do.
If these companies are already engaged in trying do do something that quite literally can't be done (again, as far as can be proven today), it's not out of line to ask them to at least try to do something that many other companies actually do in practice (pay lots of people).
It's important to be very clear that this is something that could be done, but that the AI companies do not want to even try to do.
- companies being allowed to spin fairy tales about their products' capabilities
- no real consequences to enabling scammers, copyright thieves, and misinformation factories
That stuff might be good for securing short-term investment - and it befits a society obsessed with cryptocurrency and sports gambling. But it doesn't seem good for building meaningfully smarter computers, just dumber computer users.
> AI systems are at the forefront of the technology industry, and hold great potential to achieve scientific breakthroughs, boost economic growth, and benefit consumers. As home to the world’s leading technology companies and many of the most compelling recent developments in AI, California has a vested interest in the development and growth of AI tools. The AGO encourages the responsible use of AI in ways that are safe, ethical, and consistent with human dignity to help solve urgent challenges, increase efficiencies, and unlock access to information—consistent with state and federal law.
It's impossible to understand this as a statement that AI companies are a "legal clusterfuck" or "may be entirely based around criminal activity".
edit to expand: from the headline I thought this was going to be Bonta coming out against the argument that AI training is fair use, which really would at least arguably apply to "almost everything" the companies make. But no, he's just saying not to do things that AFAICT they already all ban in TOS.
That aside, state governments are supposed to heavily augment federal laws with their own state laws and regulations.
California doesn't need or want input from a bunch of states with little to no AI (or even IT in general) industry footprint about how those companies should operate.
They don't actually have "intelligence" or agency. There's also not much we can do about the hallucination problem.
Call them what they are: errors.
I must have misunderstood you.
Only LL Cool J can call California "Cali" without derision.
https://www.reuters.com/technology/meta-used-copyrighted-boo...
After all its actually documented. Its also not likely to be fair use.
The proving now harm bit is going to be difficult, and bad on the old trumpian optics.
copyright at least forces some of his billionaires to fight with other billionares to come to a conclusion as to "what is good for America"
Its hardly a sure thing but their position does seem to be at least somewhat supported by precedent.
The tl;dr is that you have to look at the impact on the market for the use to figure out if it's transformative [1]... which means it's extremely unlikely that training for AI is going to be considered "transformative", and thus they lose the first factor. Given that AI is absolutely reamed on the fourth factor (especially now that they're paying people to use their content for training, that's basically a concession on the fourth factor), there isn't really any grounds for them to claim fair use.
[1] Yes, it's pulling the fourth factor into the first factor, it is a rather garbage opinion, but it is precedent as of 3 years ago.
Rule of law means having clear guidance on what is or isn't illegal. Vague guidance that everyone in the industry "might" be breaking the law isn't responsible, it's setting up a mechanism to trade favors and selectively prosecute enemies to reward friends.
Be better, guys. This isn't the right take.
That is an extremely basic rule of civilized governance. You can't wordplay around it.
This is not the entire legal system, but it is a critical part of it.
Rule of laws means there are known laws everyone has to follow no matter who they are.
The AG's stance on whether an action is a crime, and their policy towards prosecution, does not need to be tested in court. That is something they can communicate without litigation.
Internationally, it conflicts with Canada.
> Otherwise please use the original title, unless it is misleading..
"California AG to AI Corps: Practically Everything You’re Doing Might Be Illegal"
> Create or sell an AI system or product that has “an adverse or disproportionate impact on members of a protected class, or create, reinforce, or perpetuate discrimination or segregation of members of a protected class.“
It's a straw man to characterize that as saying disparate outcomes ALONE are why AI systems might be discriminatory. They may well can be, and likely are, actually embedding biases in the models.
The simple examples of how many systems often prefer "he" for doctors and "her" for nurses shows that bias in datasets results in bias in the models. Yes, that is a result of the dataset and reflects the dataset (and maybe even real world statistics on doctors and nurses!) but it does mean the system may treat women and men differently when there is no legal justification to do so.
All such Harrison Bergeron-style speech-policing is antithetical to a free society.
Of course it requires some actual harm before the law is involved. Perhaps AI recruiting systems favour resumes that match the "correct" gender to profession. Perhaps a hypothetical AI sentencing recommender gives stiffer penalties to people that live in high crime areas; after all, statistically they are more likely to re-offend - that's just not an acceptable sentencing factor. Perhaps a house appraisal AI accidentally recreates redlining.
These are just hard to prove, so my example is there to demonstrate that even though it's hard to demonstrate we should take the risk seriously.
Perpetuating "harmful stereotypes" reasonably falls under this language. Trying to correct the bias of the dataset, or worse, correcting the bias of reality, is a fool's errand. The same motive produced the Google AI images controversy last year and constitutes a substantial percentage of instruction fine-tuning for "alignment".
I agree that care should be exercised in rolling out algorithms for sentencing guidelines.