Microsoft, OpenAI sued for ChatGPT 'privacy violations'
theregister.com
theregister.com
I don't expect this lawsuit to lead anywhere. But if it does, I hope it leads to some clear laws regarding data privacy and how TOS is binding. The recent ruling regarding web scraping makes the case against OpenAI a lot weaker. [1] Data scraping publicly available data is legal. People didn't need consent to having their data be used, there was an implicit assumption the moment the data was published to the public, like on reddit or youtube.
I keep seeing this idea reoccur in the suit:
>Plaintiff ... is concerned that Defendants have taken her skills and expertise, as reflected in [their] online contributions, and incorporated it into Products that could someday result in [their] professional obsolescence ...
Anyone is able to file a suit, I wish people stopped assuming that a news report automatically means it has merit.
1. https://www.natlawreview.com/article/hiq-and-linkedin-reach-...
* Was it published publicly? This is basically defined in the courts as "if you make an unauthenticated web request does the data return?". This is where scraping comes in- if you make the data available without authentication you can't enforce your TOS, because you can't validate that people actually even accepted the TOS to begin with.
* Is the data able to be copyrighted? This is where things are interesting- facts can not be copyrighted, which is why a lot of scrapers are able to reuse data (things like weather, sports scores, even "for hire" notices can be considered factual).
* If it would typically be considered covered by copyright, does fair use come into play?
* Are there any other laws that come into play? For example, GDPR, CCPA, or other privacy laws can still add restrictions to how data is collected and used (this is complicated by the various jurisdictions as well)
* Was the work done with the data transformative enough to allow it to bypass copyright protections? This goes back to when Google was scanning books. Because they were making a search engine, not a library, their search tool was considered transformative enough to allow them to continue.
It's not enough to say "because it's on the internet, it's fair game for everyone to use". This is a really complicated area where things are evolving rapidly, and there's a lot of intersecting law (and case law) that comes into play.
And another complication is that OpenAI is not exposing any static data. A response is generated only after prompting. I'd argue that LLMs are closer to calculators than databses in function. The amount of new information that can be added is also limited, it's is not a continuous learning/training architecture.
I do hope this leads to more clear laws regarding data privacy, but I can't imagine the allegations of "intercepting communications", violating CFAA, or violating unfair competition law will hold.
To put it another way, it's legal for me to go to the library and borrow a DVD or a book or poems. That doesn't give me the right to publish the poems again under my own name. Whether I find the poems from scraping, borrowing the book from a library, or even just reading it off of a wall I don't get ownership rights to that data.
The same logic applies to a lot of other laws around data. If you collect data on individuals there are a bunch of laws that come up around it, and many of them don't really concern themselves with how you got the data so much as how you use it. The fact that it was scraped doesn't grant any special legal rights.
You keep making the claim that because it was scraped people can do whatever they want, as scraping is legal. That is the only thing I'm arguing against, because that is a gross misinterpretation of how the case that made scraping legal was decided. LLMs aren't relevant to that point (which is exactly what I keep saying- the method of collection doesn't magically change the legality of it).
That being said, you're still wrong. The USPO has said that the output of LLMs are the outputs of algorithms and are not creative works. Therefore you can't "own the copyright to the new work you make" because the work itself can't be copyrighted at all. No one can own the output of an LLM.
Also, just because it seems you want to be wrong on every level, it is absolutely possible that a neural network would be able to repeat data from its training set. This is an incredibly known problem in the field.
https://www.bloomberglaw.com/external/document/XDDQ1PNK00000...
That doesn't mean you grant a license to produce derivative works other than search indexes. Legally, it's different. (Germany codifies these as separate "moral rights": Urheberpersönlichkeitsrecht.)
You can usually disregard such articles as you can expect biased/incomplete reporting.
Lawsuit claim amounts have zero bearing on reality. They must be specified in any classroom, but lawyers just always specify massive amounts without justification.
Any reporting on this amount indicates ignorance in the system or intentional dishonesty.
One of the "I wonder where this will go" things with the reddit and twitter exoduses to activity pub based systems is that it is trivial for something to federate with it and slurp data without any TOS interposed.
The TOSes for these systems are typically based based on what can be pushed to them - not what can be read (possibly multiple federations downstream).
It's been a bit surreal seeing modern day Luddites come out of the wood works basically coming up with any ethical/legal argument they can that is a thinly veiled way of saying "I don't want to be automated!"
Not commenting on whether or not they are right per se, but it's weird seeing history repeat itself.
To make coal mining automation analogous to chatGPT the machinery would have had to use something the coal miner did to learn how to automate their work? I'm imagining a camera looking at all the coal miner's work and then the machine can immediately do it, but better.
I agree it is a tad different, but like with someone's coal mining which is in the public domain for anyone in the tunnel to see, likewise anything you write unprotected online is in the public domain and fair game I think?
(I should caveat that I think if they get what they want, we all lose in a big way. Not that I think this is going anywhere)
We're coming up on the outer bounds of our systems of incentives. Captialism, as a system, is designed to solve for scarcity, both in terms of resources and in terms of skill and effort. Unfortunately, one of the core mechanisms it operates on is that it's all-or-nothing. You MUST find a scarcity to solve or you divorce yourself from the flow of capital (and starve / become homeless as a result).
Thus, artificial scarcity. It's easy to spot in places like manufacturing (planned obsolescence) IP (drug / software / etc patents) and so forth. I think this is just the rest of humanity both catching on and being caught up with. Two years ago, everyone thought they had a moat by virtue of being human. That's no longer a given.
One hopes that we'll collectively notice the rot in the foundation before the house falls over (and, critically, figure out how to act on it. We have a real problem with collective action these days that may well put us all in the ground).
Why? Except for the longshoremen in the US getting compensation and an early retirement due to the introduction of containers, I know of exactly 0 (ZERO!) mass professional reconversions after a technological revolution.
Look at deindustrialization in the US, UK, Western Europe.
When this happens, the affected people are basically thrown in the trash heap for the rest of their lives.
Frequently their kids and grandkids, too.
Businesses change and adapt. Workers too — but people often don’t like change, so many choose to stay behind. Should we cater to them?
I used to do a lot of work which is now mostly automated. Things like sysadmin work, spinning up instances and configuring them manually, maintaining them. I reconverted and learned terraform, aws etc when it became popular.
Should I have gotten help from the government to instead stick to old style sysadmin work?
I don't think anyone beyond a few marginal voices are calling for a ban on job automation. What they seem to prefer is that, if they are to be automated out of a job, they should be compensated for their copyrighted works having been used in the process of doing so.
Regardless, at the very least people who are being automated should get some government support. Not everyone can easily retrain.
Suppose you're a farmer. You've been working on your tractors for decades, and have even showed the nice folk at John Deere how you do it. Now they've built your improvements into the mass-produced models, and they say you can't work on your tractors any more. Who should reap the profits?
Suppose you're a writer. You've spent a long time reading and writing, producing essays and articles and books and poems and plays, honing your craft. You've got quite a few choice phrases and figures of speech in your back pocket, for when you want to give a particular impression. Now, there is a great big statistical model that can vomit your coinages (mixed in with others') all over the page, about any topic, in mere minutes. Who should reap the profits?
Suppose you're a visual artist. You enjoy spending your time making depictions of fantasy scenes: you have a vivid imagination, and, so you can make a living illustrating book covers and the like. You put your portfolio online, because why not? It doesn't hurt you, it makes others happy, and maybe it gets you an extra gig or two, now and then. Except now, there's a great big latent diffusion model. Plug in “Trending on Artstation by Greg Rutkowski”, and it will spit out detailed fantasy scenes, photorealistic people, the works. Nothing particularly novel, but there was so much creativity and diversity in your artwork, that few have the eye to notice the machine's subtle unoriginality. Who should reap the profits?
"You build a dam that destroys 10000 homes, who should reap the profits?"
• Should we be destroying people's homes to build dams without their consent?
• In general, are people being compensated when these things happen to them? i.e., while it might be nice, does this actually happen?
The Luddites (the real ones, not the mythological bastardisation of them) continue to be sympathetic characters.
The famous: "it depends" :-)
AI most likely falls under: "they should be", IMHO.
That's the real flaw in Luddite thinking -- you can destroy the machines.
Tokugawa Japan, Qing China, many other places including in Europe for centuries.
That's too extreme.
My point is that we're reaching a point where people need to be compensated. We can't just destroy their lives, collect all the money in 2 bank accounts and call it a day.
Just because I posted something on reddit because I thought it was funny, doesn't implicitly give permission to anybody to take that post and profit from it. You're doing a disservice to consumers by acting like it's their fault for being exploited.
I disagree with you on whether it should count as being exploited. I don't see fanfiction writers professional impersonators or as inherently exploitative. I understand that some people would disagree because there is a difference in scale. But using technology to mimic and, in some sense, replace human effort is the reason it is useful.
I believe this will shift how and why people value organic media. The standard of what makes content "good" will rise in the long term. When stable diffusion first came out, I compared the generated art to the elevator music. I feel the same way about the output of LLMs. I might feel differently in a few years if models get better at the rate they currently have been, but that's not likely.
I agree that people should have more control over how their data is used, and I'd love to see this suit lead to stricter laws.
I'm not too worried about copyright issues because regardless of whatever happens with upcoming case law and legislation, any regulation against the input data will be totally unenforceable. It's nearly impossible to detect whether or not an LLM was trained on some corpus of data (although maybe there is some "trap street" equivalent that could retroactively catch an LLM trained on data it wasn't allowed to read). And even if the weights of a model are found to be in violation of some copyright, it's still not enforceable to forbid them, because they're just a bag of numbers that can be torrented and used to surreptitiously power all sorts of black boxes. That's why I'm much more worried about legislative restrictions on hardware purchases.
I hope it leads to more people realizing that a TOS doesnt override their individual rights and that the legal system works to support them.
It's codified in the fact that saying you'll do something means you're socially obligated to do it, and legally obligated if you receive something in return.
It seems to me there are a to of counter examples to this "right" you speak of. So many that it doesn't seem like it really exists.
There's sort of an exception for military service, but even soldiers have acess to military courts.
The same argument could be used to defend ubiquitous face recognition in the street though (“when going to the street, there's an implicit assumption that your presence in this place was public”) but I'd really like if we could not have that…
There's a case to be made that corporation gathering data and training artificial intelligence don't need to have the same right as people: when I go to the street or publish something on Reddit, I'm implicitly allowing other people to read my comments, but not corporations to monetize it. (GDPR and the likes already makes this kind of distinctions for personal information by the way, so we can totally extend it to any kind of online activity).
My favorite LLM analogy so far is the "lossy jpeg of the web." Within that metaphor, I don't see how anyone can claim copyright on the basis of a pixel they contributed that doesn't even show up in the lossy jpeg. They can't point to it.
https://theinnisherald.com/the-other-once-upon-a-times-a-his...
https://en.wikipedia.org/wiki/Legal_issues_with_fan_fiction
Fanfiction and fan art also tend to run afoul of the infrequently (but occasionally) litigated part of copyright - copyright of fictional characters.
https://en.wikipedia.org/wiki/Copyright_protection_for_ficti...
I came across this with the Eleanor lawsuits - https://www.caranddriver.com/news/a42233053/shelby-estate-wi... - and while I believe that that instance Eleanor falls on the "this shouldn't have been copyrightable" (took a bit to get there), the question is "what protects the representation of Darth Vader?"
In general it tends to be ignored and tacitly encouraged... but it isn't protected.
So the technology is cool, but I'm firmly of the stance that they cut corners and trampled peoples' rights to get a product out the door. I wouldn't be entirely unhappy if this iteration of these products were sued into the ground and were forced to start over on this stuff The Right Way.
If ChatGPT regurgitates verbatim or nearly verbatim, something it slurped up from OP's blog, is that not plagiarism? Where do you draw the line? Where would a reasonable person draw the line?
Often rather than claiming human aspects to the machine, they are going further, and claiming machine aspects to the human.
Using mechanistic analogies for explaining the human body or mind isn't new, but as machines become better and better at imitating humans, those analogies become more seductive.
That's my rant; the danger with 'AI' isn't so much that humans are enslaved by machines, but that we enslave each other -- or dehumanize each other -- with machines.
You are entitled to control it's distribution and use. You are not entitled to control it's influence and effects.
AIs are not massive repositories of harvested data. The models are relatively small (<20GB).
https://www.pinsentmasons.com/out-law/news/google-thumbnails...
> A US court ruled this week that Google's creation and display of thumbnail images does not infringe copyright. It also said that Google was not responsible for the copyright violations of other sites which it frames and links to.
> The Court said that Google did claim fair use, and that whether or not use was fair depended on four factors: the purpose and character of the use, including whether such use is of a commercial nature or is for non-profit educational purposes; the nature of the copyrighted work; the amount and substantiality of the portion used in relation to the copyrighted work as a whole; and the effect of the use upon the potential market for or value of the copyrighted work.
Taking copyrighted material and using it to train a model is not a copyright infringement - it is sufficiently transformative and has a different use than the original images.
Note that AI models can be used for different things. A model trained to identify objects in an image has never had uproar about the output of "squirrel" showing up in the output text.
The model also, as a purely mathematical transformation on the original source material does not get a copyright. If it needs to be protected, trade secrets are the tools to use to protect it. A model is no more copyright worthy than tanking an image and applying `gray = .299 red + .587 green + .114 blue` to it.
The output of a model is ineligible for copyright protection (in the US - and most other places).
The output of a model may fall into being a derivative work of the original content used to train the model.
It is up to the human, with agency in asking the model to generate certain output to be responsible for verifying that it does not infringe upon other works if it is published.
Note that the responsibility of the human publishing the work is not anything new with an AI model. It is the same responsibility if they were to copy something from Stack Overflow or commission a random person on Fiverr... its just that those we've overlooked for a long time - but it is similarly quite possible for the material on those sources to be copyrighted by and/or licensed to some other entity and the human doing the copying into the final product is responsible for any copyright infringements.
Saying "I copied this from Stack Overflow" or "I found this on the web" as a defense is just as good as "Copilot generated this for me" or "Stable diffusion generated this when I asked for a mouse wearing red pants" and represents a similar dereliction on part of the person publishing this content.
Would that satisfy you?
No. Both legally and practically, you absolutely do not.
The only thing copyright law gives you is an exclusive right to sell it for a limited period of time, as a whole in its original form or similar -- and to transfer that right.
Regardless of your desires, anyone can reuse it under the conditions of fair use. They can copy parts of it for parody purposes. If they're not selling anything or taking away from your sales*, they can reproduce it verbatim for private purposes. And even if they are selling something, they can summarize it, quote from it, rephrase it, and so forth.
And you don't actually get to decide any of that.
* Edit: added "or..."
> I wasn't asked and I don't really care to donate work to large corporations like that... I do get to decide what happens with it.
And I said:
> No. Both legally and practically, you absolutely do not.
You think you get to decide whether large corporations can train on your work. I'm saying the the law suggests you very much don't get to decide that.
Send some links if you see some definitive case law sorting this stuff out.
But we do? Open sourcing something with caveats is common. This code is public BUT not for commercial use. This code is public BUT you must display attribution etc.
Sure, blogposts are unlicensed (that I know) but the idea of something publicly available being held to restrictions is nothing new.
But telling me it is illegal to share what I learnt because the original source is copyrighted... doesn't sit right with me.
What protects particular solutions is patents. For example if someone were to obtain a patent for computing GCD of large integers the usual fast way, well then everyone else would have to use a different solution.
This analogy to someone reading a book, perhaps peppered with lots of legalese to the point of being hardly recognizable, will definitely be used in courts at some point. And I can't see how it wouldn't stand as a valid argument.
If you pick up a book and learn a fact, then yeah, you’re allowed to share that fact.
It’s weird that this topic keeps devolving into a form of “so what, it’s illegal for me to learn things?” Because: no, it’s not. And: You and a piece of software are treated differently under the law. You have a different set of rights than ChatGPT.
Gods, no. Where did you get that from?
Those might not be a problem regarding this specific case, but the case can easily be made that it ought to be.
I don't understand your point. Do you think it makes any difference whether I use my laptop, or a pen, or ChatGPT to violate copyright?
On the other hand, it's completely feasible to make a license that stops someone from training their model with some piece of info, is it not?
And regardless -- the problem now is that expectations of how content can be consumed are now fundamentally violated by automation of content ingestion. People put stuff up on the Internet with the expectation of its consumption by human minds, which have inherent limitations on the speed and scale on which they can learn from and reproduce things, and those humans are also legally liable, socially/ethically obligated, etc.
Now we have machines which skirt the limits of legality, and are able to do so on massive scale and without responsibility to society as a whole.
Different game now.
Then people obviously aren’t aware that bots have been indexing web pages and showing summarized information without going to the web page for three decades.
Further, almost every site has had an e.g. robots.txt which has permitted content harvesting only for certain accepted purposes for a couple decades now. So clearly people already had a sense of how they wanted their content harvested and for what purposes.
So you’re okay with Google making money off of your content. But not OpenAI?
Not on the phone yet, but on a Mac which could include iMessages.
Another part which bothers me is that I have lots of different personalities online. On most sites I use different usernames, and I wonder if there will someday be an AI which can match all the different online profile to a single person, even if different username are being used etc.
Putting it more bluntly, it is somewhere between a parasite and a slave driver.
I wonder what would an LLM trained on Google code and internal documents look like?
And it adds nothing. I'm sorry but saying "Whether that content is used to train a human mind or an artificial one is probably not up to you" may be worse than saying nothing at all.
First because it shows enough doubt on whether it's up to the authors of content (IP laws, fair use, intent of the use, and many things I ignore), while giving no laws as an example or frame of reference.
And second because it's comparing a human mind that we know exist, to an artificial one, which implies:
1. An LLM is an artificial mind, or close to one, whatever that is (again, not defined).
2. If they were to exist, they would be both equivalent and treated the same as a human one.
The amount of jumps in a couple sentences, added to the uncertainty of how copyright would/will work, multiplied by the numer of times I/we read that type of comment every single time, it's getting tiresome. And it's adding noise to the noise-signal ratio.
If you're tired of responding to these comments then stop. It's the internet, everyone is at different places in exploring topics and having discussions. Don't poo-poo on someone else's journey and instead move on with your day. There is no required reading (other than TFA) on hacker news.
If you want to prevent a web spider from scraping your blog, use a captcha or robots.txt. Copyright law doesn’t apply to this scenario.
Don't get me wrong, this is a grey area where copyright laws and general consensus haven't caught up with new techonology. But if you voluntarily stick something up online with the intent that anyone can read it, it seems a bit mean to then say "wait no you can't do that" if someone finds a way to materially profit off it.
Seems like if it's legal for a person to do it should be legal for software to do for the most part.
Surely, there is some pretty large subset of things where "if it's legal for a person to do it should be legal for software" does not hold up?
So how about the default is "not allowed"
If you ask ChatGPT the rules for D&D, the private sourcebooks are all in there.
...wait, isn't that false? legitimately asking.
or is it because it was done by a corporation that makes it illegal?
im thinking of how restaurants dont sing happy birthday and fair use restrictions etc
If I recite them to myself, in my home, it's fine. If I do it at a gathering at my house where we're playing D&D, fine. If I do it as a performance, in front of a crowd, or as a recording, now I'm no longer fine. Context matters in a copyright cases. Not to mention, to claim fair use, you do have to claim you violated copyright. Fair use is just an allowed violation.
As to Happy Birthday, that's actually ok for them to do now. The person/group that held the copyright to Happy Birthday was found to have not actually have held them in the first place. Happy Birthday is actually an older song called "Good Morning to All". Swap "Good Morning" with "Happy Birthday" and "children" with "dear [PERSON]" and you have the lyrics. This was not deemed a substantive change. And since the copyright on "Good Morning to All" has lapsed, Happy Birthday is in the public domain.
I can give you a rock that I own, which I hope we all agree is not copyrightable, and ask you to sign a license that you will keep it indoors. If you put it in your yard, you are breaking the license and potentially liable. This has nothing to do with copyright.
[0]: https://fairuse.stanford.edu/overview/fair-use/four-factors/ [1]: https://creativecommons.org/faq/#can-i-apply-a-creative-comm...
The rules of games cannot be copyrighted either. The artistic elements can be trademarked, but if ChatGpt merely explains the rules to you in different ways, that isn't infringement either.
Being non-commercial is not an automatic fair use exception. Being commercial does not preclude fair use. And rule concepts are not copyrightable, only the specific expression. Rules may have other IP protection, including patents.
That's not even true in the US anymore. You'd have to convert those rules into some sort of device, or argue that the game is a business method.
Whoever told you that is lying to you. You are not legally allowed to personally memorize and recite copyrighted works all you want, any more than you're allowed to personally memorize, write down copyrighted works, and distribute them as much as you want.
All piracy is a process of computer-assisted remembering and reciting.
I can't go write and commercialize what I learnt directly, but I'm not breaking the law by quickly seeing how some book I didn't buy ends so I can talk about it at a party - and then everyone knows how it ends which might affect whether they want to buy said book and upset the author. But, tough shit, what I did was legal. I can even use the ending as one set of input from dozens of inspirations for my own book where the end result is transformative enough where the sources are unrecognizable. And if I had learnt about the endings from a dozen books without buying those books I didn't break any laws even though I am now commercializing something in being inspired by them all to make something new.
> The lawsuit is seeking class-action certification and damages of $3 billion – though that figure is presumably a placeholder. Any actual damages would be determined if the plaintiffs prevail, based on the findings of the court.
Obtained from (check notes) public internet forums
> For the 16 plaintiffs, the complaint indicates that they used ChatGPT, as well as other internet services like Reddit, and expected that their digital interactions would not be incorporated into an AI model.
You've got to be incredibly naive if you think public Reddit data isn't used to train ML models, not least by Reddit themselves
2: You're the one who went with "invented" ;)
3: I know you're exaggerating, but I think you think you're exaggerating much less than you actually are.
It's not really important to the debate around unlicensed use of copyrighted works to train AI models, but it wouldn't surprise me at all if the majority of Reddit users have joined since 2018. It's tough to get reliable active user counts, but they seem to have risen substantially over the past five years.
It also wouldn't surprise me if the majority of Reddit users were indeed from prior to 2018, but at the very least > 2018 would be a very substantial minority.
[1] https://www.kaggle.com/datasets/ehallmar/reddit-comment-scor...
I am sure OpenAI thought all this through, so I can only assume they said "fuck it let's pull an Uber and do this anyway." We are in for lots of interesting legal headlines
They picked 3B hoping to get several million...
But that is a moral point, not a legal one; IANAL and can't say anything valuable about the legal merits.
Ideally AI makes us all redundant and the money stops mattering anything like as much, similar to how owning land stopped mattering anything like as much when the industrial revolution happened.
Regardless, I think this is a policy question rather than a legal question, even if this fight happens to be in a court.
If you're going to make a claim this strong, you should expand on it. Should software be able to have custody of children? Should it be able to kill in self-defense? Should it be able to make 14th amendment claims? Exactly what part of the case (other than the damage claim) is hard to understand?
It's legal for me to look out of the window and watch my neighbor go to the supermarket.
It's _not_ legal for me to build an automated surveillance system that tracks everybody on the street 24×7 and stores everything into a large database.
There's more deliberate action when you post something on a public online form than just existing in a place outside of your house. Especially considering you've always had the option to use reddit anonymously anyway.
Read, yes - post no.
And - you can no longer create an account that is not tied to an email...
When this happens to closedAI, it just seems like a profit grab.
Not that it changes the legality of it. Just optics.
Wonder if that matters in court.
Then they came for the writers, but I did not speak out because I was not a writer.
Then they came for me, and there was no one left to speak for me... well, except ChatGPT.
edit: corporate LLMs have pulled the "one death is a tragedy, ten thousand deaths are a statistic" ploy off fully. If you want people to question whether you're even violating copyright, make sure you violate all of them at the same time. They'll just decide that you're an act of god and not covered under earthly laws.
Not that all capital is distributed by merit, plenty of people used military might or factionalism/leaders/politics to obtain disproportionate amount of capital.
But if you are against the last 2 happening, I don't see what you expect a reorganization of society to accomplish since you are going to get a power structure of factionalism/leaders/politics taking priority. (Sorry bud, no an-com utopia ever existed, they all had factionalism/leaders/politics, thus defeating the entire purpose of removing class.)
I think most of us think we can capture/retain power easier with money, than having to climb up inter-party politics.
I don't necessarily disagree with your later points. I do, however, disagree with giving up.
At least its equitable (based on value of output), ofc there are legacy issues as well.
Some demagogue can swoon the masses and take it all if not for capital. That demagogue could be Trump or Stalin.
Know the consequences of what you are advocating for.
It's not. By definition, it's based on control of capital. That's why it's called capitalism. In other words, those aren't "legacy" issues; they are literally the system as designed.
Since utopia is impossible, its a choice between:
>Capitalism, where people can typically pull off the american dream in their lifetime.
or
>Let politics determine how much material things you get
The latter seems especially scary if you are familiar with history
Capitalism follows a very simple algorithm. In a capitalist economy, capital always accumulates, with all exceptions being precisely that: exceptions. Are you defending the exceptions or the rules?
Realize there was a very long and quite recent time when capitalism was impossible. By your logic, we should reinstate the divine right of kings.
I don't think this is relevant. If OpenAI had trained a model on just one copyrighter holder's content it would likely not be different legally, even if the model would perform much worse.
> Nuance has strict data agreements with its customers, so patient data is fully encrypted and runs in HIPAA-compliant environments
Additionally Epic seems to already be storing these clinical notes in databases and Nuance which Microsoft owns has already technically been a 'hot mic' in these same doctors office for some time. The new offering is an AI-draft note generator.
I'm personally skeptical that model output would suddenly be under different rules than the other voice-to-text AI model output?
When I worked on an expert witness report for a big law firm we just used Word.
Creator: Acrobat PDFMaker 23 for Word
Producer: Adobe PDF Library 23.3.247; modified using iText® 7.1.6 ©2000-2019 iText Group NV (Administrative Office of the United States Courts; licensed version)
So it was likely made in Word and exported to PDF. (One can anyway guess from the "look" of the paragraphs that they're not using anything like Knuth–Plass line-breaking, which rules out things like *TeX and InDesign.)The 1-28 pleading numbers on the side are annoying. They're specific to courts in California and a few other jurisdictions, and the rules of court require them. But many other courts don't have them, and they only help to cite specific lines within pages; eg "Complaint 5:4-9" means "Complaint at page 5, at lines 4 to 9". It's occasionally useful for court filings like this, but more useful for court/deposition transcripts of testimony to show precisely where a witness said something.
Related: I tried building an RNN to generate legal pleadings back around 2018/19 and gathered a bunch of docs like this from courts across the country as training data. Processing text with those pleading numbers was a pain, so I built a CNN to classify whether a document had pleading numbers or not, which affected downstream processing. Probably the wrong approach in a bunch of ways, but I was just learning.
Seems to be a blatant violation of GDPR. So I assume they’ll be fined for it sooner or later and forced to cleanup the training data anyway.
GDPR doesn’t prevent opt outs of this kind of thing.
So if they’re claiming they have the right to process data on the legal basis of consent, and they claim the absence of that cookie constitutes that consent, then they have no legal basis, and are thus in violation of the law.
We are actually working on a tool to create billion-size free-to-use Creative Commons image datasets and prepare them for training models like Stable Diffusion. There is a blogpost about it here: https://blog.ml6.eu/ai-image-generation-without-copyright-in...
In terms of implementation, I wonder about a few things:
Do models trained on more data have to pay more? LLaMA was trained on 1.5T tokens, the original GPT-3 was trained on ~300B tokens. And this is only partially related to model quality, LLaMA 13B and LLaMA 65B were trained on the same data, but the 65B model is better. What's the incentive to ever use the 13B model, if the licensing cost is 100x-1000x the model inference cost?
Who defines a word? Each model uses a different tokenizer. I'm personally amused by the idea of a government-mandated tokenizer.
What about generations that never see human eyes? As an NLP researcher, I've generated millions of tokens for training and automatic evaluation purposes -- are those subject to licensing as well?
The idea is to keep it simple, so it wouldn't be based upon the specifics of training, just whether or not it used public data. Anything else would require companies to divulge trade secrets and that won't fly. And words are defined here as, well, words -- English words. There'd be a separate fee per pixel/voxel, and then a catchall for non-language/non-image models.
2. Given that the internet is global, is every country supposed to make their own versions of this? Will I have to pay the EU tax to use models that might have been trained on data that Europeans posted online?
What if I'm not a massive corporation with millions of lines of code to train on and I want to pay for an AI coding assistant? Doesn't this make it effectively illegal for me to purchase such a product for a reasonable price when big companies will presumably be able to use it without paying the tax?
Another situation - let's say you're a company that contributes heavily to open source, but also accepts external contributions. Could Facebook train a model on the React codebase, for example, without having to pay the AI tax?
Another situation - suppose I start an LLM coding assistant and sell it to my friend. Presumably I don't have to pay the tax as a "low revenue" company. Then I get acquired or get some huge seed round and suddenly my customers have to pay the AI tax. Doesn't this just nuke all my customers?
Anyway, as a software engineer, I personally want people to use my code for whatever they want to use it for, without having to pay me for it. I indicate that by using an MIT license. Why throw that precedent out the window?
And this would not prevent you from explicitly licensing your code or writing to let people to train on it. But what it would do is say that if someone didn't explicitly license it then it is covered under the policy.
Look at all the share economy players it boils down to offload the risk, labor and debt but keep the margin.
Now that the gates are open, we'll probably be entering the "free money" cycle soon.
https://iapp.org/resources/article/us-state-privacy-legislat...
https://leginfo.legislature.ca.gov/faces/codes_displayText.x...
It isn't even turned off by default. Many sites just give you an "i accept" button or even if you want to manage the preferences, the "accept all choices" button is where the "confirm my choices" should be.
Bigger companies will just append this to their TOS and push it down the customer's throat. That if MS doesn't settle out of court and the case gets thrown together with any major oppositon to the data mining
Thankfully AI doesn't work by memorization.