There's also copyright reversion, which is a related new provision that applied to older copyrighted works. Quoting from an article I just pulled up
"...the 1976 Act created a new right allowing authors and their heirs to terminate a prior grant of copyright, the Act also set forth specific steps concerning the timing and contents of the termination notice that must be served in order to effectuate termination. The termination of a grant may be effective “at any time during a period of five years beginning of the end of 56 years from the date the copyright was originally secured”..."
But this is a red herring because the fact a model has been trained in the past doesn't mean a copyright lawsuit is "retroactive". The infringement would presumably be occuring anew every day you make it available on your web site.
Meta has been lobbying hard around that for years.
Only copyright can see through all of that, you would have to gut fair use in order to have an effective anti-scraping law.
I don’t think stack overflow is all that valuable once your model has access to github due to their good friends at MS.
The money in proprietary AI is on the top end now, open source / edge is destroying monetisation on the lower end. Top end means high quality domain specific data.
As a heavy ChatGPT user I disagree. Lack of up to date data is one of the biggest issues I face every day - technology changes fast, libraries change APIs, new tech comes out, etc.
As of my last knowledge update in September 2021, Quickwit is an open-source search engine infrastructure that is designed for building and deploying search solutions quickly and efficiently. It focuses on providing fast and scalable full-text search capabilities for applications and websites. Quickwit is built on top of the Rust programming language and leverages technologies like the tantivy search engine library.
Courts are more deliberate than you would like — no denying that. But this is a feature not a flaw. It may be that damage will be done by then. Perhaps irreversible. But I would like to think if there is a will there is a way and that if things are terrible enough the governments will be bold in their responses.
Let's say the French government decides that OpenAI must change something about their business practices if they want to continue operating in France. OpenAI says "nope", and blocks access to French users.
Suddenly French companies aren't able to use GPT-X anymore – while their competitors in other countries can. How long do you think it will take before a storm of corporate outrage forces the government to relent?
Any individual government (except, perhaps, the combined US and EU governments) is powerless against today's technology megacorporations, because they can take much more away from a country than that country can take from them. If push ever comes to shove, it will become obvious where the true power lies. So far, the corporations have barely even tried to throw their weight around.
That's one possible outcome. (ETA: You DO have a point here, but...)
The other is, you know, something like every website explicitly telling me, via an annoying popup, how much they value my privacy. Also, me not being able to access half of US news sites to this day.
The last time EU raised their finger, every technology company (FAANG included) shat their pants.
And that was simpler times, times when a cookie stored in your temp folder without websites shouting they're about to do so, was somehow the biggest concern of an EU netizen. It almost seems ridiculous, compared to the damage AI could do (the extent of which which nobody really knows).
Bof, les alternatives à ChatGPT ne sont pas si mal.
And even if the open source alternatives were far behind rather than just a bit — all this talk about corporate moats and their absence may be blind to the strengths of OpenAI's offerings, but even so it can be replaced if it must — the storms of protest in France are normally by the people, not by the corporations.
But that's not true, and people know it.
> the storms of protest in France are normally by the people, not by the corporations
Correct. CEOs of big corporations just call the ministers directly and tell them to get in line, or else.
Based on what I've seen? They're good enough to be interesting, more so than GPT-2.
They don't need to be amazing from day one to be a foundation for replacing the status-quo.
> CEOs of big corporations just call the ministers directly and tell them to get in line, or else.
I roll to disbelieve (that it works, not that CEOs attempt it); that sounds like conspiracy theory to me.
You mean corporations that wield more power than most governments, and have revenues equivalent to the GDP of entire countries?
If Universal or 20th Century Fox were to ever become a serious obstacle, Google and Microsoft are simply going to buy them. This isn't the early 2000s anymore. The power balance has shifted dramatically.
US$26.2 billion globally in 2022 according to IFPI, and US$31.2 billion according to Statista.
Other than Netflix, I think FAANG just doesn't care that much about such a small market (the market being "actually producing it", given they're already part of the previous numbers for selling and streaming it).
And of course, both A's and the N of FAANG have their own commissioned TV/film content.
Yeah, doesn't remember. Mhm...
Oh, it just can't remember the license terms of the code it "reads", so it can't comply with these licenses or help people to comply with these licenses.
Convenient.
I suspect the answer to the question "is it, though?" is one for the lawyers and lawmakers rather than for the software developers, and it may well vary wildly by jurisdiction.
People dont want to acknowledge that the LLM structure reflects rather closely what it is being trained on, but the incredibly large number of parameters suggests it is closer to a photographic fit than a true abstraction. larger models being more likely to memorize training data (Carlini et al., 2021, 2022)
The fact that the information gets mangled and somewhat compressed doesnt change this close relationship.
But that's essentially what LLMs are doing, lossy compression of the entire web
If they had announced this sooner hardly anyone on the internet would have noticed. Props to them for adding it now.