Can you imagine Uber or open AI trying to get all the proper rules out the gate? It would either be a non starter or extremely expensive. And at the end of the day it’d be a watered down version of what’s possible. The power is all in the established hands. Perhaps the AI is trained on newspapers but not movies.
By doing what they did, they showed what’s possible and how excited users are. Now there’s a balance of power and both sides can come to a compromise.
Additionally, if you ask for permission people have to imagine how things will be and for something innovative that’s hard. Can you imagine open AI explaining to the New York Times what chat gpt was going to be, before it was released? At best they’d have some kind of approximation or worst just not understand and seek to kill it.
I would’ve done the same thing as open AI because it’s the most valuable thing, there’s enough legal grey area to not be flagrantly in the wrong, and things will get sorted later. Same thing with Uber - technically it was “ride sharing”, someone is going from point a to b and you as passenger just happen to join them, but obviously quickly expanded beyond that
I agree.
> Now there’s a balance of power and both sides can come to a compromise.
The only side that needs to compromise is OpenAI (et al). They used someone else's property without permission or compensation for commercial benefit.
OpenAI could have used this ignorance to their advantage! (Prior to ChatGPT release)
They could have went to NYT pre-2023, said "Hey, can we pay you $100k to come to your offices and run some machine learning algorithm on your content library? We can even do it in an air gapped clean room because we don't need to take copies of your content home with us" -- they very well may have said "Sure, what's the harm in that, come in next week"
And with Sonarr, Radarr, Lidarr, Jellyseerr, Prowlarr, flaresolverr, VPN, and qbittorrent - you can automate a request->watch pipeline for your users.
And if you use distributed block storage, you can make this scale pretty damn big.
(Prowlarr scrapes torrent sites for viable torrents. Uses Flaresolverr to bypass cloudflare shit.
Sonarr Lidarr Radarr - categorizes and prepares downloads for movies, tv shows, and music and submits them to qbittorrent.
Jellyseerr is a request dashboard that works alongside Jellyfin that allows users to request stuff.
VPN/Qbittorrent is a combined tool thats used as a download-endpoint for the whole thing.)
The Arrs suite is definitely a thing, as are the other tools I talk about.
There's a reference docker that automates this whole toolchain https://github.com/AdrienPoupa/docker-compose-nas
Whats nice is that Prowlarr will scrape all the torrent sites for goodies to index, and your users in jellyfin can just request, and it just shows up. You can even have a discord channel and the tools will announce when stuff's done.
Piracy is 10000x better than paying for continually worsening streaming (pile of shit).
Edit: And frankly, the billionaires copy with little ramifications, and add paywalls. I'll copy as well. The big difference is I do it for free.
If you've ever gone down this path, it's a great way to see the value proposition in a service like Netflix or Prime. It's also great to highlight where the value gets diminished because of the legacy media dark patterns like migrating content licenses between providers and region restrictions.
Our current copyright regime is what happens when legislation is pay to play. There are no conscientious, well informed representatives in charge looking out for the welfare of their constituents, but they're more than happy to dance on corporate strings, as long as the strings are paid for. No rational, well informed society would consent to being exploited by an industry like American legacy media, but here we are.
Piracy is an act of civil disobedience. Nobody gets hurt, and the sooner we come around to a better informed, more rational, less exploitative legal regime, the sooner we'll have more time, money, and energy to spend on things that matter. Instead of spending it on keeping the already wealthy even more flush with unearned cash.
Imagine if Netflix or Prime didn't have to contend with the old Cable conglomerate licensing schemes and could both have copies of content available for viewing - you could choose the objectively, technically better service, or you could choose the service that you simply liked better than the others.
Instead, we're incentivized to pay over and over and over again for the same things, with no value being provided except to the gatekeepers.
If something is published to the internet that is not behind a paywall, it's been broadcast to the universe. Just like google is able to index things, machine learning models fall squarely under fair use in the technical and technological sense.
NYT and other legacy media companies don't give a flying rat's ass about the ethics of copyrights or the rights of their journalists. They desperately want a slice of the profits and are projecting their entitlement through the use of lawyers and a horribly broken and corrupt copyright regime.
Their behavior is a lot like the classic Mafia extortion trope - "it'd be a shame if something happened to your fancy new AI techology."
As far as I'm concerned, legacy media can't die fast enough. These media conglomerates mostly made sense 70 years ago. In a world with the internet, they're as archaic and outdated as quill pens and vellum.
And yes, I have gone to the library with my laptop, and rented DVDs of stuff I couldnt find. I then ripped them, and returned them.
And trust me, nobody cares about someone copying some early authored VHS or DVD of a 1950's-1980's movie. Not all of us archivalist/pirates care about current movie of the week. There's a LOT of older and really neat stuff trapped on VHS as well.
The real worry with copyright law is the super-long terms create abandoned content nobody has, and is stuck on obsolete formats. Its not all about "current movie of the week".
I was trying to illustrate (and so was the original commenter) that many things only become a problem when you start making money by doing it, especially at massive scale with millions of customers.
E.g. I have no problem with someone making a LLM (or a Netflix clone) using copyrighted material if they are doing it for pleasure and not profit.
If you’re going to make a Netflix clone and share it with your friends, cool (would probably get upvoted as a Show HN even, like Popcorn app a few years ago). But if you’re going to do that and turn it into a billion dollar business, not so cool and you should expect to be legally challenged.
Copying is not theft, and I don’t think anyone would dare to claim otherwise.
Unfair use of someone else’s intellectual property, however, is.
IP fuels innovation and progress—the very same that gave us tech and LLMs. IP is a major factor that incentivizes us to make new things, as you can subsequently profit from the value you created for the world. Take it away and you get a society where barely anyone is motivated to bother (unless they are coerced into it, with predictable results).
I notice you say distributed block storage here but don't name a service like every other part of the pipe, is there something you use? I can't imagine you're storing 1000s of 4K movies and TV series in S3 and not paying 100s of dollars monthly for the privilege
Typically, the model is trained by "reading" each document only a single time. Should you have to license in perpetuity anything you read or see online that has a copyright because it changed your brain a little? If not, why should the model be treated any differently?
What about the Phi family of models that are not directly trained on copyrighted data but instead are trained on data produced by GPT-4, among other sources.
Would totally derail faux open-source LLMs.
Genuinely open-source LLMs woukd bw fine.
Because a model isn't real, it's not alive, it's not human, it doesn't know right from wrong, it's a software machine, nothing more.
You can disagree, but legally speaking it will be interesting to see the legal strategy OpenAI takes with the NYT. I'm almost certain that OpenAI won't be using the human brain argument in their defense. Legally, claiming your machine should be treated the same as a human is a losing argument.
First, we know that personhood cannot be reduced to the brain (cf. the gut flora studies and all that).
Second, more generally that claim would require solving what’s known in philosophy as the hard problem. The arguments that consciousness arises from a certain arrangement of entities we perceive in reality (or that it doesn’t exist at all) is specious: it implies that those other things exist in absolute sense, i.e. are the underlying territory, even though the only evidence for their existence is coming via consciousness. That’s a lot of entities to grant magical existence to out of nothing. Once you assume that those things are more likely to be a map, the question turns into what is the territory then, and the spotlights shift to consciousness (a.k.a. the only thing we can be absolutely sure to exist, as it is required for the “we” part to exist in the first place).
And yes, of course, if OpenAI uses the “sentient LLM” defense, all the power to them, but the next logical step would be to consider that industry based on forced labour.
If I use BT to download a music album I have bought off Bandcamp after Bandcamp goes out of business, that should be fair. If I set up a business that makes money by helping people download whatever albums they want, not caring whether they paid for them, that is obviously illegal.
This leads to my position on derivative models, if they are based on models that themselves are trained on copyrighted works without proper licensing.
Nice.