Zuckerberg appeared to know Llama trained on Libgen
rollingstone.com
rollingstone.com
I'm more interested in how a for-profit corp decides to obtain a copy for development of a commercial product, and how they execute that ... whether they still have the data, and whether legal know about it :)
It's exactly not the kind of thing you can say you "found on a USB stick lying around in the car park"...
<glances at shelf with many, many external USB drives hooked up to a Pi 400>
Oh, really? :)
> "I think torrenting from a corporate laptop doesn’t feel right,” wrote one engineer in April 2023, adding a smiley face emoji. (A later email acknowledged that the “SciMag” data had indeed been torrented.)
It is the best at OCR though. Not many people are talking about that. It’s a very nice thing to know.
The data Google had was book scans, search engine indexing of arbitrary 3rd party content, and private email and documents they hosted.
A new chapter in "information wants to be free". Copyright was always an artificial restriction on human instinct. We now enter an age where piracy is keeping up with the Jones', and those who respect intellectual property choose irrelevance. Prosecution becomes impossible as the laundering grows ever more sophisticated. Adaptation and acceptance painful, but the only path forward .
GitHub was working on a feature which supposedly tells you which repository the snippet you just got is copied from, or IOW, which repositories have similar code, sorted by date. Effectively pushing the blame further on you by making you spend the time you just saved by investigating which repository provided the code you just got from "AI".
Supposedly, if the license is not friendly, you can delete the snippet and write your own version. :)
To be crystal, I am sympathetic with GPL and copyleft movements. I do think we've entered a time where it will be hard for license holders to enforce their rights. The laundering will only get more effective. And, per my original comment, competitive pressure will incentivize models to sail the most piratey tack.
Plus, the stack's latest version contains at least one GPL repository which their license tool failed to detect. So it's not something hypothetical in the first place.
It’s trained on 15T tokens. So how many did you provide that were genuinely novel? And how much money do you want? Like $5 from OpenAI? And $0 from meta since it’s open source?
I personally hope we can all get on the same team with AI and treat its advancement as scientific research for the betterment of humanity.
so I think $150,000 per copyrighted work ingested is fair
The range for that is huge though, it can be in the hundreds of dollars per work, or if the infringement is shown to be wilful then a judge can award up to $150,000 per work.
Are we suggesting that we should ditch creators' rights and instead value intellectual property along the lines of "I should be able to copy all your stuff as long as long as I copy lots of other stuff too, and give it all away for free or almost free...?"
Like, you get fined for speeding, but if you keep speeding you'll get you're license revoked, and if you keep driving after that you get jail time. The payment required is punitive, but the point is to stop you.
You will have to explain why Hunter Thompson copying every word of every Hemingway novel isn't copyright infringement, but a computer doing the same is.
1. Humans are not machines; arguments saying because a human can learn LLMs must be allowed to copy is not interesting.
2. Did Thompson publish the work? It sounds like you're referring to an activity Thompson did in private, to improve his skills as an author. Meanwhile, lawsuits are alleging that LLM services reproduce copyrighted materials.
3. What can be fair use at small scale is no longer fair use at large scale.
I make a high-res photo of a banknote. I can print that out at home. The bad stuff starts at the step after that...
https://www.reddit.com/r/graphic_design/comments/ah9s8n/trie...
It's fairly broken, but on balance it seems the creators are the ones getting screwed.
I did years of research in a scientific lab which resulted in <drum roll> 2 (yes two ... count them) peer-reviewed papers.
My colleagues and I did the work, wrote up the damned papers, yet to get them published we had to sign over copyright to what I'd now suggest is essentially a rent-seeking scientific publishing mafia.
All a long time ago, but I never had (and still don't have) the ability to either legally download or legally redistribute my own work...
The betterment of humanity seems to involve some parties making a ton of money while the people who provided the data apparently just need to be grateful.
You cannot make a story featuring Simba.
Just use Kimba the White Lion.[0]
s/AI/capital/.
It's painfully obvious that this is going to make material conditions worse for most people who use their minds to work instead of their hands. to these people, the "betterment of humanity" is a cruel joke.
It’s not about being paid for including their work, it’s about being compensated for having done so without permission. For crying out loud, they went out of their way to remove copyright notices from the pirated work.
> It’s trained on 15T tokens. So how many did you provide that were genuinely novel?
Then they can just take it out. And go ahead and take out every thing you didn’t have permission to include. What’s that? The model is now significantly worse? Yeah, these things compound.
> And how much money do you want? Like $5 from OpenAI? And $0 from meta since it’s open source?
No, they would’ve wanted for the work to not have been included without permission in the first place. Do you understand the world you’re advocating for? You’re arguing it’s OK for rich people to do whatever they want if they throw some scraps on the floor for you. Not everything is about money. Unfortunately there’s no other reasonable way (legal and non violent) to punish these infringers.
> I personally hope we can all get on the same team with AI and treat its advancement as scientific research for the betterment of humanity.
What you’re expressing is “I hope everyone will stop arguing and agree with me”. These moguls care about themselves, it is incredibly naive to believe they give a rat’s ass about “the betterment of humanity”.
I find it very unlikely that the commodification of knowledge work will be for the betterment humanity, I don't know if people are expecting here that just because the value of more people's labor becomes zero that we will do, what, do away with money? No, it will just mean that fewer people will have the chance to earn the right to use space and resources in a meaningful way.
There's no law to force it. So of course it won't be.
Even if there were a law to force it, how would you enforce it?
But for some reason, these companies think they don’t need to bother, and can just use everyone’s stuff.
Wait, I phrased that wrong. For a very good reason based on long precedent, these companies know that IP law is a tool to be used by big companies against individuals and sometimes other big companies, but never by individuals against company, so they know they don’t have to bother.