Zuckerberg Approved AI Training on Pirated Books, Filings Say
news.bloomberglaw.com
news.bloomberglaw.com
If u remove morals from the equation, nearly every CEO would have made that same decision if in that position.
IMHO this is a moral win on Zuck side.
There is wrongdoing and there is obvious evidence that you known what you’re doing is wrong. That really limits options on Defence
Too high? Straight to jail.
Too low? Believe it or not, straight to jail.
During the early Llama 1 days The Pile dataset was in heavy use by many. Bit later people figured out that a subset of it - Books 3 - was especially problematic.
I’m guessing all the big houses threw that piece out in later models since it’s extra radioactive
- if you are an individual then it's called "pirating copyrighted work"
- if you are a multi-billion dollar corporation then it's called "use of uncleared material for training"
Zuckerberg murdered some old library books to train a model. Zuckerberg genocided training data!
Heck, everyone who read your comment here stole it. I'm so sorry for your loss.
This is just that with lots of levels of indirection.
Will he be allowed to lead Meta if convicted as criminal?
#dataPoisoning
"Your LLM, Captain Zuckerberg."
"Oh, thank you very much!"
This isn't the 90s. Computing isn't about discovery, not in the big leagues. Its about grinding up authenticity and feeding it into a machine to convert it into shareholder value.
If they want the value, let them pay for it or release the models open source for all to benefit.
There is a theoretical implementation of copyright that is your friend.
The realities of the laws as implemented today are abusive and hostile.
Maybe. I think it means we're at a spot where I'm not reasonably convinced that existing copyright laws are actually better than the free-for-all you're describing.
I'm definitely with you that it'd be ideal if we had a way to handle direct plagiarism like described in your comment (Although if you dropped the "put my name on it" part, I don't really see much issue).
But we also have all the fun today of companies using copyright to silence critique, shut out competitors, take educational information offline, demonetize videos they don't like, and otherwise absolutely abuse the hammers copyright law has given them (often - with no reciprocal hammer to stop this type of abuse).
And that's not even getting into the discussion of whether or not 70+ year old characters and stories should be available for modern authors to reuse and reinterpret. (even more egregious when you consider the vast majority of those tales are direct reinterpretations of older stories themselves...).
Or we can discuss whether it should really be legal to sell electronics hardware that has digital locks inside it that not only am I (the legal owner) not given the key to, but for which it is literally illegal for me to even attempt to open.
----
So basically - If I had to pick between "You'll own nothing and rent everything, and all public discourse is subject to DMCA strikes or other removal" vs "no copyright"... My vote is for "no copyright".
But the reality is I think we can strike a much better balance than those extremes, we just can't do it without upsetting large profit streams for existing, very wealthy and entrenched, entities... and usually that doesn't happen without tearing things down first.
(Of course, I'm using "the 1%" rhetorically, it's really more like 0.01%)
As a society, we all clearly benefit from fair use far more than we benefit from members of the copyright cartel buying another mansion or private jet.