I created books3 to help settle the question of whether AI companies should be allowed to train on books. The outcome of "it's okay to pirate books as long as you're only training on them" was a long shot, but it would've let individual hackers train their own AI models (assuming access to sufficient compute, which you can get e.g. via https://sites.research.google/trc/about/).
Now we're in a world where you have to have dozens of millions in capital to do substantial work.
I heard at one point Eleuther was gathering public domain training data. I wonder if they ever built a corpus large enough so that training on books doesn't really matter...
A Gaylord is a type of box that fits on a pallet. There are multiple ways to palletize products, like shrink wrapping or metal banding
(an observation, not agreement)
Let them take on the liability
> Anthropic spent many millions of dollars to purchase millions of print books, often in used condition. Then, its service providers stripped the books from their bindings, cut their pages to size, and scanned the books into digital form — discarding the paper originals. Each print book resulted in a PDF copy containing images of the scanned pages with machine-readable text (including front and back cover scans for softcover books
That would be pirating. So your complaint is that they didn't do more piracy?
To be clear, this isn't a problem with the court process. Everything here appears perfectly in accordance with the law. It's just an absurd state to be in.
the people operating frontier labs are bad people they cannot be trusted in any way
the best solution to them would be to send them to monster island (even though it's really a peninsula)
This is worse than pirating books to an absurd degree, it's almost a parody - the company that slurps all human knowledge ends up not only metaphorically, but also physically destroying those books, like an information vampire.
Authors don't even receive any financial compensation if the books were bought second hand, either. There's no benefit in doing that. (Not that making one final sale of a hardcover copy would make any difference though)
If Anthropic were at least buying ebooks, this insanity wouldn't need to happen. Unfortunately there is no bulk rates for buying millions of ebooks like you have in the used book market
There is less publicly available knowledge now on the Internet than there has been 3 years ago.
There's such a thing as fair use and digitizing privately owned printed material is absolutely legal... including for corporations.
But if you don't want to ban them, telling them to buy one book of each, likely second hand, is complete pettiness that resulted in destructive scanning of millions of books, many of which were already practically available in digital form.
That's because the judges are supposed to rule on questions of law (ie. "is AI training fair use?"), not whether they think AI's good or not.
at least in Player Piano they paid the workers who made the cassette tapes that made the robots work.
our current LLM overlords demand that they be able to basically steal the sum total of all human knowledge so that they can sell it back to us at a rate they set.
they should have been shunned by society and made penniless when they first announced their goals but we have a bunch of deeply misanthropic people who have money and want to make a world where computer slaves do their bidding.
So, mind sending me your bank account information? I'll promise to make good use of it.
You get a service. The service is using their compute power to run a model and their scientists to build the model.