Would anyone agree if you replaced companies with people in that argument?
Why shouldn't a company follow the same rules as everyone else just because the scale at which they're doing it is so large?
I'd argue a company doing something like this should be forced to buy the books NEW and benefit the authors, and if they're found guilty of copyright infringement they should be punished at a scale a few orders of magnitude larger than an individual would be.
> Before you can train an AI, you must light 1 million dollars on fire
If I want to train an AI, I probably need to spend a larger part of my budget as an individual to do so than an org, should I be given the resources for free or severely discounted because I want to make money out of it?
I suppose one _could_ argue in favour of such a practice if it was going to benefit society as a whole, but is it?
The best solution I can come up with would be a digital library where one org, say the internet archive has scanned everything once, then they're charge a licence fee to these orgs to ingest a copy, and the part of the payment goes to the author, no big wastage, the information gets archived and the orgs pay their share.
I’m sure there’s lots of unintended problems with this, but it does feel like a common base set of training data like this is exactly the sort of thing the government can and should do.
Oh but the reason is that they're now making $3 billion/year, partially because of those books. I see an argument for the inefficiency behind having to rescan books that are already scanned, but not the cost. If there was a way to buy pre-scanned books from Google Books or whatever then I somewhat see where you're coming from.
I argue that there were positive effects of Anthropic having to buy and scan physical books:
* The choices people made choosing which physical books to buy and scan helped make Claude what it is. Personally I sense a difference between Claude and OpenAI and Gemini, and part of it comes down to the choices they made in training material. Sorry to go on and on, but how many choices here were made because it was a rainy day and the trains were down, so an intern went to bookstore A instead of bookstore B?
* While buying the books used didn't help the authors it helped the struggling bookstores selling their books. Literal dollars into the hands of local workers. When I fast forward to today and see how LLM companies are literally stealing the energy from the communities their data centers are based in, and polluting them with shitty power plants I can at least think of that as one positive outcome, even if it only happened once.
As far as the 7 million+ books Anthropic didn't pay for, their series B in 2022 brought in $580 million. They could have afforded those books.
It would be one thing if they were buying "used" digital copies of the books, but the fact that this is only legal with scanned physical copies makes it extremely wasteful.
Copyright has been very silly in the digital realm from the beginning and is unlikely to get less unhinged from reality absent a major overhaul that makes it completely unrecognizable.
Is there some nuance to the law that allows them to scan/copy them if they're physical but not if they're digital?
A lot of digital copies are also DRM'd to shit - to obtain raw text usable for AI training, you'd have to break DRM. Which isn't that hard, on a technical level - but DMCA exists.
DMCA is a shit law that should have been dismantled two decades ago - but as long as it's around, bypassing DRM on things you own can be illegal. Scanning sidesteps that.
If no physical copies existed and there were only DRMd digital copies of everything, the companies scanning books for AI training would be forced to work out some deal with the DRM-overlords to have it removed for their use. That (I think) would be a net benefit as hopefully the authors would get paid too.
Within the bounds of personal use, copyright holders should have no say over what people do with media after it is sold. That goes equally when the entity that buys the media is a company rather than a person. The entire reason DRM is a problem is that it subverts that principle using technical means.
I'm totally in agreement with you, once we buy something, it should be ours to do with as we wish, company or person. DRM is the sketchy technical solution that doesn't really solve a technical purpose, it's easily broken, but serves a legal one; the act of breaking it is the legal issue.
I make my stance by avoiding buying DRMd content where possible; DRM free games and digital books, but it's not always possible to avoid, if I buy a BD, I can't rip it to my NAS without subverting the DRM.
Linux is also the only OS running in my home (on computers with screens and keyboards) so I mostly can't even legitimately play those DRMd things if I buy them, whether it's a BD, or Netflix in my web browser, or whatever else if I wanted to.
I'm very, very much anti-DRM.
EDIT: Typo
how can a company be covered under personal-use?
Which is the kind of thing you would expect it to do.
I mean, demanding you pay money to the source of data in your quest to create a monopoly you are pretty much guaranteed to abuse later on while becoming filthy rich is not exactly unfair.