> But, that little bit of data is a bit more data than existed before,
No, it’s not. Undigitized data is still data. Wording, style, nuance, type, artwork, binding, metadata, attributions, citability… all of these things are permanently lost, which is a fucking tragedy because they don’t have to be. Even if you’re not willing to take the time to scan every page in a cradle scanner, as rare books should be, (and don’t tell me they don’t have the money to get a handful of library interns to do this,) you can disbind the books and store them as they did in the Caselaw Access Project at Harvard Law. They removed the pages from the binding, scanned them on a high speed conveyor belt scanner which yielded full color 600 DPI jp2 images, placed the pages back in the binding like a folio that could be re-bound if needed, vacuum sealed them, and stored them in a salt mine. It’s not like it was slow, either — we did 40k in 18 months and we did take the time to scan the rare ones with a cradle scanner. And we did it all in less than open AI probably spends in a day on inference.
> So, in that sense, the knowledge is better being spread compared to copyright where the book stays in a warehouse until it is disposed of.
That’s a false dichotomy. Libraries exist for this exact reason, and their not already having a copy does not make “you snooze you lose” a morally acceptable strategy.
I’ve been pretty cool on the direction of SV for the past decade at least, but I am absolutely gobsmacked by the unbridled hubris of these companies over the past 5 years.