The Internet Archive transforms access to books in a digital world
eff.org
eff.org
Except for that bit where somebody in power at the IA decided to ruin the future of CDL by declaring a free-for-all during the pandemic. I still want to know who it was that made a decision so mind-bendingly idiotic it put the very existence of the Internet Archive itself at risk. Of course our current copyright laws are stupid! That doesn't mean you as an organization with twenty million dollars in yearly revenue get to just ignore them and trust that having the moral high ground will protect you from the lawyers of an entire industry.
You beat me to this comment. No wonder they are now being sued. That decision may send them the way of Google Books. I hope it doesn't. Maybe in twenty or forty years we will see an accepted library like Spotify's arrival 15 years after Napster.
Archive didn't even have the moral high ground. Authors and publishers deserve to be remunerated for their effort and controlled lending supports that. When that's off the table you're just like libgen.
What surprised me the most was how many people strongly defended the IA's actions. In their heads, the IA could do no wrong. Or they failed to grasp the illogic of "public libraries are closed, so we're going to pretend we own their books and lend them out."
Even here on HN, there were some people who refused to acknowledge that what the IA did was both unfair to copyright holders and a poor strategy for trying to uphold CDL over the long term.
There are people here who refuse to acknowledge that what copyright holders do is unfair to everyone.
Copyright allows you to decide how your code is distributed. Software patents restrict innovation in a manner that was never intended to be possible.
Yes. All human intellectual work comes down to discovering a really big number. This labor has immense value. The data itself can be trivially copied and distributed worldwide at essentially zero cost.
Society needs to find new ways to value this human intellectual labor. New ways that don't depend on forcing artificial scarcity on the result.
> Not only is it an utterly helpful notion
Helpful to whom? To the monopolists, sure. Meanwhile everybody else has been robbed of their fair use and public domain rights.
> enforcing copyright in the 21st century helps us to protect our computing freedom
Absolutely not. The number one reason why computers are no longer under our control is copyright. It is impossible to stop copyright infringement if we can create and run whatever software we want. They must have mechanisms that stop us from running unapproved software in order to have any hope of enforcing anything. Otherwise it's trivial for computers to share data with each other, it's as natural as a river flowing.
If you unlike their software or hardware, then just stop buying it and buy or support free software and hardware. Freedom is not free.
That's a half-measure with limited impact. Why spend considerable resources freeing up a small number of works? The right thing to do is to abolish copyright straight up. At the very least, fix copyright by significantly reducing duration so that copyright ownership is reassigned to the public automatically at zero cost regardless of what monopolists want.
> just stop buying it
People won't have a choice. They'll either buy locked down hardware or they won't be able to consume anything. Streaming services have already started enforcing DRM rules that prevent owners of free software and hardware from consuming their content and it's illegal to circumvent this even if you pay for it.
But you can both be in favor of reducing the maximum copyright term AND believe that it is unfair to authors to just steal their works outright, which is what the IA effectively did through its "Emergency Lending" scheme.
> it is unfair to authors to just steal their works outright
It would have been stealing if those were physical books. This is data. There is no stealing, there is only copying.
They're not guilty of copying. There's nothing morally wrong about copying. They're guilty of not compensating authors. Creators deserve to be compensated but copyright is not the way it should be done in the 21st century. There needs to be a new way.
But don't find yourself actually needing to get in touch with a human there. I have a friend that inadvertently doxxed themself on a website they own; they removed the pages, added a robots.txt to disallow crawling and is trying for months to reach someone over there to just get that data removed from The Wayback Machine, to no avail.
Perhaps archive.org doesn't have a lot of staff or they simply don't care but I imagine these kinds of things happen often enough that they should at least have an automated process in place (and continuing to crawl a website that specifically forbids it through the standard robots.txt is a bit icky).
Opinions and practices are very different across different projects in the space. Wikipedia and WP commons, for example, have detailed policies for personal data. Commons is so careful about copyright, it sometimes can’t host a file that lots of commercial publishers freely use, such as (photos of) ads or logos that lack explicit licenses.
Both approaches have their pros and cons. I can’t help but think both projects would be improved by moving a bit into the direction of the other.
See: https://archive.org/details/softwarelibrary_msdos_games
> But don't find yourself actually needing to get in touch with a human there. I have a friend that inadvertently doxxed themself on a website they own; they removed the pages, added a robots.txt to disallow crawling and is trying for months to reach someone over there to just get that data removed from The Wayback Machine, to no avail.
That sucks. I've always been able to get a hold of a human the few times I emailed them - but nothing due to something like this.
> Perhaps archive.org doesn't have a lot of staff or they simply don't care but I imagine these kinds of things happen often enough that they should at least have an automated process in place (and continuing to crawl a website that specifically forbids it through the standard robots.txt is a bit icky).
Also, FYI they stopped honoring robots.txt directives: https://teleread.org/2017/04/24/the-internet-archive-will-so... — so if you're a webmaster that would like to exclude your website from the internet archive, you'll have to email them directly to make that request.
Good. It shouldn't be as simple as "put a robots.txt on your domain" to erase the entire publicly archived history of any sites ever hosted under that domain. This happened a lot when domains would expire.
Surprisingly, both Google and Bing removed the data within 1 week of making the request.
I find the idea you, in control of a server, can send someone data and then demand they erase it a bit ridiculous. If you want to control distribution then control it. Require a login to access any data you don't intend to be public.
The CDL program rests on a shaky legal theory that's already proven wrong by Capitol Records v. ReDigi. That is, it seeks to argue that as long as we have some DRM technology that lets us recreate the circumstances of something protected by first sale or copyright exhaustion (which, btw, is separate from Fair Use), then we should have first sale rights.
I like the spirit of this argument, but it will not persuade a court; because the current law doesn't work this way. First sale works specifically because the act of physical resale does not create a copy and does not tread upon the copyright monopoly even a little. Copying something and then destroying the original (as in ReDigi) does. If you want digital first sale you have to ask Congress for that.