Archive.org and California to start a data sharing and preservation project
blog.archive.org
blog.archive.org
How feasible would a reliable, distributed archive be; given how massive amount of data Archive.org has? After all, it was created precisely because the already-decentralized web was too ephemeral and unstable. I don't think decentralization is a panacea in this case.
Some problems are best solved by institutions.
The way to solve this is to provide the Internet Archive with enough resources to build out a globally distributed storage system. Could you hack something together using their torrent tracker for every item served? Yes. But you don't hack together something made to preserve digital human culture in perpetuity.
> The way to solve this is to provide the Internet Archive with enough resources to build out a globally distributed storage system.
Yeah, I agree. I do see a space for other appropriate institutions (such as the Library of Congress, British Library, etc.) to pool resources and facilities with the Internet Archive to achieve that goal.
Ultimately, it'd be awesome to see each national library run a semi-autonomous IA copy that synchronizes with all the others, but can continue to operate independently (scrapers and all), if need be.
2016: http://990s.foundationcenter.org/990_pdf_archive/943/9432427...
more: http://990finder.foundationcenter.org/990results.aspx?action...
Getting the content coverage people sometimes assume we already have is another matter. Additional funding (thanks for you donation!) go towards additional crawling and keeping up with the endless treadmill of media types and protocols. Eg, headless browser crawling development and deployment to capture javascript-heavy sites (https://github.com/internetarchive/brozzler); this is much more expensive than "classic" crawling.
For more on increasing storage costs and the under-funded state of web archiving in general, I recommend David Rosenthal's blog, eg:
https://blog.dshr.org/2018/05/longer-talk-at-msst2018.html
https://blog.dshr.org/2014/03/the-half-empty-archive.html
Far more effective and robust than hoping the archive is "suck it up for us" is to upload snapshots/dumps/exports yourself! Anybody can create an archive.org account and upload content (recommend https://github.com/jjjake/internetarchive over the HTML form), within reasonable limits. Obviously, care needs to be taken to remove sensitive (and personal) information first.
The University of California is a part of the government of the State of California established in the State Constitution, whose governing body is comprised of 18 members appointed by the Governor and confirmed by the Senate, plus seven ex-officio members, three of whom are State elected Constitutional officers (Governor, Lt. Governor, and Superintendent of Public Instruction) and one of whom is the Speaker of the Assembly.
(That said, it is unusual and potentially misleading to refer to UC as “California”, but not because UC is actually separate from the government of the State.)
Though I've heard credible complaints from "copyright" holders vs archive.org.
But both data-privacy and copyright they try to create ownership of information and must do so through intrusive legal measures because physical nature makes is against it.
The project aims to demonstrate how members of a cooperative, decentralized network can leverage shared services to ensure data preservation while reducing storage costs and increasing replication counts.