Internet Archive Scholar
scholar.archive.org
scholar.archive.org
I've started saving the html (including the css seems like too much overhead, and often it's incomplete or relies on downloads still - and screenshots are not searchable) of every interesting article I find online, downloading quite a few videos too with yt-dlp. I'd long copypasted all interesting comments into a txt file, but now it seems like data hoarding's the way to go - at least in moderation, focusing on things I'll actually refer back to.
I remember 15 years ago, discovering pdf dumps on random sites like a kid in a candy store. Perhaps it'll be like that again, with people presenting museums of their favorite old pages.
There's also plent¥ of other similar tools:
This would be fantastic, being able to browse a curated internet made of accumulated lists from other trusted users on the net, similar to how ad blocking lists work today.
You are genuinely trying to steer the internet into what it used to be: a museum of knowledge and expert discussion.
Edit: Ah but wow, Polyform license. Huh.
>22120 archives content exactly as it is received and presented by a browser, and it also replays that content exactly as if the resource were being taken from online.
I've been looking for this for a really long time.
https://www.reddit.com/r/linux/comments/coazye/what_does_rli...
First, my licensing protects my rights as the creator of Diskernet. I want to ensure that my hard work is not used or modified without my permission, especially in commercial settings. I believe this is a fair and reasonable request.
Second, my licensing allows individuals to use Diskernet for free for personal use. This means that anyone can download and use the tool to improve their own browsing experience, without any cost or obligation.
Third, my licensing allows businesses and organizations to purchase a license and use Diskernet for their own purposes. This allows me to continue improving and supporting the tool, while also providing businesses with a valuable tool for archiving and organizing their online content.
I understand that the Polyform Strict License may not be the most common licensing approach, but I believe it offers a good balance between the interests of the open source community and my rights as the creator of Diskernet. I hope you will consider giving my tool a try, and I believe you'll find it to be a valuable addition to your online browsing experience.
The readme clearly links where to buy a license for that.
If they brought back MHTML saving support it'd be a great win.
(In one case I was able to give back: bundling up the several scans an author had of a half-century old paper from their student days into a single, hopefully cromulent, PDF)
Edit: recall also that accepting that links are one-way and might be dead was the key simplification that allowed HTTP to take off after prior attempts at hypermedia had failed.
2 spinning rust drive can store the library of congress. ~ 2,000 drives would store the web (1). How many millions of these drives get manufactured per year? Our technology systems are failing us - all those words are being lost, like tears in rain.
(1)Back of the envelope estimation: https://www.worldwidewebsize.com/ ~ estimates 50 billion websites, with some estimates ~ 6 pages of information per website. Let's say 1mb per page on average. So ~2,000 drives would store the entire web.
But stuff goes missing at Wayback because people don't agree to their pages being backed-up. Copyright, whatever. So it's like Global Heating, the tech is there, but people just can't agree. So 'pirate' backer-uppers go to jail. And island-nations and expensive ocean-side properties are being submerged. So it goes.
1. https://addons.mozilla.org/en-US/firefox/addon/markdownload/
When I did a major college project in 2003, I made sure to make pdfs of any academic article that I referenced. It actually saved me, because some articles disappeared by the time I went to revise my references.
I do similar. I've had https://github.com/ArchiveBox/ArchiveBox bookmarked for a while as something to try better organise all that, but like a great many things I haven't go around to it yet.
So probably not one for my use case.
How can I go about creating this? Are their off-the-shelf solutions, will I need to say combine scrapy with elastic search? The links in this thread look promising.
Somehow it's the IA's job to fix problems that we all know are problems, sadly.
Your Donation Will Be Matched 2-to-1! [...] Right now, we have a 2-to-1 Matching Gift Campaign, tripling the impact of every donation. (from the home page)
A huge percentage of the operating budget is from small donors. The funding is preposterously small compared to other public-service public interest such as Wikipedia.
A lot of us take it for granted and assume there is e.g. support from FAANG companies proportionate to the degree they lean on it.
This is 100% NOT THE CASE.
Please advocate for recurring institutional donations from your firm. The audience reading this has a lot of voice in a lot of organizations who could without a though sign up to make annual 10K, 100K, 1M donations...
...and essentially, none do.
Please help change that!!!
Internet Archive Scholar - https://news.ycombinator.com/item?id=26419782 - March 2021 (3 comments)
Internet Archive Scholar: Search Millions of Research Papers - https://news.ycombinator.com/item?id=26401568 - March 2021 (47 comments)
Internet Archive is a great resource, however, it should get state funding as it provides a fundamentally important archival service. It's too bad it has to rely so much on private philanthropic donations (although state support comes with possible political interference, i.e. censorship of material that some politician doesn't like, maybe that's less of a problem with private donations, although then you could have some billionaire doing the same thing).
Indeed scholar.archive.org does not currently use citation count in search rankings. We have a decent citation graph, which we are working to expose in scholar (it is visible in fatcat.wiki today). Would probably only ever use citation count as a weak boost in search rankings (eg, "any citations at all", "more than 25 citations" as boosts, nothing beyond that), don't want to create too strong a feedback loop influencing future citations.
scholar.archive.org specifically was partially funded by the Mellon Foundation (and partially through donations and other service revenue). IA overall has diverse funding, including grants and service revenue from the USA (Library of Congress, IMLS, etc); other national governments (paid crawl services); foundation grants; universities and libraries (crawl, preservation, and digitization services); and of course general donations. The last category of course has the fewest strings and lets us pursue new projects which might be hard to get traditional funding for. Remember that the whole premise of web archiving was considered radical and quixotic at the beginning!
(source: I work at IA on scholar)
Google Scholar on the other hand brings me to paywalls, even though the articles are so old they should be out of copyright.
[1] - https://scholar.archive.org/search?q=%22Frederick+G.+Scovel%...