How to download all of Wikipedia onto a USB flash drive
planetofthepaul.com
planetofthepaul.com
As I recall, there were several clever things that the app did to reduce the size of the dump; many stub/redirect articles were removed, the formatting was pared down to the bare minimum, and it was all compressed quite efficiently to fit in such a small space. Patrick gives more technical detail on an earlier version of the app's homepage: https://web.archive.org/web/20080523222440/http://collison.i...
In retrospect I do kinda miss _not_ having cell reception on vacations, as it was easier to disconnect from stuff.
In retrospect I do miss _not_ having the internet or cell service at all. It's part of why I like to watch shows and movies from the 80s or 90s. It's funny because I love technology as much as anyone else on HN, but at the same time my idea of the perfect retired life is one that is almost entirely offline.
(As I check HN from the bathroom, which would’ve been a childhood dream)
I asked the train staff about WiFi onboard, and they said that they didn't have it and the preferred it that way. People take their train, not to get from A to B, but to disconnect and meet other people, read books, or watch the beautiful landscape going by. If people had Internet access, they'd be glued to their devices and wouldn't meet fellow travelers, and that's the magic of this line.
(They even said that their corporate management wanted WiFi in the trains so that the staff could digitize a lot of their paperwork, but the crew was resisting it because they believed it was exactly the lack of connectivity that keep people taking the train for vacation.)
What's preventing you from not taking a phone with you on vacation?
At the time i was super into fan stories about Ultima Online like PK Ghost and things and would download that and we would all pass the Handspring around to read them.
Dang this actually brought up some good memories for me thinking about.
in fact i invested in avantgo right before the crash. i like to think it was the $500 or so i invested that truly precipitated the original year 2000 dot com bubble crash.
Useful for when I have been traveling and need GPS nav and there isn't cell service.
I wish it was a little more obvious how to search it, or what all the variations mean, but it looks like a valuable resource.
It is worth noting that Kiwix works on multiple OSes and on phones and has a wifi hostspot version (that you might run on an raspberry pi, for example). Internet-in-a-box similarly works as a wifi hostspot for ZOM archives.
Lastly, it is worth mentioning that there are tools for creating your own ZIM files; it looks like the most straightforward way is to take a static website and use a utility to convert it into one self-contained file.
I'm looking for a way to archive all the websites from my browser bookmarks and then download them for offline use.
ArchiveBox[1]: Pretty much a self-hosted wayback machine. It can save websites as plain html, screenshot, text, and some other formats. I have my bookmarks archived in it and have a bookmarklet to easily add new websites to it. If you use the docker-compose you can enable a full-text search backend for an easy search setup.
WebRecorder[2]: A browser extension that creates WACZ archives directly in the browser capturing exactly what content you load. I use it on sites with annoying dynamic content that sites like wayback and ArchiveBox wouldn't be able to copy.
ReplayWeb[3]: An interface to browse archive types like WARC, WACZ, and HAR. The interface is just like browsing through your browser. It can be self-hosted as well for the full offline experience.
browsertrix-crawler[4]: A CLI tool to scrape websites and output to WACZ. Its super easy to run with Docker and I use it to scrape entire blogs and docs for offline use. It uses Chrome to load webpages and has some extra features like custom browser profiles, interactive login, and autoscroll/autoplay. I use the `--generateWACZ` parameter so I can use ReplayWeb to easily browse through the final output.
For bookmark and misc webpage archiving then ArchiveBox should be more than enough. Check out this repo for an amazing list of tools and resources https://github.com/iipc/awesome-web-archiving
[1] https://github.com/ArchiveBox/ArchiveBox [2] https://webrecorder.net [3] https://replayweb.page [4] https://github.com/webrecorder/browsertrix-crawler
I'm going to explore all the solutions and start building my setup soon.
That and the TeXlive PDF manuals can get me through anything.
In those cases having local access to Wikipedia (and not necessarily just en; Kiwix has archives for all the languages) can be a great learning resource and reference.
https://wiki.openzim.org/wiki/ZIM_file_format
This all lets you squeeze English Wikipedia into 90 Gb. But it also makes it much more difficult to edit in-place, and, of course, no MediaWiki means that it cannot possibly work like git pull requests.
I've downloaded the entirety of wikivoyage for example.
But next time this happens I will have an USB stick with all the necessary knowledge. The definitions for voltage, current and frequency should however be printed out in case my laptop battery charge is insufficient for accessing the USB stick.
It worked offline and had images and I traveled to Peru with it and learned so much. Does anyone remember this sort of thing?
I’ve tried wix formatted copies and they do work but the experience on an offline ipad was simply better. Thanks in advance.
I believed it morphed into "Wikipedia for Schools" ^0 - possibly this ^1 is a comment about it?
0: https://en.m.wikipedia.org/wiki/Wikipedia:Wikipedia_for_Scho...
1: https://www.speedofcreativity.org/2008/11/11/wikipedia-to-go...
Edit: nope :)
[^1]: https://github.blog/changelog/2021-09-30-footnotes-now-suppo...
all the links always appear plaintext for me
by the way, do you still have an olpc? i never got to use one but remember seeing them. my one weird piece of similar era tech is a cr48, the early chromebook google gave away. I remember on the form for requesting them it asked what you would do with it. i responded “install linux on it” and they gave me one.
My neighbor years ago used to always chuckle at me using it with an Happy Hacking Pro keyboard because of the price difference between the two.
I used a similar approach for https://wikiscroll.blankenship.io
1. kiwix dump
2. unpack to HTML
3. process with cheerio to create json files
4. Create git repo and push to github pages
Works well for infinitely scrolling content, it's just Math.random on top of static files.
> Its file compression uses LZMA2, as implemented by the xz-utils library, and, more recently, Zstandard. The openZIM project is sponsored by Wikimedia CH, and supported by the Wikimedia Foundation.
From https://www.kiwix.org/en/documentation/
File size is always an issue when downloading such big content, so we always produce each Wikipedia file in three flavours:
Mini: only the introduction of each article, plus the infobox. Saves about 95% of space vs. the full version. nopic: full articles, but no images. About 75% smaller than the full version Maxi: the default full version.
IIRC the articles were rather easy to download and convert even on my early 2000s netbook. The media (pictures, video, audio) though was painful to deal with, and it didn't take long to find out that Wikipedia without diagram s and figures was not a great experience.
I remember there was some work done to categorize articles like with the Dewey system, but so far, you can't really reduce the size of those exports.
Of course it would require a lot of work. Maybe it's already possible to categorize articles of they belong to a "portal".
But yeah, it doesn't seem the Wikipedia foundation really care about those kind of problems. To be fair they lack money.
Is TV Tropes available as a single file ZIM download?
they are not translations, they are completely different articles under the name brand and platform of Wikipedia
an entry that may be just a blurb in English may be one of the most comprehensive and fully fleshed out and researched entries on the site in German, for example
I don't think that makes any sense. By that logic any currently working device should be assumed to last another $currentlifetime. My 20 year old car is not gonna last another 20 years. My 10 year old laptop won't last another 10. If my car somehow did last another 20 years, it would not then make sense to assume it would still be running in another 40.
Makes more sense to look at all objects of the same class. If 75% of laptops are dead in 10 years and 95% are dead in 15, and your laptop is 10 years old, you can infer that 5 out of 25 surviving laptops will make it another 5 years, or 20%. (These numbers completely made up, just an example.)
The problem with the "half life" rule is that it corrects for this in the dumbest possible way, not only providing an inaccurate estimate for most of the object's life, but even getting the first derivative wrong for most objects. Usually, lasting longer does not make the expected remaining years of service go up, but the rule implies it does!
Take people for example. At birth, a woman in the United States has a life expectancy of 81. If she makes it to 60, she can now expect to make it to ... 85. Not a big change! Every year she lived (even her first), her remaining life expectancy went down, not up. See this chart I made comparing the life expectancy of people versus a theoretical "half-lifer": https://0x0.st/otZ_.png
You might want a device like that to have offline maps as well, especially as those are more likely to be immediately useful. The easiest way to get there is the OsmAnd app - like Kiwix, it does a number of tricks to compress things, so it's quite feasible to have a complete offline road and topographic map of US in your pocket.
(Note that Google Play Store availability on the device is immaterial, since Kiwix and OsmAnd are also available as as downloadable .apk, and are also listed in F-Droid store.)
[...]
>"The current Wikipedia file dump in English is around 95 GB in size. This means you’ll need something like a 128 GB flash drive to accommodate the large file size."
Great article!
Also, on a related note, there's an interesting philosophical question related to this:
Given the task of preserving the most important human knowledge from the Internet and given a certain limited amount of computer storage -- what specific content (which could include text, pictures, web pages, PDFs, videos, technical drawings, etc.) from what sources do you select, and why?
?
So first with 100GB (All of Wikipedia is a great choice, btw!) -- but then with only 10GB, then 1GB, then 100MB, then 10MB, then 1MB, etc. -- all the way down to 64K! (about what an early microcomputer could hold on a floppy disk...)
What information do you select for each storage amount, and why?
?
(Perhaps I should make this a future interview question at my future company!)
Anyway, great article!
https://library.kiwix.org/?lang=eng
I was looking at what other sites are available, and seems there are quite a few. Are there any specific ones apart from Wikipedia that HN readers would recommend?
Bandwidth is getting cheaper and cheaper, and arguably if you can afford to get that initial 100 GB Wikipedia dump, you can afford downloading it more than once (and vice versa, if you can download multi-gigabyte differential updates periodically, you can afford the occasional full re-download).
One application where I could see it making sense is a related project [2] which streams the Wikipedia over satellite: Initial downloads at this point probably take several days of uninterrupted reception.
[1] Google has once implemented a custom binary diff optimized for Chrome updates, but I'm not sure if it still exists. [2] https://en.wikipedia.org/wiki/Othernet
ZIMply on branch `version2` has worked pretty well for me [1]. The search works a lot better and it's really nicely formatted.
Kiwix divides the data into chunks and adds various indexes and stuff to allow searching data and fast access, even on slow CPU devices. But if you can live with slow loading, you can probably halve the storage space required, or maybe more.
So the best algorithm to use from there is starlit, with a compression factor of 8.67, compared to lzma in 2MB chunks which can only achieve about 4:1 compression.
The only reason it isn't eligible are compute constraints (and I don't think the hutter prize allows a GPU, which nncp needs for any reasonable performance).
Once you've done all that work, I'm not sure what you've actually accomplished. By the time any sentient being gets around to visiting JWST, I wouldn't be surprised if an unshielded commercial drive would be rendered totally unusable by radiation.
What good would an "offline backup of Wikipedia" do in that situation?
Wikipedia is good for one thing, and one thing only: getting some cursory knowledge on a topic you're unfamiliar with. It's the tourist map to the "sum of all human knowledge." If you expect to use it for anything else, you're asking too much of it.
Presumably you have seen a science fiction film before, use your imagination.
There are 14 people on the ISS. If they were the only ones left, they would certainly not have the breadth knowledge of a Wikipedia dump.
I like the idea of periodic Wikipedia moonshots, although the storage format is kind of an open question, I've wondered for awhile if a DVD made from e.g. quartz, platinum, and titanium might be up to the job.
A full backup would fit on 12 double-layer, single-sided disks; I'm being conservative and not using Blu-Ray numbers, since density and longevity are always somewhat in tension. Probably more expensive to put them safely on the moon than to manufacture in the first place.
In space you probably don't have to worry as much about minor surface scratches and oxidation, though. You just have to worry about outgassing and meteoroid impacts. Some of them you can stop, and some you can't. On the bright side, they're very rare.
I think common media formats like DVDs are designed with a lot of emphasis on speed, both of reading and of duplication. This compromises both density and longevity. If you, instead, allow yourself the luxury of FIB milling to write and an electron microscope to read, you can manufacture in any vacuum-stable material at all, and you can engrave your archival message with, say, 50-nanometer resolution. At one bit per 50 nanometers square, you get five gigabytes per square centimeter.
I think that with e-beam photoresist cross-linking followed by etching you get about 500 kilobits per second, and I think FIB milling is a little slower, so it might take a few weeks to make the copy — obviously unacceptable for a consumer CD burner but fine for periodic Wikipedia moonshots.
No, but FAT32 does. Exfat, on the other hand has a file size limit of 16 exibibytes. That, combined with exfat's cross-platform mounting (NTFS has a lot of limitations in this regard) makes it a superior formatting system for flash based offline file transfer.
On a network? Use zfs+ or something.
Would be interesting to see a graph of usb size easily available vs. Wikipedia dump size.
When I buy a storage device I usually have an intended purpose for that storage and would not like to have to delete all of the files that some manufacturer thought would be useful information but I would have to delete to make room for what I want.
* Getting a flash drive
* Formatting a flash drive (which includes a subsection on not formatting it but buying one that's already formatted instead, while there was a separate section just before this one on buying a drive)
* Waiting for a file to download
At the same time downloading both Wikipedia and Kiwix are in the same section. Then, installing Kiwix is in a section called "You're done" which isn't next to the section on downloading Kiwix.
All possible dumps: https://dumps.wikimedia.org/other/kiwix/zim/wikipedia/
https://wiki.kiwix.org/wiki/Content_in_all_languages
Why I know that? I wanted to travel as system administrator in some antartica base with a whole copy of stackoverflow with me.