Open-source app releases first complete copy of English Wikipedia with images
xowa.sourceforge.net
xowa.sourceforge.net
Just to be clear, XOWA isn't an installer for MediaWiki, but it's own app. This allows it to avoid the dependency on the entire MediaWiki tool-chain (apache, php, mysql, MediaWiki). Unfortunately, this means that XOWA has to reproduce the same logic, which is quite a challenge...
I was already pretty far along with my own parser before Parsoid was usable though. (and my parser has its own DOM / hooks)
If you're interested in going this route, you can look at http://www.nongnu.org/wp-mirror/. This should build a local MediaWiki instance with one click. Keep in mind that it's a bit slow: it takes two days to build simple.wikipedia.org with images. In contrast, XOWA sets this up in about 30 min
And I thought ah yes 'Outer Space' that could be another location an offline version could help.
> Space required when completed: multiply the dump file by 2.5. For example, for English Wikipedia, the dump size is 10 GB. When done, your English Wikipedia will be 25 GB.
Ouch, looks like I won't be trying this here any time soon.
http://en.wikipedia.org/wiki/List_of_Pok%C3%A9mon_(252%E2%80...
In the end, you're still going to need about 25 GB for English Wikipedia. If you want something smaller, you can try one of the other wikis (for example, Wiktionary, Wikiquote, Wikisource, etc.) Each of these are generally about 5 GB.
Hope this helps.
[1]: http://aarddict.org
EDIT: Just read further down the comments and noticed the mention of Kiwix, apologies for the duplication.
For one thing, it has an Android app, and it's easy to put the whole thing on your external SD card. It also provides an index for full-text search.
However, I'll point out that Kiwix has not updated English Wikipedia since January 2012. Also, XOWA works directly with the Wikimedia dumps (http://dumps.wikimedia.org/backup-index.html) so it's (a) always up to date and (b) can work on any wiki (Kiwix needs to release the zim file first)
Also, XOWA can run from an external SD card (including FAT32 formatted ones)
Kiwix is a lot more polished than this, especially the user experience.
This looks interesting becuase reproduces mediawikis very well. The text-only dump imported pretty well once I figured it out.
The kiwix ZIM files are a pain to prepare, and use the public facing API that can be pretty unreliable, especially if your mediawiki hosting is of poor quality. It also forces you to hand curate what you export.
Using the direct XML dumps prepared server side is much cheaper, but then you lose the curation. The image situation is unclear here. There are some really confusing instructions about it, but I wasn't able to figure it out.
I'll keep using and recommending kiwix for now, but this could be promising.
Regarding images. There is some assembly acquired, but I tried to make the instructions as simple as possible. If you look at http://xowa.sourceforge.net/setup_simplewiki.html, then there should be two steps:
* Download the .7z file from from archive.org: http://archive.org/details/Xowa_simplewiki_2013-10-30_images... * Unzip the .7z file to your XOWA directory. If you're on Windows and have C:\xowa as your folder, you should get a file called C:\xowa\file\simple.wikipedia.org\fsdb.main\fsdb.abc.sqlite3 as well as many others
enwiki is a little more difficult, but only in that it requires downloading more files.
Let me know if you run into other issues. I'm going off to work now, but I'll check again later.
EDIT: I forgot to add that if you set up ImageMagick and Inkscape (installation instructions are on XOWA's Main_Page), you can download images dynamically for each article (i.e.: you don't need to download the entire image dump first)
I tried tar'ing up my images directory from the server, and unpacking them in a few locations on the filesystem that looked like likely places, but that didn't work. The filesystem layout was kinda confusing with the "user" and "wiki" separation.
How would one prepare a similar image database for an unsupported wiki? I expect this is a custom thing you prepared as opposed to the xml text dump which is a standard mediawiki dump format.
As to the imagemagik part, it doesn't work for an unsupported wiki. Also, it would be impractical for me to manually crawl my whole site triggering downloads of images, and even if I did that, it is unclear how to package and deploy it. The deployment needs to be completely offline becuase there is no Internet at the prison.
Overall, setting up one of the well known wikis is probably pretty smooth, but the private wiki requires a lot of technical knowledge about implementation details that make this tool impractical for unskilled users. Right now, kiwix deployment it is close to ideal. I just need to instruct the unskilled user to replace the ZIM file.
There is one small deficiency in the kiwix deployment in that the automatic index files are user specific unless prepared in advance and recorded in the libary.xml file, so in practice I had to prepare a script to make sure the index and library were right. The actual deployment is "copy zim files to this dir, then double-click on this script"
To answer your question, yes: the image databases were prepared with expectations of a standard Wikimedia wiki. These wikis have a standard file layout of wikipedia/wikidomain/thumb/hash0/hash01/name_of_file/thumbnail_file/; EX: wikipedia/commons/thumb/9/97/The_Earth_seen_from_Apollo_17.jpg/270px-The_Earth_seen_from_Apollo_17.jpg.
If you're using a MediaWiki installation, your files should be laid out similarly. You can change the XOWA config file to explicitly specify this WMF layout. XOWA allows the user to work directly with the WMF tarballs, so this should work for you as well. You can look at this thread for another user's attempts: https://sourceforge.net/p/xowa/discussion/general_archived/t... If you have questions, feel free to ask / post.
The other alternative is that XOWA should have the ability to read from a non-Wikimedia directory. Another user asked for this for his own private wiki: https://sourceforge.net/p/xowa/tickets/159/. In this scenario, you'd have all your files in some root directory (C:\images) and XOWA would index the directory and look-up the file by filename. You would probably need imageMagick and inkscape installed though.
Regarding your other point: I will probably centralize all the directories, instead of spreading them out between /wiki/, /file/, /user/. I had a reason for this layout, but it's causing confusion among a few users. You could always zip the files with relative paths, and instruct the users to unzip the zip. For example, the XOWA wikiquote package is one zip file: https://archive.org/details/Xowa_enwikiquote_2013-11-19_comp.... If you unzip it in the /xowa/ dir, it will automatically put all files into relevant folders
In the end, if you have a routine set up for kiwix, you're probably best sticking with it. Keep in mind that XOWA does offer some other nice features that you may / may not need. (editable wiki pages; Wikimedia Lua code). It also offers a lot customization. For example, one of the users added Mathjax to XOWA on his own. (he then proceeded to add a lot more: sorting / collapsing, wikidata skin, redlinks, etc.)
Let me know if you're interested, and I'll see what I can do to help. Otherwise, thanks for the use case scenario. It's definitely something I'll consider supporting in the future!
However it suffers the number one torrent issue: they do not tolerate change. This means that
- When an article changes, you need to generate a new torrent - When a new torrent describes the archive, it needs to be downloaded from scratch by all peers, so that the maximum number of peers are available for a newcomer.
I hope you'll understand that this is not the official way to distribute archives...
Edit: Kiwix looks nice too, thanks!
(I tried googling the topic, but I'm very lost as to even where I should start. The main question that I'm interested in could be summarized as, "How do we ensure our knowledge as a species/society is not lost in an unintelligible format?")
But this is useless to the post-apocalyptic hunter gatherer, civilization would have to be reestablished by then. And hopefully they at least understand the idea of writing words on paper and don't just worship it as a religious thing.
There has actually been work on doing that, attempting to mark radioactive waste dumps in a way even someone in a completely different culture from a distant future could understand. It's really interesting and there is pdf on it here: http://prod.sandia.gov/techlib/access-control.cgi/1992/92138...
If you are going to communicate with a completely alien civilization, that is one that can't even understand images, expressing information is even more difficult. My best guess is that you send a message with a really obvious pattern to it, then use that pattern as a basis for sending more information.
For example, send a ton of examples of simple code in a simple programming language, and their output. They can figure out what it means. Then encode your messages in the programming language somehow. Send a simulation of 3 dimensional space and little objects in it interacting, for example.
Somewhat related: http://lesswrong.com/lw/qk/that_alien_message/
It could possibly be made future-proof by using a flipbook format, assuming we can find a suitably long-lived material to print on.
I kept thinking about the "how to tell future people about radioactivity" example and I think it's hard to convey the actual effects, while it may be easy enough to convey that it's dangerous.
A movie would be much more apt at explaining what was there and what the consequences of irradiation are that a few stills.
Also, there was another page along the lines of what you're thinking: https://en.wikipedia.org/wiki/Wikipedia:Terminal_Event_Manag.... Although it was an April Fool's day joke, it does give some idea of what's involved.
For less advanced future generations, given the lack of a universal language I think the best you can do is store copies in many languages and/or translation dictionaries, and hope that at least one of them survives. http://rosettaproject.org/disk/concept/ is an interesting take on this.
In an odd way, we're going back to days where encyclopedias came in CDs. There was Encarta on a single disc, now we can have a lot more for around 80Gb ( http://xowa.sourceforge.net/requirements.html )
It would place more responsibility on the user to maintain their copy of the dump though.
It's a bit like maintaining your copy of an OS. You can stick to the "stable" branch or, if you're feeling adventurous, you can switch to "release". If you're really into the bleeding edge, you can go with the "nightly" build.
All-in-all, I really like this.
One concern I have is the possible increased bandwidth load for WP. Maybe you can include a small icon or notification to support it by donations. Couldn't hurt to have one there for yourself as well.
Anyway, thanks for the food for thought as well as your suggestion. I added donation links for archive.org and wikipedia tonight.
Since March 2013, inmates of the Bellevue prison in Gorgier, Switzerland can request access to an uncensored offline copy of the French Wikipedia, based on the Kiwix software. Internet access is severely restricted for the prisoners, most of whom serve long-term sentences [1]
[1] http://meta.wikimedia.org/wiki/Wikimedia_Highlights,_June_20...
- low-bandwidth availability, particularly in less-developed regions of the world
- censorship evasion
- security concerns. some users want to access Wikipedia without exposing their machine to the internet
There are probably a few others I'm missing....
Unfortunately, having your own copy of Wikipedia could also be used to enable censorship. For example, a fundamentalist school could have their own version of Wikipedia from which they've purged all articles about evolution, etc. Then they could configure their firewall to block the real Wikipedia.
By and large, for most private individuals, an offline app would allow them to evade censorship. I'd hope that this benefit outweighs the risk of the other's abuse.
A local HTTP service or desktop app, DVCS and indexer would do a fine job of this.
An abridged version would be a couple of gigabytes perhaps which isn't beyond the realm of possibility. That'd fit nicely on a smart phone/tablet and could be taken somewhere with less than adequate data connections (read most places on this planet).
Origin of birds: 8,506 in the last month.
Ogden L. Mills (secretary of the US Treasury under Herbert Hoover): 399 views in the last month.
What is in the public interest is not the same as what the public show an interest in. Page views won't necessarily help you filter Wikipedia...
https://en.wikipedia.org/wiki/Wikipedia:CSD
Just edit the page and {{db-nonsense}} {{db-test}} or {{db-vandalism}} as appropriate. :)
Thanks. :)
Though it is a reflective LCD and text-only.