So it's not like you need to crawl the sites to get content for training your models...
So it's not like you need to crawl the sites to get content for training your models...
It's not clear which files you need, and the site itself is (or at least, was when I tried) "shipped" as some gigantic SQL scripts to rebuild the database with enough lines that the SQL servers I tried gave up reading them, requiring another script to split it up into chunks.
Then when you finally do have the database, you don't have a local copy of Wikipedia. You're missing several more files, for example category information is in a separate dump. Also you need wiki software to use the dump and host the site. After a weekend of fucking around with SQL, this is the point where I gave up and just curled the 200 or so pages I was interested in.
I'm pretty sure they want you to "just" download the database dump and go to town, but it's such a pain in the ass that I can see why someone else would just crawl it.
More recently they starting putting the data up on Kaggle in a format which is supposed to be easier to ingest.
Also make a cname from bots.wikipedia.org to that site.
1. Go to https://dumps.wikimedia.org/enwiki/latest/ (or a date of your choice in /enwiki)
2. Download https://dumps.wikimedia.org/enwiki/latest/enwiki-latest-page... and https://dumps.wikimedia.org/enwiki/latest/enwiki-latest-page.... The first file is a bz2-multistream-compressed dump of a XML containing all of English Wikipedia's text, while the second file is an index to make it easier to find specific articles.
3. You can either:
a. unpack the first file
b. use the second file to locate specific articles within the first file; it maps page title -> file offset for the relevant bz2 stream
c. use a streaming decoder to process the entire Wiki without ever decompressing it wholly
4. Once you have the XML, getting at the actual text isn't too difficult; you should use a streaming XML decoder to avoid as much allocation as possible when processing this much data.The XML contains pages like this:
<page>
<title>AccessibleComputing</title>
<ns>0</ns>
<id>10</id>
<redirect title="Computer accessibility" />
<revision>
<id>1219062925</id>
<parentid>1219062840</parentid>
<timestamp>2024-04-15T14:38:04Z</timestamp>
<contributor>
<username>Asparagusus</username>
<id>43603280</id>
</contributor>
<comment>Restored revision 1002250816 by [[Special:Contributions/Elli|Elli]] ([[User talk:Elli|talk]]): Unexplained redirect breaking</comment>
<origin>1219062925</origin>
<model>wikitext</model>
<format>text/x-wiki</format>
<text bytes="111" sha1="kmysdltgexdwkv2xsml3j44jb56dxvn" xml:space="preserve">#REDIRECT [[Computer accessibility]]
{{rcat shell|
{{R from move}}
{{R from CamelCase}}
{{R unprintworthy}}
}}</text>
<sha1>kmysdltgexdwkv2xsml3j44jb56dxvn</sha1>
</revision>
</page>
so all you need to do is get at the `text`.I know there are now a couple pretty-good wikitext parsers, but for years, it was a bigger problem. The only "official" one was the huge php app itself.
I wrote another Rust library [1] that wraps around `parse-wiki-text-2` that offers a simplified AST that takes care of matching tags for you. It's designed to be bound to WASM [2], which is how I'm pretty reliably parsing Wikitext for my web application. (The existing JS libraries aren't fantastic, if I'm being honest.)
[0]: https://github.com/soerenmeier/parse-wiki-text-2
[1]: https://github.com/philpax/wikitext_simplified
[2]: https://github.com/genresinspace/genresinspace.github.io/blo...
Crawling is more general + you get to consume it in its reconstituted form instead of deriving it yourself.
Hooking up a data dump for special-cased websites is much more complicated than letting LLM bots do a generalized on-demand web search.
Just think of how that logic would work. LLM wants to do a web search to answer your question. Some Wikimedia site is the top candidate. Instead of just going to the site, it uses this special code path that knows how to use https://{site}/{path} to figure out where {path} is in {site}'s data dump.
Sounds like the problem is not the crawling itself but downloading multimedia files.
The article also explains that these requests are much more likely to request resources that aren't cached, so they generate more expensive traffic.
A torrent of all images updated once a year would probably do quite well.
For those unfamiliar: The images that are marked NonFree must be smaller than 1Megapixel. 1155 X 866. In practice, 1024x768 is around the maximum size.