Extracting data from Wikipedia using curl, grep, cut and other bash commands
loige.co
loige.co
Start with: https://en.wikipedia.org/api/rest_v1/page/html/List_of_Olymp...
Run this js one-liner: [].slice.call(document.querySelectorAll('table[typeof="mw:Transclusion mw:ExpandedAttrs"] tr td:nth-child(n+2) > a:nth-child(1), table[typeof="mw:Transclusion mw:ExpandedAttrs"] tr:nth-child(3) td > a:nth-child(1)')).map(function(e) { return e.innerText; }).reduce(function(res,el) { res[el] = res[el] ? res[el] + 1 : 1; return res; }, {});
The result is an object with the medalists as keys, and the count as values. JS objects are unordered so sorting is left as an excercise for the reader.
I can't help but notice a small bug.... Driulis Gonzalez for example has medalled 4 times, but your script gives his count as only 3. Similarly Gévrise Émane isn't listed by your script. Something to do with split tables cells I suspect.
Still, it's inspiring. I've often used bash and perl to scrape data from web pages. I'll definitely consider JS in future.
The only remotely sane way to do this is to use the Mediawiki API [3] to get the pages you want, then use an actual parser like mwlib [4] to extract the content you need. Wikidata and DBpedia are also promising efforts, but both have a long way to go in terms of coverage.
[1] http://stackoverflow.com/questions/1732348/regex-match-open-...
[2] https://www.mediawiki.org/wiki/Help:Extension:ParserFunction...
Parsing HTML with regexps is fine if you're just curious roughly how many images are in a page. It's great for quick command line experiments. It's just not good when you need to be "doing it properly".
But on the other hand, this is often how long term "proper" solutions are born — evolved from something cobbled together in couple of hours.
I learned that one after mocking up some UI screens using VB6 and then having to explain over and over again that no, just because we showed you some buttons on a page doesn't mean that the program (which had to be a Java applet, mind you) was "almost done."
Most alternative parsing libraries (at least all the ones I looked at, and that includes Wikimedia's own Parsoid!) don't bother with implementing all that complexity. Which means that they can be often tripped by sufficiently tricky markup; and this turns out to be a quite low bar.
Parsing MediaWiki markup properly is an insanity on a level comparable only with TeX and Unix shell scripts. Even PHP is saner.
FWIW, every sane extensions reports its existence to [[Special:Version]] which at least includes a list of extension tags at the bottom but you'd need to implement all of them in an alternative parser.
dbpedia is a triple-store that allows us to perform simple queries against wikipedia data like listing music bands based on a particular city:
SELECT ?name ?place
WHERE {
?place rdfs:label "Denver"@en .
?band dbo:hometown ?place .
?band rdf:type dbo:Band .
?band rdfs:label ?name .
FILTER langMatches(lang(?name),'en')
}
or queries that involve multiple subjects, categories etc.https://query.wikidata.org/#SELECT%20%3Fhuman%20%3FhumanLabe...
SELECT ?human ?humanLabel ?count WHERE {
{
SELECT ?human (COUNT(*) as ?count) WHERE {
?event wdt:P31 wd:Q18536594 . # All items that are instance of Olympic sporting event
?medal wdt:P279 wd:Q636830 . # All items that are subclass of Olympic medal
?human p:P1344 ?participantStat . # Humans with a participant of statement
?participantStat ps:P1344 ?event . # .. that has any of the values of ?event
?participantStat pq:P166 ?medal . # .. with the award received qualifier of any of the values of ?medal
}
GROUP BY ?human
}
SERVICE wikibase:label { bd:serviceParam wikibase:language "en". }
}
ORDER BY DESC(?count)
LIMIT 100
EDIT: The dbpedia search should be something like: http://dbpedia.org/sparql?default-graph-uri=http%3A%2F%2Fdbp... SELECT ?human ?count WHERE
{
{
SELECT ?human (count(*) as ?count) WHERE {
?event rdf:type dbo:OlympicEvent
{
?event dbo:bronzeMedalist ?human .
} UNION {
?event dbo:silverMedalist ?human
} UNION {
?event dbo:goldMedalist ?human
}
}
GROUP BY ?human
}
}
ORDER BY DESC(?count)
LIMIT 100This is a bad attitude to have for working with data processing, where QA is necessary and the accuracy of the output is important. A 50 LOC scraper with comments and explicitly-defined inputs and output from functions is far preferable to a 8 LOC scraper that those without bash knowledge will be unable to parse.
And the 8 LOC bash script is not much of a time savings as this post demonstrates; you still have to check each function output manually to find which data to parse / handle edge cases.
I have so many `zgrep ... | awk ... | sort` scripts that I legitimately couldn't tell you what they do anymore, just what the correct output is.
I really like the feeling of proving to myself that I know how to use my command line well, but in most cases I end up spending longer trying to remember the differences between OS X built-in `sed` and the `sed` I know and blah. Lots of wasted time.
I've started keeping a "scratch" git repo with scripts for one-offs, so the next time I say "oh I want to go through this CSV and run X on each line with Y" I can look at my old code and replace the necessary parts.
echo x |\
# Comment here
cat | # Or like this
cat
An interactive shell with the option interactive_comments set (the default) ignores these comments, which is useful if you wish to copy+paste+execute bits of scripts.I recommend mk (original make replacement for Plan 9, available for different OSes through Plan9Port [0,1]). It has a bit more uniform syntax, and can check that the dependency graph is wellfounded.
mk is not there yet.
and even if that is the case, I'm still better being a vim user at a system that only has vi and ed, then any other editor's user.
Well, that depends. Shell scripts can be commented, and they can be built progressively and interactively by building a pipeline. That's a great choice for a one-off task, and in my experience much faster than many other approaches.
It also works well for tasks close to the system. For example, our users are able to download large archives of data, and we keep over 99% of such downloads indefinitely. We delete downloads > 100GB after they're more than 6 months old.
With a shell script run by cron that's achieved with find + rm + curl|jq (to tell the API the download is deleted).
Make -I dir
or Make --include-dir=dir
Choose whichever one makes most sense.anyone who uses single letter options on a script should be punished.
single letters are for one time typing.
.
is portable, whereas: source
is not... which was a surprise. :)I've gotten surprisingly far using just curl/cat/cut/sed/awk/wc & friends. When I need to build on top of that, I go write a real program, but the UNIX-fu tells me what I need to build on top of.
> 8 LOC scraper that those without bash knowledge will be unable to parse.
And someone without knowledge of Node.js will be unable to parse the 50-line JavaScript monstrosity.
> And the 8 LOC bash script is not much of a time savings as this post demonstrates; you still have to check each function output manually to find which data to parse / handle edge cases.
It only took that long because he detailed every step as a tutorial. If you have basic literacy of the shell, it's no time at all.
To refute you, I decided to solve it myself, using only the "hint" at the beginning of the article to use ?action=raw to work with the source wikitext (I had not yet read the rest of the article; I had not seen his solution).
It took me literally 2 minutes: (Setting in url='https://en.wikipedia.org/wiki/List_of_Olympic_medalists_in_j... for readability on HN)
[2016-08-15 17:52] curl -s ${url}
[2016-08-15 17:52] curl -s ${url}|grep flagIOCmedalist
[2016-08-15 17:52] curl -s ${url}|grep -o '{{flagIOCmedalist.*}}'
[2016-08-15 17:53] curl -s ${url}|grep -o '{{flagIOCmedalist.*}}'|cut -d'|' -f2|sed 's/[[\]]//g'
[2016-08-15 17:53] curl -s ${url}|grep -o '{{flagIOCmedalist.*}}'|cut -d'|' -f2
[2016-08-15 17:54] curl -s ${url}|grep -o '{{flagIOCmedalist.*}}'|cut -d'|' -f2|sed 's/\(\[\|\]\)//g'
[2016-08-15 17:54] curl -s ${url}|grep -o '{{flagIOCmedalist.*}}'|cut -d'|' -f2|sed 's/\(\[\|\]\)//g'|sort |uniq -c
[2016-08-15 17:54] curl -s ${url}|grep -o '{{flagIOCmedalist.*}}'|cut -d'|' -f2|sed 's/\(\[\|\]\)//g'|sort |uniq -c|sort -n
You can see the only place I fumbled a bit was with escaping brackets inside of brackets in sed, which is admittedly a little wonky.Sure, for something that might need to run repeatedly, it probably doesn't handle future edge cases that might arise. But it's not production, it's a one-off exploring the data script. Not every one-off program you write needs to be production quality.
And being worried about "those without bash knowledge"... don't be afraid to use your operating system!
That said, this line of solutions has a big edge case: it relies on the editors of Wikipedia being consistent and formatting each row as one line in the source. If I were to solve this totally on my own, choosing to ignore the "hint", and work with the rendered HTML, and just do most of it with a nokogiri one-liner.
Knowing that there's an ?action=something to get just the page HTML without the navigation and such. I spent about 3 minutes finding the "render" action (documented here: https://www.mediawiki.org/wiki/Manual:Parameters_to_index.ph... ). Then anther 3-ish minutes poking around the DOM in my browser to get an idea of what I'm working with, and visually inspecting the layout of the article; each cell with a medalist has two links, the first to the medalist, and the second to the country. Then it took me a whopping 5 minutes to hammer out the rest of the one-liner.
(Similarly, url='https://en.wikipedia.org/wiki/List_of_Olympic_medalists_in_j... for HN readability)
[2016-08-15 18:14] curl -s ${url}
[2016-08-15 18:15] curl -s ${url}|nokogiri -e '$_.css("table.wikitable td a:first-child").each{|a| puts a.text}'
[2016-08-15 18:16] curl -s ${url}|nokogiri -e '$_.css("table.wikitable td a:first-child").each{|a| puts a.text}'|grep -vE '^[0-9]{4} '
[2016-08-15 18:17] curl -s ${url}|nokogiri -e '$_.css("table.wikitable td a:first-child").each{|a| puts a.text.sub!("\n", " ")}'
[2016-08-15 18:18] curl -s ${url}|nokogiri -e '$_.css("table.wikitable td a:first-child").each{|a| puts a.text.sub("\n", " ")}'
[2016-08-15 18:18] curl -s ${url}|nokogiri -e '$_.css("table.wikitable td a:first-child").each{|a| puts a.text}'|grep -vE -e '^[0-9]{4} ' -e '^details$'
[2016-08-15 18:19] curl -s ${url}|nokogiri -e '$_.css("table.wikitable td a:first-child").each{|a| puts a.text}'|grep -vE -e '^[0-9]{4} ' -e '^details$'|sort |uniq -c|sort -n
And most of it was fumbling around with thinking that the year/details links were one link with a newline in them, when they are in fact two separate links.For a grand total of 11 minutes; generously.
My point is: Don't be afraid to play around with your tools and data! Have fun! Not everything needs to be production quality! And being written primarily in bash doesn't necessarily mean that it isn't production quality!
https://github.com/wikimedia/parsoid
That is MediaWiki's official off wiki parser that can turn wikitext into HTML or HTML back into wikitext. It would be reasonably simple to hook into its API and use it for data extraction instead.
sed -n 's/.flagIOCmedalist|\[\[\([^]|]\).*/\1/p'
https://en.wikipedia.org/w/api.php
The article text is a raw blob of wikitext you have to process, but you don't have to go to stupid lengths trying to parse HTML without a browser.
It also didn't require having a JVM preloaded to make startup times acceptable during development (naming no other tools).
I do use shell tools to process data, a lot. They're particularly good for exploratory programming and initial analysis of new datasets.
Using the wikipedia Python library (to search for oranges :)
https://jugad2.blogspot.in/2015/11/using-wikipedia-python-li...
And there maybe libraries for other languages too, since the above library wraps a Wikipedia API: