* http://www.entitree.com/en/family_tree/Elizabeth_II
* https://family.toolforge.org/ancestors.php?q=Q187114
Tools found on this page: https://www.wikidata.org/wiki/Wikidata:Tools/Visualize_data/...
---
Some SPARQL queries: https://www.wikidata.org/wiki/Wikidata:SPARQL_query_service/...
---
Out of topic: I wish wikipedia would provide an API to get the infoboxes (made using Lua or wikidata).
The easily parsable Infobox data can probably already be found in Wikidata (assuming there is a property).
The other way round seems better, but obviously too late.
Some wikipedia infoxes are based on wikidata. I can't find an example, but here some links:
* https://commons.wikimedia.org/wiki/Commons:Wikidata_infobox_...
* https://commons.wikimedia.org/wiki/Template:Wikidata_Infobox
* https://en.wikipedia.org/wiki/Template:Infobox_person/Wikida...
There are lexeme too, and it is not based on wiktionary. Search the prefix "L:" (without quote).
Example:
* https://www.wikidata.org/w/index.php?search=L%3Acat&search=L...
* https://www.wikidata.org/wiki/Lexeme:L7
Also, there are a lot of tools on toolforge.org. One is reasonator which produce sentences from a wikidata item: https://reasonator.toolforge.org/?q=Q1339
Perhaps, but I already know how to scrape HTML and I know the data I wanted to pull out was in there. I have no idea how to query wikidata and it could have ended up being a blind alley.
Also, it was only my reading your comment just now that told me wikidata was even a thing.
When I was analyzing Wikipedia about 10 years ago for fun and, later, actual profit. I did the responsible thing and downloaded one of their megadumps because I needed every English page. That's what people here are concerned about, but it doesn't matter for your use case.
To be fair, the original comment just made a valid observation in a casual way, he didn't criticize the approach of the OP, nor was he impolite.
But I know it's pretty common to see haters nitpicking things all around ;)
I ended up loading the full nightly db dump and filtering it streaming from the zip instead. Faster and it actually worked.
The code to do that is at https://github.com/boxed/relatedhow
Currently it doesn't support some SPARQL features, but I've found it to generally be quite a bit faster for most queries.
I only thought of it myself because you mentioned the problem with deducing which parent is the mother and which is the father, and I remember in wikidata those are separate fields.
I know javascript and had the pages at hand.
I looked at wikidata and some pages about, but still had no clear idea how to use it and no motivation to digg into it. Because js just worked with a small custom script, to retrieve some pages and data.
superior RDF triples are like martian language to millions of humans
over
> Instead of redirecting their efforts to a more general graph model which has actual hype and use by developers
neo4j is basically this. You can also load RDF into neo4j using neosemantics and query it using Cypher instead of using a conventional triplestore with SPARQL, which is nice.
RDF was designed primarily for data interchange and there's nothing that beats it at that.
And for the model: property graph. But yeah, enjoy your Stockholm syndrome with your model where reification is required to annotate an edge. Also even your nickname is an aknowledgment of RDF failure: named graphs (n-quads) were created because RDF triples aren't good enough for modeling data.
You're right about bioinformatics, but lets do a quick check on http://sparql.club/ on who else is looking for RDF/SPARQL specialists. Oh look: automotive industry, finance, publishing, medical, research etc.
In real life you use tools for both.
Yes, I made some improvements ( https://www.wikidata.org/wiki/Special:Contributions/Mateusz_... ).
But overall I would not encourage using it, if I would know how much work it takes to get usable data I would not bother with it.
Queries as simple as "is this entry describing event, bridge or neither" are requiring extreme effort to get right in a reliable way, including maintaining private list of patches and exemptions.
And bots creating millions of known duplicated entries and expecting people to resolve this manually is quite discouraging. Creating Wikidata entries for Cebuano Wikipedia 'articles' was accepted, despite that Cebuano botpedia is nearly completely bot-generated.
And that is without unclear legal status. Yes, they can legally import databases covered by database rights - but they should either make clear that Wikidata is a legal quagmire in EU or forbid such imports. But Wikidata community did neither.
So far I have not found way to achieve this without laboriously maintaining my own database of errata, and new exceptions keep appearing.
However, having read the article, they didnt have an easy time with scraping Wikipedia either.
So I'd probably still recommend people look into wikidata and SPARQL if they want to do this kind of thing.
Theres a few tools that generate queries for you, and some cli tools as well:
https://github.com/maxlath/wikibase-cli#readme
It makes Wikipedia better too, in a virtuous cycle, with some infoboxes like those that he scraped being converted to be automatically populated from wikidata.
The Wikidata folks are well aware of the limits on their SPARQL service. They just posted an update the other day:
https://lists.wikimedia.org/hyperkitty/list/wikidata@lists.w...
The SQL endpoint is at https://quarry.wmflabs.org/ however it doesn't have the actual data so much as metadata (mostly) so its not super useful.
Still, it was a cool article and a good example of scraping information.