Wikidata
wikidata.org
wikidata.org
Wikidata is one of the main data sources in my Conzept encyclopedia project: https://conze.pt (https://twitter.com/conzept__)
Is this an LOD based solution where we can plug in arbitrary KGs?
Did you fork the UI from some existing Knowledge Graph viz tools, or is it made from scratch?
Its just me (Jama Poulsen) currently. I hope at some point to be able to work with more people on this project and also to have a base for other archives / institutions (as a B2B product and service) to use the "Conzept UI" for their knowledge base. The Conzept framework is already pretty generalized, multi-lingual and customizable, but more needs to be done. I'm aiming to be able to do a first pilot at the end of the year (or later, depending on the progress). Feel free to send me a message if someone is interested in this.
There will be better user documentation coming (the current guide is a bit dated). Still thinking how to make that more modular, UI-integratable and maintainable.
The whole main app design and UI is from scratch, but many embedded apps are developed by others (I've been donating to some of them for their great work, but more needs to be done here IMHO). No papers or anything on all that currently.
Is the source under an Open-Source license? Can it be seen anywhere? Do you accept contributions?
The user base is still small, which is fine, as there are still many small issues to fix. I'm having a lot of fun developing this and hearing the feedback of users.
I started the project mostly for myself, to have a more integrated topic exploration tool, but things grew to also be of use to others. I think the main use case is just learning about topics of personal interest. I would like to experiment with some social features in the future (eg. ephemeral audio chat for topics).
Anything you would like to see?
I will readily admit this is due to my own ignorance; I just wanted to post a point in a Reddit comment and investing a lot of time to properly learn SPARQL wasn't worth the effort, but I wish the data was more easily accessible without a fairly large time investment. WikiData is basically the only structured source for these election results that I could find (the government publishes them of course, but in a PDF that's always different).
And I'm a programmer by trade! Non-technical people will have an even harder time.
WikiData is pretty good, but I feel it could be truly fantastic if the data was more easily accessible.
I made a script to get the highest mountains in norway and get them on a map in Jupyter - looks nice, includes pictures if there are etc, but it doesn't pick out all the highest mountains. Most of them, but some are missing.
Kind of applies to everything, no? :-)
But yeah, I don't think GraphQL is "bad", it's just hard to get started with (same applies to SQL really, although we're all used to that so we've forgotten about getting started with it). I want to use WikiData maybe ... once a year? Or less? I will probably have forgotten a lot of stuff next time I want to use it.
This is a Hacker News thread from March 2012 titled "Wikidata: The first new project from Wikimedia Foundation since 2006" that is interesting to revisit:
"A gentle introduction to the Wikidata Query Service"
https://blog.google/products/search/introducing-knowledge-gr...
"DBpedia (from "DB" for "database") is a project aiming to extract structured content from the information created in the Wikipedia project."
I think they acknowledge that info-box information is no longer good enough (both in coverage and completeness) and have handed over the moniker of the open-access KG to Wikidata. To their credit, I think its the right call.
src: I worked in a lab lead by DBPedia's founder (and one of the contributors to the original infobox to KB code).
[0] https://blog.liu.se/olafhartig/2019/01/10/position-statement...
^^This enables, amongst other things, fact's validity to be quantified. Instead of saying <barack obama> <president> <united states>, you can now say {<barack obama> <president> <united states>} <from> <DD-MM-YYYY>; <to> <DD-MM-YYYY>
[1] http://downloads.dbpedia.org/current/ (from 2018); http://downloads.dbpedia.org/ (see the folder titled 2016-10)
Please read it carefully before creating new items. If the item you want to create is about yourself, your new business or the song you just wrote, it likely does not meet the notability policy.
To me the greatest value of Wikidata was making me aware of RDF and SPARQL.
In most cases, if you are relying on data business needs, it would be best to maintain your own RDF dataset and host it either just on HTTP, or on something like https://dydra.com/.
WikiData deseperately needs RDF ingestion, and if this is made available (can be done outside of Wikidata) then it would be easier to periodically sync datasets with Wikidata.
On that note however, you could export all Wikidata triples you need and just host that on your own SPARQL server (e.g. Jena) or use it with RDF tools like rdflib.
It seems like magic, but it is possible to entirely outsource the infobox to wikidata. See for example https://fr.wikipedia.org/wiki/Mart%C3%ADn_Abadi whose infobox is created with the sole `{{Infobox Biographie2}}` line.
location(San Francisco (Q62), Geolocation{-122.4183, 37.775})
This requires that Geolocation be a special type like integer or float. VERY BAD. There would be a massive proliferation of such things.Alternatively Wikidata could do this as
Q95 (or whatever) = new object
longitude(Q95, -122.4183)
latitude(Q95, 37.775)
location(San Francisco (Q62), Q95)
This is the standard reified approach used by binary semantic networks and is ugly as sin, unnecessarily polluting the object space with little object poops. The clean way to do this would have been to declare location to be ternary, so we could say: location(San Francisco (Q62), -122.4183, 37.775)As for types, “GlobeCoordinateValues”[0], i.e., locations, form one of the types allowed for values, and consist not only of two integers, but indeed of four distinct values: latitude, longitude, precision, and the reference globe, since Wikidata does not limit coordinates to the Earth. There is also no “massive proliferation”, the data model[1] knows 12 types[2], of which four are four different kinds of texts (untranslated, monolingual, multi-lingual, and list of multilingual texts).
[0] https://www.mediawiki.org/wiki/Wikibase/DataModel#Geographic...
[1] https://www.mediawiki.org/wiki/Wikibase/DataModel
[2] https://www.mediawiki.org/wiki/Wikibase/DataModel#Datatypes_...
- "Borane" has two unlinked entries: Q127611 is "any chemical compound composed of boron and hydrogen atoms only" while Q15634214 is specifically boron trihydride. Both correct, but the latter should be labelled as an instance of the former.
- "Anomalocaris" has Q37395 for the "extinct genus of radiodon" (instance of "fossil taxon") but species (e.g. Q49557506 Anomalocaris cranbrookensis) do not link to it.
- There are no lists of links, for example Q936518 ("aerospace manufacturer") doesn't have a list of its instances (e.g. Boeing, Q66, or Arado Flugzeugwerke, Q624899).This is intentional, you can construct these via the query service. There's also an optional "gadget" registered users can add to their configuration, that can do this automatically when visiting a page for every instance of some "inverse" property.
You can query the database using SPARQL: https://w.wiki/3XgD
No, it should be labelled as an subclass of the former. Instance would be a specific piece of borane.
Perhaps some kind of synchronization would be feasible, where with standardized infoboxes relations between known items could be extracted when specified on Wikipedia, and plopped into Wikidata.
For myself, I looked around for a service that could present infobox-data from pages in a category as a table, with filters and sorting. But ended up just firing API requests and parsing the boxes.
Main website: https://www.openstreetmap.org/ Wikidata tagging documentation: https://wiki.openstreetmap.org/wiki/Key:wikidata
In general, Wikidata is not made up of discrete "datasets". Rather, each identifiable real-world entity, event or concept has a unique Q-identifier and a listing of "properties" that apply to that entity. So, to construct a dataset, you'd query for entities that participate in some set of properties you might care about.
understanding which tags from the ontology to use and then building the query.
I'm using Wikidata to automatically categorize visited websites and used programs for time tracking purposes.
For example, here's a query to get the entity (e.g. company) that has a specific domain (news.ycombinator.com), and get the categories that entity is in that that are a descendent of the "service on internet" category:
SELECT distinct ?service ?website_url ?outer_category ?outer_categoryLabel WHERE {
?service wdt:P856 ?website_url.
optional {
?service wdt:P31 ?inner_category.
?inner_category wdt:P279* ?outer_category.
?outer_category wdt:P279+ wd:Q1668024.
}
VALUES ?website_url { <https://news.ycombinator.com/> }
SERVICE wikibase:label { bd:serviceParam wikibase:language "en" }
}
Returns- social news website - online service - website
Relations have to be written as `wdt:P279` even though they have names [1], and the query engine just silently accepts lots of stuff it shouldn't. For example if above you do `"url"` instead of `<url>`, it will just not return any results because URLs is a separate type and URLs never match strings. And if an entity doesn't exist it will also just not return anything. Then there's this hacky "label service" thing that implicitly creates new output variables (outer_categoryLabel) to make it actually return text.
The UX of the query site [2] is also pretty bad.
I feel like Wikidata would be much more used if it had a easier query language. Or better, real bindings to common languages with full intellisense for properties etc). Something like this:
const service = wikidata.variable();
const inner_category = wikidata.variable();
const outer_category = wikidata.variable();
const results = await wikidata
.filter(service.official_website("https://news.ycombinator.com"))
.filter(service.instance_of(inner_category))
.filter(inner_category.subclass_of(outer_category).anyDepth())
.filter(outer_category.subclass_of("service on internet").anyDepth());
console.log(results.map(row => row.get(outer_category).getLabel({language: "en"})));
It's also really hard to find good answers about the SparQL language without reading hundreds of pages of dry documentation.[1]: https://phabricator.wikimedia.org/T196450 [2]: https://query.wikidata.org
Just a Speaker's Corner to say whatever one wants regarding Wikidata?