Semantic MediaWiki
semantic-mediawiki.org
semantic-mediawiki.org
Today, the first thoughts that comes to my mind aren't what my work could do with this information, but more about sustainability, accuracy, and power.
A bunch of huge companies are going to hoover up all the data you've compiled, and use it in such a way that most people using it will never see your Web site.
So won't be prompted to contribute back wiki-style, won't get all the benefit of your quality control, won't see your donation appeals, won't know you exist, won't see the other things that you would like to see, won't give you analytics on your impact, won't fund you with ad impressions, etc.
Which might all be OK for your use case, but over time it could break most of the current useful percentage of the Web.
— Iain M. Banks, Excession
Despite being fictional, it seems at least as correct as your saying.
Also applies to corporations.
But had they been adopted for what they are: well defined metadata according to vocabularies and ontologies that following accepted, interoperable standards, the development of ground truth and knowledge graphs would have been much more advanced that the current state.
So my guess is that projects like semantic mediawiki, wikibase etc will eventually see renewed roles and gradually link with the more statistical and algorithmic processing of information.
Yes, this is very true.
As an addendum, I think many people in industry don't quite realise how prevalent RDF and RDF-based ontologies are in academia. They're one of the main ways researchers integrate data. No one cares about lofty ideals, it's about the usefulness of RDF graph model and the various associated standards built on top of it.
I have been wondering why it got adopted there and not elsewhere. Its not clear. Adoption or not of software and information tools is a mix of cultural and economic reasons beyond pure technical aspects.
CSV, json etc are information black holes. Eventually people will realize that the systematic annotation of data within a metadata framework is not fanciful but an enabler of more confident and robust information exchange and processing
If you check out the common ontologies in use you will get an idea which areas are covered: https://lov.linkeddata.es/dataset/lov/
With LLMs those tags have suddenly become a lot cheaper to make. Arguably their cost was the largest barrier to the semantic web vision. But now it may already be practical to semantically markup wikipedia content automatically as submitted. But does AI moot the semantic web, or enable it? Is the semantic layer redundant with an AI interface? I can't predict.
Depending on the context, read / write model, and how inference costs go, inferring on edit to classify and tag the content could prove more expensive than having users do it client-side on read.
Semantic technologies or let's say "structured data entry" can benefit from having better auto-suggestions and validations from machine-learning.
The part that I find real interesting is how we put the users in between, so they're in full control and can use AI as an "assistant", but improve on it. And that the "structured data entry" helps the users to give him structure (that will later be helpful to the end-consumers).
It is used att places like NASA as a knowledge base [0] and some other places.
I guess what is hindering wide adoption is that it is still quite a bit too much conventions and caveats to learn to become productive with it, that it has a hard time reaching outside niche use cases.
(Basis for speaking my mind on the topic: Developed the RDFIO plugin for RDF import/export in SMW [1])
Big difference is that wikidata is an entirely separate wiki from client wikis, where smw is much more closely integrated into "articles". This allows wikidata to abstract across multiple language wikipedias. In practise that means that wikidata is often used to store factual data outside of articles that than gets pulled in via templates. SMW tends to more get used as a backend to turn mediawiki into a user programmable CRUD/workflow app (particularly with PageForms extension)
SMW is less concerned about scalability and performance. Wikidata puts anything of questionable scalability separate from the system (e.g. sparql api is isolated from rest). SMW would be unlikely to scale to anywhere near wikipedia levels in its current form.
Where extensions like SMW really shine is for wikis which discuss highly structured content, like content in video games. Another extension which I've seen used in this space is Cargo [1].
It was, in a nutshell, not great: there's a fundamental contradiction between wikis, which are very loosely structured and edited by humans who do all sorts of wacky unpredictable shit, and hierarchies, which need to be carefully pruned and maintained.
So this use case didn't work, and I think it's telling that (AFAIK) not a single Wikimedia project has adopted it.
SemanticMediaWiki has some questionable scalability when it comes to large sites. Wikimedia was never going to adopt it for technical reasons.
The advantage of SMW imho is still that it brings the semantic / structured data features right into the MediaWiki itself and that there's a big ecosystem of related extensions around it (like Page Forms). So far I don't know of an open-source information management system which comes close in capability to this. Setting this all up and maintaining it is difficult, though - I even had to start developing custom tools to assist me with that once it got more complex.
However if we're talking about Wikipedia, any scaling concerns at all basically make it a non-starter.
https://wiki.guildwars2.com/wiki/Guild_Wars_2_Wiki:Reporting...
We never bothered doing anything like concepts, detailed types, etc...we only used it for basically adding tags to pages that you could query as part of a search. It worked okay with that. We had some issues where we wished a property could be an enumeration or nullable, which I don't think ever happened in one of their many releases.
Adding the metadata represented a bit of a challenge. We eventually had it generated from templates that would usually get filled out anyway.
Still amazed at that video when I rewatch it.
[fyi, the company developing that extension was bought, and Project Halo disappeared in limbo]
Eventually I presume it all goes down to Michael Erdmann [@emwiemaikel on Twitter] having the right to maintain the codebase.
https://evolvingtrends.wordpress.com/2006/06/26/wikipedia-30...
It took more decades and the arrival of Transformers and self-attention (see Attention is All You Need) for the "End of Google" (the search engine, not the company) to become a reality.
I was able to find Marc on Twitter: https://twitter.com/marcfawzi
Seems to have picked up his old idea of a geek-run VC fund, but using AI: https://evolvingtrends.wordpress.com/2006/06/21/geek-run-gee...
[Or maybe Notion, i never used it]