The Semantic Web Is Dead. Long Live the Semantic Web
blog.blikk.co
blog.blikk.co
There is no better way of expressing why "the semantic web" as it was originally conceived failed, and will always fail. This is not to say that automated processing to extract higher-level information of the kind the article argues for might not be both helpful and possible, but it will always be awkward, difficult and imperfect. The awkwardness and difficulty will go down as AI gets better, but the imperfection has a fairly high minimum level because the kind of meaning humans extract from language varies enormously between individuals.
As a friend who works in geology put it: "If I send a bunch of geologists out to survey an area and combine their data naively, I can tell who mapped where but not what anybody mapped." The illusion of shared meaning is powerful, and with enormous effort we can get communities of shared meaning that are large and powerful enough to be extremely useful, but we do that either through top-down control (politics, corporations, militaries) or intense bottom-up interaction (the sciences), not the kind of loose, informal, disconnected mechanisms that the Web enables.
I think the Semantic Web doesn't try to solve this problem and doesn't need to. There are other approaches, such as probabilistic programming, that are meant to reason with uncertain data. But to do that someone must make the data available to you and tell you how uncertain it is in the first place. I think that's what the Semantic Web, or whatever you may call it, can do. Make data available, but leave open the interpretation.
My take is that it is a bunch of ideas to make it easier for search engines to crawl and use your data about all kinds of things. On one hand, that is great.
However, the huge problem with using it is that you end up creating your own search engine to crawl data sources to do anything useful. That is to say, you have to crawl the data, store it somewhere, and then build up your own systems for querying or doing anything useful with the data.
The use case many people have is that they want to use and API to get some particular bit of data out and that's it. Like, say you want to do a search for a list of tweets on Twitter with the hashtag #HackerNews. The sane thing to do is to be able to hit a twitter api endpoint and have it send you back a list of tweets with that hash tag.
Now, imagine if instead you had to index all of twitter and filter yourself for tweets with #HackerNews in it. Is that better for that particular use case? No, it sucks.
There are certainly cases where you DO want to crawl data and do your own data analysis on it. But, that is a much more limited use case for many developers and there isn't as much value in that as people seem to believe.
A better solution would be something like REST with HATEOS on a much larger scale. You'd be able to index things nicely, but still have the benefits of smarter API calls.
Unfortunately, I don't see this happening anytime soon, despite the interesting things you could build with it.
> The use case many people have is that they want to use and API to get some particular bit of data out and that's it.
In theory, that's what SPARQL and SPARQL endpoints are for. Plus you get things like federated queries and an open data model (RDF), that allows to combine multiple data sources without "wrapping" schemas.
But well, this is kind of utopia and yes, I doubt it will ever be "a thing".
The industry track exists for the record. Much of semweb companies were funded using EU money and went bankrupt when the money went out. I can't believe the EU continues to inject plenty of money for the semweb in H2020 despite having wasted more than 1B in the previous decade (publicly admitted by the EU).
At ISWC 2012 in a workshop that took place the day before the conference, some guy (don't remember his name) asked the speaker "would this building explode, do you think the semweb would still exist ?". The speaker tried to find examples of people doing semweb in the industry, but did not manage to convince anyone, not sure if he did it for himself. This was (and is) a nice summary of the situation of the semantic web.
[2] http://103.1.187.206/core/1338/
[3] http://arnetminer.org/page/conference-rank/html/All-in-one.h...
[4] http://academic.research.microsoft.com/RankList?entitytype=3...
It's like asking Einstein back in 1905 how is his work going to be used in industry.
Semantic web is not an industry or engineering field. It's theoretical and academic and it has its purpose because it lays strong theoretical foundations.
Physicists study law of nature, there is no such thing in the Web which is a deeply human field. It's quite an insult for Einstein to be compared to the Semantic Web ...
In all seriousness, 'Semantic Web' has always felt like a SciFi inspired version of AI intelligence. A concept that sounds cool, but in reality can't ever work.
Take movie ratings & "suggested viewings" for example. Jim, Bob, and Steve all watch the same movie. Jim thinks it's funny because of the physical gags. Bob thinks it's funny because of the dialog & jokes. Steve likes it because the hot new actress is naked. Dave likes the director and cinematography. All 4 guys give it a rating of 4/5.
With this one data point Streaming-movie-place.com cannot ever 'guess' what to suggest to these guys to offer more movies for them to watch. The hope is that once these guys start watching and rating other films a pattern will emerge. That pattern can then be marketed and offer valid suggestions.
BUT reality is too different. We like different things for different reasons, and no algorithm can ever get it 100% right. How many of you have a Netflix queue of things that you want to see, but the suggested movies are full of crappy suggestions? most or all of us I bet.
Which brings me back to my point; it's a sci-fi illusion. It can never exist in real life. Humans are too damn fickle. (Which brings me back to KDE; I wish they would give up on neopunk/symantic desktop crap. it's bloated, slows the system down and offers nothing in return. Or I am just using it wrong.) </rant>
In particular, scraping the View is pretty much a last resort. If you have to, it's likely that either the source doesn't have the manpower to become semantic anytime soon, or it's hostile to providing easy to parse data anyway.
1. Most of the implementations and formats and everything really boil down to a great way to publish graph data on the web, and query it using a pretty nice graph query language, and for most general cases everything just works.
2. A lot of the web is implementing incomplete select parts of SW technologies even if they don't realize it, and then promptly putting it all behind an API key and shared secret. When SW takes hold, that will all have to be different, IE a common mechanism for authentication / authorization, and some kind of way to quantify what is supported by a service, and all of that exists but again, every service is different it seems right now, and they all fear unfettered access.
3. You can use all of it today if you want, and the library ecosystem is very rich, IMHO. Plop it all into Neo4J / Jena / rdflib (etc etc) and have at it.
When implementation is left to a large population, there will be differentiation as that is our nature. The semantic web wanted it easy, it wanted the implementers to organize. Organizing it with all the differences is how it will have to be. Unless a system implements the standard for you and each author upon creation, there will be differences, deltas and no standard.
Probably the best semantic web / metadata system that has been built does do this and that one is at the NSA.
But a working Semantic Web is a huge deal (think: bigger than Google) and will happen, make no mistake. But when it happens, it will be via a different approach.
Edit: spelling (thx Joshua)
Based on products like Siri or Cortana it may seem like we are coming closer to NLU, but the reason those are working at all is because they are very topic-specific and tiny subproblems of what NLU is really about.
One of the great things about the semantic web, to me, was always that you could process it on a cheap VPS or your own desktop, creating new systems with little or no knowledge of math or machine learning. (I know projects that do this today.) That's not going to happen for a long time with natural language processing.
Why do you think this? Is it just a feel or is there some company that is really close. As another user mentioned in this thread natural language understanding seems to be the same thing as solving strong AI. Solving strong AI is a huge deal. So big that the people in control of it will probably become the most powerful people in the world.
I doubt that a sufficiently well-defined notion of natural language understanding that does not specifically include strong artificial intelligence in its definition would require strong AI. Constructing such a definition is left as an exercise to the reader.
Thinking that strong AI is required for natural language understanding may end up being similar to how it was once thought that beating humans at chess would require advanced AI. Brute forse can do wonderful things, as can weak AI.
I'd love to find out more about what you have in mind here. Can you elaborate?
THIS is why it failed: you guessed what the thing was from the name instead of actually learning about it.
Also: "death throes"
You might find this interesting: https://vimeo.com/92351230
When anyone can query anything, the SW will build itself -- a sort of new web, but of data, not text.
ps x | grep "nginx" | wc -lJust like many other research fields it lays strong theoretical (and at times also practical) foundations.
This is very poorly explained; people tend to think it means something about computers understanding the data.
Love, RSS
Edit: Not sure why this got downvoted, just trying to illustrate what I thought the term meant.
Using semantic HTML (rather than presentational HMTL) for content is somewhat related to the Semantic Web, but not equivalent to it -- the Semantic Web is/was about open, composable ways of extracting and processing meaning from content on the web. HTML itself (even HTML5 or the current state of the HTML Living Standard) -- even when used in a cleanly semantic way -- isn't expressive enough to do much with that without building additional ontological structure on top of that, hence things like RDF and Microformats.