An Intro To The Semantic Web: Why You Need To Know About It Sooner Than Later
webcentralstation.ca
webcentralstation.ca
No. This will not happen.
Look, before Google, internet search relied on people adding meta tags to their pages so that the search engines could know what a page was about. You had two languages, one for humans, and one for computers.
But people don't really care about having correct meta tags or headers or other invisible markers containing instructions for search engines, which is why Google steamrolled the market when it appeared. Google ignored all the instructions and instead analyzed the human language on web pages, and inferred relevance based on that.
If semantic search engines or analyzers or other technology wants to become widespread, it cannot rely on extra markup, it cannot rely on people adding descriptors of meaning to the data it indexes, it has to determine that meaning from the data directly.
Google are now at least partially on the Semantic Web bandwagon. They, as well as Yahoo, have been detecting, parsing, and using RDFa and microformats data - when present - for a while now, and using it to provide enhanced search results. See:
http://www.google.com/support/webmasters/bin/answer.py?hl=en...
People don't give a shit about metadata.
The important thing here is to not commit the fallacy of the excluded middle... There's a place for the Semantic Web between "it's a total pipe dream and no part of it will ever come to fruition" and "it will be fully realized in exactly every detail as originally conceived by TBL."
People are already using Semantic Web technologies to some extent, the open question now is "just how prevalent will this stuff eventually become?"
Granted, I'm not sure what role the semantic web, as defined, will play in this over the long term. And, to achieve optimal costs, I would tend to agree that the great majority of data will need to be annotated by machines.
However, I think it is important to point out that human generated annotations like markup and ontology can play an important role in 2 scenarios: 1) places where there is not enough data to make meaningful machine inferences, 2) vertical domains that have very well-defined structure where the cost of human annotation is actually less than machine annotation.
It's seven years later, nothing happened in terms of real-world adoption, it's safe to declare it DOA. In fact it's been safe for several years.
And never mind the fact that there is real world adoption. Google and Yahoo both embraced RDFa a couple of years ago, and have you checked LinkedData.org lately? There's a constantly growing body of data out there in semantically interoperable formats: http://linkeddata.org/
[1]: http://webbackplane.com/mark-birbeck/blog/2009/04/20/rdfj-se...
http://data.gov.uk/apps - same thing here, driven by the semantic web and linked data.
It's only one small part of the overall Semantic Web, but it's an element that's here and being actively used already.
Perhaps the metacrap rant (http://www.well.com/~doctorow/metacrap.htm) was right?
Seems that way at the moment.
Additionally there seems to be a whole lot of reinventing the wheel. The best semweb people are aware of all the past research in logic programming and automated reasoning, but most semweb enthusiasts seem to be hardly aware of prolog let alone that rdf triples are just another way to express what clauses do in prolog.
If we're really trying to solve the 'problem' that semweb addresses we'd be seeing more articles titled "intro to logic programming, knowledge representation and automated reasoning"
"It doesn't take economics into consideration. most people/companies aren't that interested in sharing their documents in an anonymous way.
WWW is a web of documents. SemWeb is a graph of data.
More reading: http://www.w3.org/People/Berners-Lee/ http://www.w3.org/DesignIssues/
I think this is a better intro for technical folk.
A better idea is extracting latent semantic information from the existing messy web. However, "meaning" is extremely difficult to characterize, and attempting to encode it in an interoperable manner inevitably leads to a lowest common denominator approach. That will probably still provide tons of value and be much more ubiquitous than structured data, but will ultimately be shallow and fall far short of the vision most semantic web proponents evangelize.