Semantics and the Web: An Awkward History
lists.xml.org
lists.xml.org
Second, this is not about the Semantic Web, despite 'Semantic' and 'Web' appearing together in the title. It's about the history of semantic markup, which is where markup is meant to describe what text is. (Since it's about "meaning", the word semantic is used.)
Be sure to see the video (https://www.youtube.com/watch?v=77qDvd5uOx8), presented by Simon St. Laurent at 2021 Balisage, a conference for markup specialists, if you are interested in the evolution of markup from SGML to HTML to XML to XHTML to HTML5. To whet your appetite for what the video is really about (not the Semantic Web), Simon starts like this:
Hello, this a story where we, the fans of meaning conveyed by markup, mostly lose after a long winning streak. To soften the edges a bit, I'm telling the story with Playmobil figures...
It's very well written and cleverly presented -- worth a watch if you're at all interested in the historical arc of markup.
That's helpful, but I was linking to the thread index since that post spurred the majority of discussion in September, not just the direct responses bearing the post's title, so be sure to read all responses. In particular, I found interesting the tension between those seeing XML as a general-purpose serialization format (and wanting to simplify it further) vs those who're unhappy with XML having lost out on the web and sticking to SGML, HTML, markdown, etc. as formal document formats, a recurring topic on xml-dev.
As far as I know those tools were either immature, nonexistent or proprietary/expensive.
I thought that signaled the death knell of XHTML.
Youtube video: https://www.youtube.com/watch?v=77qDvd5uOx8
Transcription: http://simonstl.com/balisage/TRANSCRIPT09042021.txt
Pairs well with Brian Kardell's History of the Web series (2015), which covers some of these bits & is a delightful enjoyable read as well: https://bkardell.com/blog/Brief-ish-History-of-The-Web-Part-...
Also, note, very little about "semantic web". The word semantic does not appear in the transcript (other than as the title).
I really enjoyed watching Simon's talk myself - a slightly different perspective of the last 30 years is also in this talk I gave https://youtu.be/dFsRz1PGjdw?list=PL_BRVuWxk8srToZ69z-_5vBJI... .. It's got no playmobile figures, but a lot of gifs :)
I think there are a lot of small factors that create a big headwind for semantic markup, but I think the most important are:
1. Semantic markup can make it easy to build bigger/better levers, but a lot of these aren't universally useful. (They aren't compelling reasons for most people to suffer.)
2. There's no great starting point you can recommend to anyone from which they'll be able to incrementally improve their documents and start building levers without large up-front investments.
Is there anything better than RDF?
1. The simultaneous emergence of walled gardens
2. The concepts were too challenging for most developers to fully grasp and too distant to inspire common interest.
You have to remember the web prior to 2005. Nobody had heard of Facebook and Twitter did not exist. Google was just a search engine and data auction. Most of the web displayed static content dynamically generated by either ASP, PHP, or TCL. Most of the people on here have probably never even heard of TCL.
Back then all the value of online businesses were some form of payment processing, think ecommerce, or application processing like data mining. The idea that the data itself had value aside from the products and services it represented was known but not fully realized. This wasn't even deliberate.
Emerging online services needed to generate revenue to repay their investors and in most cases the only thing that stuck was online advertising. You can show ads to anybody, but the more precisely targeted those ads became and the more they followed users across third party sites the more valuable they became. You have to understand that in most cases these are high quantity but nearly worthless transactions so anything that could raise the value of a transaction is a really big deal. This is how the walled gardens happened.
This would have been obvious to anyone trying to build even a trivial application. Just try to build a todo list, recipe application, calendar, etc. It will be like walking in quicksand. It will start out great and then every step will be slower than the last until you're expending huge amounts of effort to just make the next step and in the end you'll see what you've got and realize that you expended a ton of energy to make someone else's life easier and even that won't necessarily be clear. Performance will probably be dismal and there's little hope that it will scale at all.
In the rest of the world nobody cares. It works or it doesn't. The end. Most software developers want to believe that they are really smart people, and they just might be, but that doesn't equate to effort. Learning and extending the world of XML takes effort.
My experience sitting on both sides of that fence tells me that unless you have a domineering personal investment the only thing that matters in software is employability and the semantic web did not crack that code. Employability is a primitive binary equation often reduced to a few lines on a resume.
Eventually Google's natural language-esque search just got big enough that no other interface to the web really mattered anymore, and their speciality was in parsing unstructured text & messy HTML into simple phrases. Better semantics might've helped other spiders, but by that point nobody cared about other spiders anymore.
For things like social sharing, OpenGraph had the commercial support of Facebook and Twitter and was far simpler to implement. And other communities, like MediaWiki, used RDF only for the low-hanging fruit (Wikidata) while the most valuable info was still locked behind freeform blobs of text (Wikipedia).
The semantic web took more effort to implement than the crap it usually describes. Most humans just don't really want to waste time classifying stuff. The communities that do (science, pirate communities, libraries, etc.) already have their own classification schemes. There was just no need for another web-only classification scheme, no popular desire for it, no sufficient commercial interest behind it, no end-user advantage of using it over HTML... is it surprising that it failed? It tried to solve a problem nobody really had, using a solution that was quickly eclipsed by machine learning. And for the few actors (search engines, social networks) who actually wanted effective classifications, their own algorithms were both more effective and more private, not relying on/enabling their competitors. Open classification excites librarians and archivists, maybe, and nobody else.
Just wondering what information do you consider valuable here?
I don't think it was anyone's vision that rdf replace prose text entirely.
It's not that prose is any more or less valuable than classification, but that classification takes a lot of work and doesn't come naturally to most knowledge producers. It's a speciality unto itself that most people aren't super interested in. In the case of schema.org and RDF, that effort had no visible impact on anything, so people just stopped trying after a while. In the case of Wikidata, it's still ongoing, but nowhere near as popular as the more natural (to our species) output of Wikipedia.
FWIW, I'm saying this as someone who spent time classifying articles on Wikipedia, contributing to Wikidata, marking up our pages with microdata, managing databases for a living, etc. I'm not philosophically against better classification. Just observing that efforts to automatically describe human output via algorithms and ML seem to work a lot better than asking human content editors to self-classify their output with anything more complex than WordPress tags. People don't tend to enjoy or be effective at complex classifications, especially when there's no visible benefit from it. All that microdata never really surfaced anywhere and there was no useful mechanism of discovery or sharing.
Really? I didn't see a funeral?
[Except maybe the podcast usecase]
Sure, maybe podcasts use it as an underlying protocol, but users' experiences depend more on the front-end client (some app) than the playlist protocol. It's not really a discovery protocol either; turns out having a central index like Spotify or the podcast apps do works better anyway.
I think it's irrelevant in the sense that it could be trivially replaced by app-specific implementations, and few people would even notice anymore.
But they tend to come up with their own classification schemes.
The idea of a interoperable semantic web wasn't a bad one, there just was no real presentation layer for it, and no popular demand for it without one.
However Amazon, Spotify, libraries, etc. classify their stuff internally, to their eventual end-users, the data is transformed into a HTML or app frontend of their design. Users don't see the underlying classification schema, just the product pages or book search or whatever.
Sure, you could probably use XSLT to transform a semantic document into pretty HTML, but most companies went the other way of using custom code to read data and output unstructured markup, because it was easier and faster. HTML and JS "standards" evolved way quicker than RDF could keep up with.
Beyond that, had there been a web-wide search engine that could easily let you search by classifications, microdata might've been helpful. But that never became popular. Closest thing I know of is Wolfram Alpha, and that uses its own algorithms and classifications too.
shrug
What do you think?
Consumers hardly ever get given the tools to curate content apart from upvotes, downvotes, and maybe comments which they can put hashtags in, and even those often get taken away or rendered useless.
The very fact that some people often do it for free when the value is mostly captured and extracted by mainstream corporations suggests otherwise.
Those same corporations have a strong preference for passivity, exhibited by their preference for the term "consumer".
I thought the most hilarious demonstration of this tendency to try and train users in to passivity was exhibited by the flamboyant VC-inspired suicide of Digg.
https://developers.google.com/search/docs/advanced/structure... and https://schema.org/
The only reason it was tried at all is because a large chunk of the Web standards community believed they could ignore economics and rule Web authors by decree. This belief is so attractive that even today you can regularly find HN commenters who hold it ... the ones who complain about the way the Web is and argue that browser vendors or standards bodies should make Web devs behave differently.
One contributing factor was that the semweb community really needed more working engineers. I tried a few times to implement things and you could very quickly end up in cases where the documentation in 5 places was inconsistent and the lone example on some random W3C page didn’t work with the one available tool or hadn’t been maintained in so long that some other random XML spec had mutated incompatibly. Things like that really pumped up the cost of what was already a dubious proposition.
I mean RDF is alive and well and Schema.org is widely adopted.
The closest we have are things like schema.org, which are really just translation layers that people run their own proprietary data models through in order to produce something interoperable. But we're certainly not using the interoperable representations everywhere; only when Google or some other big company has the financial incentive (and monopoly power) to mandate it.
If, instead of a big inscrutable blob of HTML/JS/CSS, Amazon gave me the same shaped data for a product listing as Ebay does, then as a user I would have a lot more power over how my browser represents any generic product listing, across the entire internet.
And Amazon / eBay / Twitter / Facebook brand identity would be worth so very much less.
Amazon doesn't want you to think about a product on their website as interchangeable with a product from eBay or WalMart, they want you to think about their product listing i.e. the listing of the product surrounded by the Amazon Experience.
Likewise, Twitter doesn't want you to think of a tweet as a bit of content equivalent to or interchangable with a toot from mastadon or post from Facebook; Twitter wants you to see the Tweet (the content within the context of using Twitter).
It is a bit like calling Disney World a theme park. Sure, it has roller coasters and stuff, but describing a thing is not experiencing it.
Sounds great!
The reality is, it won't happen, because average consumers don't care and the big producers and suppliers don't benefit from it at all- quite the opposite.
The apocryphal lessons of the betamax vs vhs wars seem to never be learned- no matter how elegant or technically superior something is, it is the "whole product" that matters.
This is why we have governments: to enact legislation that requires companies to not run roughshod over the public interest. Regulation is the magical fiat wand. I know that HN's views of GDPR are all over the place, but the reality is that it is possible to enact meaningful legislation to restrict the powers of tech companies, and magnify the power of the end user.
You could make the same argument that the average user doesn't care about the data protection measures that GDPR provides. And the corporate interests definitely chafe at those requirements today.
That would certainly be ideal, but uniformity is not a prerequisite for more powerful, user-serving interfaces.
It doesn't always take special intent on the part of the site operator to end up providing a rich source of data; many sites end up doing it by accident. This is what's so great about HTML's class attribute. People often mistakenly think of classes as inherently having something to do with CSS. CSS is a carrot that often results in developers being encouraged to mark things up whether they buy in to the semantic argument or not[1].
Take a look at the video listings on pbs.org for an example of rich structure that can be readily mined by the UA. What's lacking is a generic tool that sits at about the same place as Excel on the power/expression spectrum and which allows the user to create and "train" an adapter (in a matter of seconds) for an arbitrary Web property. For example, given a list of videos and associated details (title, thumbnail, description, length, etc), point to two or more titles and your UA can, in an ML-free way, use those samples to examine the structure to create an adapter. To really make things tighten it up to an acceptable margin, it can even provide an editor that exposes the raw class names as UI so you can explicitly reconfigure the model as you please. This is a system that would not take an expert to use, at least not any more than the "expertise" that is required to recognize that, say, the table widgets in many apps allow you to sort data or rearrange/resize columns via drag and drop.
1. It's not perfect, because you have garbage-out fads that can come along and knee you in the stomach, like Tailwind CSS and compilers that generate meaningless class names, but there are short-term and long-term mitigations.
As the parent poster says, an effect could be "then as a user I would have a lot more power over how my browser represents any generic product listing, across the entire internet. " - from the perspective of any major site providing product listings, that's not a good thing, it quite clearly says that it would transfer power from the site to the user, and the sites do not want that.
For them, becoming a source of "generic product listings" would be a horrible strategic failure, destroying a foundation of their business model and so it would be worth spending a lot of money to ensure that their site does not work with any standard promising this effect and/or that the standard fails to exist in the first place.
So the key concept is that if you want to improve interoperability and the proposal to get there (e.g. Semantic Web) requires content providers to do some work, then it's inevitably going to fail, because the major content providers do not want interoperability - the up-and-coming content providers might want interoperability to gain users from the major existing ones, but only until they gain marketshare themselves.
The implementations I saw were basically just a case of hoping they’d guessed at what other people needed well enough and waiting to see if someone used it. Changing the serialization format wouldn’t have changed that dynamic.
The reason you can tell that it's not the format is that this hasn't changed as API-only services became commonplace. The problem preventing collaboration wasn't that people were putting structure in HTML but that businesses didn't see an advantage in doing so.
this is the problem, but it's magnified by the fact that you put responsibility in the hand of content makers and web developers, of course they are not interested in this. If you reframe this around making inteligent systems talk to eachother, separated from the web stuff, then the incentive is clearer to whoever has it, those who want to make themselves visible to mash-up services, aggregators, e-commerce, booking systems and what not