W3C’s transfer from MIT to non-profit going poorly
twitter.com
twitter.com
I would take a skeptical view of his take on what's happening. The w3c is a very dysfunctional organization and there has been a lot of turmoil internally. Jeff Jaffe who had been CEO for more than a decade quit in November. There are power plays behind the scenes to fill this vacuum.
https://www.theregister.com/AMP/2022/07/01/w3c_overrules_obj...
[1] https://www.eqar.eu/qa-results/synergies/european-blockchain...
Maybe I'm misunderstanding, but the HTML 5 spec is entirely the work of Ian Hickson/WHATWG (financed by Google), and has been for over ten years. Until about 2018 W3C has merely snapshotted, rubberstamped, and editorialized WHATWG's specs, and has since stopped doing even that, and now W3C's HTML5.2 spec has a banner retroactively redirecting you to WHATWG's current head [1]. W3C is involved via the CSS WG and ARIA still. But yeah, as a standardization body, W3C has failed spectacularly. There's no standard as such; WHATWG's "living standard" is merely a stagnant (yet still unversioned) memorandum of understanding of extant browser vendors (Google Chrome and Google-financed Mozilla plus Safari) to implement features, with a large part of browser APIs not implemented by FF and Safari due to fingerprinting concerns.
w3c's XHTML was machine understandable. It took great care to introduce rigid and reusable semantics. You could spider the web and directly extract facts without ML or heuristics. It spiritually brought HTML closer to RSS and RDF, and there would have been a lot you could do with that.
Google didn't want anything to do with that future. Search is their moat. They pushed a format that is easy for humans to author, tolerably messy, and knowledgeless for machines. If the format was difficult to extract knowledge from in a scalable and reusable way, Google could keep their search kingdom.
XHTML was a ladder rung into the semantic web. A distributed knowledge graph that could be mined, remixed, and extended with little effort. It was marginally harder to author (the mimetype issue was ridiculous), but extending it was low hanging fruit.
This isn't necessarily how it went down, but the incentives do align. I do think we'd have more automation in the world today if we'd have adopted XHTML.
Now with AI/ML and NLP I think we'll get to this kind of web without a format to rigidly specify semantics. I still have to think that the ideas of semantic web and P2P could have delivered decades ago if Google and Facebook hadn't overwhelmingly pushed the envelope on centralized resources.
Google even cooperates with other search engines in promoting unified formats to "extract knowledge from in a scalable and reusable way", namely the schema.org standards.
It would require people to use extra effort to author a table of "address information" (or whatever). Maybe that would not have happened, but the format also suggested the building of reusable components, browser integrations for semantic understanding, and much more.
The world it alluded to was much brighter for machine understanding and reusability.
/s
That would be interesting to me, and I wasn’t aware of it until your comment.
Edit: sounds like it’s not a thing:
> The DOM, the HTML syntax, and the XML syntax cannot all represent the same content.
https://html.spec.whatwg.org/multipage/introduction.html#htm...
> For example, namespaces cannot be represented using the HTML syntax, but they are supported in the DOM and in the XML syntax. Similarly, documents that use the noscript feature can be represented using the HTML syntax, but cannot be represented with the DOM or in the XML syntax. Comments that contain the string "-->" can only be represented in the DOM, not in the HTML and XML syntaxes.
I.e. the big thing you'd fail to capture when doing an automated HTML -> parsed DOM -> XML trip is noscript tags. No big loss. Others I can remember are similarly tiny, e.g. some issues around textarea elements containing only newlines as their initial contents. Some discussion at https://html.spec.whatwg.org/multipage/parsing.html#parsing Ctrl+F "roundtrip", although note that most of that is actually about what happens if you construct weird DOMs with JS and try to do a DOM -> serialized HTML -> reparsed DOM roundtrip.
I think it's pretty safe to say in general that the procedure of "parse HTML to DOM, serialize to XML" will preserve everything interesting about a document. Especially if that document is already valid HTML.
You can have an element called "foo:bar" in HTML, but you can't in XML-with-Namespaces-in-XML; likewise, you can have an element called "a\u0300" in HTML, but you can't in XML (and the set of possible element names is even more different if you're looking at an implementation of XML 1.0 4th Edition, which most are, rather than the much later 5th edition).
You can't have a comment that contains "--" or ends with "-" in XML. (This is potentially the thing that comes up most often in real world content: people like doing <!------ DON'T TOUCH THIS ------>, which in HTML produces a comment whose value is "---- DON'T TOUCH THIS ----", which you cannot represent in XML.)
The first attempt at that "knowledge graph" thing was the "keywords" meta tag. Its failure should be taken as instructive regarding the difficulty level of the project as a whole.
> It was marginally harder to author (the mimetype issue was ridiculous), but extending it was low hanging fruit.
Silly me, I thought it was more about mandatory strictness vs user-generated content, rather than what mime-type the server had to label it as.
Knowledge graphs are basically just semantic networks, which arguably date back to 1956 (or 300 depending on how you define things)
It did absolutely none of this. It was an presentation layer built on top of XML, and it did not really touch semantics at all. The W3C's proposed direction for the future of the web was for everybody to write data as XML, and then apply XSLT transforms to build XHTML as the structural presentation layer, with CSS providing styling information on top.
XHTML added very few additional semantics, other than loose concepts that were worse than useless.
> It was marginally harder to author (the mimetype issue was ridiculous), but extending it was low hanging fruit.
It was difficult to author (there were days you could run random websites through the "XHTML Validator", and despite them trying, find plenty of errors). The mimetype issue was that Internet Explorer would not recognize a proper XHTML mimetype, while other browsers wouldn't turn on XHTML mode when the mimetype was text/html.
https://webkit.org/blog/68/understanding-html-xml-and-xhtml/ -- this blog post was written as XHTML's last dying breath, and I think was the nail in the coffin for it.
There was zero need for XML; recall XML is just a proper subset of SGML by its original developers, which was used to specify HTML until version 4. And even today, SGML is the only game in town to parse HTML (including version 5, up to minor trivialities) based on an international standard, whereas Hickson's/WHATWG's procedural parsing spec has become unmaintainable and isn't covered by a test suite, precisely because it doesn't follow a formal model, but started from a prose description of what an SGML parser does but lost track. Not a great outcome for a markup language used daily by billions of people; especially since HTML the markup language hasn't really changed for a very long time, whereas everything around it (CSS, JS) had to change drastically to make up for HTML's stagnation and first W3C's, then WHATWG's fuckups.
W3C has lots of failures and bad standards. I don't think the base XML spec was one of them, even if i think JSON is better for most usecases.
It's now maintained by Dominic Denicola, Anne van Kesteren, Tentek Celik, and a few others.
The HTML Standard has not been maintained by Ian since 2015. It has since then been maintained by myself, Anne, Simon, and Philip. Affiliations during that time are, respectively: Google, Mozilla then Apple, Bocoup then Mozilla, and Google. (Indeed, mostly only browser companies have been paying people to be editors for web specs!) https://html.spec.whatwg.org/multipage/acknowledgements.html...
But the contributor pool is much larger. We're quite proud of the vibrant (not stagnant) community we've created, which is continually evolving and improving the spec. Contributions come from all corners; many are from people employed by browser engine companies, but others include students, web developers, consultancies like Igalia and Bocoup which various companies hire to work on web standards, W3C staff and members, representatives from server-side runtimes like Node and Deno, and so on. Some interesting pages to peruse might be https://github.com/whatwg/html/commits/main , https://github.com/whatwg/html/graphs/contributors?type=c , and https://blog.whatwg.org/ .
Overall I think it's pretty exciting you can run a successful standards organization like this, getting such high engagement levels despite employing no full-time staff and with operating costs being entirely server bills. Including, no membership fees. Which is perhaps relevant to the OP. https://twitter.com/Hixie/status/1603917371214729216
The WHATWG only includes features in our standards which have multi-implementer commitment: https://whatwg.org/working-mode#additions . This is different than some places where anyone can publish a "standard", even if the target platforms have no intention of implementing it. I don't think this reduces to "merely a memorandum of understanding"; I think it's best when standards reflect reality. Other SDOs can disagree, and that's fine; it's healthy to have a marketplace.
As for the scare quotes around "living standard", you might enjoy https://whatwg.org/faq#living-standard and the follow-up questions.
https://www.coindesk.com/markets/2020/10/15/filecoin-launch-...
That said, don't get me wrong, I'm always down for a quick pitch fork roast on the internet.
However, I can quite easily see this happening on the MIT side. Some mid-level bureaucrat who doesn't even know what W3C is will be losing budget, so they're playing hardball assuming the usual level of scrutiny. They're going to get a surprise when they get dumped on by their managers because this suddenly hit a lot of eyeballs and is garnering negative PR for the entire university.
I have no idea why anyone of any level of technical sophistication or containing halfway decent communication skills makes the attempt. Choose a free blog, write something more substantive, and write a succinct Twitter post to get people aware of it. Or at least do that at the same time you post a balkanized “thread” like the author here and link to the more substantive post in the process.
Otherwise why bother writing up research either? People might just read the abstract. Why bother watching a full baseball game? You can just get the condensed highlights and winner afterwards. Why bother watching a movie? You can watch the trailer for a lot of the great bits and then read a synopsis online.
Plenty of people actually do those things and I guess that’s okay too but there value and good reason to put out the whole product too.
https://mastodon.social/@robin/109524929231432913
(I'm not sure we should even be linking to the Twitter copy, given that the author's name there is giving attribution to him as "@robin@mastodon.social", so surely that should be considered the authoritative version, and it doesn't require delegating to any third-party aggregator hack.)
MIT is playing hardball with people's jobs and W3C assets.
W3C is playing hardball MIT's reputation.
I think the fact it's reached the point they're publically talking about this means they're is very little chance MIT is going to be backing down. The real question for me is would US officals allow W3C to move aboard. Could they prevent it? I have a feeling MIT's lawyers have thought alot of this out already.
It's more like softball, if we're being honest. 99.9% of the public doesn't care, and of the small portion of the public who is familiar with both MIT and W3C... I'll just predict that nobody is going to show up and protest, or bring torches and pitchforks, or anything because of twitter threads. Nobody is going to cut MIT's funding because of this, and they'd have to really cut in order to make MIT reconsider dumping what must be a money-loser for them already.
Really playing hardball with MIT's reputation would involve getting Tim Berners-Lee in front of the mainstream press to talk about this.
> MIT is playing hardball with people's jobs and W3C assets.
That is hardball.
The thing is, the 0.1% that do care are the people MIT care about what they think. MIT don't care what most people think. They're not giving them money, they're not giving them status, they're providing MIT with nothing. Which is kind of why those people don't care. But the people do care are giving them money and status and other stuff.
I think moving abroad would simply massively backfire on W3C - it would turn them from an org struggling to stay relevant into a completely irrelevant org immediately.
For that matter, what liabilities are we talking about here? Hosting a website? Maybe i am just naive, but what else is there?
This is obviously one sided, but assuming most of this is factual… not good.
My understanding is that besides its reputation and the fact everyone knows about it, MIT is fundamentally not really different than any other great tech university. And many big universities are starting to turn more into businesses. The "admins have seized the Ivory Tower" (https://news.ycombinator.com/item?id=33856624&ref=upstract.c...) applies to MIT as well.
At least 20 years ago, MIT's greatest strength was its student body and its culture. You got the sense everyone was striving to learn as much as they could, and most students reveled a bit in their nerdiness. In high school, I took classes at a well-regarded state school, and didn't get the same sense of intellectual hunger. IHTFP (simultaneously I Have Truly Found Paradise and I Hate This F'ing place) summed up culture pretty well. You got the sense that you and everyone else had lined up to drink from the fire hose, and were going to struggle through it together, and come out the other side better for it. I have several friends who got grey patches in their hair during undergrad from stress, that went away shortly after graduation and didn't show up again for another 15 or 20 years.
I hope that pressure cooker feeling isn't actually necessary for rigor. I hope MIT has found some way to keep the rigor while being a bit more easy on the mental health of the students. MIT ensured every month had at least one holiday by inserting one fake Monday holiday in each month without a holiday, as a mental health break. I heard the mental health breaks were a result of the high suicide rate in the 1980s. Thankfully, none of my friends committed suicide, but a few friends of friends committed suicide in my time.
It’s a bit tragic that a lot of times this comes at the cost of mental health, though MIT has gone a long way to improve that.
They also made freshman year courses pass or fail.
The university in 2019 signed a five-year extension of its lucrative partnership with the Russian technology research institute, which has long raised espionage fears among foreign policy experts and the FBI. The extension came just three months after the federal government announced it was investigating MIT’s compliance with reporting requirements for the Russian money it had received in connection with the project.
The article notes MIT only ended the cooperation after the invasion of Ukraine.
https://www.wgbh.org/news/local-news/2022/02/25/mit-abandons...
What would be the impact of the USA part of the team shutting down? The big USA companies will still be there and will keep advancing their agendas. What the rest of the world can do?
Or perhaps MIT is offering a bad deal on purpose to sink negotiations?
- Mit: $26.4B
- Yale: $42.3B
https://en.m.wikipedia.org/wiki/Yale_University_endowment https://news.mit.edu/2022/endowment-2022-1007
Also, MIT isn't Ivy League, technically.
Seems like not a neutral thread, but posturing and propaganda in its own right from the W3C.