XML, Java, and the Future of the Web (1997)
xml.com
xml.com
Instead we have HTML5 which contains nothing that couldn't have been easily expressed in XHTML but have none of it's strictness that would have enabled much simpler and better browser implementations.
The arguments against XHTML amounted to "but wahh it should still work even if the syntax is wrong". Thing is none of the major proponents of XHTML were advocating for exterminating the existing doctypes and their associated quirks modes. People could continue to use that but XHTML would be there, ready for the advent of non-shit programming to hit the web.
People are starting to come around the ideas of strictness in programming languages again with Typescript and co rapidly gaining in popularity. It's just such a massive shame the entire Web community forced this massive wheel-reinventing cycle on us for no damn reason, completely ignoring the lessons learnt -literally everywhere- else.
Just like MongoDB gaining mindshare - because webscale - only for people to realise those relationships and structured types were actually quite good at keeping your data clean.
No it doesn't. The HTML5 parsing algorithm is well-defined, and parsing it according to the spec is not an order of magnitude more difficult than properly handling XML according to its spec. What makes browsers difficult does not come from the part that involves ingesting the initial response payload and building it into a tree in a spec-compliant way—what makes things difficult is all the other stuff involved with creating and maintaining a browser.
"Invalid" HTML5: Please refer to section X.Y for the heuristics to apply in this scenario.
"Invalid" HTML4: Uhh, just do whatever? But probably you need to do whatever the other browsers do.
Implementing the other behaviours is more work than not implementing them, and may be more difficult if your internal data structures don't match those of the browser for which it was originally introduced as expedient or even emergent behaviour.
SGML has no opinions on rendering, since it's just a generic markup language. Nor does it have opinions on "what do with a hr that's a direct descendant of the table despite not being permitted", to pick the first example from the spec (it should be treated as a preceding sibling).
That said, i'm also stuck in the middle of migrating older versions of frameworks (Spring) to newer ones (Spring Boot) and it's proving to be extremely cumbersome and feels like it'll be impossible to get right without spending multiple months on the migration. Most of the older technologies are considered "dead" nowadays for a reason, though at the same time i agree that they can hold a certain wisdom.
Remember in 1994 just before Java, when Sun hired TCL's developer John Ousterhout, then unilaterally announced that TCL would be the official scripting language of web browsers?
https://en.wikipedia.org/wiki/John_Ousterhout
http://www.softpanorama.org/People/Ousterhout/index.shtml
And that triggered RMS into posting his controversial "Why you should not use TCL" announcement, which kicked off the Great TCL War of 1994:
https://news.ycombinator.com/item?id=12025218
https://vanderburg.org/old_pages/Tcl/war/
"Java is a DSL for taking large XML files and converting them to stack traces" -Andrew Back
Isn't Web Components an attempt at something like that?
Could you elaborate a bit more on that? By my understanding, order of elements in XML is significant, and child elements are kept in an ordered list.
I believe you mean tag/ as every parse I know of would choke on tag\.
if it is tag/ the XML spec says there is no difference between these two scenarios, thus most tools and libraries will serialize as tag/ (I say most because maybe somebody silly somewhere did differently)
I am optimistic for XMLs future though, as the rise of stuff like TypeScript and Rust shows that people do appreciate strict languages nowadays, and WASM means that modern versions of XSLT can be utilized on browsers even if the browser vendors cannot be bothered implementing them.
Most websites are currently blobs of div. Why ? Because the idea of a web page being mainly a document completely missed the part were artists fiddling with messy code would produce things that customers would like better.
It's the same reason we currently see giant images on home pages and 10mo static assets to load. Because in the end, the best technical decision doesn't win. The one that gives the result the customers end up preferring wins.
If you have 3 hours to get a design done, you have the choice between making it look great, or having a clean semantic and maintainable markup, you will have to drop one. And the market will select according to the result.
So one ends up with a soup of divs, made pretty with CSS and then replicate the behaviour of a beatiful dropdown fully in JavaScript, ignoring the browser builtin infrastructure.
Naturally because this breaks down in every browser, it is then extended with hacks for every corner case.
Depends on whether accessibility is a priority. Depending on where you live it may be mandatory by law, for larger organizations.
The situation were a website is made accessible from the start is such a tiny rare occasion it cannot be used to explain why xhtml failed.
I wish the author had enumerated some examples here.
I started building a custom JSON representation of HTML templates with embedded functions for importing fragments and generating markup etc and soon realized I was re-inventing XSLT [0] but kept going anyway. I haven't really seen another template language that can translate one DOM tree to another. It seems like XSLT is the eventual implementation of DSSSL in TFA, I'd be curious if anyone can speak to why it never caught on.
Once you understand apply-templates though, it becomes much easier.
https://gist.github.com/spiralx/dcb7e5fa4e5dedd800e292848ed2...
Working with lists in scheme is pretty great. I stopped using any other xml tools for my own projects.
- Ericsson's Erlang/OTP
- Sun's "The Network is the Computer" and "Utility Computing", IoT - "Jini" (Internet-connected toasters;), and non-PC devices - smartphones, set-top boxes, smart cards - JavaCard is still popular.
- Bell Labs/Lucent's Plan 9 and Inferno OS (neither is gained adoption)
- Semantic Web (which never took off).
IMO, XML and Enterprise Java were regression, but so is REST, JSON and dynamic scripting languages. The most recent example is GoLang.
EDIT: the XML itself is a good family of specs (especially XSD, XPath and XSLT), the problem was with abusing XML by forcing it onto everything: from configuration to logs to over-the-wire format like in XMPP. I remember IBM even sold XML Accelerator appliances ;)
From what? (especially XML as im not even sure what E-Java means)
I think (as you also mention yourself) that XML is quite good at some things (it's schema based, can mix schemas, has namespaces, etc). Just (as you mention as well) it was overused at some point.
Given we agree so much: what is it that XML was a regression from?
JSON is a regression from XML wrt schemas/strictness; but then it was very low overhead (which was needed as is it gain momentum used in places where the client (JS browser app) and the server (serverside of said browser app) are both maintained by the same organization).
IMO Java was a dumbed-down C++ for Enterprises and an attempt by Sun to escape Wintel monopoly. We could've had a much better mainstream programming languages now than Java or C#. Better PLs means higher bar for software quality.
It was the lack of understanding with NeXT that made it fall apart and eventually Oak made into Java.
EDIT: Here are some references,
https://cs.gmu.edu/~sean/stuff/java-objc.html
https://en.wikipedia.org/wiki/OpenStep
https://en.wikipedia.org/wiki/Distributed_Objects_Everywhere
Part of Inferno OS lives on what Go took from Limbo, and Go still isn't capable of using dynamic libraries like Limbo did.
In Limbo/Inferno one could either download the module and run the code locally or execute the code remotely without downloading it.
The migration into variable_name type_def also had already taken place, although the Limbo does require : as separator.
E.g. it's very successful in things like layout/UI description, which is naturally language-like. And not that successful in data exchange applications, which is not language-like. (To exchange data with a language-like medium is like integrating programs by constructing their command lines; possible, but way less convenient than by using a library.)
If we are to compare it with something it would be Loglan or Lojban, not JSON. The difference is that it's a working and widely used Loglan. Something like this will have to be central to navigate a vast sea of documents, which is what Web was supposed to be. Web has changed quite a bit, but those documents didn't go anywhere and the problem is as acute as ever. More acute, I would say; in 1997 there wasn't that much content there.
SAX was the hammer applied to every problem.
Though I think that namespaces are pretty trivial to add to Json. At least in a primitive form. Just add key “namespace” to some object.
And there’re multiple efforts to make Json Schema. But so far most of APIs that I saw, do not use those.
The SWIFT network is looking to migrate to XML in the next few years, and to my knowledge FedWire is looking to do the same.
Oh... typescript... you...!!!
- JSON allows arbitrary integers to be represented as numeric values. That works fine on its own, but it causes serious problems if you trust the name of the format, "JavaScript Object Notation". You can't do that in JavaScript objects.
- JSON has no comments. Or version designators.
The big problems in XML are named closing tags and the weird distinction between children and attributes. JSON did those things better, by not having them at all. But XML is a long, long way from "like JSON, but worse in every way".
(Another big problem in XML-considered-as-an-ecosystem is the prevalence of people who want to deal with it by using regular expressions. As far as the technology goes, this is a non-issue -- it's a problem with the user, not the technology. But I do have to admit that the JSON equivalent appears to be loading the data into a parser that doesn't work, as opposed to loading the data into a non-parser that doesn't work.)
JSON doesn't mandate the use of binary numbers, parsers can keep number literals as strings and let the user choose whatever binary storage.
>The big problems in XML are named closing tags and the weird distinction between children and attributes. JSON did those things better, by not having them at all.
JSON has this design freedom too: should a collection of values be an object or an array?
So it is a coin toss whether a perfectly valid JSON file can be processed (read/written - say pretty print) by a perfectly valid JSON library without its contents getting trashed? Quality software engineering right there.
What are the practical problems one runs into due to this mechanic? Isn't that just, any number literal scheme in any language ever?
It doesn't, numbers not in the range of a double have implementation-defined behavior.
* good interoperability can be achieved by implementations that expect no more precision or range than IEEE754 double precision
* for such implementations, only numbers that are integers and are in the range [-(2*53)+1, (2*53)-1] are guaranteed to represent the same number on all of them.
> An implementation may set limits on the maximum depth of nesting. An implementation may set limits on the range and precision of numbers. An implementation may set limits on the length and character contents of strings.
But those are not good ideas, and neither is rejecting numbers that are explicitly allowed by the grammar, but happen to be bigger than 9007199254740992.
There also appears to be a contradiction between these directives:
> A JSON parser MUST accept all texts that conform to the JSON grammar.
> An implementation may set limits on the size of texts that it accepts. An implementation may set limits on the maximum depth of nesting. An implementation may set limits on the range and precision of numbers. An implementation may set limits on the length and character contents of strings.
Helpfully, the RFC itself specifies which one should win:
> The key words "MUST", "MUST NOT", "REQUIRED", "SHALL", "SHALL NOT", "SHOULD", "SHOULD NOT", "RECOMMENDED", "NOT RECOMMENDED", "MAY", and "OPTIONAL" in this document are to be interpreted as described in BCP 14 [RFC2119] [RFC8174] when, and only when, they appear in all capitals, as shown here.
The discussion of limiting number values glosses over a pretty big logical hole:
> This specification allows implementations to set limits on the range and precision of numbers accepted. Since software that implements IEEE 754 binary64 (double precision) numbers [IEEE754] is generally available and widely used, good interoperability can be achieved by implementations that expect no more precision or range than these provide, in the sense that implementations will approximate JSON numbers within the expected precision. A JSON number such as 1E400 or 3.141592653589793238462643383279 may indicate potential interoperability problems, since it suggests that the software that created it expects receiving software to have greater capabilities for numeric magnitude and precision than is widely available.
Of course, integer values of more than 54 bits are also generally available and widely used. A JSON number such as 36028797018963968 does not suggest that it expects receiving software to have any greater capability for numeric magnitude or precision than is widely available. The actual reason for the mention of integer value restrictions is not discussed, but it is mentioned earlier in the RFC:
> JSON's design goals were for it to be minimal, portable, textual, and a subset of JavaScript.
These goals were not achieved, but they are still corrupting the discussion of numeric values. This RFC is a dog's breakfast. Writing "arbitrary noncompliance will be allowed" makes a mockery of the idea of a standard, and according to its own terms, the document does not even allow the restrictions it claims to allow. This document doesn't do anything except attempt to legitimate any and all existing or future "JSON parsers". There is no JSON data which, according to the lowercased restriction allowances, can be guaranteed to be accepted by a "compliant" JSON parser.
Regarding limitations on text size, you can always return an out of memory error. As long as it's not a _parsing_ error, it's technically fine, you are still accepting all texts that conform to the JSON grammar but telling the client that there's not enough memory to store the parsed output.
Which has proven to have zero real-world use case.
> - JSON allows arbitrary integers to be represented as numeric values. That works fine on its own, but it causes serious problems if you trust the name of the format, "JavaScript Object Notation". You can't do that in JavaScript objects.
That's not a problem with JSON, that's a problem with standardized JavaScript. I believe JSON was named before there was an official standard that required JavaScript implementations to be terrible.
> - JSON has no comments. Or version designators.
Indeed, and it's all the better for it.
> Indeed, and it's all the better for it.
JSON5 attempts to fix the comments problem: https://json5.org/
A lack of comments is never good, especially when you want to provide a workable example of what an entity looks like, while annotating its contents, but at the same time allowing it to be parseable. Furthermore, i wouldn't scoff if i ever saw comments about non-trivial fields in web APIs actually being returned, to better explain how to use them. Of course, at the same time i believe that something like that would be better suited as a part of WSDL, WADL or XSD schemas, but JSON and the technologies around it have essentially done away with strict schemas, which makes using them about as reassuring as dynamic languages - i.e. unreliable.
Far too few people do properly versioned APIs and far too few people do OpenAPI specs that are generated from code automatically and are publicly available for the above to be a moot point. Whereas with XML based services, i could open the WSDL/XSD file for a 15 year old API and know what's going on within 10 minutes. It seems like the industry has lost that bit of wisdom somewhere along the way of chasing agility - the same way that knowing how to generate code and operate with metamodels and models has also been tossed aside.
It also attempts to fix some other problems. I'm fond of "numbers may be hexadecimal".
I can't resist observing how much easier the fixes would be, if JSON data included a version designator.
You can't expect a number from an API and get a boolean back, without your system breaking. There should be contracts between any two parts of a system, or any interlinked systems.
It's inexcusable to have breaking changes without doing something to give the users the ability to react to them before breakage: be it changing the signatures and deprecating the old methods for libraries which would be picked up by CI and would not build the code before these being addressed, or changing a WSDL/WADL/OpenAPI service description, which would then propagate into failing integration tests before new versions would be deployed.
You'd get a JSON schema of v14 and then another of v15 which would introduce breaking changes, but all of the downstream systems would see the changes and essentially figure out that they cannot use this new API (ideally, in a scheduled and automated process).
So essentially, it would be like this:
- you have an old API version that is used
- for example: your-app.com/api/v14/pictures/cats/bambino?size_x=640&size_y=480&page=5
- the new API version would get released, which would change paging semantics
- your-app.com/api/v15/pictures/cats/bambino?size_x=640&size_y=480&count=10&offset=40
- the CI system would pick up changes from /api/v14.json and /api/v15.json service descriptions
- it would detect breaking changes and developers would be alerted (either automatically as a part of integration tests, or when trying to update integrations)
- they'd update the API integration to address these issues, before the old would eventually be sunsetted
Alternatively, even a header about deprecation being returned on the current API endpoint would be better than one day just discovering that production has broken: https://tools.ietf.org/id/draft-dalal-deprecation-header-03....Of course, sadly most companies out there aren't interested in versioning their APIs or even providing any sorts of service descriptions, because all of that costs money and time. So in the moniker of "Move fast and break things" the focus ends up being on the second part.
Nope. The ECMAScript standard released in 1999 already specifies:
> In ECMAScript, the set of values represents the double-precision 64-bit format IEEE 754 values including the special “Not-a-Number” (NaN) values, positive infinity, and negative infinity.
https://www.ecma-international.org/wp-content/uploads/ECMA-2...
Care to explain more? I though they were over engineered compared to bare bones JSON. For e.g. JSON schema is still catching up to XML schema and the ecosystem it had back then (XSLT, XPATH) etc.
XML is of course heavy weight, but we should compare YAML against XML, not JSON - JSON doesn't have for e.g. comments, CDATA etc.
In fact, Pandoc's internal format is exportable in JSON
It's self-evident that XML is better for the mixed content problem (which makes sense, given that mixed content is something that it was consciously designed for). Anyone arguing otherwise is, ironically, approaching the issue by narrowly considering only the things that JSON is good for and then evaluating XML against that—e.g. without it ever occurring to them that they should consider mixed content, as in the example given—or they're deluding themselves/being dishonest.
In your imagination, maybe, but otherwise, no, it's not. The claim being prosecuted is that "Everything XML wanted to be, JSON did better." (It would be bad enough to lose sight of this once, but I explicitly repeated it for your benefit in the comment that you you are directly replying to. This iteration makes #3.)
Thankfully we got rid of it and replaced it with data formats that got rid of the anti-features like comments.
Now YAML is mostly replacing JSON in places where you want human readability, and JSON is only used for data exchange.
"The Web was a huge success, there must be a reason. Oh I know, it must have been HTML. We need more of that, and bigger, extensible! Behold, XML!"
I remember shaking my head in wonder during the height of the XML hype. What were all those supposedly very smart people thinking?
To me, HTML was already a huge mistake, the Web succeeded in spite, not because of it. What could you do in HTML/XML that you couldn't in s-expression, only simpler, more readable, and more writable? How many man-years have we collectively wasted having to read/write/generate/process/validate/... the mess that is HTML/XML, as opposed to a simpler, more sensible format such as s-expression?