Better Than JSON?
wiki.alopex.li
wiki.alopex.li
> Not sure anyone really knows how XML happened. It’s
> basically the W3C’s fault, I think? It’s okay for some
> things but in the end I’m not sure it’s something anyone
> actually wants to use, it’s just going to be one more of
> those mistakes of the past.
Look, I was doing web dev when XMLRPC was in. For simple API stuff, JSON ended up being worlds better. No fiddly XML preamble, no schemas, no envelopes, just enough structure to be able to quickly pull out a couple of fields and we're set. Beautiful.But let me tell you how much I appreciated XML and schema validation when we were working with fifteen different partners all using a standardized XML format with a schema. These are cats who were putting XSDs into their IDEs, clicking generate to get a class, and then just not caring.
So when our integrations broke because they were sending XML payloads that failed schema validation, our conversations were so easy.
Before you say "yes but technically you could have done the same with $OTHER_THING" -- sure it's possible, but XML is ubiquitous. Shit, our tools were in Ruby.
XML can be a bear, but sometimes you need a bear.
I do agree with the XML criticisms around writability and complexity, though I think if you made a Venn diagram of the actual complexities inherent in managing a multi-client document-exchange system and the complexities of XML+Schema, you'd find that they're similar--and that many people who start building such an exchange with JSON documents end up stumbling into the same amount of complexity except in an ad hoc fashion.
It's funny, I wonder if other people have the same surreal experience: as the mass zeitgeist moved away from XML I was more or less like, okay, gotta go with it, as the driftwood goes with the currents. It's interesting how the zeitgeist functions. You can tell it's happening when people look at your functional system and are like, "Uhh, why aren't you using X...?" as though it's self-evident that you should drop everything and rewrite. Where X was, over the years (dating myself) {"Java", "PHP", "Node", "Rails", "Jquery", "REST", "Angular", "React", "Thrift", "NoSQL", "Protobuf", "Hadoop"}. Some of which I quite happily used, BTW, so the interesting part is the penetrative nature of the conversation, e.g. how the quite-functional-but-not-hot technology becomes the sick gazelle falling behind the herd, even as it is not empirically sick.
And then the ratchet turns, and the pendelum swings in reverse, and suddenly that thing sent from god was instead discovered to be sent from satan... and you're still trying to explain that tradeoffs exist
XML can be a bear, but sometimes you need a bear.
Phrase heard at conferences in the 90's/2000's: "XML is like violence. If it doesn't work, you're not using enough of it."That same saying would also work replacing XML for soap (but not SOAP).
This coming from someone who hates XML, avoids it at nearly all costs - there are times it's the right tool for the job. And you list that time. If you have a 1-1 API contract, don't use XML. If you have 1-N, don't use XML. But if you have N-N API contracts? Where you're talking to half a dozen other services, and they're all talking to you? XML is an excellent choice.
Good times.
Fine, but with for example, Protobuf, you share a .proto file and so there's zero percentage chance of this happening.
[0] https://github.com/edn-format/edn
[1] https://github.com/cognitect/transit-format
[2] https://github.com/Datomic/fressian/wiki
[3] https://www.youtube.com/watch?v=JArZqMqsaB0&ab_channel=Cloju...
Binary formats win for the following cases:
- payload size
- serialisation/deserialisation speed
- simplicity of client / server code (e.g. replacing a full blown HTTP server with a simple ZeroMQ one)
Sure, having payloads that are human readable is great for initial debugging/verification, but once things are put live it's just an unnecessary expense.
HTTP, by itself, is one of the more complex layer-7 protocols known to man. UTF-8 is not simple. Converting a decimal string representation of a floating point value into IEEE 754 is not simple (particularly compared to... memcpy). In a bizarre example of the tail wagging the dog, the ARM ISA has added an instruction specifically for parsing JSON...
There's really not a lot of value in having different serialization formats, so the "even if it doesn't make sense" requires a specific context for determining what makes sense.
It is nothing to do with JSON.
technically, "arbitrary precision decimal"; floating point is a storage representation, to which JSON is formally agnostic, though RFC 8259 recommends that limiting to the range and precision representable with IEEE 754 doubles is a good idea for interoperability.
I totally hear you on serialization/deserialization speed, though it is amazing how inefficient some of the "machine readible" implementations can be, and how efficient people have been able to get JSON parsing, its horrible design disadvantage in this area is lethal.
The simplicity argument is an amusing one to me. I think a lot of people build simple client/server code for JSON, but these "simple" implementations have all kinds of rough edges that prove problematic long term. Once you start doing things right, JSON proves to be far more complex.
I really mean 'typically useless' in the scenario where the communication is being done internally, not when you're actually communicating over the public internet. There can be quite a bit of wastage in this scenario. I really don't care about the 'Date' header or the 'Server' header, but the HTTP spec does, so they're there.
All I'm saying is that while JSON is horribly inefficient for some payloads, the reality is that most serialization formats tend to be efficient for specific kinds of payloads, but still have plenty of common cases they don't attempt to be efficient with. For example, machine readable formats like protobuf, thrift, avro, etc. employ some variation on varint encoding to keep small integers compactly represented, they encode strings relatively inefficiently; They encode field type & number, followed by a varint for length, followed by uncompressed UTF-8 encoded payload... and I've found that if anything it seems once you go through the compression ringer it often isn't much different in terms of space consumption vs. JSON... and sometimes it can be less efficient. Given how often fields end up being encoded as strings (even though they should be enums, ints, etc.), this has resulted in a lot of surprising cases for teams who have tried switching away from JSON to reduce payload sizes. If you cared about efficiency, you'd probably have a special encoding mechanism for short strings, and then a reasonably compact (but fast to parse) representation for longer strings, probably using something like SCSU encoding instead of UTF-8.
JSON Schema http://json-schema.org/ is widely used for things like Swagger / OpenAPI.
https://diziet.dreamwidth.org/6568.html
SBE (Simple Binary Encoding) also makes some interesting performance claims.
My pet peeve with both protobuf and capnp is the lack of an option type; they seem to be designed for languages with type systems that include null, rather than more modern languages like Rust or Kotlin.
Btw JSON is often a popular choice, but although this is not explicitly mentioned in the main article as a con which is an oversight imo, the actual JSON standard does not support NaN/+-Inf for floating point. That's insane, which is why most implementations have an option to allow that anyway, but not all, which can be quite the showstopper.
The idea of JSON is to offer two important ways to aggregate, which are different enough that they deserve separate treatment - ordered array and associative array. They are economically represented, with [] and {}, and internal structure is also made with small costs - commas (,) between elements and colons (:) for keys in associative array. We can argue that both of those features are required.
Chosen primitives - numbers, strings, booleans and null - are also selected from what's regularly and consistently used, and the choice is supported by two decades of JSON application. Roots of this selection are in design of JavaScript, which in turn relied on common practice for basic primitives.
If JSON is considered as what I think it is, structure out of basic primitives, it's at the optimal point. Crockford's decision to avoid JSON versioning makes a good sense.
Applications of JSON - human-readable texts, performance-optimized communications - can suffer from e.g. lack of built-in comments, non-extensible "type system", non-optimal bandwidth usage (if direct ASCII or even UTF-8 is used for JSON encoding, not some other compressed approach). They however can be fixed by building on top of JSON. For example, parties can agree upon reserved keys and compression schemas, constraints in form of grammars (JSON-Schema). This flexibility stems from the fact that JSON is at the very foundation of data structures, so "better than JSON" is to an extent like "better than two's complement number representation".
People forget this, but ASCII (and therefore UTF-8) already has characters in it for field/record/etc. separation. One increases complexity significantly (and introduces inefficiency) by overloading other characters with their semantics.
There are much more simple & efficient ways to demarcate ordered arrays in particular (e.g. simple length prefixed structures).
Nulls are highly problematic for any number of reasons that are widely documented, and JSON unfortunately compounds this by having two ways of representing them (explicitly and by omission).
One might quibble with the notion that a format that didn't even exist in 2000 doesn't really have two decades worth of applications working on it, but setting that aside, the design constraints of JavaScript and the challenges that have emerged from it are far from ideal. The work arounds that have emerged (the whole Number fiasco, the "standardizing" on UTF-8 while using UTF-16 when escaping, line terminators, the backslash escapign of control characters, the inconsistencies between JSON & JSONP, the binary64 "good interoperability" rule, and comments being disallowed, and of course the various JSON Schema efforts).
While I agree that an ideal serialization standard would not have versioning, and Crockford's decision to avoid JSON versioning might consequently seem to make good sense... in practice this has lead to significant inconsistencies between implementations, many of which lead to interoperability problems that could at least be somewhat mitigated by versioning. (http://seriot.ch/parsing_json.php)
"Fixing" a bad abstraction by layering on a bad abstraction is a common house of cards meme. It is highly problematic to say the least. It's how "simple" JSON rapidly evolves to be at least as complex (and at least as error prone) as ASN.1.
This made me laugh more than it probably should have.
The kiss of death. It's a graveyard of things from a previous era, pure maintenance mode. I think library authors get tired of maintaining for free then hand it off to these life support organizations. The new way for open source software is a decentralized ecosystem where people fix software they actually use, so the cream rises to the top on its own and you don't need the sponsorship of these digital museums.
So there's a ton of utterly crap XML applications out there with no rhyme or reason, and these crap XML enterprise apps are most developer's understanding of "real world XML". Similar to how terrible C++ code being all over the place is making people hate C++ the language.
-------
XML can be used very well however: its use in Jabber / XMPP is clear and concise. XML Namespaces are needed due to the extensions of XMPP, and various programs can interact with the protocol. Specifications are clear, the ability to "stream" data is proven, different parser strategies have been implemented. Its clear that XML in this instance works, and works pretty well.
I mean if there's enough of it, it kinda becomes "real world XML". Right or wrong human's ability to actually use the tool or tech well is often where the bar is set. Could be a great idea but if folks can't do it... won't fly.
I can't count how often I've seen XML that looked like this:
<container name="containerName">
<attribute name="attrib1" value="value1" />
<attribute name="attrib2" value="value2" />
<attribute name="listAttibute" value="pipe|delimited|values" />
<attribute name="containedAttribute" value="<container name="otherContainer"..." />
...
</container>
Which was, technically, XML (it even conformed to a schema!)... almost like "malicious compliance".I think that's the main problem of XML it's not meant to be read by humans but people thought it was.
So you have an already complex document parser (which you are using only a fraction of the full functionality!) with the added complexity of a schema system layered on top. This is why JSON took off so fast, it ditched the whole document part and just gave you the data.
<items>
<item id="1" attr1="value 1" attr2="value 2" attr3="value 3" />
<item id="2" attr1="value 1" attr2="value 2" attr3="value 3" />
<item id="3" attr1="value 1" attr2="value 2" attr3="value 3" />
<item id="4" attr1="value 1" attr2="value 2" attr3="value 3" />
<item id="5" attr1="value 1" attr2="value 2" attr3="value 3" />
</item>
It's not something that an IBM committee would have come up with but it fits the needs of an application. If XML has an "intrinsic fault" it's related to DTDs and namespaces which allow for the creation of horrible, complex monstrosities. XML can be every bit as stripped down and to-the-point as JSON, trading the curley braces for angle brackets. JSON's best point is that it doesn't support any of this needless complexity but it's still easy enough, with enough nested objects and arrays, to make something that's effectively not readable by a human.Items can be order-dependent in XML. So the "id=1" thing is still superfluous. See xs:sequence. (In contrast: xs:all declares something to be order independent).
If you were doing XML for your own apps, I think xs:sequence items simplify parsing grossly.
Ex: if you have Address, maybe you always want "Street", then "Zip Code", then "State".
<Address>
<Street> 1234 Blah Street </Street>
<City> Springfield </City>
<State> Whatever </State>
<Zip> 55555 </Zip>
</Address>
By declaring that Street -> City -> State -> Zip must happen in that particular order, you can grossly simplify the parser's job when reading such data in. (Or writing it out).--------
If there are multiple People at this address, you could (and probably should) just list them out in order.
<Address>
<Street> 1234 Blah Street </Street>
<City> Springfield </City>
<State> Whatever </State>
<Zip> 55555 </Zip>
<People>
<p> Marge </p>
<p> Homer </p>
<p> Lisa </p>
<p> Bart </p>
<p> Maggie </p>
</People>
</Address>It's also highly redundant (so it basically can't be hand written), undesrpecified (do I put this value as an atribute, value, or sub-element?), underpowered (how does it represent numbers? Byte streams?), complicated (have you ever created a DTD by hand?), and just out badly designed.
Of course, JSON and YAML share some of those same problems, and add some different ones. But the complaints about XML aren't some superfluous opinions about how it looks.
Oh there are lots of ways to beat XML into looking like it can do it but they’re all bad, break all the actual semantics of XML and turn a nice deserialization story of a tree where each node is a complete thing into mess.
Turns out that very few people actually needed a tree of heterogeneous string k-v pairs.
For comparison, time formats in JSON are all over the place.
In XML, there was always some motherfucker trying to stuff CSV into an attribute.
And the difference between ID as specified and ID as implemented lead to a lot of problems, reaching its zenith (or maybe nadir?) in the XML Signature spec.
If you need three values use an array (or in xml, child element)
I haven't tried multiline strings but it allows comments and trailing commas
You don't have to worry about maintaining numeric field tags, can remove fields, can make previously required fields optional (by promoting them to unions with null), or even change their type (by promoting them to unions of <old type> and <new type>)
You need to pass your schema around with your data somehow, but there's a file format specified for that, and you can still just dump to JSON if you can't be bothered.
The schemas are specified in JSON too... which in my view makes it more robust than the JSON + JSON-Schema combo
[0] https://avro.apache.org/docs/current/spec.html#Schema+Resolu...
It's both binary and text, and supports all the common data types, and doesn't require a schema or extra compilation steps.
> “CAR”/“CDR” are not a part of the expression syntax
If you were implementing them for a serialization format, you'd leave out cons cells and car / cdr.
But in late-binding LISPs, which is where they're typically used, (a . b) is part of the syntax. It's worth calling out for why you don't want it: it's badly typed since it could be a 2-tuple, or it could be a list depending on what 'b' is.
So S-expressions suffer a similar problem to XML: there is not a clear way to represent many data structures.
Welcome to Node.js v12.18.4.
Type ".help" for more information.
> new Date(-10000)
1969-12-31T23:59:50.000ZSee: http://erlang.org/doc/apps/erts/erl_ext_dist.html
and for an overview:
https://medium.com/@niamtokik/serialization-series-do-you-sp...
I've been using it lately but i'm i'll qualified to review it to any meaningful degree. It's neat, (reportedly) fast, but has seeming zero traction and thus makes me uneasy.
(I'm currently (configurably) using it in place of bincode, for a content addressed store. for some context)
We'd also have a use case where we think about using bincode, this seems worth to check out/compare as a possible alternative.
From an aesthetic point-of-view, XML is ideal. It reads like S-expressions or Lisp but with enough features to be a robust serialization format or even represent code. Validation alone puts it ahead of JSON. Too verbose? What kind of editor are you using that can't add the closing tag automatically or show you validation errors as you go. If type checking is good, then XML is good.
The main reason XML is hated now is that it's too difficult to learn.
People that care about performance will use protobufs/capnproto, which has all the safety of XML and more. People that want something to work out of the box everywhere without too much thought will use JSON. There is no room for caring about correctness for its own sake.
Doesn't instill confidence in that content.
fyi: XML was subset by W3C (the SGML "extended review board" specifically) from SGML to become the generic syntax for web vocabularies (replace HTML syntax by XTHML, and define new SVG, MathML vocabularies). Didn't work out on the web, though.
It is pretty clear by the language used that it's all written tongue in cheek.
(jk)
Very interesting article and I like the style in which it was written.
Another factor to consider is whether the serialisation code is intrusive or not. Being intrusive can be an issue for legacy codebases, and IMO is not paricularly welcome regardless.
Anything using a schema that is agreed upon between participants isn't like JSON at all.
There are more. I don't know if java dot-properties file is worth including.
All of the mentioned formats are now insecure by default, because they deserialize objects. (Just msgpack not, which has other problems). Even JSON, the most secure format originally became insecure in its 1st update. All the binary JSON variants, like BSON went bonkers.
I am not sure what the takeaway is from it aside from the author's personal opinions on various technologies
JavaScript is better than JSON because it has loops, conditionals, comments, modules, functional programs, typing, etc etc
- Fluentd & Fluent Bit
- Microsoft SignalR https://docs.microsoft.com/en-us/aspnet/core/signalr/message...
The best part is that these are a superset of JSON and could easily be added if there would just be more buy-in (lots of unofficial support already exists).