The Rise and Rise of JSON (2017)
twobithistory.org
twobithistory.org
The litany is not short and I'm sure I've overlooked a couple.
But if you add too many features, it may end up like XML.
Perhaps CDDL:
For numeric types, if you handle integer larger than 2 billions or floating numbers where the precision is critical (that's where json fails), these can and should be represented as string. I don't think there is any common format that guarantees large numbers and/or arbitrary precision to be encoded and decoded properly across languages and platforms.
They're really bad for a data interchange format though, because people inevitably start putting important data in them and you end up with two different ways to write strings, one of which isn't supported by every parsing library.
Thus each discussion of JSON ends up being two groups talking past each other, the people using it as a configuration format who lament the lack of comments, and the people using it as a data interchange format who celebrate it.
The solution: since it's too late to add comments now, don't use JSON as a configuration format.
I don't have a single competitor to recommend, but JSON5, TOML, JSONC (mentioned in a sister comment), or something else along those lines might be better. I'd probably just go with whichever of those is popular in your community.
(and yaml is TERRIBLE)
For me, readability is king both for myself and because I sometimes want non-developers to be able to hand-edit files.
Any curly brace format is a non-starter. Too much clutter, too easy to create invalid files that aren't obviously invalid at a glance.
Significant white space does have the advantage of meaning exactly what it appears to mean and being intuitively understandable to most people.
For data interchange formats comments do presents major problem.
Any feature or tool can be abused and misused. Comments are useful for configuration files, period.
The core project implements a JavaScript module [1] but it looks like there are also experimental JSON5 libraries for Go [2], Python [3], Ruby [4], and Rust [5].
[1]: https://www.npmjs.com/package/json5
[2]: https://github.com/yosuke-furukawa/json5
[3]: https://pypi.org/project/json5/
[4]: https://github.com/bartoszkopinski/json5
I noticed that JSON5 doesn’t include a Date type. There have been discussions about this in both the JavaScript [6] and spec [7] repos but with no resolution yet.
<script src="//unpkg.com/json5/dist/index.min.js"></script>
to every HTML page as extra JavaScript instead of a native browser implementation.Per Douglas Crockford, the creator of JSON, comments were an anti-feature:
> I removed comments from JSON because I saw people were using them to hold parsing directives, a practice which would have destroyed interoperability. I know that the lack of comments makes some people sad, but it shouldn't.
> Suppose you are using JSON to keep configuration files, which you would like to annotate. Go ahead and insert all the comments you like. Then pipe it through JSMin before handing it to your JSON parser.
* https://web.archive.org/web/20120507093915/https://plus.goog...
Discussion on the post (2012):
* https://news.ycombinator.com/item?id=3912149
Remember: JSON was designed primarily as a data exchange format between computers.
I think this is a bit of 'over reach' in what JSON was intended to do, so it's perhaps not surprising that there may be some things 'lacking' for that purpose.
This allows trailing commas and comments. This can also reads strict json of course. Libraries are available for all popular languages.
However, JSON is hitting the same limitations and problems that XML faced, and is following in their shoes (namespaces, schemas, x/jpath, implementation drift across libraries, etc).
With JSON, you have to go looking for this sort of trouble if you want to find it. An out-of-the-box JSON parser is all but guaranteed not to coerce a schema, fetch remote resources, apply namespaces, and so on -- if you want to do those things, you have to go out of your way to do them, and your decision to do so is likely motivated by considering tradeoffs rather than cargo-culting because the decisions are A) explicit and B) made in a cultural context where schema-free and remote-ref-free JSON is common practice rather than a sneered-upon possibility.
With XML, you don't have to go looking for trouble, the trouble finds you. XML's ecosystem absolutely did push everyone in the direction of using schemas and remote-refs. DTDs were both common and encouraged with a heavy hand. All the Big Protocols used them, and instead of emphasizing the differences between Big Protocol use cases and duct-tape use cases, the culture tended to get preachy about DTDs being "the right way" to do things. If you chose differently, you were amateur scum not worthy of the Enterprise. This encouraged cargo-cult application of DTDs in places where they weren't appropriate (e.g. maximally economic / flexible glue) and for purposes that weren't appropriate (e.g. as a substitute for documentation). It also lent "moral license" to the wrong party in validation-related toe-stepping. Party A turns on validation, Party B sees their integration stop working, Party A argues that if Party B hadn't been writing morally degenerate XML they wouldn't have been caught out and can point at an abundance of preachy XML gospel to back them up. Party A gets away with arguing that an API change isn't an API change, or that a new internet/VPN dependency isn't new. Party B gets trod upon and remembers that XML was complicit in it. Sure, the XML ecosystem could have avoided these problems without a single code-change or standards-change by creating a culture of acceptance around schema-free XML, but they didn't.
When JSON promised both cultural acceptance and appropriate defaults for the common low-rent use case, people flocked to it, and rightly so. The rest is history.
And is it following in XML's shoes? I've not had to work on or integrate with any systems that did this Enterprise™ silliness either.
Now granted, I'm fairly separated from the world of IBM and SAP type companies, and I'm sure there's all sorts of unholy abuse of JSON going on in a dark server room somewhere, but the answer in most cases is just to choose technology stacks that don't do that to you. These problems are generally self-imposed rather than imposed by the technology.
Yes. Some folks, even at my relatively small company, are using JSON and also pushing for extensive schema use.
JSON - just like XML - allows you to use its vanilla form. But the additional tools are available if/when you need them, because the need exists.
I think that's why JSON won: it was first and foremost and engineer-driven standard as opposed to one that incubated inside a bunch of enterprises before being dropped on an unsuspecting world.
S-expressions have the advantage of not containing the redundant object type (there's no need for an object, map or dictionary when one has alists or plists), and the even greater advantage of elegantly representing code.
Also they are easier to write parsers for.
I can see why Lisp style syntax is bad for human comprehension. It’s because of the lack of an explicit delimiter, like a comma. I don’t think human eyes are very good at using a single space character, as a delimiter.
And with this format, you’d have to agree upon the data exchange structure. There would have to be a master key somewhere, maybe as the header. Albeit, this format is far more efficient than XML.
But, XML probably won out because it was explicit, and allowed multiple levels of nesting, but at the expense of overly wordy delimiter tags.
And JSON made it simpler, by forcing it to be a lightweight key-value pair.
It’s too bad something like this didn’t take root and become more popular. This might be a very useful advanced data exchange format for some scenarios. Although the headaches and problems associated with trying to understand, and work with it, might far outweigh its efficiency benefits.
Unless of course you'd put the comments in the data itself, but that's kind of ugly.
Or do you want to send different parts of e.g. a map in separate streams?
If I got multiple items, they're in an array. If such a reader blocks reads to an element until it's finished, I've won nothing if I have to wait until that array finished.
The utility here is if you can process each element one at a time. If you need everything to get a meaningful result it offers no benefit to you.
Then you parse elements one by one. While your particular library may not support it, it's not a very hard thing to implement
["a", 1, "b
A parser could at this point give us the "a" and the 1, but not the "b" since the string could have more content.You probably want the parser to give you an object that behaves as a collection you can for-each loop over, and that blocks when there is no more data available.
Where do you think it's different from streaming XML (SAX)?
Plenty of software that uses regular JSON also strips unused data when writing back to a file it read because internally it transforms the document into a data structure with no space dedicated for unknown data.
It would be up to the software, or a very strict standard or library, to preserve unknown/unused data.
JSON is normally "good enough" as a general wire interchange format, and it's human-readable and whoever comes after me will already be familiar with it. But if there's a time when it's not good enough and its failures are in expressiveness, I'd totally consider JSON5 or some other alternative instead.
{
"_COMMENT": "default; override with --s3-bucket",
"s3_bucket": "foo-bar"
} [1, 2, 3, 3.14159, "comment: Pi!!", 4, 5, 6, 7, 8]
over: [1, 2, 3, 3.14159 /* Pi!! */, 4, 5, 6, 7, 8] [1, 2, 3.5 /* faster, use 3 for more stability */, 4]JSON was first designed as a data exchange format between computers (that was lighter weight than XML (which was lighter weight that SGML)), so that fact that it's being used in persistent fashions is getting away from its primary purpose.
I suspect that a lot of the popularity is piggy-backing on the popularity of JavaScript.
But what the heck is .CSV doing on an incline?!
I always thought it was chosen since it was a seemingly sensible delimiter at first glance and then you only realize it's less useful well after you see others using your spreadsheet app for storing complex strings.
It’s hard to find references of the intentions of the creators for a format this old, which evolved over decades of different software having similar input formats, starting in 1972. Most of these early descriptions accept both spaces and commas as separators, hinting at a manual input process.
The closest I can find is a 1983 reference¹ which indicated that the format “…contains no other control characters except the end-of-file character…”, which I take to mean that the format is easy to handle and process with other text-handling tools, but still, no definite reference of direct intent.
https://archive.org/stream/bitsavers_osborneexeutiveRef1983_...
It does seem like it was made to be easy to visually parse but it still feels like the delimiter choice has been something of historical baggage (to the point where just changing the delimiter itself, such as TSV, makes it easier to parse).
Sometime between 8-12 I 'invented' csv when I need to save data for some simple game I built. It seems like the most obvious solution to come up with: store data in a text file, separate the fields with a comma (or other character) and the 'entries' with a newline.
Once you're using a binary format, you might as well use something like parquet or hdf5 (or ASN.1 I suppose) that give you other goodies in addition to airtight data/organization separation.
It is still the most common interchange format for tabular data. Tabular data is too verbose/unreadable when represented in standard JSON. Databases and Data Science are also on the rise; .csv continues to ride this wave.
In other words, people thinking they can just join(",") keeps them from reaching for an actual serializer (JSON, XML, etc), and if they realized they already have to bring in a CSV library anyways, they might consider using another format. Comedy gold, huh? Though I'm only half-joking, I've done it before and we've all consumed "CSV" from people who had the some presumption. :)
The downside is that you can’t process a row at a time.
T̵h̵e̵ ̵l̵a̵t̵e̵s̵t̵ ̵v̵e̵r̵s̵i̵o̵n̵ ̵o̵f̵ ̵J̵S̵O̵N̵ ̵w̵a̵s̵ ̵a̵p̵p̵a̵r̵e̵n̵t̵l̵y̵ ̵r̵e̵l̵e̵a̵s̵e̵d̵ ̵l̵e̵s̵s̵ ̵t̵h̵a̵n̵ ̵a̵ ̵y̵e̵a̵r̵ ̵a̵g̵o̵.̵ [EDIT: I was mistaken: the latest standard was RFC 8259. However, this was still published on 2017-12-13, after the article was written.] There have been three different RFCs alone, all defining JSON. (And let’s not even get into the thing called JSON5.)
* https://tools.ietf.org/html/rfc8259#appendix-A
Not very Earth shattering revisions IMHO.
<complexType name="MyMethodResponse">
<sequence>
<element name="A" type="string" />
<element name="B" type="string" />
</sequence>
</complexType>
and you later add a "C", existing clients will fail, so you will have to create a new method along side the old one to remain compatible.Similar to defaulting to error-ing on extra tags, I never got the point of making everything an element and adding extra redundant attributes. That's really just the result of XML that's automatically generated.
Like the `<element... type="string">` in your example of classic XML, why not just use `<a>3</a>` or `<a int="3"/>`. XML proper styles still suggest the classic long/autogenerated form. Writing XML in html5 style like `<my-a data-value="3" />` makes it much friendlier. Or as I tend to use it for various internal protocols to wrap multiple CSV data in a file:
<some-data type"csv" columns="A, B, C">
1.3, 3.4, 5.6
1.3, 3.4, 5.6
</some-data>
Technically That's really just html5 in a file I guess. :-) But html parsers also don't tend to complain about extra tags too.I recently used GraphQL on a project and it had some nice advantages. I love the idea of protocol buffers but have never had the chance to use them in anger. But if I'm honest, the boring option is JSON and it is what I would use for just about any API I had to expose nowadays.
Previous discussion: https://news.ycombinator.com/item?id=17832936
I had to add XML and CSV, because "nobody would integrate with a JSON API"
In 2011 we moved from an PHP app to an SPA where JSON came in handy.
It's a format that still bears excessive decoration (what's the purpose of quotes around field names? what are all those commas for?) yet it's limited in the types of data structures that it's able to express (natively). I'm not particularly fond of Clojure specifically but a format like EDN would have been superior in just about every way.
The datastructure complexity being limited is also a pretty significant key to its success. More complex datatypes means greater chances for JSON handling libraries to lack compatibility.
The only substantial shortcomings of JSON I see are shortcomings associated with any textual serialization format. Optimizing for human readability in a use case that's 99.99% of the time not read by a human.
That's literally how low you have to go.
"The only substantial shortcomings of JSON I see are shortcomings associated with any textual serialization format. Optimizing for human readability in a use case that's not 99.99% of the time not read by a human."
There are a few other shortcomings that are reasonably substantial/significant, but yeah, that's the gist of the problem.
JSON looks good compared to what we'd be using instead of JSON, which is nothing so nice and structured as XML. The competition to JSON is something infinitely more ad-hoc, probably without a distinct parser, such that a generic library to generate or consume it is impossible, and getting usable error messages is equally impossible.
> "The only substantial shortcomings of JSON I see are shortcomings associated with any textual serialization format. Optimizing for human readability in a use case that's not 99.99% of the time not read by a human."
I agree with this and disagree at the same time: Optimizing for human readability means optimizing for the weird case, the 0.01% (but it seems to be more often than that) of the time you need to go beyond the tools you have to fix something. Saying that's rare is true but inapt: Seatbelts are only used in rare cases, too.
It's like, "hey, we came up with a simple way to do it, but to make it easier to deal with this little edge condition that creates complexity, let's significantly up the complexity and number of edge conditions so they're endemic to the space, and then we're all good go".
The irony is, I invariably end up needing to use a computer to help me read JSON anyway.
If you've got good built in tooling for payload visualization, then you might have minimal overhead to debug from a text like format. Both protobufs and flatbuffers (not to mention BSON), have good tools that spit out JSON equivalents.
In reality however, there are some cases where protobufs, flatbuffers, or BSON are superior to JSON, but there are a lot of cases where they aren't. You'll have to weigh the pros and cons for each situation. And a lot of the time, there's not time to benchmark everything, so you kind of have to guess how the elements of the system are going to interact.
The one element that every system has is humans, so it's a fairly safe bet that humans will have to read whatever format you use.
I spend probably 15-45 minutes a day just in Postman, testing JSON calls. If something goes wrong, I'm inspecting requests/responses in Chrome. When we integrate a new team member, we don't have to have them install any tools--they're included in the browser they have installed. We don't have to write any schemas. When I start a new project, I don't have to install any libraries: they're included in my language(s). When we integrate with a partner company, we hand them sample requests/responses as text.
How many JSON documents per day do we have to send to get the payoffs you're claiming?
And my company is not unique: in fact, the stack I'm using is one of the most common stacks on the market.
In reality, now, JSON opens in everything from Chrome network inspector to Vim, and protobuffers/flatbuffers don't.
Chrome also decompresses gzip and understands TCP/IP, but it doesn't handle LZMA or SS7.
Sure, thanks to our cult-like following of bad principles, we've made support of on a pretty broken stack with lots of terrible consequences ubiquitous. I'd argue that's bug, not a feature.
I make a solid income solving problems that people pay me to solve. Why should I abandon that and devote my life to achieving 9% size and 4% availability time increase[1] that no client has ever asked me for?
And to be clear, it's not that I don't care about performance. It's that if my automated tests notice an endpoint loading slowly or if a client complains about performance, I can almost always achieve order-of-magnitude performance gains by optimizing a SQL query or twiddling some cache variables, which almost never happens by switching serialization formats. I have used protobuffers in a few cases, where profiling indicated it as a solution, but this has not been the norm in my experience.
The first optimization is getting it to work, and the second optimization is whatever profiling tells you it is.
[1] https://auth0.com/blog/beating-json-performance-with-protobu...
But I think you're right that as long as we stay this course there's going to be more problems and so you'll be able to make more money solving problems that didn't need to exist.
You're absolutely right about the first optimization being to get it work. You're just discounting the reality is that you're making it far more difficult for that to happen.
You can make tools that present data in any format in a way that is easy for humans to digest. Letting a very small and trivial part of the problem drive what is a much larger problem space is pretty flawed.
{"time":"2020-07-22T10:59:14.95406-04:00","message":"{\"level\":\"debug\",\"module\":\"system\",\"time\":\"2020-07-22T10:59:14.953909-04:00\",\"message\":\"Running MetricCollector.Flush()\"}"}
this is a very moderate example of what I deal with daily, all because JSON includes quotes around fields.
I wonder if it would’ve been better for them to Base64 encode their message. Of course, this itself presents other problems.
https://amzn.github.io/ion-docs/
Superset of JSON with many extra features that people in the comments here desire from their data-language, such as support for comments, timestamps and s-expressions.
So everyone gets to suffer for his beliefs.
Maybe something like TOML? It's very human-readable and simple. But it's not a very good serialization language.
Thankfully JSON came along and got back to the simplicity of early XML and XML-RPC and stayed there.
Deserialization and remote code execution vulnerabilities all over the place. That was brutal.
Who thought this was a good idea to pass arbitrary function names and arguments for the remote servers to resolve and execute blindly? The regular vulnerabilities in the XML parsing libraries themselves was the nice cherry on top.
I'm not sure who would would to that but I certainly didn't. Ultimately XML-RPC is no different from REST/JSON except it's in a different format. What you did with that format is a totally different issue.
<?xml version="1.0"?>
<methodCall>
<methodName>examples.getStateName</methodName>
<params>
<param>
<value><i4>40</i4></value>
</param>
</params>
</methodCall>
The thing is meant to call arbitrary functions with arbitrary arguments. It doesn't take long until there is a straight up exec functions exposed or some accidental command injection.It's strange to look at it 20 years later. The adoption of JSON really got developers to stop shipping RCE vulnerabilities every other week. Yet nobody must have thought of that when deciding what to use.
{
"methodName": example.getStateName
"params": [14]
}
Although you'd probably instead have an REST endpoint contain the method name and the entire JSON body is the parameters. But the difference is minor. There's no reason this allows arbitrary execution than anything else.Methods directly exposed to the web is how 99% of all MVC frameworks work.
My favorite quote about XML-RPC (from the plan9 people):
Some part of me desperately wants to believe that XML-RPC is some kind of elaborate joke, like a cross between Discordianism and IP Over Avian Carriers
I have exactly the same feeling regarding the JSON "protocols" and whatnot.
“The good thing about reinventing the wheel is that you can get a round one.”
XML is a mess and a chore to work with (at least in any language that isn't Java I guess, but even then).
Yes, it has some rough edges. Yes it could be better.
But overall it's good. Not too complicated and not too hard. Works fine for most stuff.