JSON with Commas and Comments
nigeltao.github.io
nigeltao.github.io
JSON is something we can (almost) all agree on for dumping loosely specified, human readable representations of data structures. It lets users and client application developers be lazy and not have to learn a new library for our preferred data format. The lack of schemas or validation means we can often get away with wishy-washy, hand-written, plain English specifications, or specifications "by example". It's convenient.
For anything intended to be robust, for serious data interchange, use something else. Anything with support for machine-readable schemas, validation, and robust encoding rules. Binary is preferable but a JSON serialization is still a must-have escape hatch.
And here's a shocking idea... provide your clients with an SDK. It's pretty easy to build one around a multi-language data serialization framework and maintaining an SDK is still easier than maintaining a spec that your clients will probably implement incorrectly.
https://github.com/lloeki/ruby-skyjam/blob/master/defs/skyja...
gives that:
https://github.com/lloeki/ruby-skyjam/blob/master/lib/skyjam...
I would certainly enjoy having a DSL to write descriptive code to validate using JSON schema, but it would be even better if the Ruby definitions could be generated and persisted in Ruby files using that DSL.
Also, storing things in basic hash/array types works, but having dedicated types is useful, so that one can ensure not shoving one kind of hash in place of another unrelated kind of hash.
As for types themselves in general, there's RBS and Sorbet. One could have type definition generation as well for even deeper static and runtime checks.
It's actually 2020-12, which is two versions after Draft 7 (they shifted from Draft n to YYYY-MM after Draft 7, and since then have had 2019-09 and 2020-12.)
And that's true of most languages, though there is some 2019-09 support. (It really doesn't help that there is also OpenAPI which baked in a variant—“extended subset”—of Draft 5 JSON Schema.)
Existence of a schema definition file and checking responses against is signalling that I can trust an API vendor to be at least aware of the requirements for clients. (Whether they randomly change the schema definition or ignore it is a second question, but at least somebody once thought about formalising and it's not an complete adhoc dump of today's internal data representation)
Apache Avro has support for parsing and utilizing schemas at runtime, even in C++.
For Apache Thrift you have things like thriftpy: https://thriftpy.readthedocs.io/en/latest/
I'm not aware of a type-safe mechanism for Flatbuffers or Protocol Buffers.
- strongly typed
- ad-hoc or schema (your choice)
- no code generation step
- edit in text, send in binary
- supports ad-hoc data structures or schemas per your preference
- supports all common types natively (doesn't require special string encoding like base64 or such nonsense)
- supports comments, metadata, references (for recursive/cyclical data), custom types
- doesn't require an extra compilation step or special definition files
- Has parallel binary and textual forms so that you're not wasting CPU and bandwidth serializing/deserializing text. Everything stays in binary except in the rare cases where humans want to look or edit.
"Amazon Ion is a richly-typed, self-describing, hierarchical data serialization format offering interchangeable binary and text representations. The text format (a superset of JSON) is easy to read and author, supporting rapid prototyping. The binary representation is efficient to store, transmit, and skip-scan parse. The rich type system provides unambiguous semantics for long-term preservation of data which can survive multiple generations of software evolution."
E.g. XML has all that, and yet the average use of xml is much more fragile thsn the average use of json. Similarly asn.1 has all that, but does anyone actually like the fact the x509 certs use it? (To be clear, json would definitely not be appropriate for that case either).
- HORRIBLE schema format.
- Very few FOSS libraries, most of which are terrible. I used Lev Walkins asn1c[0], which I patched slightly and coupled with some semi-generic code that walks the generated data structures to emit JSON (for debugging purposes).
It was very handy to be able to use jq on the data.
There is no official JSON schema RFC, there is no official linking between json-schema.org and the JSON RFC. There is no push to make every language that supports JSON to also support a JSON schema system.
Finally, there isn't even a guarantee that all human-readable JSON document specifications can be expressed within a JSON schema in any sensible way.
I never got anything out of xml validation either. It just lead to this ridiculously verbose garbage xml that was even less readable with all the namespace, namespace declarations, etc. Very tedious to write manually as well. Also lots of weirdness doing e.g. xpath or xsl transformations against that. The whole SOAP / web services bubble eventually imploded when people figured that you could send tiny json objects instead via REST.
I have many issues. People sending invalid json to my APIs is not one of them. You get a bad request if you do. End of story. Don't don that if you don't like bad requests. It's not a problem that justifies a lot of over engineering.
Json, yaml, toml, hocon, etc. are basically all just variants of attempting to send human editable blobs of information. Fine for apis where all of the requests are created by programs. Unfortunately people also abuse them for things like DSLs that are authored by humans. The real problem there is not using a strongly and statically typed language that simply does not allow illegal things. A schema is just a stop gap solution when you don't have that.
Kotlin is actually great for creating proper DSLs. You get IDE auto-completion and red squiggly lines when you do it wrong. I think Rust also has some nice syntactical constructs for creating DSLs. Typescript might also emerge as a language that is very suitable for that (with maybe a few more features borrowed from Kotlin). Using a non compiled language with weak typing kind of defeats the purpose. Hence the endless ruby but not quite ruby like DSLs for things like puppet.
S-expressions might have caught on for this purpose, but they lacked a killer app (jquery, and later rails). Like JSON, they are easy to parse and easy to generate, but being more loosely-specified than JSON, it's less clear how to map a given S-expression to a native type. As a glue format, JSON feels just right, even if it's a bit picky sometimes (e.g. trailing commas).
For anyone who wants extra features like comments, YAML is the oldest popular format I am aware of that is a superset of JSON. For any new format to succeed, IMO it needs to sufficiently distinguish itself from both YAML and JSON, and not just support a feature set that happens to lie somewhere between the two.
I was doing mostly Perl and Javascript when it caught on, and to this day I have very mixed feelings: it's wasteful of space but still doesn't allow comments; its type system is basically a technical-debt generator; and for all that people still get it wrong pretty often. On the other hand, it's more or less human readable for simple data structures and it's more or less everywhere.
My hunch is that for a new format to take off it wouldn't so much need to not fall between YAML and JSON: it would need to be the default format of something with such super exponential growth that even us oldies would have to use it.
I believe it's hard to explain JSON's popularity and wide adoption without talking about javascript. With javascript, JSON was right from the start an ‘eval’ away from being parsed. The barrier to entry to adopt it simply was never there. Once you start to expose JSON APIs to clients, other servers also start to need to consume data from those servers. Rinse and repeat until you reach mass adoption.
They would not have been loosely-specified if they were specified, like JSON was :-) I mean this is taking things a bit backward. When tools use a data format based on S-exprs, they define more clearly what is or isn't valid (OCaml Dune, Guix, etc.)
Yes, but they all implement the JSON RFC slightly differently, or implement it in a totally non-compliant way.
It's very easy to do JSON wrong, to not even be aware of doing it wrong, and do it wrong in a way that limits its interoperability (see: the default behaviour of the json library in Python). This is not theoretical either, just google for 'json nan github' to see the hundreds of production programs accidentally emitting JSON that cannot be ingested by RFC-compliant parsers.
And regarding integer ranges: Any user has restrictions on top of the generic format. Some fields have to be present for the application to work, some fields have to be a string, others an array. Some integer has to be between 0 and 100, some array has to have 5 elements. Given that 99% of integers in practice fit in a signed int32 there is little problem (97% are probably 0, 1% are 1, 0.5% are -1 and only the rest other values ...) If you are on the edge you have to know and work-around ...
The problem is that you are not guaranteed to know, as a JSON library user, that the library has not mangled the numbers it received on the wire prior to your application receiving it - so you don't know if you having .a set to 42 is the result of 42 being sent over the wire, or an implementation dropping bits that are outside the range of you library's support. The RFC does not mandate what a JSON implementation should do in case of numbers outside its' support.
> Given that 99% of integers in practice fit in a signed int32 there is little problem.
Until you hit that 1%, and you find this out the hard way, and you have no way of solving it because you don't control the emitting side. Again, this has happened in practice to me, when a Python library was emitting large numbers as numbers (as the RFC permits), while the receiving side silently casted to 32-bit floats, losing data (as the RFC permits). Both sides are right per the spec. Technically, everything worked as per spec. Practically, the product was broken.
99% of the time you might be okay. But the 1% of edge cases makes it that you can never rely on JSON, which makes it a bad interchange format if you care about reliability and safety. You _can_ make it work if you severly limit yourself and are deeply aware of all the possible issues that using JSON has (and there's a lot more). Or you can just pick some other standard that solves these basic things for you (eg. Protobuf).
Random side note: ECMAScript has no integer type, but only Number, which is a float, thus when dealing with such numbers in a JavaScript frontend you are in Problem Land anyways ... which again shows that boundaries have to be thought of, independently from the specification of the data exchange layer.
Had to look that up. Was under the impression the rfc only specified the syntax. But here it is
> This specification allows implementations to set limits on the range and precision of numbers accepted.
I don’t think the implication is that it must be done by silently loosing precision though. A loud error would fit that description just fine. So in this case I would blame the implementation not the spec.
However, we're coming from XML, which was better defined, and there are also binary "quasi standard" formats like protobuf, so those problems had already been solved.
Yet still "the world" has moved to JSON, which seems to indicate that those problems probably were not all that important for most use cases.
As a case in point, the python parser breaks the JSON spec already with regards to Infinity/NaN, but then has flags to configure this. See https://docs.python.org/3/library/json.html#module-json
> The RFC does not permit the representation of infinite or NaN number values. Despite that, by default, this module accepts and outputs Infinity, -Infinity, and NaN as if they were valid JSON number literal values:
I wish something would done in this area relatively soon though because JS has added a number of features that would make a new JSON much nicer like multi-line strings and BigInts.
I think JSON + comments, commas, template literals, BigInt, NaN, Infinity, and BigDecimals (if/when those land in JS) would be very useful. (It'd be nice to include dates, but that's tricky w/o a literal and because dates)
Current parsers aren't uniform here. Since the JSON spec is silent on what post-parsing format is used for numbers, each parser is free to do whatever makes sense in the context of the host language. I reckon you'll find some JSON parsers use bigints already, especially in languages with first-class bigint support (such as various Lisp dialects)
> And I'm not sure anyone wants a format where the result may change types based on the size of the number
That's nothing to do with the format, that's to do with the parser. Some parsers already do exactly that – use an integer type for numbers that are integers, use a floating point type for numbers that contain decimal points
Yep. Python is one such language. I've seen this catch people by surprise when they discover their serial number (which granted, should have been a string in the first place) doesn't survive a trip from Python to JSON to Javascript, among other languages.
Sometimes I debate collecting stories like this.
A new mime type doesn’t really help because the parser doesn’t check the mime type, it assumes the programmer did that. I’m not against making a new mime type or specification, but there are already so many and making more doesn’t seem to help. Yes it would be nice if PHP and Python implement the logic for ignoring trailing commas the same way, but 99% of the time that isn’t all that important. Yes there will be cases where it matters and bugs are introduced, but since there already significant differences between JSON parsers I don’t see these kinds of things as making it much worse.
I think the most correct way to deal with this problem is to get IETF and ECMA to update the JSON standard first. Honestly, comma-after-final-element and comments are such important quality-of-life features that I never understood why they weren't part of the original spec.
On the other hand, it might be easier to just adopt TOML and forget about JSON.
The whole reason JSON is text-based is to make it human-readable while also enabling data exchange, but the lack of comments works against that goal.
Wasn't there just one Stalin, though?
Stdin. New phone, forgot to disable autocorrect.
I remember seeing some conversation from the JSON authors around comments and specifically not allowing comments into the spec because they did not want people to use them as extension mechanisms, so the whole no comments in JSON thing was very much intentional.
- There is no minimum/maximum number/integer value (or sizes) defined in the RFC [1] - you have no guarantee that a number you emit will be readable by all RFC-compliant parsers, so the only safe option is to emit all numbers as strings. There is also no behaviour mandated for parsers that encounter numbers/integers outside of their supported size range, so you can't rely on any fail-safe behaviour.
- There is no standardized behaviour for repeated dictionary keys (which are allowed but 'discouraged' per spec), and different implementations will treat them in different ways. This is especially painful when trying to build JSON middleware that does a parse/check/modify/emit of arbitrary data.
- Implementations in some languages (eg. Python) are non-RFC compliant by default (`python.dumps` will emit NaN/Inf/-Inf, even though the RFC forbids that), and generally all implementations are similar-but-different-enough to trip you up [2].
All three of these have bitten me in the past when trying to interface with something that spoke JSON, and as such I refuse to design new systems that use JSON as any source of truth.
https://developer.mozilla.org/en-US/docs/Web/JavaScript/Refe...
>Trailing commas in objects were only introduced in ECMAScript 5. As JSON is based on JavaScript's syntax prior to ES5, trailing commas are not allowed in JSON.
json was largely `eval`d when it first hit the scene. It was easy to parse because you just received it and `eval`d it directly to parse. worked in every browser. security concerns over time lead to this being rightfully replaced with the JSON object and its associated encoder and decoder.
however, since early javascript didn't allow for trailing commas, neither could JSON if it wanted to be able to be `eval`d.
The only real reason why comments would be problematic is that they are a pain to preserve in a consistent way when editing a file, and thus would require extra work in parsers / serializers. Still, it would be worth the cost imo.
Actual JSON-reading applications must ignore comments because they aren't allowed to care about them. On the other hand anything useful can be placed in a proper field, a comment would be a hack to pass information to human readers but not to JSON consumers (saving some time and memory). Such a technique could only be considered if people are supposed to read the JSON files and some information is useful for them but not for the reading application: a niche within a niche.
Preserving comments would actually be a complicated special purpose feature, reserved for something like structured text editors that can perform nonstandard parsing with comments included in their special object model (of the large Javascript subset/superset/variant they choose to support and roundtrip, not of JSON).
I honestly don't want to bash JS, but typing is not it's strength.
(Not affiliated)
On the other hand, there is an RFC for JSON where there isn't for Markdown, so it ought to be possible to make use of the usual standards process to formalise a successor to the original JSON which might include some of the very common additions some people always want (trailing commas, comments, proper dates) - that this hasn't happened over the last 15 years suggests really that the problem is actually due to lack of agreement.
I would like to see why the author chose to do yet another format, vs adopt JSON5.
For example, allowing identifiers as object keys.
Starlark to solve this exact problem, and is a Python subset so that people don't have to learn something totally new either. It's still turing complete, but at least guaranteed to not have any side effects, or even be able to access anything but an explicitly given execution context.
It also makes it easier for other tools to generate, validate, or output config that feeds into your system. They can do whatever processing they want and emit plain old JSON.
I did this to support hand-writing game data in JSON (e.g. monster stats, campaign dialog scripts) and then converting it to MessagePack at build time [2]. This gives you very high simplicity (no schemas) and excellent performance.
As for trailing comma the only real issue in practice is noisy unified diff output. But for JSON word diff or similar works better in any case when the trailing comma is not an issue.
Comments would only serve to make humans edit the file by hand more than they do now, and to maintain it like a snowflake. You should not be maintaining your data in JSON, it should only be a momentary stop along the way to a more robust data storage system [with a schema].
Optional end commas would only encourage people to use shitty hacks to try to craft JSON without a real parser or encoder, which would 1000% end up with broken JSON all the time, which would then force vendors to support shitty broken JSON.
There's a reason it's designed like it is. Cutting corners is not going to result in better outcomes.
Computer-generated data never needs to include trailing commas or comments.
Really the principal use case here is JSON as human-edited configuration files. In which case I suppose it's nice to have a name like "JWCC", but really it's just two flags for JSON decoding libraries to add (or a single nice combined flag).
So that really would be great IMHO, if you could just call json_decode($json, JSON_ALLOW_JWCC). I don't think we need a new MIME type or anything. But a file extension of ".jwcc" would be a nice convention too.
Also, NaN numbers are an entire features, which they talk about above.
Even if one ultimately needs to keep json document’s human structure (like key order, comments, whitespace nuances), they may create and use more syntax structure-aware library to do that.
Sure. I’m particularly thinking about mvn upgrades or “npm update”, which modify pom.xml/package.json files to upgrade the libraries, after checking rules (non-breaking changes or not, vulnerabilities or latest, etc).
For mvn, libraries never succeeded to modify the pom in-place without wrecking the file format, so mvn upgrades never became a thing. It’s also a demonstration that a DOM with comments in memory doesn’t ensure we can output the file as-is ;)
For programatically upgrading JSON config files, it's not just comments but whitespace, indentation, etc. that need to be preserved. Honestly that needs an entirely separate tool/library from normal JSON encoding and decoding -- e.g. special jwcc_insert() and jwcc_delete() functions that guarantee the file remains untouched (exact formatting preserved) except for the modified part.
It's really more akin to when your apache.conf file gets modified programatically -- everything is preserved exactly except for the specific lines that get touched/added.
The text "deliberately removed comments from JSON" links to https://web.archive.org/web/20150105080225if_/https://plus.g... where Doug not only explains the reason, but also a solution which works with the existing standard. It's imperative that the JWCC author explain what they find lacking in Doug's solution - stripping comments before handing off the a JSON parser. A new format is a really weird way to go about fixing a problem that already has a solution.
A baseless reason is not a compelling reason. Some imagined eventual use is not relevant to the utility and convenience. Interoperability (in this greater undefined sense posited) cannot be maintained via syntax anyway.
> stripping comments before handing off the a JSON parser.
Putting an extra processing step before use of JSON is impractical.
> I removed comments from JSON because I saw people were using them to hold parsing directives
> Putting an extra processing step before use of JSON is impractical.
Care to explain?
Java has not been destroyed, nor any other language that uses injected behavior. Ironically, applying parsing directives are basically another tool being run (similar to a linter) so I'm not sure what he's even getting at when suggesting running yet another tool. This just complicates the ecosystem and is part of what makes javascript seem so primitive in use and syntax. Features that are counterproductive for poorly considered reasons.
Similar thing if you want to inspect a JSON dump into a file.
Without a standard format for comments, you have no reason whatsoever to expect the ad-hoc comments in a JSON file made by anyone who isn't you to be styled like JavaScript comments. And if your response is "then everyone needs to use JavaScript-style comments", well, TA-DA, you've just added comments to the spec. You can't have it both ways.
> It's imperative that the JWCC author explain what they find lacking in Doug's solution - stripping comments before handing off the a JSON parser
If the syntax rules for adding comments are unspecified, then the the syntax rules for what to strip out are unspecified too!
Indeed.
> Without a standard format for comments, you have no reason whatsoever to expect the ad-hoc comments in a JSON file made by anyone who isn't you to be styled like JavaScript comments.
But, since comments are not supported in JSON, you can expect JSON from external sources to contain no comments whatsoever.
That’s precisely Crockford’s point: he didn’t want the format to have a comment syntax that works across different parties so that the only way to use comments is internally, i.e. with a custom comment syntax that you strip before interacting with external sources.
So instead you have people coming up with a million different bad hacks to support comments in non standard ways (e.g. "fake comment" properties, duplicate keys, etc.) Coupled with the fact that since it's not in the standard, you never really know if the file you're creating may be read by something that won't strip comments. It's the worst of all possible worlds in my opinion, with 0 benefit.
No, because his goal was to make JSON a highly-compatible interchange format, which it is! There are roughly zero incompatible preprocessor directives, and roughly zero problems using JSON to exchange data between parties. Seriously, when have you received JSON that you couldn’t immediately parse with your standard JSON parser of preference? The fact that some people might devise their own tools to store incompatible JSON on their end with things like comments, and strip those things out when transmitting them in order to make them compatible JSON, does not qualify as an incompatible preprocessor directive. On the contrary, thats precisely the intent of the choice to not have comments in JSON.
Every single time for any non-mediocre definition of parsing. JSON has no way to transmit the meaning of values, no datetimes, no units, no sets, nor any other semantic attribute. That means your "parsed" result is always useless and wrong on its own, literally a lesser-dimensional projection of the original information, until you interpret it again in an entirely ad-hoc custom manner. Wouldn't it be nice if you could store instructions inline for how to do that.
Creating schemas for declaring types is literally a hack for adding type! I think its very bizarre to say "You dont need to do X because you can just do X."
> Also how hard is to write {"type": "datetime", "value": "2021-02-23 12:46:37.07"}
"datetime" isn't a universal format specification. 2021-01-10 could be October 1 or January 10. What time zone is this in? Local to the sender, UTC, local to the recipient, somewhere else? Is this a 24 hour clock or something that gets a modifier for the other half of the day? You're making some classic mistakes that people make when they don't think about the complexity of the problem domain of communicating information unambiguously.
It is a "communication protocol" - base of any messaging system. Message has header with protocol name and version: HTTP, TCP, DB connection, SOAP, any message queue - everything working that way.
> "datetime" isn't a universal format specification. than use another (see the comment above)
self-contained message is anti-pattern
As recently as 2017, the Google Translate API used to return completely invalid JSON with empty array indices instead of nulls ([,,,] instead of [null,null,null,null]). This broke several JSON parsers until they patched around Google's cavalier abuse of the language. Your experience of everything always working perfectly doesn't mean that everything always works perfectly.
When they accept it, it becomes defacto valid. That's how acceptance works.
See also http://seriot.ch/parsing_json.php#41 for a big table of "JSON-encoded data and JSON parser that didn’t work perfectly together". Read the whole page though. It's quite enlightening.
That's neither here, nor there. The narrow vision many/most tools were created with is laughable compared to the actual creative uses people put them into.
Heck, the internet wasn't created for collaborating, socializing, shopping, reading, listening to music, etc., anyway, it was created to have a war-proof network for army use, yet here we are...
It's like the inventor of gif declaring it is pronounced 'jif'. Who cares, basically nobody else says it like that.
Would you consider your use of fire (e.g. for cooking) "illegitimate" if the creator of fire said so? ("No, must eat food raw! Fire is meant for heating only! Ugh!").
Quoting https://en.wikipedia.org/wiki/ARPANET :
> It was from the RAND study that the false rumor started, claiming that the ARPANET was somehow related to building a network resistant to nuclear war. This was never true of the ARPANET, but was an aspect of the earlier RAND study of secure communication. The later work on internetworking did emphasize robustness and survivability, including the capability to withstand losses of large portions of the underlying networks.[51]
See also Licklider's work with https://en.wikipedia.org/wiki/Intergalactic_Computer_Network leading up to ARPANET, or his "The Computer as a Communications Device" - https://signallake.com/innovation/LickliderApr68.pdf where he describes how the network might be used:
> You will not send a letter or a telegram; you will simply identify the people whose files should be linked to yours and the parts to which they should be linked-and perhaps specify a coefficient of urgency. You will seldom make a telephone call; you will ask the network to link your consoles together,You will seldom make a purely business trip, because linking consoles will be so much more efficient. When you do visit another person with the object of intellectual communication, you and he will sit at a two-place console and interact as much through it as face to face. If our extrapolation from Doug Engelbart’s meeting proves correct, you will spend much more time in computer-facilitated teleconferences and much less en route to meetings.
and has a section titled "On-line interactive communities":
> Available within the network will be functions and services to which you subscribe on a regular basis and others that you call for when you need them.In the former group will be investment guidance, tax counseling, selective dissemination of information in your field of specialization, announcement of cultural, sport, and entertainment events that fit your interests, etc. In the latter group will be dictionaries, encyclopedias, indexes, catalogues, edit-ing programs, teaching programs, testing programs, programming systems, data bases, and—most important—communication, display, and modeling programs.
Collaboration and reading were surely some of the main goals in the vision that became ARPANET.
Nonetheless, according to Stephen J. Lukasik, who as Deputy Director and Director of DARPA (1967–1974) was "the person who signed most of the checks for Arpanet's development:
"The goal was to exploit new computer technologies to meet the needs of military command and control against nuclear threats, achieve survivable control of US nuclear forces, and improve military tactical and management decision making."
> "The ARPANET was not started to create a Command and Control System that would survive a nuclear attack, as many now claim. To build such a system was, clearly, a major military need, but it was not ARPA's mission to do this; in fact, we would have been severely criticized had we tried."
History's complicated, isn't it?
So, we look at other things: Licklider became Program Director at ARPA in 1962 and ARPANET started in 1966 (which is when Lukasik joined as Director of Nuclear Test Detection before becoming A.D. the next year, then Director in 1971). And as Lukasik writes in his paper, the late 1960s were a different funding era than the early 1960s when Herzfeld's "foundling" started.
Recall that the 1968 Mansfield Amendment prohibited military funding of research that lacked "a direct or apparent relationship to specific military function" - far different than the Ruina years where office directors and program managers had significant autonomy and funding authority. An effective ARPA director after the Mansfield Amendment was passed is going to be someone who is good at viewing ARPA projects through that military support lens, yes? Which might be different than the lens used earlier?
Next, quoting https://en.wikipedia.org/wiki/Robert_Taylor_(computer_scient... :
> Taylor hoped to build a computer network to connect the ARPA-sponsored projects together, if nothing else, to let him communicate to all of them through one terminal. By June 1966, Taylor had been named director of IPTO; in this capacity, he shepherded the ARPANET project until 1969.[11] Taylor had convinced ARPA director Charles M. Herzfeld to fund a network project earlier in February 1966, and Herzfeld transferred a million dollars from a ballistic missile defense program to Taylor's budget.
It therefore seems very much like collaboration, at the very least, was indeed part of ARPANET's goals when it started in 1966.
FWIW, as Martin Campbell-Kelly and Daniel D Garcia-Swartz point out, at https://citeseerx.ist.psu.edu/viewdoc/download?doi=10.1.1.58...
> ARPANET network was one among a myriad of (commercial and non-commercial) networks that developed over that period of time – the integration of these networks into an internet was likely to happen, whether ARPANET existed or not.
They further consider it "Whig history" to think ARPANET plays a critical role in the modern day internet.
And I agree with that assessment.
Yes, but in the end it's he who pays the bills who decides. Without them it's lights out.
Lukasik came in later, and had no signing authority or supervisory position on the initial ARPANET work.
Following your guideline, the ARPANET therefore wasn't "created to have a war-proof network for army use"; that was a design goal which came later, after the collaboration goal.
You can expect it, and then you'd be suprised when they do.
See people really want/need comments, and they're gonna implement them anyway in dozens of little parsers/libs, and you're gonna have to deal with files having them anyway...
He didn't want. That's great, but using JSON doesn't magically eliminate the need for documenting what things _mean_. Oh, this text isn't just text and is actually a datestamp? Oh, this array isn't allowed to have repeat values? Oh, the valid values for this field are "red", "yellow", "purple", and 5? You could have had that documentation inline. Instead you're forced to have it somewhere else because the need for documentation didn't magically vanish when Crockford waved his wand.
All that eliminating comments accomplishes is needlessly limiting the utility of an otherwise mostly fine format. JSON could have been a good configuration language. Instead it isn't and apparently people have to fight to justify fixing that.
The only reason JSON is as popular as it is is because JavaScript parses it natively (JSON.parse is defined by the JS spec).
I mean I remember using XML for serialization back in the day and how it made we want to kill anyone who was involved in its creation. JSON won because it’s a great and simple standard and yes because it maps 1:1 to js. It was very popular with people who hated js when it came out too - just because of how much better it was than the other options
It's not unlike saying "actual code shouldn't allow comments - you can have a separate preprocessor in your non-code file that strips those out." Sure, you could do that, but all your line numbers would be meaningless and everyone would hate the language from day 1.
(Also, as gets pointed out every time this topic pops up on HN, including multiple times on the submission already, the main demand isn't people wanting to add comments to the JS flowing across the wire, but to config files sitting on disk. Responding to people saying "feature X would help use case Y" by noting how many people are using it for use case Z even without feature X is logically incoherent.)
And many people have. You're literally posting this in response to someone submitting the new spec they created to solve this.
I don't understand what you're trying to argue here.
Just because he had a reason to give doesn't mean it wasn't a bad reason. His stated reasoning failed with hindsight.
(a) People already do what he thought was preventing (custom parser behavior) by removing comments, (so his choice failed to prevent what he wanted to prevent)
(b) people have already created dozens of JSON parser variants that accept comments, because we really want those. And not just some niche devs - Microsoft and dozens of other big companies have tools that accept JSON + dangling commas + extra comments (so his choice was second-guessed and bypassed anyway, just in ad-hoc ways instead of a better universal one).
It wasn't his place to define whether we get comments or not based on some potential parser abuse. But he did it, and now we're stuck with that decision...
People still do it, and they still call it JSON. Heck, many JSON parsers, who otherwise work fine with regular JSON still accept it.
Even parsers using a different name like JSONC still allude to the JSON connection - and still are an argument that people found a lack in JSON.
Etymology/originology never produced much benefit over pragmatic non-prescriptive examination of how people use things. If anything, it helped confused the situation with pedantic objections.
He was writing the specification, of course it was his place to do so. Who else would be in the right "place" to write his own specification that no one was enforced to use.
Stripping comments before handing off to a JSON parser is not a full solution. For example, if the file has to be augmented with additional data and written out again, the comments would be lost.
{<object>} {<object>}
That isn't valid JSON, is it? The Qt JSON parser doesn't handle it.
There are libraries[1] that support parsing it and it's not too hard to do yourself, either. Some fairly popular projects use it to represent multiple responses in a single response body[2].
[1] http://ndjson.org/libraries.html [2] ElasticSearch uses it for msearch response bodys, for example.
This format works great when using Amazon Athena (Presto) against log files written with one JSON object per line.
In fact I've been using JSON with comments and trailing commas in my own projects a lot already, and I imagine lots of other people have converged to the same idea since these are really the only two pain points when handwriting JSON. It's about time we give it a formal name.
{ [1, 2]: "a", [3, 4]: "b" }
If it's a map I'm a bit more unsure how you'd check to find the object (quickly). You're basically getting into the how to store a struct/class as a map key. Either way its a bit more involved than just having json parser return an immutable array and map, and the more I think about it the more edges cases there are.
I'm here to tell you that `Object.freeze()` works just fine on arrays.
I guess you could have this json+ convert into some custom javascript class that allows arrays/objects as map keys rather than the normal javascript object that only accepts strings as map keys.
Sure - why not? JSON doesn't have to map directly to JavaScript objects.
'Welcome to Node.js v15.7.0.
Type ".help" for more information.
> { [1, 2]: "a" }
({ [1, 2]: "a" })
^
Uncaught SyntaxError: Unexpected token ','
Thar' be dragons! https://stackoverflow.com/questions/32660188/using-array-obj...However, it's equally true that a better language and object notation would support non-string scalar keys.
However, I don't think it is useful to prod JSON (or a variant) to go in this direction.
People are more interested in JSON derivatives like JSON5 or Ion than more complex languages like Dhall in my experience.
Dhall is new to me. Would you tell me more about other configuration languages you seen? I often use TOML myself.
ob = {...}
somemap[ob] = value
// vs
somemap[ob.id] = value
In languages which allow weak keys in such maps, a garbage collector may even reclaim whole key-value pairs once ob falls out of existence. But to my opinion, in json it is barely useful for a number of reasons. It is more programming technique than data transfer format, and these features (and reference loops) should be encoded at a higher level than json.That would mean your API would need to serve different variants based on the Accept header, and your backend code would need to generate the two different variants.
And then you'd have to question if sending API responses with unused things like comments is actually a good idea given it's entirely wasted data in a production environment.
I suspect few people would use a newer JSON format until it's been available for a long time, and even then they wouldn't use it on big sites.
https://en.m.wikipedia.org/wiki/TOML
Not that there's anything wrong with adding comments and commas to JSON. Call it JSONv2, keep the .json extension and let people bump around upgrading for a bit. It's hardly much of a change but has significant benefits, even outside the config file use. For instance it can sometimes be useful to annotate raw data, and it'd be nice to have that built in. Certainly the commas is a no-brainer.
What’s missing?
ISO 8601 can represent a simple offset from UTC, but it can’t do a lot of the things that the IANA time zone database can do.
Adding an IANA library to the ISO8601 library (often part of the same lib) has so far solved that problem for me. I’m sure it starts struggling with dates before the early 1800’s though.
Also doesn’t work well if you’re dealing with relativistic effects I think.
It absolutely has every benefit you cite.
It's the only ISO spec I know the name of without having to look it up.
To be honest I think you're better off just not encoding the timezone in your main timestamp string at all and instead adding the full IANA timezone name (e.g. "Europe/London") in to a separate field. In sensible formats all the dates within a single JSON object or object will have the same timezone anyway, and to save bytes you can simply say in your spec "If absent, assume UTC".
Another problem with ISO is the specs aren't freely available, so people will generally guess or refer to old versions, drafts, or RFCs (like RFC 3339)
That means your json parser either will accidentally parse some strings as dates, or will have to be told which strings are (or even might be) dates, introducing a partial schema, or it will have to leave it to the application to convert strings into dates. Workable? Somewhat, but not ideal, just as not supporting numbers, and leaving string-to-integer-conversion to the application (after all, JavaScript will happily convert strings to integers in eval) would be ‘not ideal’.
Also, are you going to accept 2021-02-23T08:10:19Z, too? Leaving out the seconds, minutes,…? for interoperability, you would have to specify that.
The practical approaches are to represent it as time since the Epoch (in floating point) or a string where both sides somehow agree what "03/04/05" means.
All these are allowed according to the ISO standard.
If your data field represents a time (such as the selection of the time when a daily should recur each day), then you should use the ISO-8601 format with of a time.
If your data field represents a specific point in time (such as the time an action happened), then you should use the full representation of the date and time.
> you want the T separator or space?
When concatenating the date and time, the 'T' separator is used. Although, this is an implementation detail I've never had to touch myself, since every date library I've ever used has been able to format dates into an ISO format, and parse them from an ISO format without me having to consider the separator.
> Is time zone represented?
Again, this depends on what the data being represented fundamentally is. If it's a point in time that an event happened, yes. _Usually_ if you're representing a full date & time the answer is yes, you should represent the time-zone (I always serializing these dates in UTC and include the TZ).
The TZ should only be forgone if the data represents the abstract notion of a specific date and time, rather than a specific point in time.
To me (other than the "T" separator), these are all fundamental data questions, and not serialization questions. You need to know whether a field is a DATE, a TIME, a DATETIME and whether is a Zoned Date Time or a Local Date Time regardless of your serialization format.
ISO-8601 conveniently supports all these uses-cases and more.
A data format is a way of structuring or serializing data. The whole point of the format is to read and write arbitrary data. It may have a simple schema, or features designed for the loading and unloading of data. But all those features are designed to assist the machine, not the human operator.
A config format is designed specifically to assist a human operator, not a program. Humans should not have to consider data types when they write a config, or when they feed it to a program. Humans should have useful features to make their lives easier, like variable substitution, pattern matching, inheritance, namespaces, simple newline-separated whitespace-indifferent commands, etc.
Programming languages are just elaborate configuration formats. And that's where the problem begins: how much "power" do you give the configuration format before it gets unruly? It's difficult to find a balance. People end up using simple data formats because the parsers are widely available and they can "fake" advanced features by making programs interpret a specifically-crafted data structure as a configuration instruction. But there's very little thought put into how this can be extended to make more complex configurations easier.
Most web servers and other complex software [usually run by sysadmins] have a real configuration format or configuration language. Poorly-written software pretends a data format is a programming language, or even worse, forces you to use an actual programming language. These designers fundamentally don't understand or don't care about the user.
The single/double square bracket thing is also non-obvious at first.
There's a ton of ways to represent booleans, a ton of ways to represent dates, a ton of ways to represent numbers of a ton of different bases, a ton of ways to represent strings of various line-endingness'. It really has caused a significant number of easily avoidable issues in my experience.
All the complexity related to anchors and such is unfortunate, but that’s a problem for parser writers.
The link in the main article convinced me otherwise: https://noyaml.com/
I wanted to see if CameronNemo would spell it out.
I'd like to point out the HN Guidelines:
https://news.ycombinator.com/newsguidelines.html
> Please don't comment on whether someone read an article. "Did you even read the article? It mentions that" can be shortened to "The article mentions that."
Even putting aside what "F" stands for, this means that comments like TFA or RTFA are not in the spirit of the discussion here.
That's exactly what the parent wrote, so he's totally fine with respect to the HN Guidelines.
To quote: "The author reviewed most prior art in TFA.".
The term is very easily confused, to state the obvious.
In my opinion, I think it would be better to avoid it. I'm sorry if one negative meaning deprives people of the joy of using a nicer one, but such is the nature of language.
Please review https://news.ycombinator.com/newsguidelines.html.