JSON parsers that can accept comments
douglascrockfordisnotyourdad.technomancy.us
douglascrockfordisnotyourdad.technomancy.us
JSON in its original form is fine if you want to pass data from a server to a browser's fetch request. No need for comments.
(If you're passing data from one process to another, and neither of the processes is a browser, have you considered protobuf?)
For humans writing config files, such as the VS Code one where you sometimes need to go in and do things by hand, both the ability to temporarily comment stuff out and do put comments above particular settings, are extremely useful if not essential. The "comment: '...'" approach doesn't work here or anywhere else where you have a schema to validate against.
Comments are not part of how JSON was originally defined. It's technically correct (the best kind of correct, some say) that comments are not part of the spec. But comments are sure how most people who edit JSON as a human [want to] use it. If the spec says A but everyone is doing B because that's what best solves their problem, then either you write a new spec and call it JSON++ or JSON-C or something and watch most people switch to that, or you adapt the spec in the first place.
We've seen this with markdown (ok, commonmark) and HTML already, among other things.
Not only it supports comments and trailing commas, it allows to avoid commas at all (they are simply ignored by EDN). Commas indeed are just redundant and incur visual noise most of the time.
[1, 2, 3]
[1 2 3] ; much better
The spec: https://github.com/edn-format/edn/blob/master/README.mdExamples at github (261k files found at the moment) : https://github.com/search?q=path%3A*.edn+&type=code
Github search only shows the first 5 pages of results. To extract more results split the search with more specific path qalifiers:
path:*a.edn
path:*b.edn
path:*c.ednI would argue semicolons are and only sometimes commas. Personally I prefer [1, 2, 3] much more over [1 2 3]. With long numbers and dots and strings and whatever, the seperator is helpful for me.
https://layer8.space/@douglascrockford/113595316189101091
> If the comment is important enough to be put in the text and remembered, then make it an explicit part of the data structure. That will also make it easier for tools to find it and process it in a useful way. {"comment": ...}
> Use a preprocessor tool such as jsmin to remove the commentary before passing it to a JSON parser.
I kind of wish JSON had trailing commas though
either it's a data transfer format and comments should be stripped before transmission
or the comments are part of the data model
makes you have to think clearly whether the comments are just "notes to self" at authoring time, or something relevant to the consumer
OTOH of course there are plenty of 'greyer' cases
Both JSON (as defined in the RFC) and JSON5 have a nice property of being well-defined, meaning that you can use different libraries in different languages on different platforms to parse them, and expect the same result. "JSON but parser behaves reasonably (as defined by the speaker)" does not have this property.
"Despite the clarifications they bring, RFC 7159 and 8259 contain several approximations and leaves many details loosely specified."
And Gruber wouldn’t give Jeff Atwood permission to call his variant <something> Markdown, or it seems anybody else, so we ended up with CommonMark, and GFM.
Json5 is good for JSON at rest, as others have mentioned already.
JSON as currently spec'd is honestly quite bad at both jobs, but the most rational defense of its use as a data format is that it's (mostly) human readable. Given that that's its main value proposition, what exactly is the reason for saying that JSON-as-data-format should not have comments? What do we lose if we allow them?
Because JSON originally did have comments, and people were putting pragmas into them, and so different parsers would act different depending on whether they understood them or not. Comments ended up being an anti-feature in JSON because people were abusing them.
Source:
> I removed comments from JSON because I saw people were using them to hold parsing directives, a practice which would have destroyed interoperability. I know that the lack of comments makes some people sad, but it shouldn't. […]
* https://web.archive.org/web/20190112173904/https://plus.goog...
> Suppose you are using JSON to keep configuration files, which you would like to annotate. Go ahead and insert all the comments you like. Then pipe it through JSMin before handing it to your JSON parser.
Doesn't seems he is that against the idea
I would call out portability instead, which is not dependent on the byte ordering or endianness issues of binary data formats.
sort of like: javascript is portable code, json is portable data.
But there are dangers there - look at how horribly comments get abused in code:
* doctests are nonsense, just write tests. (doctests like rusts that just validate example snippets are the closest thing to good I've seen so far, but still make me nervous).
* load bearing comments that code mangling/generation tools rely on (see a whole bunch of generated scripts in your linux systen - DO NOT EDIT BELOW THIS LINE)
* things like modelines in editors that affect how programs interact with the code
* things like html or xml comments that on parsing affect end user program logic.
Comments can be abused, and in something like JSON on the wire I can see systems which take additional info from the comments as part of the primary data input. Often a completely different format... and you end up with something like the front-matter on your markdown files as found in static site generators.
Point being, comments are not a purely benign addition.
these are mostly a warning sign for humans, to be read as "if you need to modify the script below this line, a) you gotta be knowing what you're doing, we are not held liable for support if you change stuff around there b) please contact us to make sure we didn't miss a legitimate need or c) you're trying to do something in a bad way and there's better ways to do so".
I don't see anything intrinsically wrong with doctests. I also can't see a better way to do "load bearing comments," and I'm not eager to go back to "Step 2: Edit your .bashrc to include foo."
Why are doctests not tests?
> doctests like rusts that just validate example snippets are the closest thing to good I've seen so far
Rust's doctests don't seem to be fundamentally different from Python doctests, which is the language I've seen most commonly make use of doctests.
Also languages with some progrommatic capabilities like cue, dhall, jsonnet, nickel etc.
Non of them are perfect, and some are less suitable for certain use cases than others. But IMO pretty much all of them are better for human editing than json, and in many cases yaml.
- doesn't require to quote everything
- has lists/dictionaries
- uses indentation and new lines instead of commas and brackets
- doesn't have 1000 unnecessary features like YAML
Also, you don't need all types from JSON.
JSON has a very minimal set of types and I regularly use all of them. I guess you could argue that integers and numbers could be combined, but I think that's it.
That said, I don't like it as a config file read/written by humans.
Seeing as I can only see the use case as a file format to be read/written by humans in the loop, then maybe the conversation should be about compiling the file format to a data format for compatibility outside of the user tooling.
The idea is that this forces better formats.
How well this works? Well, then I got an "x-comment" property or non-standard comments. Nonetheless. If people see the need to hack some extension in, they'll find a way.
Why did they bother making it text-only ASCII then ?
I like to solve problems - or at least bringing them to me doesn’t result in a loss of status for either party. People notice this about me and bring me problems. Someone recently described to people what is essentially my process: the likelihood of the cause divided by the difficulty of verification. Partially sort and just start checking off assumptions.
A lot of cheap but low probability options get shuffled higher, and just sending the wrong data is a common enough problem, especially with caching. And if it’s nearly free to look at the payload, it’ll get checked. If it isn’t people will try everything else to avoid it.
JSON is notable for making UTF-8 encoding a hard requirement.
…which was pretty ballsy back in the mid-2000s. We were still fighting with Shift-JIS and Windows-1252. Excel didn’t add proper support for UTF-8 until depressingly recently.
I don’t remember when I started pushing for utf-8 everywhere but it was “early” by most people’s standards, so I know what you mean.
And one of the things that makes me dislike MySQL is that they have a field type called utf-8 that isn’t. And they didn’t fix it, they introduced a new type instead. So that footgun was still there for all to trigger. So mad.
> Previous specifications of JSON have not required the use of UTF-8 when transmitting JSON text. However, the vast majority of JSON-based software implementations have chosen to use the UTF-8 encoding, to the extent that it is the only encoding that achieves interoperability.
1. comments are metadata (specifically Human/LLM-readable metadata vs machine-readable metadata)
2. general-purpose data formats should support metadata
What no-comments saved us from was stuff like this in our data interchange:
{
"count": 123 // bigint
"price": 10.99 // @precision=2
"date": "2024-08-12" // @format=YY-MM-dd
"data": /* !transform(rot13) */ "uryyb"
"storage": 5 // Unit(TB)
}
And who knows what deeper layers of hell we avoided.Frankly, VSCode shows that all this time people were complaining about no comments in JSON config and how hard it was to write config in JSON, they could have just written their apps to strip comments at read time.
So we do have the best of both worlds.
JSON is awful for writing manually because it requires typing too many quotes, commas etc. I think JSON is meant to be machine-generated and machine-read and therefore doesn't need any comments.
Having a mess where JSON parsers sometimes do, and sometimes don't allow comments is a bad outcome.
JSON5 is advantageous because it's explicitly separate - it has a different file extension, it needs different libraries. Also it supports a few other utility features people want (like trailing commas) without bringing in the whole dangerous kitchen sink like YAML does.
That looks like a worst of both world solution IMHO.
You are probably better of with a TOML -> JSON mapping and would be better of with YAML too if YAML hadn't had really stupid ambiguity pitfalls.
- JS syntax compatibility (low cognitive overhead)
- Decent balance between machine-readable and human-readable
- High familiarity for developers (if they know JSON, which they likely do, they can work with JSON5 with near-0 learning curve)
Plus, JSON5 support is _somewhat_ widespread (maybe ~50% of tools I use support json5 for config?)Files shouldn't be labeled as compliant with a standard and then not be. Full stop.
The standard that is JSON does not support comments. Don't call something JSON that isn't.
Use one of the many existing JSON extensions or create a new standard. DO NOT however just adhoc crap like this suggests, that's a road to hell.
https://learn.microsoft.com/en-us/aspnet/core/fundamentals/c...
//
/* ... */
#
--But the point is, if you are in control of your own parser, you can just use any one of those.
Yet YAML, which is somewhat related to JSON, uses
# This is a comment
I'm neither pro-comment or anti-comment, but I'm all for standardized comment syntax
to avoid parsing hell.It's a bad idea to write JSON parsers that can accept comments. Comments are not part of JSON.
There are plenty of extensions to JSON, and they have parsers. Use them. Or, use a language that is already an extension to JSON, like JavaScript or YAML.
I would tend to agree, but I don't think it's such a problem if it requires an option to be enabled and the docs make it clear that incorporating comments means it's not really JSON anymore.
(Naming it my-custom-format parser would circumvent ambiguity)
https://nigeltao.github.io/blog/2021/json-with-commas-commen...
Eh, who knows? But I’m skeptical that the omitted comma actually saves any space in common situations.
Because of that it didn't become XML or YAML or markdown or protobufs or *RPC or microsoft config file format or all the rest. All those formats have "reasons", but they are harder to understand, harder to parse or not as portable or on and on...
(that said, I wouldn't mind comments. JUST comments).
const obj = JSON.parse(jsonText.replaceAll(/("(?:\\"|\\.|[^"])"|[^\/])|\/\/.|(\/)/g,"$1$2"));
The above should be multiline JSON safe and give you single-line comments to end-of-line, it matches first strings (and checks for escaped double-quotes and escapes to get them correctly) or all other characters except // sequences (outside of strings) and passes that straight through via group 1(anything outside of string not starting with /) or 2(single / even if it's outside of JSON spec), if a // sequence is found it's not passed through and disappears.
Yes, regexps can be abused. This one should be fine though but only use it for config files you control :), use as CC0 and keep an keen eye if translating to another language since escapes will differ if the regexp goes into a string instead of a regexp literal like in JS.
Remember how JavaScript used to reside inside a HTML comment?
As has been pointed out multiple times in the thread so far, a random string that is ignored by the receiver is semantically equivalent to a comment, and people totally stuff random crap into strings to extend the data format in sometimes-horrifying ways. Banning comments doesn't stop us from abusing strings, it just means we don't have a good way to indicate that a part of the file should be ignored by the parser because it's not actually data.
Abusing the string is the correct place to put poorly planned data because it's just data and it has a stable position in the data (and syntax) tree.
But consider this example:
{
"a": /* directive */ 42,
"b": 42, // directive
// directive
"d": 42,
/* directive */ "d": 42
}
What does the code look like that has to reconcile the directive with the data it modifies?That's a completely different and worse problem than someone "abusing a string" with something like `{ "a": "42 directive" }`.
I think JSON made the right call. People just confuse that with it being somehow forbidden to use comments in JSON config, but they're wrong. Just look at VSCode's config. It was possible all along despite decades of people whining in blogs.
We didn't need comments in JSON over the wire. We didn't need a new format. We didn't need another blog post about it. Just write your tool to accept comments in its config.
You can just make a parser that ignores comments, while still allowing them in the syntax. You don't need to store the comments in the AST, or to actually parse them. You just need to skip them.
Imagine you’re a HTML parser library author in 1998. You’ve been happily skipping comments. Now you get complaints that your parser doesn’t see JavaScript. Turns out that everyone has agreed that these new script tags should be embedded inside comments.
Should you keep skipping comments and tell your users you’ll never support this HTML that everybody else now considers valid?
There are parsers which allow for comments if one wants, so if someone wants to engineer an insane system no one is stopping them, provided they can ensure the parser on the other end of it.
This is a data format meant to be readable by humans. As such, it's natural to want to support things like configuration via end-user editing values in the text as a use case. Data is occasionally going to need comments to explain valid options or add context for people to edit. This is a reasonable thing to have.
"like this"
never // Like this
Bad Things will surely befall us if we break the taboo.We already have XML. Stop trying to make JSON another XML.
I'm now knee-deep in a gig that uses all sorts of w3c standards, which enforce json-ld, json-schema, json-whatnots and so on. It's a terrible mess of unreadable complexity that has LONG been solved in XML. Decades ago!
I keep thinking that if you need one of those "fancy" features, why not just use XML? Why bolt them onto JSON, a format that was deliberately (?) set up to not have all these features?
I think that if you don't need these features: by all means, use JSON. It's simpler, cleaner, easier. But that it's only simpler, cleaner, easier because it's not XML. And that all these attempts at making it XML, end up with a JSON that's harder, more complex and often messier than the XML that we had for decades and that JSON tried to "solve" by being simpler. So it's one step forward, two steps back.
Then I discovered a handful of YAML parsers in various languages support code execution on parsing. Which basically killed YAML for me.
Have you ever written or tried to write a YAML parser?
Ok, we're being pedantic, perhaps it depends on how much information you put in a Yaml file, I usually see it used for small config files.
Also JSON + comment that's JSON5.
I don’t think anyone proposes switching to YAML for your API endpoints
I like JSON lines format where I can strip comment lines out.
Given its importance, we should take this into the standard body. Namely, the IETF should be pressurized enough to update RFC 8259 for every nitty-gritty about comments. It is clear that this new version of JSON needs to address multiple use cases which are often mutually exclusive, so we probably need to introduce some short indicator for profiles at the very beginning of the document. (Yes, that means you should declare in advance that you want comments and other humane syntaxes. I don't think that's very hard when you do want comments.)
It is often useful to express a JSON object as a single line of text. Comments that start from a symbol ("#" or "//") and continue to the end of the line can't be joined into a single line of text without throwing away the comments. /* */ is distinct by having an explicit end-of-comment symbol that allows meaningful text afterwards.
(I don't like the C preprocessor but in some situations it's the easiest solution.)
{"data":
#include "/home/myservice/secrets.json"
}// is unnecessary, though. It adds nothing besides an extra character to type.
It's the same as whitespace. You don't expect it to be round-trippable, it's just to make the source file easier to read. Likewise I would interpret any expectation of the preservation of comments the same as an expectation of the preservation of whitespace: weird, unreasonable, and probably arising from a misunderstanding of how JSON works.
XML preserved whitespace, too! I dislike XML intensely, but there was a method to its madness. There’s a method to JSON, too: simplicity and usability.
You start adding things to JSON, you lose that.
It’s unfortunate how many APIs are documented with a single json example, but annotating them with comments (and maybe types?) would help immensely.
Ambiguity is just one of such downsides
> We don't need a new format, we just need parsers for the existing format that are written by people who understand that Douglas Crockford is not their dad.
That's a new format, you just don't want a name change
Edit: I read the fine print on the page after posting this... Turns out that I'm not the first to come up with jsonc. The author just isn't a huge fan and wants JSON parsers to accept comments. Good luck with that!
{"comment" : "This works everywhere doesn't it?"}
{
"A": 1 // Don't remove this important but undocumented value
}The whole point of JSON is that it is an object notation: JSON parses into a kind of lowest-common-denominator object, either a number, a string, an array, a hash table, a boolean or a null. Every language supports those, one way or another. You can read an object into your runtime, mutate it, then print it back out and only the mutated part will change (this is round-tripping).
If comments were added, then to preserve round-trippability they would need to be readable. For that to work, they would need to live somewhere in the read-in object. But then, what would the result of reading this be?
3.14 # pi
It can’t be 3.14; it has to be something like [3.14, #<COMMENT: "pi">]. What even is a comment object?Where would the comment go in an example like this?
{
"foo": "bar" # baz
}
Maybe the returned object is something like {"foo": #<COMMENTED-VALUE: "bar" " baz"}? But what is that type?It’s a huge can of worms!
Sure, you could just give up on roundtripping. But it’s important!
XML of course has support for comments, but they are objects within its infoset: when you read in an XML document, you can access the comments as part of it. One of the points of JSON is that one doesn’t need all the complexity of XML to get stuff done.
In the vast majority of cases, it’s not. If removing comments breaks roundtripping, then so does altering the indentation or adding/removing whitespace, even ever-so-slightly.
EDN parsers strip comments when read – it’s not an issue in practice.
https://xkcd.com/927/ applies
On the other hand, people are adding comments to formats like JSON.
Comparatively, writing an alternative JSON parser is probably a bit wiser.
It's not a binary choice. I am more in favour of JSON comments than I am for CSV comments, but if I were to need them, I would probably use the Emacs style:
https://github.com/emacsmirror/emacswiki.org/blob/34edac6f86...
https://hey.hagelb.org/@technomancy/statuses/01JEKS1B4Q5NS34...
Doug, please adopt me!
As a file format, just use jsonc or json5, if you want to use comments.
if you have human managed configuration do not use JSON, it's a pretty bad choice for it and adding comments is not (fully) fixing it
or in other words
IMHO if you have JSON you shouldn't have a need for comments, if you do there is something wrong
char str[] = "That's like saying this is a C comment";No, he's your drunk uncle who sounds smart when you're young. When you learn to form opinions of your own, you realize that his opinions are worthless hot air.
Many years ago I went to a session where he promoted asynchronous Javascript as the "one true way." In the Q&A, I pointed out that callbacks made code much harder to read and maintain compared to threaded code; and then asked if there was a cleaner way to do it.
His answer was quite rude. A few years later, we got the "async" keyword, which solved the problem.
And they don't solve the core issue of async control flow that async/await solves, so it's not an alternative to async/await much less a superior one.
A classic example is when you want to conditionally do something asynchronously like B() in this case.
function process(id, callback) {
A(id, (result) => {
if (result === 3) {
B(() => {
C(result, callback);
});
} else {
C(result, callback);
}
});
}
Versus: async function process(id) {
const result = await A(id);
if (result === 3) {
await B();
}
return C(result);
}
Add a couple more layers of this and the async/await function stays simple and flat, and the callback version grows significantly more complex and nested.No serious person is going to claim that any form of callbacks is superior to async/wait
Note that your question didn't challenge this point (Async javascript is the "one true way", given that blocking in browser context always degrades the user experience).
The way you've bitterly dwelled [0][1] on something Crockford once said about async during a Q&A for 15 years or something is maybe a little bit weird.
There, I've explained myself, and I think the explanation is relatively straightforward. Now you. Why do you obsess bitterly over something Crockford said to you fifteen years ago during a Q&A? Perhaps more importantly, have you discussed Crockford with a therapist?