I don't think your take makes sense. Comments were purposely left out of JSON to mitigate the problem described by Hyrum's law.
Smaller interface means a smaller surface to allow abuse. You're just arguing to facilitate abuse because you feel the abuse that was indeed avoided by leaving comments out is unavoidable, which is a self-contradiction.
On top of that,think about the problem for a second. Why do you feel it's reasonable to support comments in data interchange formats? This use case literally means wasting bandwidth with data that clients should ignore. The only scenario where clients are not ignoring comments if exactly the scenario that it's being avoided: abusing it for control purposes.
If you want a data interchange format to support that, you should really adopt something that is better suited for your whole use case.
> and your suggestion doesn't help with eg syntax highlighting tools for your config that will treat comments as syntax errors (...)
That's the whole point. This is by design, not a fault. Why are you pretending a language supports a feature it never supported and explicitly was designed to not support?
If you take your law seriously, this is irrelevant because the surface of abuse is on the same scale of practical infinity, so it doesn't matter that one infinity is technically smaller.
For example, based of the example in the quote: you could stick those directives from comments info #hashtags in stringy values, with the same effect that there is no interoperable way to parse values (or if you add "_key_comment" - there is no interoperable way to even parse keys as some of them need to be ignored)
So the designer has achieved no benefit by removing a valuable feature
> abuse that was indeed avoided
Nothing was avoided, you can have the exact same abuse tucked into other elements
> Why do you feel it's reasonable to support comments in data interchange formats?
Why does the author of the quote sees the obvious which you don't see even after reading the quote? Go convince him his comment makes no sense because of "data interchange"
Obviously it's not only used for data interchange in cases where every byte matters (reminder: this is a TEXT-based format) and also comments matter to humans working with this data
> adopt something that is better suited for your whole use case.
And this discussion is literally about a format supports that? But also, how does this in any way mitigate the flaws in the designer's arguments?
> Why are you pretending a language supports a feature it never supported and explicitly was designed to not support?
Same thing, why is the author of the quote makes this senseless suggestions then? Go convince him first
Hyrum's law is not mine. It's a statement of fact that goes by that name already for a few decades and is noted by anyone who ever worked on production services that exposes an interface and is consumed by third parties.
> (...) this is irrelevant because the surface of abuse is on the same scale of practical infinity, so it doesn't matter that one infinity is technically smaller.
It really isn't. JSON does not support comments, thus they aren't abused in ways that sabotage interoperability. The only option you have is to support everything at the document schema level. You can't go around it.
> For example, based of the example in the quote: you could stick those directives from comments info #hashtags in stringy values, (...)
Irrelevant. You're specifying your own schema, not relying in out-of-band information.
That's exactly how a data interchange format is designed to be used.
There is no way around it. If you understand the problem domain, the fact that leaving out comments is an elegant solution to a whole class of problems is something that's immediately obvious to you.
If specifying the schema is important then why not an interchange format with strict schema enforcement like XML?
If minimal feature sets to avoid abuse are so important, then why not a binary format like protobufs which also have strict schema dictates?
And if interoperability is important then why not ditch REST/JSON all together and instead favor SOAP which strictly defines how interoperability works?
That's why I don't buy the "comments might be abused" argument. JSON doesn't have a single problem domain and it's not the only solution in town.
I don’t have anything against XML by the way, it was purely horrible to work with because people were using it in so many weird ways. Personally I prefer toml, but I guess we all have our preferences.
JSON wasn't some magical, made-on-the-fly format that makes Crockford some kind of genius. It was simply the standard Javascript object literal notation with some added constraints. I think some of those constraints make sense (i.e. are there any other languages that support both single and double quotes for string literals?), but funnily enough, some of the biggest issues with JSON interoperability is it is very underspecified in the areas that matter, such as the type and width of numeric literals, what to do in some edge cases like duplicate keys, etc. Just did a quick search, and here is a post that outlines some of the real security risks this underspecification leads to: https://bishopfox.com/blog/json-interoperability-vulnerabili...
Yes, quite a few like Python, PHP, Fortran, COBOL, Lua, R, Ruby, Perl, Bourne Shell, Dart, Groovy, etc.
Though some only interpolate with one or the other.
Sure it’s derived from JavaScript and it plays a major part in frontend development today. It was really the other way around though, when Ajax picked up people started realising that they could use json to make frontends work with “just” html and JavaScript.
This actually would put JSON in a worse place than XML - while XML has an overly complex infoset, that infoset is actually defined and standardized. Representing "a property with a comment on the property name and one before and after the property value" so that information is not lost in parsing would explode the complexity needed for an "interoperable" JSON library.
if someone wants to create some sort of scheme where they do "createdAt$date" as a property name to indicate the value is to be interpreted in some agreed-upon date format, that at least doesn't lose data if the JSON data doesn't understand that scheme, or require a new custom parser in order to properly interpret that data, compared to something like /* $type: date */ "createdAt" :...
This doesn't explode anything and you don't need to lose any data, so the monstrosity of XML still has no benefit, and neither does "createdAt$date", which would need a custom library anyway, so it doesn't matter where you insert your types
A bit off topic when the only consideration here is comments, but XML allows for type definitions and structures data that simply isn't possible in JSON.
People have found ways to attempt to add types to JSON but they aren't part of the spec at all and are just convention-based solutions.
If you need additional fields in the json to hold comments, why not add the fields however you want? And if you need meta-data for the parser, you could add it the same way. In a project, i am working on right now, we simply use a attribute called "comment", for comments.
e.g. use "_" as a prefix to mark comments, and then tell you applications to ignore these attributes.
{
"mystring": "string123",
"_mystring": "i am a comment for mystring",
"mynum": 123,
"_mynum": "i am a comment for mynum",
"comment": "i am a comment for the entire object"
}That may not be a huge problem for you, you see the "comment" key and know its just a comment and can ignore when a parser treats it like a normal string field. It could be an issue though, for example I could see any code that runs logic against all keys in the object becoming harder to maintain.
> If you need additional fields in the json to hold comments
I don't need additonal fields, I need comments
> why not add the fields however you want
Because I can't do that either, there are noticeable limitations
You can make special key names that are really directions for something.
You can make enrire k v pairs that are never used by anything that actually parses the json normally.
Argument was invalid as far as I can see and calling it "sorry it makes you sad" is, wow I don't even know where to begin with that.
Having annotation happen in a dedicated place designed for it is better than having it happen where it was not designed to be, end of math problem.
For any spec there are people who want something spec doesn't do and people writing the spec need to say no to requests that they consider not in scope as much or more often than they say yes
Comments are not some weird thing one person wants for their weird reason that no one else needs to care about. It's like leaving out a letter from the alphabet.
Just acknowledging someone wants something out of scope is not an apology
Saying "I intended to make a defctive thing doesn't" doesn't change the fact that it's defective.
A car without windshield wipers is defective, or at best inexcusably limited. Saying "I only designed it to use in the sun" doesn't make it suddenly perfectly useful.
If you want to try to say that json was actually intended to only be used in special conditions like an exotic car with no roof, then it's fair for everyone else to say "Ok, well that intentional design is a crap design of limited use. People actually need a car WITH windshield wipers."
The simple fact that it's a text format invalidates all the attemped arguments that it's just a machine to machine data format never intended for humans to mess with. People aren't "holding it wrong".
CSV has no comments and is probably more popular than json.
If small parser was the most important, that's what binary formats are.
Or if the data needs to pass through a text handler that can't handle binary, we already had csv and a few different standard and cheap encodings.
csv's excuse for lacking some of the bare minimum features is that csv was first. Csv is from the time of typewriters and handwritten data. What would eventually become known, the most common needs, just hadn't been encountered yet.
A single person did not create csv after 40 years of seeing what is needed, and decide that csv shall not have an expected functionality in it's spec. There essentially isn't even any spec for csv. No one explicitly authored it. It's a typewriter format.
csv continues to get used after that point because of inertia. Once it's used many places, it continues to get used in new places because most new things have to interoperate with the existing ecosystem if you want to sell to the most possible customers.
Also csv, unlike json & similar, IS a pure data format, not used for things like configuration, because it's really only barely human usable, because the columns don't line up and there's no form of annotation other than a header record. Not to mention every part of the format, quotes, commas, even newlines, are valid data that may appear in a field, and so the only way to read the file is to manually reproduce the streaming state-machine in your head.
If something as limited as csv is so good and doesn't need anything else, then why did he invent json when csv already existed for jobs that narrow in scope?
csv is not remotely proof that text formats don't need annotation.
Your two examples are just two examples of why we don't need comments for data interchange: yes, you put the data in a trivial, stable position in the parsed data that all parsers can parse rather than write some sort of DSL that has to be extracted from comment nodes that may appear in various places in the tree.
Turning this:
{ "key": "value" /* directive */ }
Into this: { "key": { "value": "value", "info": "directive" } }
Is the whole point. The more "abusive" you imagine the contents of "directive" to be, the more reason it should exist as the latter data, not more reason we should accept the former.Most json users are not writing all of the software that both generates and consumes the json.
I think roundtrip with comments is not feasible in general. Most code just expects a hashmap, which it edits and then serializes back. You would need some really clever handling for comments to do it.
Yes the comment free spec forces to normalize what would be comments into data spec and it is frustrating to put more effort into things
Syntax errors and erroneous highlighting are not even item 10000 on my list of JSON concerns.
Dare I pull a tired cliche and say “you’re using it wrong”