Parsing JSON Is a Minefield (2018)
seriot.ch
seriot.ch
It would be much simpler if all primitives were strings, and it'd probably save a few people from accidentally doing the wrong thing while dealing with prices.
There is something very nice and expressive about the existing JSON types. Just 6 types (null, boolean, string, number, array, and dictionary) are enough to cover a ton of use cases, and as you suggest, one can always fall back to "stringly typed" alternatives by implementing one's own serialization and deserialization for extra types.
CBOR features are almost one-to-one with JSON, except that the encoding is more size-efficient, it supports a few additional types (e.g., integers and floats are separate), and it allows semantic tags.
Expectation: tags in CBOR allow you to pass semantics. Reality: multitude of tags, and absence of strict rules for the tags make it pain in the ass.
However, some of the things I mentioned above, do have benefits for interoperability with JSON, although they aren't good for a general-purpose use; I think that it would generally be better to make a good format rather than trying to work only with the bad ideas of other specifications. (Fortunately, I think what I described above could be implemented using a subset of CBOR.)
However, using these formats (whether CBOR or JSON) is often more complicated than should be needed for a specific use anyways.
For configuration formats, I 100% agree with you. I do not want any data type except a string and a hashmap (maybe an array if you're being luxurious). Not an int, not a float, not a boolean, not a datetime (looking at you, TOML). For configuration formats I am always immediately feeding those files into a language with a richer type system that will actually parse them; my program and its embedded types are the schema. (Users of dynamically-typed languages may reasonably disagree.)
However, for the serialization use case, I'm not so sure. There's an argument that having a schema against which to do lightweight validation at several points in the pipeline isn't the worst idea, and built-in primitives get you halfway to a half-decent schema. I'm ambivalent at worst.
They are not. Configuration is a very tiny subset of a more general problem that you also mention: serialization.
Your config file will be de-serialized by your program and parsed into some specific types. Including numbers (tons of edge cases), dates (tons of edge cases), strings (tons of edge cases) etc.
It becomes worse when your program is used by more people than just you: which field is a date? In which format? Do you handle floats? What precision? What's the decimal separator? Do you do string normalization? What are valid and invalid characters, if any?
You can't pretend that your config is "just strings". They are not
Human input is full of tradeoffs, that’s why it’s bash and not typescript in your shell path column. And you’ll meet a great resistance from users if you make your config fully typed and require to refer to schema dtd ns or whatever bs xml had.
Bash is there purely for historical reasons. And it sucks.
> And you’ll meet a great resistance from users if you make your config fully typed and require to refer to schema dtd ns or whatever bs xml had.
That schema can and will help editors to validate and autocomplete things on the fly, and can also serve as a reference for what actual data the config accepts.
You can't build a generic schema validator that will accept exactly the valid configs for some program and nothing else anyway, so forget the half-assed type checking attempts and just provide the hierarchical structure. It's up to the application to define the valid grammar and semantics of each config option and parse it into an application-specific type.
But different languages interpret different strings in different ways by default.
This leads to major bugs.
One of the great strengths of JSON is that parsing a number is well-defined.
The way you're suggesting would lead to people emitting JSON with leading zeros sometimes, and then some languages end up interpreting certain numbers as octal.
No thank you.
JSON numbers are far more restrictive than strings and carry precisely defined meaning in a way that arbitrary strings don't. They're only "just certain strings" in the same way anything can be serialized to a string, which doesn't really mean anything.
What does jq do to them?
echo 1.4e99999999999999 | jq
1.7976931348623157e+308
While I agree that the meaning of json numbers exists, I'm not sure which JSON standard you're referring to that contains this meaning. json.org certainly does not contain it, and links to ECMA-404, which just says "JSON is agnostic about the semantics of numbers."I've never gone back to formalize the grammar or otherwise mature it. But it's served me well as-is, and it's been easy to convert "up" to JSON or YAML or XML or what-have-you, once the case for an interface beyond plain text proves worthwhile.
TCLON?
https://news.ycombinator.com/item?id=12796556
Parsing JSON Is a Minefield (2016) - https://news.ycombinator.com/item?id=28826600 - Oct 2021 (173 comments)
Parsing JSON Is a Minefield (2018) - https://news.ycombinator.com/item?id=20724672 - Aug 2019 (178 comments)
Parsing JSON is a Minefield - https://news.ycombinator.com/item?id=16897061 - April 2018 (246 comments)
Parsing JSON is a Minefield - https://news.ycombinator.com/item?id=12796556 - Oct 2016 (292 comments)
Note, just so nobody reminds me, don't parse JSON with eval for security reasons. I'm just curious how it would work from a parser completeness point of view.
(the only possible problem is if you are designing a security system, but even then, since all the ambiguity is whether to reject the string, it will cause DOS at worst)
- The numbers is floating points, but cannot be Infinity and NaN. It is not a integer type, so long integers might not work properly. (There are other problems with numbers too, as mentioned in that article.)
- The strings is Unicode. Non-Unicode (including binary data) doesn't do properly, and even Unicode can have problems (some of which are mentioned in that article, but there are others too).
- Keys are only strings, not numbers.
- Syntax convenience isn't so well, e.g. doesn't have comments, optional trailing commas, etc.
- The format is difficult for reasons explained in that article, too.
One possible alternative would be a format based on a subset of PostScript (instead of JavaScript), e.g. (a part of a example from Wikipedia):
<<
/first_name (John)
/last_name (Smith)
/is_alive true
/age 27
/phone_numbers [
<<
/type (home)
/number (212 555-1234)
>>
<<
/type (office)
/number (646 555-4567)
>>
]
/spouse null
>>
PostScript also has binary format, comments (with a percentage sign), hex string literals, etc. (And, commas are not used, so the problem with trailing commas also does not apply.)(Nevertheless, I did write a JSON parser (and also a JSON writer) in PostScript.)
It is also possible to use binary formats, CSV, etc, depending on what exactly is needed by the program; for many reasons, one format cannot solve everything.
I personally hate the usual interpretation as float and see it as a common but extremely-implementation-induced failure. It's far better interpreted as an arbitrary precision numeric type, not float or int. The spec even says as much and only says that implementations mostly suck so watch out. IMO precision myopia is why we end up with e.g. Python's refusal-by-default to (de)serialize from/to Decimal.
edit: I didn't mention integer keys, because object members canonically start with a letter.
I never understood these two choices in the spec as they are totally against the “human-readable” goal…
This is not true, JSON numbers are simply signed decimal numbers. They might be parsed into floating point (as is the case with JavaScript), or any other numeric type, which makes them unreliable without additional constraints beyond what JSON specifies.
>- The numbers is floating points, but cannot be Infinity and NaN.
The numbers are in fact real. Infinity and NaN are not reals.
It does at least support comments though. Biggest flaw in JSON by far.
I dunno, it matches up reasonably well with languages that have nestable custom types.
XML labels nodes, and JSON labels edges. They both have pluses and minuses.
https://docs.rs/serde-xml-rs/0.6.0/serde_xml_rs/#caveats
Look at how much more complex this is than the equivalent JSON code, which requires none of these annotations:
Would you mind explaining why you thinks it is an apple and oranges comparison?
Feedback:
> I wrote yet another JSON parser (section 6)
Link defunct.
One day a student came to Moon and said: “I understand how to make a better garbage collector. We must keep a reference count of the pointers to each cons.”
Moon patiently told the student the following story:
“One day a student came to Moon and said: ‘I understand how to make a better garbage collector...
```