Unintuitive JSON Parsing
nullprogram.com
nullprogram.com
For example, instead of faithfully implementing the grammar from the specification, allow numbers with leading zeroes and then produce an error for them.
Another situation where this comes up is parsing language keywords. Instead of writing a separate lexer rule for every keyword, write a single rule for keyword-or-identifier, and then use a hash table lookup inside of that to determine if it is a keyword or identifier.
Using the approach in TFA, can the laser handle tokens that are prefixes of other tokens? Or even tokens that share prefixes?
Lexing "truefalse" as two adjacent tokens "true" and "false" seems slightly crazier than just lexing it as one (meaningless) token.
edit: I guess because "01" isn't a valid token.
The parent was talking about js.
But that's the problem. The tokenizer doesn't talk to the grammar parser (and vice versa)
The tokenizer could understand numbers with leading zeroes and throw an error there.
Something to think about: do languages - not json - interpret -1.2 as [MINUS][NUMBER] or just [NUMBER]. Or how does languages deal with 1.0-2.0 compared to 1.0+-2.0
As for the leading "-", in languages that have expressions it is common to parse the "-" as a prefix operator because that covers both negative numbers (-1.0) and negating variables (-x).
But in JSON there are no expressions like 1.0-2.0 so the leading "-" is parsed as part of the number.
Expr: ‘-‘ Expr {
return -1 * $2;
}
| Expr ‘-‘ Expr {
return $1 - $2;
}
| ‘(‘ Expr ‘)’ {
return $2;
}
| ...Of course there is no logical reason why the parser shouldn't have this concept just because the spec doesn't require. IMO, beyond basic correctness, user friendly error messages are the main differentiator between excellent parsers and crappy parsers.
The message in Firefox is: JSON.parse: expected ',' or ']' after array element at line 1 column 3 of the JSON data
In Chrome: Unexpected number in JSON at position 2
The case discussed in the article would benefit of a display with spaces between tokens.
Example: Source: [01] Compiler error: [ 0 1 ] ----^ SyntaxError: JSON.parse: expected ‘,’ or ‘]’ after array element
Such situation of two tokens without a separator character becomes much more obvious.
Notice also the character showing point of error. OCaml has been doing for a while ( https://ocaml.org/learn/tutorials/common_error_messages.html now shows underlined parts) then clang, cf. https://clang.llvm.org/diagnostics.html . This is much quicker for the human to communicate, than "line x column y".
Maybe there's a way to re-parse with a friendly parser with there's an error and dev tools is open.
In the case of the above, if last is ], previous is number and before is 0, then put a nice error message where octal notation is not allowed.
This is not “pure” from a language theory sense but in practice if you add a handful of common ones, it makes for a much better experience.
Helpful error messages for humans writing JSON is a special use case, not "the main differentiator" marking an "excellent [parser]".
Exposing a "helpful" JSON parser to the internet is a bad idea, and probably a waste of electricity, since the error messages will likely go into a black hole. An "excellent" parser might even reject some technically valid forms, for not being regular (indicating a suspicious or malfunctioning client).
I'm guessing from this that a "real JSON emitter" is one that perfectly implements the JSON spec and has absolutely zero bugs? Does such a thing exist?
> Exposing a "helpful" JSON parser to the internet is a bad idea, and probably a waste of electricity, since the error messages will likely go into a black hole.
There's a lot of software where any kind of unexpected error just gets repeated back to the user. I'm sure you've seen a modal popup like:
Unexpected error, please try again
(SyntaxError: JSON.parse: expected ‘,’ or ‘]’ after array element)
This is immensely unhelpful to basically everyone involved. If the message at least hints that it's due to a problem with the input itself -- for example: "SyntaxError: Invalid value ‘01’" -- then it's much more likely that an end-user can figure out which value is causing the problem and work-around it, and report a much more meaningful bug to the developer (thus allowing it to be solved significantly faster).Taking the argument that "helpful" JSON parse errors are pointless to its logical conclusion, there should only be a single possible error "JSON parsing failed", and I can't see how that's anything but a recipe for making everyone absolutely despise your API/product/etc, especially were your JSON parser ever to have even a single bug that caused that response erroneously.
If the JSON Serializer has a bug then the only option is to fix the bug instead of adding a human that is manually correcting mistakes in the generated JSON output. The only purpose of human friendly errors is because humans are involved and since they aren't we have to unnecessarily add them to make your complaint valid. (Don't tell me you're feeding the human friendly error messages into an ML algorithm. That's just bullshit.) The days of office workers doing pointless busywork of this type are long gone.
While I agree fixing the bug is the only option, it still has to be identified first. A bug that is based on data can appear to happen "randomly" and if the error is a useless one like the OP or even just "invalid JSON" it can take a long time before isolating the exact cause (especially if it's not obvious which piece of software is even the source of the error), and in that time end users, support and ops people lose trust in the system.
It think parser quality is a matter of correctness, error messages, speed, and resource use. How much each of these are to be prioritized depends on application. I'm entirely with you that most of the time error messages matter a lot and are often worse than they should be.
That is, it's not terribly difficult to deal with the tradeoffs if your description of the problem is at all correct
In the past I've encountered JSON lexing that only considers token boundaries on "special" characters i.e. ",}]:" and whitespace. This will return a lexing error when it sees "01" (equivalently "truefalse").
One could imagine first tokenizing only based on whitespace, then only starting to figure out what the tokens are. Which means parsing them individually. Which means another parsing step.
I think this would match human more closely: structure is more obvious based on visual separation than detailed analysis.
I guess it wasn't done that way because the current way of operation means one parser to rule all sources, and that parser can handle more complicated cases. That kind of design decision is more surprising later, but is kind of understandable when you draft a language as the same time as your first parser.
https://developer.mozilla.org/Web/JavaScript/Reference/Error...
See for yourself:
> 'use strict'; 01;
SyntaxError: "0"-prefixed octal literals and octal escape sequences are
deprecated; for octal literals use the "0o" prefix instead {
"path": "/foo",
"mode": 0644,
"contents": "bar"
}> A number is very much like a C or Java number, except that the octal and hexadecimal formats are not used.
and image:
https://json.org/img/number.png
as shown literally on the JSON home page:
I am all for good error handling, but at some point you do have to blame the user.
You can silently beat your child everytime he makes a mistake, until he accidentally does the job correctly (and doesn't get beaten), but it seems to me that making use of our ability to communicate can be much more efficient (and significantly less painful for the child).
And json is merely a (very innefficient, and somewhat problematic) protocol for information exchange; it's not something you should expect people to have read the spec for, especially when its whole popularity stems from it being "intuitive" -- that is, you don't really need to read the spec to deal with it effectively
JSON.parse("0o10") === 8?
I get SyntaxError: Unexpected token o in JSON at position 1
Essentially you have Netscape and IE. Netscape added support for "octal", IE did not, that meant that you had code like `x = 017` that had different values in the two engines. Given the early JSON parsers essentially just called eval() on the string that wasn't ok behavior for a data interchange format.
Then you have the absurd behavior of the Netscape octal implementation, which leads to such wonders as `018-017==3`, which make it a super terrible footgun.
Sensibly modern syntax makes the difference between octal and decimal very explicit with a 0o prefix, just like 0x, 0b, etc. I wish I knew why it was originally decided to not use 0o when 0x was in use.
In what sense is: {0}{1} better than [{0},{1}]? Presumably, if a few bytes are a major concern, you aren't using JSON anyway.
Concatenated JSON is ambiguous, so you need to put some whitespace between any JSON texts where both aren't arrays or objects.
This is a widely used technique.
To elaborate:
{0}{1} is better than [{0},{1}] because, when streamed, {0} is still valid, while [{0}, is not.
Also, making a language X parser accept anything that is not in fact language X is nothing but a terrible idea. If there is one thing that standards are good for, it's interoperability. And if there is one thing that hurts interoperability, it's having different implementations of supposedly the same standard accept and reject different inputs. That's how you get websites that work in one browser, but not another, because one browser was so helpful to make up some meaning for your creative markup instead of rejecting it with an error message, which obviously helps you absolutely nothing with the next browser that is of a different opinion. If you think the spec is stupid, you have to change the spec, if you don't manage to do that, you still should implement the spec, because interoperability is more important than whether your program can read some input that isn't JSON and that therefore no other JSON parser is guaranteed to understand anyway.
It seems honoring this type of technical correctness matters a lot. For example, imagine if ECMA added a new feature (e.g. 0-prefixed octal literals) in 2020..
Another issue: security. Imagine a hacker figured out that you used a mix of JSON parsers on your application (e.g. V8 and jq), and they produced different output.
For a vaguely related example, consider that some URL parsers interpret N (U+FF2E - fullwidth latin N) as ".", meaning you can sneakily add a ".." to the URL with NN (see https://www.blackhat.com/docs/us-17/thursday/us-17-Tsai-A-Ne...)
My question was more "for inputs not defined as being valid by the spec, is the result undefined (a la C++ UB where anything and everything is legal in response) or is it required to reject said input".
The sibling response says extensions are allowed, but that wouldn't come into play if an input is specifically called out as disallowed (vs simply not taken into account whatsoever).
“Numeric values that cannot be represented in the grammar below... are not permitted”
Regardless I agree we should not do such things.
This is false. In JavaScript, a leading zero, unless accompanied by a lowercase oh ('o') does indicate the number is written in octal.
08 === 0o10; // true
Here, the left side is still base 10, while the right side is base 8.
https://developer.mozilla.org/en-US/docs/Web/JavaScript/Refe...
Javascript for many years has assumed a leading zero means an octal prefix, and it's only recently that behavior has changed.
EDIT: Also your example does not disprove this. Here's another example that you should try running in the console:
011 === 11; // false
011; // 9EDIT: Also, yes, I the mistake with my initial example -- as '8' doesn't exist in octal, we have no choice but to interpret the leading '0' as padding and '08' as base ten, whereas '11' can be interpreted as base eight.
It is, at best, no longer true.
Of course, until recently JSON wasn't a strict subset of JS but that was an oversight rather than by design.
What am I missing out on? Why are they included in modern languages like JS?
I assume JS has it because so does everyone else.
Is JS a modern language? It was made 24 years ago as a prototype for a scripting language loosely mimicing Java. Presumably octals would have been still used often on recent machines.
So far I haven't ever had that need it seems.
> have an abbreviated form of binary
Yeah ok, but why octal over hex? After all, hex maps better to the underlying storage.
> pack decimal in a way that's easier to reason
How'd that work? I know about BCD but I don't see how octal improves the situation, being base-8.
> divide a number in half (down to 1) without getting fractions
Huh?
> represent file permissions (a grouping of four octals)
Ok, I get that for C and such, but how often do you do that in JS?
> Is JS a modern language?
Compared to C, where octal support is understandable, I'd say yes.
I found this to be a rather interesting article, and it'd be a shame if the discussion around it centered on such a well-tread topic.
https://en.wikipedia.org/wiki/Round-trip_format_conversion
The reason JSON can't support comments the way XML does, is that comments aren't part of the JavaScript "DOM". They disappear when you parse them, so they can't round-trip, and there's now way to stringify comments back out.
Standard XML parsers and XML/HTML DOM APIs give you all the comments and whitespace, and it's up to you to ignore them if you don't care, and they're not lost when you parse and re-serialize. Comments and whitespace are part of the standard DOM/SAX API, but you can't just nail those onto another model like JavaScript polymorphic objects/arrays after the fact.
Because parsed JavaScript objects and JSON structures provide no way to access the comments the parser threw away.
Although of course you could implement a parser that saved the comments, but it would need to support another more complex API than directly accessing JSON objects, which could somehow describe where each comment was in relation to the parsed object (since multiple comments can appear anywhere), and what kind of comment it was (// or /* */), as well as where all the whitespace the parser ignored was. (Although XML parsers typically don't tell you about whitespace inside of tags, so that can't round-trip.)
That is theoretically feasible with a JSON API implemented from the ground up in any language, like JSON.Net for example. But it's not practical in JavaScript itself (or Python or any language that parses JSON into pre-existing polymorphic arrays and objects), because you parse JSON into actual JavaScript (or Python) objects, whose API and implementation isn't under your control. So JSON being able to round-trip with any language other than JavaScript (and Python) itself would have been silly.
JSON would not be as powerful and useful a format if parsing then serializing JSON lost information. JSON was meant to be a round-trippable format, so there was no other choice but to leave out comments.
But maybe there's a use case for a "lossy" compressed JSON format like JPEGSON. ;)