YAML is not a superset of JSON
patrickstevens.co.uk
patrickstevens.co.uk
The YAML "Scalars" section[1] says:
> A few examples also use the int, float and null types from the JSON schema.
And includes these examples:
canonical: 1.23015e+3
exponential: 12.3015e+02
So, is the "+" required here or not? Is a YAML parser buggy if it doesn't parse all JSON numbers as numbers?Edit: Ah, further on, it says:
Canonical Form
Either 0, .inf, -.inf, .nan or scientific notation matching the regular expression
-? [1-9] ( \. [0-9]* [1-9] )? ( e [-+] [1-9] [0-9]* )?
The example 1e2 clearly matches this regex, so his YAML parser is broken.Edit edit:
In YAML 1.1, there were separate definitions of float[2] and int[3] types, where only floats support "scientific" notation, and must have a ".", unlike JSON.
So this article is talking about YAML 1.1, while the other article is talking about YAML 1.2.
[0] https://news.ycombinator.com/item?id=41498264
[1] https://yaml.org/spec/1.2.2/#23-scalars
Most YAML parsers default to 1.1 for compatibility reasons, because if they default to 1.2 then existing YAML documents expecting 1.1 behavior will be parsed incorrectly.
YAML is a difficult language to parse if you care about getting the correct data.
1e2 does not match this regex. 1e+2 or 1e-2 would, though.
> The content of a mapping node is an unordered set of key/value node pairs, with the restriction that each of the keys is unique
This is not about semantics, it's about grammar. While it's fair to say that JSON "usually" is valid YAML, it's still good to be strict about it, because the existence of a single counterexample can be used maliciously.
Precisely. The article is really noticing quirks and limitations libyaml, the library doing the heavy lifting behind PyYAML, not YAML-the-spec proper.
Granted, in practice, library limitations are probably what you want to know about. AFAIK, libfyaml[0] (not libyaml) is the most spec-compliant library around. It's a shame more downstream languages aren't using it.
- YAML is for files written or edited by humans - e.g. configuration files for servers with lots of comments and explanations, but generally quite simple key value pairs or basic data structures
- JSON is for files written and consumed by machines. It allows for complex, nested data structures and types, but requires technical knowledge to use.
The problem arises once you start confusing these usecases. I'd argue that once you start writing `'{"a": 1e2}'` in YAML you're quite far outside of its ideal use. I appreciate that feature creep might lead to overly complex configuration files (I remember editing XML config that allowed you to specify for-loops in my earlier days), but really, at a certain point it might be worth taking a step back and reflecting if you're still using the right tool for the right job.
A minor change of nesting is easy to accomplish in JSON, while YAML practically requires support from a text editor. JSON _can_ be rendered in a pretty print fashion for easier human editing, and it should still parse correctly irrespectively of how additional non-printing-space is added.
I have always hated YAML, and still to this day, because I cannot write a yaml file, the indentation makes no sense to me and the list syntax is black magic (you actually have several ways to write those, once again the indentation implication is obscure). So while agreeing on the goal to be written/edited by humans to my perspective it fails at it.
Also 1e2 might not be the best example as this is just 100, but as someone who had to pass a lot of neural network training hyper-parameters : passing 1e-3 and so on is definitely a use-case. I am on the negative values Xe-XX (and YAML 1.1 would parsed it OK) but I guess other domains could also use the positive side to pass values (maybe as upper limits like 1e5 or so).
I think the YAML format should have parsed those number formats from the beginning. If this is fixed now, good job, hopefully the default yaml parsers are going to be "fixed". I would still use TOML over YAML any day, waiting for a human-json (some already exist) to be popularized one day.
These are completely different serialization formats.
But the yaml docs said it’s accidental
«The YAML 1.18 specification was published in 2005. Around this time, the developers became aware of JSON9. By sheer coincidence, JSON was almost a complete subset of YAML (both syntactically and semantically).»
The SAP article definitely states it, that's the first time I have seen it described that way.
It's one of those beliefs that seems like it should be true, but isn't for obscure technical reasons.
It appears to still be true in at least Ruby and Python, which are probably the two most popular languages to write YAML-consuming programs in:
$ irb -v
irb 1.3.5 (2021-04-03)
$ irb
irb(main):001:0> require 'yaml'
=> true
irb(main):002:0> YAML.load '{"a": 1e2}'
=> {"a"=>"1e2"}
irb(main):003:0>
and $ python3
Python 3.10.12 (main, Nov 20 2023, 15:14:05) [GCC 11.4.0] on linux
Type "help", "copyright", "credits" or "license" for more information.
>>> import yaml
>>> yaml.safe_load('{"a": 1e2}')
{'a': '1e2'}
>>>
-- > The spec specifies it should assume 1.2 and 1.1 should be opt-in.
The spec for 1.2 says that, but that's the reason parsers can't upgrade to 1.2, because changing the default version will cause backwards-incompatible parsing changes. Without the version directive there's no way for the parser to know which version was intended, so it defaults to 1.1, so people writing YAML will write YAML 1.1 documents, because that's what the parsers expect.The only way YAML 1.2 is going to displace older versions is in a greenfield ecosystem that has all its tools using YAML 1.2 from the beginning, but that requires an author who both (1) cares a lot about parser correctness, and (2) wants to use YAML as a config syntax, which isn't a large population.
> YAML can therefore be viewed as a natural superset of JSON, offering improved human readability and a more complete information model.
If you must use some kind of configuration language, at least use Dhall or something more sane. Thank you in advance!
There are - similar to the JSON insanity - multiple YAML standards.
YAML 1.1 and 1.2+ are the important ones, as the "superset" argument is only valid since 1.2.
HOWEVER PyYAML is a YAML 1.1 parser: https://pypi.org/project/PyYAML/#description
This also can be responsible for many security problems, as ppl will assume things about JSON and YAML, but don't worry about which of the 8 different JSON standards / YAML implementations they use.
> The content of a mapping node is an unordered set of key/value node pairs, with the restriction that each of the keys is unique
The same is true in YAML 1.2, so it's not just a legacy thing, either.
What the ever loving hell.
For almost all intents and purposes, if you are asked to create a YAML file then you can choose JSON as your syntax instead, because your file will be understood by the YAML parser. The benefit being that JSON has far fewer quirks and edge cases.
It's comical that when people get confused with YAML (which is often) they convert their YAML snippet to JSON to see what's really going on. YAML is horrible for humans to write. Let's just use JSON, the sane syntax, instead. A few extra parents and quotes is really no big deal, and it's far easier to read unambiguously.
If we want human readability we should accept not being language agnostic or simply direct use a programming language for data as well (lisp teach), if we want a textual language agnostic data exchange format than we should ignore human readability and maybe stick with XML.
I guess there is a place for all... but I do wish s-expressions got more love. :-)
Great for cutesy human-readable applications but nothing workhorse. Just as I'd happily use Papyrus for invitations to a neighborhood barbecue but not a resume.
The problem is that it parses to different values in those two syntaxes.
No it is not?... A valid JSON document cannot start with ` or with '. You could argue that the poster added ` in the hopes of getting it formatted, but not the '.
{"a": 1e2}
There are no single quotes or backticks in it. Those are an artifact of posting on Hacker News.The old spec of YAML 1.2 section 1.3 explicitly said:
YAML can therefore be viewed as a natural superset of JSON, offering improved human readability and a more complete information model. This is also the case in practice; every JSON file is also a valid YAML file.
The revised spec 1.2 revision 1.2.2 (2021-10-01) no longer contains that sentence; but still says, in section 1.2:
The YAML 1.2 specification was published in 2009. Its primary focus was making YAML a strict superset of JSON.
and in section 6.8.1:
Note that version 1.2 is mostly a superset of version 1.1, defined for the purpose of ensuring JSON compatibility.
Given all these claims, Patrick Stevens’ observations that YAML really isn’t a superset of JSON, because YAML can’t handle all JSON number literals, and tabs as whitespace, really is surprising. At least to me.
When previously JavaScript/ECMAScript 2018 was found not to be a JSON superset, at least it was about unescaped occurrences of little-used characters U+2028 LINE SEPARATOR and U+2029 PARAGRAPH SEPARATOR in string literals. And even that got fixed (by allowing the unescaped characters) in ECMAScript 2019.
[YAML 1.2]: https://yaml.org/spec/1.2-old/spec.html#id2759572 [YAML 1.2 revision 1.2.2]: https://yaml.org/spec/1.2.2/ [ECMAScript 2019 feature Subume JSON]: https://v8.dev/features/subsume-json