Even if they were, it would still be easy to parse!
There's a whole spectrum of unofficial parser "helpfulness" here, with HTML 4 being an extreme case of parsers filled with hacks to deal with existing broken data, protobufs being an extreme case of parsers doing the One and Only True Thing, and JSON mostly toward the same end of the spectrum as protobufs, but a bit less so.
JSON hits a sweet spot of being very easy for computers to deal with almost all the time, while also being reasonably easy for humans to read and write.
I was going to add “if you started with that as the spec, it wouldn’t be hard to design something better than JSON” but real examples like YAML are pretty awkward, so probably it’s a harder problem than it seems.
But they’re not, so we have to resist the temptation to use them!
https://github.com/google/protobuf/tree/master/conformance
Here are some of the quirks of protobuf:
- non-repeated fields can occur multiple times on
the wire -- the last value "wins".
- you have to be able to handle unknown fields, including
unknown groups that can be nested arbitrarily.
- repeated numbers have two different wire formats (packed
and non-packed), you have be able to handle both.
- when serializing, all signed integers need to be sign-
extended to 64 bits, to support interop between different
integer types.
- you have to bounds-check delimited fields to make sure
they don't violate the bounds of submessages you are
already in.
I do think protobuf is a great technology overall. But it has some complexities too; I wouldn't want to oversell its simplicity and have people be unpleasantly surprised when they come across them later. :)FWIW, I think it'd be awesome if the only wire integer format were a bignum, in order to support full interoperability between integer types. Maybe even do that same for floats, too …
PHP serialization is better here, everything is type:value or type:length:value, although strings do have quotes around them, because their byte length is known, internal quotes need not be escaped. You can still have issues with genrating and parsing the human readible numbers properly (floating point is always fun, and integers may have some bit size limit I don't recall), but you don't need to worry about quoting Unicode values properly.
Protocol buffers have clear length indications, so that's easier, but it's not a 'self documenting' format, you need to have the description file to parse an encoded value. The end result is usually many fewer bits though.
Protobuf and similar are binary formats so don't have this limitation.
Canonical S-expression are both human-readable & length-prefixed. They do this by have an advanced representation which is human-friendly:
(data (looks "like this" |YWluJ3QgaXQgY29vbD8=|))
And a canonical representation which is length-prefixed: (4:data(5:looks9:like this14:ain't it cool?))Certainly no data format for data exchange between systems, especially untrusted sources.
That said, I've still been bitten by the Python implementation on the Mac acting differently from the C++ implementation on Linux, although I can't remember exactly what the issue was right now.