Thanks for your reply! It's a really good proposal, and I'll definitely consider it next time I'm serializing something.
> NUL is a tricky beast
One alternative I found interesting is Modified UTF-8 (https://en.wikipedia.org/wiki/UTF-8#Modified_UTF-8), where NUL is encoded as 0xC0 0x80. It's already in use in Java, apparently. Disclaimer: I've never heard about it before, and I don't know how well supported it is.
> Line breaks and commas
I'm from a country that uses comma as decimal separator, and I still prefer dots when programming. I dread ambiguous numbers like "1,001". I'm confident this is true for almost all technical people. And I really don't see non-technical users editing this kind of file.
If they can be trusted to read it, and to modify it without introducing syntax errors, they can be trusted to use dots and an editor that shows LF as linebreaks (i.e. anything but `notepad`).
My choice would be dots only for decimal separator, accept LF and CR LF on reading, but prefer LF when writing.
> Entity references
You can always define manual ids:
c1
{
"alice"={friends=["bob" "charlie"]}
"bob"={friends=["alice"]}
"charlie"={friends=["alice"]}
}
Of all data types that could have native encodings (colors, IP addresses, lat/lon coordinates, enums, tags, markdown, hashes, file permissions, OIDs, DOIs, ISBNs, etc), I think arbitrary object graphs bring too many downsides for a serialization format.
You don't want reader CVE's because of Billion Laughs, and you don't want to force every programmer writing a simple traversal algorithm to correctly handle cycles.