* Text AND binary so that humans can edit easily, and machines can transmit energy and bandwidth efficiently.
* Carefully designed spec to avoid ambiguities (and their security implications).
* Strong type support so you're not using all kinds of incompatible hacks to serialize your data.
* Versioned, because there's no such thing as the perfect format.
* Also, the website is 32k bytes ;-)
+ Avoids ambiguities.
- The format seems to feel the need to support everything, including things I am not sure are actual usecases (what's the point of Markup element for example? What does Metadata save us compared to just including it in document, given that parsers must parse it anyway?). This must make implementation most complex and costly, and makes reading the text format more difficult.
- Not a fan of octal notation. At 3am not sure I can't confuse 0 and o given certain fonts. Does anyone even use it these days?
- Unquoted string were discussed in the thread, I'd like to point out that it's very easy to make an unquoted string not "text-safe" (according to the spec) without noticing it, at which point document is invalid.
Just add white-space (maybe a user pasted a string from somewhere without noticing whitespace at the end or forgot the rules), a dot, an exclamation or a question mark. Having surprises like that is IMHO worse than a consistent quoting method.
Basically all the things I don't like are about the format supporting a bit too much. YAML 1.1 should teach us more is sometimes less.
I put in octal because it was trivial to implement after the others. The canonical format when it's stored or being sent is binary, and a decoder shouldn't be presenting integers in octal (that would just be weird). But a human might want octal when inputting data that will be converted to the binary format.
Markup is for presentation data, UI layouts, etc, but with full type support rather than all the hacky XML+whatever solutions that many UI toolkits are adopting. Also, having presentation data in binary form is nice to have.
I saw there's a 'Media' type in the spec. It's seems the type is actually for serializing files. But there's no "name" (or we can call it "description") field. Of course we could accomplish this with a separate field - but than again the entire type's functionality could be accomplished with a u8x array and a string field. So if you're specifying this type at all, might as well add a name field to make it useful.
The media object is a way to embed media data directly into a document such that the receiving end will have some idea of how to deal with it (from its media type). It won't have or need a "file name" because it's not intended to be stored in a filesystem, but rather to be used directly by an application. Yes, it could be built up from the primitives, but then you lose the canonical "media" type, and everyone invents their own incompatible compound types (much like what happened with dates in JSON and XML).
- I'm removing the metadata type. You're right that it's not really gaining us anything.
- I'm changing strings so they always must be quoted. This actually simplifies a lot of things.
Thanks for the critique!
Any reason for not using RFC2119 keywords in the spec? Using them should make the spec easier to read.
Unquoted strings are much nicer for humans to work with. All special keywords and object encodings are prefixed with sigils (@, &, $, #, etc), so any bare text starting with a letter is either a string or an invalid document, and any bare text starting with a numeral is either a number or an invalid document.
> Any reason for not using RFC2119 keywords in the spec? Using them should make the spec easier to read.
I use a superset of those keywords to give more precision in meaning: https://github.com/kstenerud/concise-encoding/blob/master/ce...
An important feature of RFC2119 keywords is that they're always capitalized (ie. the keyword is "MUST", not "Must", or "must"). This makes requirements and recommendations stand out amid explanatory text, improving legibility. For example, RFC2119 itself uses MUST and must with different meanings.
Because strings can contain whitespace and other structural characters that would confuse a parser.
> Having two representations for the same data means you can't normalize a document unambiguously.
The document will always be normalized unambiguously in binary format. The text format is a bit more lenient because humans are involved.
The idea is that the binary format is the source of truth, and is what is used in 90% of situations. The text format is only needed as a conduit for human input, or as a human readable representation of the binary data when you need to see what's going on.
> An important feature of RFC2119 keywords is that they're always capitalized (ie. the keyword is "MUST", not "Must", or "must").
Hmm good point. I'll add that.
So everything else needs some kind of initiator and/or container syntax to logically separate it from the other objects when interpreted by a human or machine.
- A proper editor was never around.
- Closing tags were verbose.
- Attributes vs tags was confusing.
- It didn't map "naturally" to common data types, like lists, maps, integers, float, etc.
XML is serialization. I hardly believe you was concerned about serialization while posting comment or thought about attributes-tags distinction.
This page utilizes request to server for multi-user editing. But it is easy to build truly serverless (like a file) document with same interface:
data:text/html,<html><ul>Host: <span class=host contenteditable>example.com
Change it, save it, done. Web handles input of lists, maps, integers, float and much more. 636f 756e 7472 6965 733a 0a2d 2047 420a
2d20 4945 0a2d 2046 520a 2d20 4445 0a2d
It is not practical to edit Excel documents in plain text: <?xml version="1.0"?>
<Workbook xmlns="urn:schemas-microsoft-com:office:spreadsheet"
xmlns:o="urn:schemas-microsoft-com:office:office"
xmlns:x="urn:schemas-microsoft-com:office:excel"
xmlns:ss="urn:schemas-microsoft-com:office:spreadsheet"
xmlns:html="http://www.w3.org/TR/REC-html40">
<Worksheet ss:Name="Sheet1">
<Table>
<Row>
<Cell><Data ss:Type="String">ID</Data></Cell>
Tim Berners-Lee browser was browser-editor. Can't you see parallels?"You need this special tool to work" immediately and instantly rules out "easy to edit". Or makes the debate irrelevant: every format is easy to edit if you have "a convenient UI" to do it for you.
And plain text editor is a "widely deployed special tool to work". Actual data is
countries:\n- GB\n- IE\n- FR\n- DE\n- NO
Or 636f 756e 7472 6965 733a 0a2d 2047 420a
2d20 4945 0a2d 2046 520a 2d20 4445 0a2d