NestedText, a nice alternative to JSON, YAML, TOML
nestedtext.org
nestedtext.org
…only because NestedText does not support numeric types at all. That seems like throwing out the baby with the bathwater.
It's less of a problem if you're using it for configuration files, where a program knows what key's values need to be cast to an integer or float.
But it seems disastrous if you wanted to use it for storing or transmitting data, above all between applications. You're immediately throwing out the possibility of being able to serialize and then deserialize data in basically any programming language.
I shudder at the idea of an API that accepted NestedText, where I'd need to worry about whether my floating-point output was compatible with its floating-point string parser. Yikes. I want the serialization format to handle that. Isn't a major criticism of JSON that it doens't have a built-in datetime representation?
As for stringification, JSON's data types are mismatched with pretty much everything that isn't JavaScript to some degree. If you need to serialize a 64-bit integer to JSON, you serialize it as a string because the parser on the other end is probably going to try to parse it as a double-precision floating point number. Once you've started serializing numbers as strings anyway, it's not too far to "serialize every scalar as a string".
Indeed, not being able to have lists of dictionaries, or lists of lists, is very restrictive. Seems to be for very simple configurations only. E.g. a set of preferences, but not a set of monitor calibrations. (Inventing arbitrary dictionary keys seems pretty hacky.)
I don’t think this is true. None of the examples have lists containing non-string objects, but the documentation doesn’t seem to draw a distinction between lists and dictionaries wrt what can be placed in them.
(Both lists and dictionaries are initially described as only containing strings, and later this description is expanded to include nesting; this counterintuitive arrangement may explain the confusion.)
Frankly, in most languages, this is better because you don't have the types of objects randomly change based on user input. (In a few languages, with a few libraries, you can specify the type of the document to the parser and have it fail to parse the entire document if it can't deserialize to the right type, in which case this is a little weaker. But you can still do that with NestedText, just one step after the parser - have your own function that takes a ComplexStructure<..., String, ...> and returns a ComplexStructure<..., int, ...> or throws an error.)
Only in untyped or dynamically-typed languages. But even in JavaScript one may write +obj.version instead of obj.version to make it numeric. Evading this is a straight way to hell.
In a statically typed one, conversion is typically generated from description and type checking applies just at reading.
The problem with vaguely specified format is in more simplex cases. Shall we accept 45x as number (and what it value will be, 45 or 0)? 045? 045x? What date is 1/2/3, 1-2-3? And so on.
I think the point is that the answers to those questions may well be application-specific. In which case it is better to not bake them into the file format.
- There's no way to say that you want an integer; you get a floating-point value.
- See https://news.ycombinator.com/item?id=24676484 , you can't reliably accept integers over 2^53 without taking them as strings.
- Someone can always specify something of the actual wrong type. (Imagine changing YAML "version: 1.9.1" to "version: 1.10". You can't just stringify 1.10, you'll get "1.1"!)
So, in a practical data format, the schema for your document needs to say something like "This is a number, which must be an integer between 0 and 2^16" or "This is a string, make sure to quote it" or whatever, and a generic statically-typed JSON- or YAML-parsing library isn't going to handle that for you. And telling your users "the input format is JSON" doesn't answer that question: you must make it explicit to users.
Fortunately, you can handle it just fine in a statically-typed language in one of two ways. One is to accept an object from your parser that consists of variant types and pass it through your own function that validates it against a schema, and then either returns a more-restrictively-typed object or throws an error. Such a function could easily do string conversion too if given NestedText input, as I mentioned. The other is to pass some information into your parser saying, don't act like a generic JSON/YAML parser, instead interpret these particular fields in this particular way and accept only things with this structure. If you're doing that, you can easily tell the parser to use this particular string-to-integer function on the strings in NestedText and then return an appropriately-typed object containing an integer to you.
One should remember that any sane application will be parsing the config file into internal data structures and validating it anyway so it gets little benefit from the numbers being already “parsed”.
There are also issues when something looks numeric but doesn’t parse (eg 1.2.3, 3/2, 12in, 4h30m2s, 2:30, 2020-02-29, etc). One way to deal with these is a tokenisation rule like in Common Lisp: if it is a valid number syntax then treat it as a number, otherwise it’s a symbol, but this can lead to issues (eg you would need to know that when your number needs more than float precision or otherwise doesn’t follow the rules, it should be in quotes. It seems crazy to pass that detail on to the poor sod who has to write the config file).
The benefit of standard numeric and boolean types is that different tools can exchange data in a well-understood way.
Getting rid of yaml's 30 ways to write "true" and "false" by making everything is a string just means that you now have 30 tool-specific ways to write "true" and "false".
The "everything is a string" approach already exists in shell scripts and TCL and it's not really that great.
Unless you have actual type annotations/tags (eg xml, jsonld, graphql), everything IS a string. There's no assumptions otherwise.
http://p3rl.org/guts#Magic-Variables
perl -MDevel::Peek=Dump -mTie::Scalar -e'
//g; Dump $_; tie $c => "Tie::StdScalar"; Dump $c; Dump \%ENV
' 2>&1 | grep MAGICPerl tied-variable magic just means there are (effectively) getter and setter properties attached to the variable. "Magic" is just the name that was chosen in the implementation, and it stuck.
It's used to implement variables with special, automatic meanings, like $$ for "current pid" and $! for "last error".
It's also used to implement variables with user-defined behaviours on access, which is quite handy for a lot of abstractions.
A lot of modern languages support both of these things, because they are useful, but it's not called magic in those languages, it's called something like "watchers", "proxies", "getters and setters" or "hooks".
No, the criticism of YAML-style "magic" is that it leads to entirely surprising behaviour from innocuous input. Perl magic is not that kind. If you're using a special variable, you already know why.
In fact in Perl, you can opt for longer, readable, lexicon over the terse single character variables; and that’s literally how modern Perl should be written.
Whereas the problems described with YAML is where it can automatically alter your data based on what the parser “thinks” the data should represent. Which is generally what people mean when they talk about “magic” in IT: systems that don’t honour your input and instead automatically convert it into something else. Perl doesn’t do this even in spite of it looking like executable line noise to many.
If you take a look at how string handling works in most programming languages there is a lot of "magic" going on there. Which isn't a bad thing necessarily, because most programmers don't want to deal with the intricates of strings unless they really have to. The key is that this magic doesn't get in your way and doesn't do too magical things nobody ever asked of it.
https://yaml.org/spec/1.2/spec.html#id2805071
Of course, even though YAML 1.2 is a decade old, there are still many parsers that accept YAML 1.1.
YAML has actual type annotations (tags).
I've never seen typed yaml, this is wild.
negative: !!int -12
zero: !!int 0
positive: !!int 34
Can't say I love the notation, but indeed that is type annotations. I guess neither "yaml type hints" nor "yaml type annotations" are the right query. Had to search explicitly for "yaml tags".https://docs.aws.amazon.com/AWSCloudFormation/latest/UserGui...
In Python:
>>> import json
>>> json.loads('{"x": 9007199254740993}')
{'x': 9007199254740993}
In my browser's JavaScript console: > JSON.parse('{"x": 9007199254740993}')
{x: 9007199254740992}
(Consider what happens if you try to send a tweet's ID, a perfectly normal number like 205052027259195393, through JSON. Or if you try to serialize a stack trace on a 64-bit system, where addresses are also perfectly normal numbers.)Javascript number is a floating point
Number.MAX_SAFE_INTEGER
// 9007199254740991
Number.MAX_SAFE_INTEGER + 1
// 9007199254740992
Number.MAX_SAFE_INTEGER + 2
// 9007199254740992
no automatic promotion to BigInt 9007199254740991n + 2n
// 9007199254740993nThe JSON standard doesn't place restrictions on size or precision of numbers, instead just noting that implementations can vary their treatment of and limits on numbers. While JS uses doubles for all numbers, many other languages emit an integer type for a JSON integer. So, once you go beyond the range where a double can accurately represent all integers, you run the risk of a mismatch in how the number is interpreted by different languages parsing the same JSON.
Of course the spec also allows you to create way too big or too precise numbers that would be problematic in most languages as well; it's just that this is a somewhat common bugbear.
I wouldn't necessarily call it a flaw in JSON though, more an issue with JSON.parse or really just a fact of life when dealing with numbers in JS. Alternatives to the built in JSON.parse exist to read large integers as strings or bigints.
https://developer.mozilla.org/en-US/docs/Web/JavaScript/Refe...
Described issue is not problem of a JSON but engine which parsed it and language which stands behind the parser. Any config format will eventually have same result and same issue.
So forcing programmer to parse every single piece of data for sake of "it's his responsibility" is not a case here.
I also disagree this is in any way programmer responsibility to create standardized way of creating parser for everything. This format gives you nothing but indentation so you are forced to create documentation for everything field, what type it's and what kind of values it takes. Lots of extra work for nothing when you have any other format.
I mean, there's a pretty clear argument here: Twitter themselves used to return these numbers as numbers in the API until they realized they were about to hit this problem. https://developer.twitter.com/en/docs/twitter-ids
That's an incredibly weird complaint when the real problem is that javascript's JSON.parse doesn't use BigInt for large numbers.
I don't think that's what this tool is for. This tool is for humans to read and edit data. That's a different use case from automated programs exchanging data, for which I agree you should be using standardized numeric and boolean data types and not making everything a string. But how that standardized data gets determined from data that humans enter should be up to the individual application.
> A key that requires quoting must not contain both single and double quote characters.
You can't really serialize user data with that restriction.
Maybe we don't need it, but it often helps.
Also, this data format isn't necessarily for humans to exchange data with other humans, but for humans to give data to applications in a format that's much easier for humans to use.
> if this is insufficient to communicate it to the machine
Not at all. Each specific application can easily parse this data format according to its own needs. What this data format doesn't specify is a single translation into application data that is the same for all applications. But applications don't need or want that, because they have different use cases.
If you make everything a string then the interpretation of "no" as boolean true or false is left to each tool, and there are even tools which have different interpretations of "yes/no" for each field.
Most likely if their input is meant to be machine interpreted they would need to be trained to provide specific inputs anyway. I like that NestedText doesn't hide that problem. It lets the user organization decide how it wants to manage that problem, and what symbols or words are understood by the people authoring the files.
ISO8601 joins the chat.
The nice thing about that is it solves the problem rather than hoping it doesn't matter or assuming each program's validator will think to note all of the possible data types not compatible with the program natively. If a u32 is defined in the file and you've only got doubles to work with it's a given you'll have to deal with it in your tool specific validation. For everyone else it's well defined.
The downside is it's a bit more verbose and if you have all of that info already it's pretty easy to jump to just using a binary format which will be more efficient anyways.
The spec requires that a full 64 bits of signed integer be parsed and understood, that hexadecimal, octal, and binary values be integers, and that floating point values be parsed as doubles.
It doesn't support hexadecimal float, however, which is a pity: having a guaranteed bit-identical format is a nice affordance.
Too many types could get overwhelming, but I like where amazon's Ion is [1]. It actually supports multiple number types, with decimal being the default for values with a dot.
> you would need to know that when your number needs more than float precision or otherwise doesn’t follow the rules, it should be in quotes
Not really. The configuration value should either be a number or not, which is determined by the application reading the config. As a config writer you only care to make the type match (so, in json, if the application uses number you make sure you use number, and if it expects a string you use that)
(disclaimer: I work for amazon, but have nothing to do with Ion other than having used it. Opinion is my own, not my employer's, yadda yadda)
Only if this format is intended for use-cases that never need to deal with numbers.
> One should remember that any sane application will be parsing the config file into internal data structures and validating it anyway so it gets little benefit from the numbers being already “parsed”.
That statement couldn't possibly be more wrong.
Number parsing (and encoding!) is a decidedly non-trivial problem. You need to concern yourself with -- at a minimum -- all of the following:
- Unsigned 64-bit numbers.
- A series of digits that would be bigger than a 64 bit whole number. Convert to float? Truncate in some way? Error?
- NaN
- Infinity
- Negative zero
- Denormal numbers.
- Differentiating between decimal/currency types and floating point numbers. Not all decimal values can be exactly represented as floats!
- Efficiently encoding floating point to use the minimum digits without losing precision.
- Parsing those minimal numbers with perfect "round-tripping".
- Doing the above efficiently.
- Securely too! Efficient parsers cut corners on sanity checks. I hoped you fuzzed your parser...
The above can easily amount to many kilobytes of extremely complex code. Look up "ryu" as an example of what Google came up with to make JSON number parsing reasonably efficient.
Meanwhile, reading a fixed-length number from a binary format can be done in a single machine instruction. One. It might not even take an entire CPU clock cycle! Okay, two, if you need to bounds-check your buffer, but there's ways to avoid that.
Afterwards, the bounds check is again literally just two machine instructions in complexity. That's not the difficult bit!
The difficult bit is the parsing.
1. Declare numbers as numbers in the configuration language. E.g. "decimal(1e1000)".
2. Parse declared numbers with a lossless format like Python's decimal.Decimal.
3. Let users decide at their own risk if they want to convert to a lossy format like float.
That doesn't matter at all. The author's aims will be ignored if this format is used for anything even vaguely important. Eventually it'll need tooling to both read and write it.
DevOps pipelines, applications with GUIs, or something will need to both parse and generate this format in a consistent way.
There is no such thing as a human-write-only format in widespread use.
Even programming languages are regularly generated by tools such as RPC API codegen tools, LINQ-to-SQL and the like.
One example you provide is decimals for currency values but I claim you would want such values to look like $1234 in config files so that when they are reviewed or written, the person reading the file knows they are looking at a dollar value and can be concerned if it is too large.
I’m not suggesting that applications write their own number parsing. Just do uint64::parse or parseInt or Double.of_string, or whatever else you need to access your language’s number parsing routines.
> Just do uint64::parse or parseInt or Double.of_string, or whatever else you need to access your language’s number parsing routines.
Okay, so the computer is doing the parsing.
Those functions are notoriously inconsistent in their behaviour, particularly across different programming languages. If you're not careful, you'll end up accidentally using the internationalised versions of those functions. Even if you're careful, other people won't be.
Remember, data formats are for interchange. They have to be language agnostic. They have to be well-defined, and it should be possible to write a parser for them without having to guess at the precise details.
The harmful consequences of the Robustness Principle are now well-recognised in computer science: https://tools.ietf.org/id/draft-thomson-postel-was-wrong-03....
Some things need to be done properly, nor not at all.
I am far more worried about localisation issues than language issues. If you are storing something central to multiple applications I'd argue a text file is the wrong tool
And - it is certainly OK in many instances to have fixed-width, fixed byte-order binary encoding as the format's basis. It comes with the twin downsides of wholly different categories of errors cropping up, and with the lack of a universally agreed upon tool for human entry.
Perhaps text was a fashion, though. I definitely have had thoughts in that vein lately. And in that case we shouldn't always be rushing to use it as the source of truth when we have many good, machine-level agreements about numeric formats.
It seems you're mixing up the language definition with implementations that try to follow the language definition.
If different implementations have different results then either they are buggy or the language has some important holes in the specification.
Either way,the solution to this problem is not less validation.
> One should remember that any sane application will be parsing the config file into internal data structures and validating it anyway so it gets little benefit from the numbers being already “parsed”.
That's the whole point, isn't it?
I mean, if you already acknowledge the fact that this parsing and validation is a basic requirement, why handle it as an afterthought and force developers to add their own hand-rollef absurd and unnecessary type checks and type coversions?
Wouldn't it simply easier to let the language and the parser do that already?
I mean, no one ever complained that JSON had string types. In fact, one of JSON's main complaints is that it doesn't support enough types, such as timestamps.
name(str): Dave
age(int4): 22
dob(date): 2020-02-01
photo(base64):TWFuIGl....But YAML is really quite complicated, and JSON (which shouldn't be used for config files at all) and TOML (which I love and wish it would gain more popularity) aren't exactly alternatives to YAML. So, I would be actually totally ok with "YAML, but better", as a way to deprecate YAML.
Now, it is clear from the start that this cannot deprecate YAML, because it doesn't even have booleans and numbers. But, surprisingly, I can accept this as well: ok, let's just assume that being good at dealing with strings may be enough.
The problem is, it isn't clear at all from the docs, if this is better than YAML at anything. It raises dozens of questions. I'll start with the most basic ones (using [] as a wrapper/delimiter): how do I represent values [ a], [a ], ["a"] and [""] in this file format?
Why should JSON never be used for configuration? It is sufficient for declaratively expressing anything I have encountered. Do we really need references or other stuff from YAML? For configuration this seems unnessecary, provided that the program, which interprets the result of parsing the JSON is well written.
I don't understand the qustion. They are supposed to be exactly that, [<whitespace>a] and [a<whitespace>]. I assure you, I've encountered many situations, where whitespace at the beginning or the end of the value is actually meaningful for reasons you (creator of the app) have no control over.
> Why should JSON never be used for configuration
Many reasons, actually, but the most important (IMO) being that original JSON specification doesn't support comments, nor most actual parser implementations do. Configuration file that doesn't support comments is trash and causes very real inconvenience for users. Using additional key values for comments (even if such atrocity doesn't bother you conceptually) isn't a solution in many cases (for example, when your intention is to comment out a list item).
Just for the record, JSON initially had comments and they were later removed according to Douglas Crockford.
[<whitespace>a,<whitespace>b]
How will this be interpreted/parsed? Will there be a whitespace before "b" after parsing? That would mean, that I am not able to separate visually more clearly, by adding a whitespace between list elements, which is widely considered to be a good practice in programming languages. The reason is readability.
The next thing is, that on the same line whitespace is added, but what about multiple lines defining a list? Here we do not add the line breaks and indentation to the string. It's not consistent in this way.
So I personally would never write string like that. I would always make use of quotes in such situations and probably in YAML in general, simply to make it clear, that I do wish to have the leading whitespace in the string, and it is not simply a typo, resulting from removing a former first element from the list.
JSON is limitting, but for configuration I think it's kind of limitations are often good. With comments in JSON I am still not sure, because sometimes I'd like to write them there, but would not like to include another dependency, only to be able to parse away the comments from JSON. Then I better write good docs elsewhere.
You can use HJSON which is the json with comments. It's fully compatible with json so easy to introduce into anything that does json. https://hjson.github.io/
Triple-quoted multiline strings like HJSON would be great, too.
this:: right
this: left
Will my program be able to tell its right from its left?That was my intention behind this, too:
https://github.com/crdoconnor/strictyaml/
The general structure of YAML is fine I think but its feature set grew a little bit out of control.
The "cleanliness" of the format leads to one of its inherent weaknesses - syntax can't be used to encode type information so you either need a schema to encode type information (strictyaml approach) or have magic conversions (yaml approach) or to assume strings (strictyaml w/o schema/nestedtext).
The interesting thing I discovered about schemas building this is that it kind of pays to make them extensible and build them in a turing complete language. Schema validation done using a non-turing complete language (e.g. jsonschema) allows for cross language usage but it ends up being a kind of blunt object.
So far as I can tell, this doesn’t use [ or ] in its syntax at all, so all the values you give would be represented exactly as strings without any problem.
Yes, that's why the GP chose those characters to delimit their example strings. I'll try again using ` characters to delimit the example strings.
> so all the values you give would be represented exactly as strings without any problem.
But what would those string representations hold? If I parse the below strings in a javascript context, what would I get?
` a`, `a `, `"a"` and `""`
Do I get this? " a", "a ", "\"a\"" and "\"\""
Or do I get this? "a", "a", "a" and ""
Or do I get something inbetween?> I will mention something else. The section about the "Norway problem" is not quite accurate. Some YAML loaders do in fact load no as false. These are usually YAML 1.1 loaders. YAML 1.2's default schema is the same as JSON's (only true, false, 'null and numbers are non-strings).
> Any YAML loader is free to use any schema it wants. That is, no loader is required to to load no as false. Good loaders should support multiple schemas and custom schemas. The Norway problem isn't technically a YAML problem but a schema problem.
> imho, YAML's biggest failing to date is not making things like this clear enough to the community.
> Note: PyYAML has a BaseLoader schema that loads all scalar values as strings.
I want to know beforehand what I can put in a config file and I want a fast and hard failure if what I put in there is not good.
And this should be implemented at the file format parser level, with hooks for apps to add on top of the default behavior, so that every app that implements this format gets these things almost for free.
It would be nice if we had such tools.
message Config {
repeated Server server = 1;
}
message Server {
string address = 1;
int32 port = 2;
bool standby = 3;
}
And then you use the text representation in a config file: # main instance
server { address: "127.0.0.1" port: 4567 }
# backup instance
server { address: "127.0.0.1" port: 9876 standby: true }
And load it into a message instance: Config config;
google::protobuf::TextFormat::ParseFromString(input, &config);It looks oddly like HCL. I wonder...
I know cap’n’proto also has fantastic support for using the schema for config files. You can just compile any constant as a stand-alone serialized message that you mmap into your code in a safe way. It can’t do complex math and things (at least yet) but you can express lists, dictionaries, and reference other constants, so as a config file replacement I love it. I’ve also found the format to be far more regular and consistent than you get with things like text protobuf (you’re still using the schema language instead of another format)
Are you suggesting using a binary format for your config files? I think most people would find that more trouble than a decent text format.
> ... than you get with things like text protobuf
You can just use protobuf's canonical JSON representation (thought the lack of ability to use comments is annoying).
Cap’n’proto also has plain text and JSON serialization formats if you really want to have your deployed config file be directly human-editable and deserialize from that. I was just noting a very cool feature of having your config written in cap’n’proto and it’s what Cloudflare uses to maintain a bunch of config internally if I read Kenton’s allusions to it correctly.
You can then compile it into whatever format (JSON, plain text, binary) that you want for actually reading it from disk.
JSON was a huge step backwards in the name of simplicity. And now when we are going to add similar functionality to JSON, something else is going to come out in the name of simplicity (like NestedText).
There may be room for an argument that Magento did XML badly (it did many things badly), but I don't believe I've ever seen XML done well.
In a minute I can read and write json from most languages I use.
In the same amount of time, I'm still wondering if I should use a tag or attribute in xml. cdata? expat?
It's not that xml isn't a good technology. It's that it's not appropriate for general use, especially in comparison to simpler alternatives.
The fundamental issue with XML is its impedance mismatch with common data structures which forces using Object to XML mappers (whether explicitly or implicitly). It's more or less solved with XML Schemas or DTDs, but if you're looking at just XML, you can't tell whether some element is an array or a single node. Thus JSON is better suited for serialization.
That is really not what attributes are for. I feel a bit of a fraud posting that because I'm not an XML expert and so not really clear what they actually are for. (This reenforces the parent's point: you need to be an expert to know what such a fundamental feature is for.) I remember it's something like "something used to help interpret the actual value" e.g. units of measurement. But most of the time, even if it's non-repeating with no children, you're supposed to use elements rather than attributes.
One problem here is that attributes are so much more compact (and so often easier to read) than elements that it's tempting to use them in places where you ought to use an element (and many people over time have given in to that temptation). Another problem is that the distinction between attributes and elements is almost never useful. That was the parent comment's point by the looks of things.
> The fundamental issue with XML is its impedance mismatch with common data structures
That's probably part of it, but I think at least as problematic is that it has many features that most of the time you don't need and don't want to have to care about. Things like CDATA (also mentioned by the parent comment), custom entities, external entities, DTDs (which can be inline in XML files so you need to know all about DTDs to understand XML properly). That's why there are all sorts of weird XML vulnerabilities that JSON doesn't have. Did you know you can make an XML file that reads your /etc/passwd file when it's parsed? That is not an issue with JSON.
Confusion arise once a human observer is lost.
I had been thinking that all of these extra features that XML have are just a case of massive overengineering that no one would ever need. In fact it's a case of taking something fundamentally meant for text documents with extra markup, as the name implies, and misapplying it to config files and IPC messages which are just not the original domain at all.
I think we should draw on XML strength points. People read articles in browser, not plain text. "Add to cart" is just a POST request with id
curl -d id=foo
yet we have forms and interactivity. Like in literate programming text and data live together, interactive application like a Smalltalk image.In XML we can separate data from presentation.
<?xml-stylesheet type="text/css" href="foo.css"?>
<?xml-stylesheet type="text/xsl" href="bar.xsl"?>
<root>...
Machine receives data, human receives application with documentation, builder. That's exactly what we have today except UI can be plugged to any stored document. To good to be true.I think XML was killed by poor usability. Plain text XML, XHTML and XSLT authoring is not fun.
I am trying to uncover it from DOM perspective [1], so far I like it more than Markdown. XHTML and HTML is just a serialization format. HTML is not a good one [2], [3], [4]. XSLT may have nice GUI or compact syntax like RELAX NG.
[1] http://sergeykish.com/live-pages
[2] http://sergeykish.com/script-style-is-cdata-in-html
[3] http://sergeykish.com/pre-newline-ignored-in-html-test
[4] http://sergeykish.com/content-after-html-appended-to-body-in...
No you don't... the parent commenter explained to you what it's for in a simple and concise manner... you chose to not accept that even though you're not an expert in this, and then complains you need to be an expert to do it?!?
I don’t doubt they meant to be clear, but reading it they were not and raised more questions than were answered.
As an example:
Wouldnt attributes be better served as details about the current element?
Wouldn’t elements be better served as “I am a child of the parent”?
Why would I use an attribute as a “non-repeating child” when semantically that doesn’t make sense when looking at the document? The attribute is inside the element’s definition, and seems to me attributes should be used to further describe the element being presented itself, and not be structural or describe itself as a child in any way.
(The true difference is explained in sibling comments to yours, by sergeykish and tannhaeuser, if you're interested.)
Not only can SGML (but not XML on its own) read /etc/passwd, it can format it into fully-tagged markup and then render it into eg an HTML table. Demonstrating what SGML/XML is actually designed for: encoding and authoring semistructured text. This can't be overstated in discussions like these where use cases for config formats, service payload formats, and actual text authoring are all thrown into the same basket when they shouldn't.
Btw: you can parse and canonicalize this new config file format into markup using the same SGML mechanism you'd be using for CSVs like /etc/passwd, namely short references
Btw2: you can skip/ignore markup declarations in XML, including whole declaration sets (DTDs) since these can be recognized using plain greedy regexpes, though you can't ignore entity declarations when actually used in your XML body text
Both of these solutions and all other known solutions to this problem are, as I'm sure you can see, just awful.
You can't just paste XML in XML because of the <?xml?> thing, because of entities, and because of half a dozen other misfeatures of XML.
You put XML fragments inside a parent XML document using namespaces.
This is very well supported, and used extensively.
Trying to "escape" XML to nest it in a parent XML document is Wrong with a capital W.
could you post or link to an example? i'm not very familiar with advanced XML features
or for a simple example: what would it look like to put `child` into `parent` using namespaces?
# parent-doc.xml
<parent>
<!-- embed here -->
</parent>
# child-doc.xml
<child x=3 y=5/> <!-- The special XMLNS attribute binds a short alias to a long name -->
<p:parent xmlns:p="urn:some:unique:string">
<c:child xmlns:c="urn:some:other:child:name" x=3 y=5>
<c:subchild> <!-- No need to repeat the fully qualified unique name -->
<p:tada>You can even interleave!</p:tada>
</c:subchild>
</c:child>
</p:parent>
Note that while this is possible to write by hand, typically namespaces are for documents generated and processed by tools. The XML Schema Definition (XSD) format has full support for namespaces, so you can define documents based on modular chunks. E.g.: you can "import" the SVG namespace into a diagramming XML document format namespace, but restrict its usage to only the child nodes of an "img" tag. Or MathML as the children of "graph" nodes. Both SVG and MathML can potentially import a shared "font" namespace. Or whatever.In the XML Reader API, each element has a "fully qualified" name that includes the long namespace prefix. If you use the API correctly, your tool can handle nested documents, or gracefully ignore them if it's appropriate.
The fiddly part is making this efficient, i.e.: avoiding a full string comparison against a long URI or URN. You typically have to "register" the namespaces you're interested in, and the API gives you some sort of efficient token instead of a string to use from then on.
I'm not saying it's perfect. Nothing is in XML. It was designed by committee, it brought too much of the legacy SGML baggage with it, but its namespace capabilities are a lot better than nothing at all, in much the same way that C# or Java don't have perfect type systems, but they're superior to loosely typed languages.
function escapeXml(unsafe) {
return unsafe.replace(/[<>&'"]/g, function (c) {
switch (c) {
case '<': return '<';
case '>': return '>';
case '&': return '&';
case '\'': return ''';
case '"': return '"';
}
});
}
Or you convert to the same encoding, strip XML declaration, expand entities. In short work with adequate tools.And yes, I'm being sarcastic.
Well, except for handling complex content documents like in all ebooks and, in sgml form, all webpages like this one.
It would be true if XML was not full of all this SGML debris like "entities" (really, uncontroller macros), if real schema formats was flexible enough (I needed <c> inside <a> and <c> inside <b> when they totally different), etc.
But when a config reader tool has to deal with 40+-year legacy of enterprise guys wanting to embrace the universe, but all this doesn't allow to control contents without external measures like regexp checking... that simply shuts up facing real world.
It's much simpler to use than XML Schemas, and arguably results in cleaner data models, since it doesn't have anything analogous to XML namespaces that allow for arbitrary mixing of schemas.
I can see this work perfectly fine in typed languages like C#: `NestedText.Deserialize<T>("nestedtext")` where the deserialize method handles the actual mapping of nested text objects to `T` by providing the deserializer a class / classes that handles the string -> scalar(s) mapping for the given T. That would, sort of, function as a Schema.
I think the only thing, from glancing over the project, that would need to be supported to make this really useful is nested lists/dictionaries. I don't see how this can be done but maybe I'm missing it.
And the problem with stringly typed systems is that everything is underspecified
I should have mentioned that I want something simple and readable.
I wrote a json/kotlin-serialisation library once and purposely restricted some json-features to achieve that:
1. Fields can arrive in any order - this is standard
2. Field names are matched case-insensitively - so keyA and keya are the same, because who would use two variables differing only by case. Serialization keeps the original casing of the name.
3. Missing fields throw an error. if they are nullable, they have to be explicitly set to null - so that you can be sure the serialization side upgraded to the latest version of a protocol if a field was added, and things don't just work by chance.
4. Nullable strings are not coerced to empty strings or anything like it. Kotlin is null-safe, so if it's a string, it has to be "". If it's, for whatever reason, a nullable string, you can set it to null.
5. Enums are also serialized case-insensitively - so you an write "keyA": "eNumVaLuE" if you want - typos should not break the code here, no on would you two enums differing only by case. IIRC booleans could also be TRUE, tRuE, truE etc. (but NOT t or f, or yes or no, or 0 or 1 or empty).
6. Superfluous properties are silently ignored.
These rules were a great tradeoff for quick development, mixing languages and having fail-fast behavior with a stable protocol.
(https://medium.com/@fabianzeindl/generated-json-serialisatio...)
The problem with human friendly formats is that the thing that typically makes them human friendly is removing things that make reading and editing difficult, but make disambiguation possible. If the format ever needs to be read by a machine, something has to do that disambiguation.
If it’s not provided by the format, you’ve turned every usage into a potential source of bugs that would otherwise be restricted to interchange/stack implementation incompatibilities. In other words, now your format can have a different set of expectations even on the same system.
The natural response to that problem will be to bolt on validation, types, and documentation that is provided arbitrarily (and with varying quality).
IMO, efforts in human friendly formats should focus less on stripping out funny characters, and more on which minimal set of funny characters provide:
- Good readability
- Good editability
- Clarity of structure
- Clarity of data types
- Reasonable tolerance and flexibility for variance in arbitrary formatting/style preference (particularly in delimiting long form/multiline text and annotations), because no one can agree what good readability or editability means
- A flexible type system that allows machines and humans to know what a given datum is without variation or surprises
- Maybe humans should just use a GUI?
A great comparison for this is CircleCI's config syntax and that used by GitHub Actions. The Circle format is extremely error prone; about half the time when I'm modifying a Circle config, I'll end up pushing a broken config, even though the YAML syntax itself is valid. With the GitHub Actions format, I almost never screw it up. I don't think it's a coincidence that if you convert a Circle configuration to JSON, it looks twisted and bizarre, whereas if you do the same with a GHA config, it looks perfectly ordinary and sensible.
If you think of YAML as "a prettier version of JSON", and design as if your users will work primarily with JSON, you can do fine with it. If you think of it as a medium for building your own configuration language, you'll make something awful. The problem is that any human friendly format is going to inherently encourage the latter.
Also one more goal is twin binary and text formats that are 1:1 compatible, so that you can write it in text and transmit in binary.
I'm still finishing up the reference implementation, and then will start on the schema.
What can an ICT professional claim at the end of its career? “Hey, I’ve argued about shit all my life!”
And it also diminishes your value as a team member. If you can't convince others, means the reasons you present are weak and nobody would be interested in listening to you and therefore there's not much reason in having you around.
Jeez, you extrapolated a personal observation all the way to a character assassination and firing letter paragraph.
Standups and 1on1 with you must be a blessing... a joy
My reasoning is the following:
People would only listen to you if you can prove what you say is right, because nobody is interested in hearing wrong things or unexplained things, they just aren't helpful.
"We should use nodejs!" "Why?" "I don't want to argue, we just should." Is that helpful?
If you don't have the reasoning skills to convince others, you can't present constructive ideas and back them up with an explanation.
If you can't do that, literally, what is your value to the team? Blindly and quietly execute the will of other team members? That would take too much energy from those people, to direct you on every step of the way.
When you hire engineers, you expect them to give more than they take, otherwise they're a drain on the team resources.
Collective problem solving is impossible without arguing. Arguing is trying to improve something, identify mistakes, logical contradictions, basically you're doing the work of a compiler that checks your program for correctness. Would you want a compiler that always agreed with you, whatever you fed into it? Don't think so. Same thing with engineers working together to reach a common goal. You're checking and improving each other's ideas.
Edit: moreover, if you lack reasoning skills to convince others, means you lack reasoning skills themselves. How are you going to solve problems in the first place?
Please.
Thankfully we’ll never work together so can we just continue our existences as we did before, blissfully unaware of each other?
> arguing about slight variations of the same mundanities: XML, XSD, IDL, ASN.1, Avro, JSON... Emacs vs. Vim, Weakly vs. Strongly typed, and so on...
There could be a lot of valid and important arguments around these things. Except maybe vim and emacs, who gives a shit about that.
It'd be easy to employ ... as a document separator / end indicator that could be checked for.
I'm really struggling with this assertion; IMHO one of the problems with JSON is the lack of more sophisticated scalar types.
That being said, this appears extremely readable, so my concerns could definitely be alleviated by a decent schema.
In Rust you do something like this:
#[derive(Clone,serde::Deserialize)]
struct {
an_int: u64,
a_float: float,
ordered_map: linked_hash_map::LinkedHashMap<String, chrono::DateTime<chrono::Utc>>,
unordered_map: std::collections::HashMap<i32, String>,
}
So there is no need to worry about if maps are ordered or if a value is an integer, real or string in the format itself. Ironically for Python (which the reference implementation is in) it does seem much more annoying to have to manually call `int()` on each element.I'm just a little sad that tabs are disallowed. I really think the best rule for indentation-sensitive languages is that each line must either have the same indentation in which case it is the same level, same indentation plus any amount in which case it is the next level, or the exact indentation of any previous level in which case it is a dedent. These "solutions" which just forbid tabs are half-assed and ones that try to convert tabs to a set amount of spaces just lead to confusion.
Additionally it would be nice if there was an example of a dict inside a list. I think it would work like the following but can't confirm from reading the site.
-
key: value
-
key: value
other-key: other-value [{'key': 'value'}, {'key': 'value', 'other-key': 'other-value'}]The trouble is, open source cannot do good GUIs. If a problem is best expressed with a GUI, open source consistently blows it. See Gimp, Blender, Inkscape, FreeCAD, all of which are notably worse than their commercial competitors.
What about browsers? Firefox and Chromium are open source (and fit under any reasonable definition of "GUI")
Improve the format of org-mode. Make it primarily easy for humans to interact with but also easy for cheap scripts to parse and manipulate. Create a super-fast CLI for it which ships with the ability to read keybindings from a file. Ship with emacs keybindings as a default but also a file with the spacemacs keybindings. Add the ability to run the CLI as a daemon that can be started from neovim.
Keep the format open so someone else can write some npm package for including an editor in VSCode or a webapp.
After dealing with 500+ line kubernetes configurations, this is a bad idea.
That assumption is... not applicable to most scenarios I come across and will likely lead to issues being pushed downstream and introduction of subtle bugs and defects.
> I once disabled our product for the entire country of Norway for a day because `NO` in YAML evaluates to `false`
There should be at least some support for some standardized representations (basically JSON + ISO8601 datetime + some encoding for embedding whatever stringified serialization, eg. just how HTTP uses chunks and unique boundary tokens).
In my opinion, JSON is best used as a wire-protocol. It is awkward as a configuration language.
YAML works for short configs, but becomes unmaintainable for longer configs. I think the primary problem is that the indentation is significant. I also think the language spec is far too complex.
INI format works for short configs, but also becomes unmaintainable for longer configs. Ironically I think this is because INI is too primitive, the opposite problem of YAML, but has the same effect.
I am not familiar with TOML or DHALL, mostly because I stopped looking after I implemented the JSONNET system and liked it so much.
Addendum: I have used text-formatted protobufs in limited situations with good results. But I don't think that protobufs is a good general purpose configuration language.
Addendum2: The amazing thing about simplicity of the INI file format is that I was able to write a "single line" sed program to parse it in a bash script. The following finds the value of the $key in the $section in the $config_file INI file (definitely works on GNU sed, I think it works on MacOS sed too, not 100% sure though):
sed -n -E -e \
":label_s;
/^\[$section\]/ {
n;
:label_k;
/^ *$key *=/ {
s/[^=]*= *//; p; q;
};
/^\[.*\]/ b label_s;
n;
b label_k;
}" \
"$config_file"An editor problem, perhaps? We don't maintain office documents using vim; why edit structured configuration files using a plain text editor, if doing so is arduous?
My text editor[0] abstracts the underlying hierarchical data format behind a tree-based widget[1]. Whether YAML, JSONNET, NestedText, CSON, XML, or TOML backs the widget becomes an implementation detail.
[0]: https://github.com/DaveJarvis/keenwrite
[1]: https://dave.autonoma.ca/blog/2019/07/06/typesetting-markdow...
a) Do not have values of different types occupy the same fields
b) Have a schema defined (explicitly or implicitly as part of the parsing), especially because you're likely working with a type system
This isn't good if you want `x = parseJSON(blob)` kind of API, but that's definitely not what you want for any kind of human-editable config.
It seems simpler than TOML, I'd give it a try.
Protobufs is similarly a serialization protocol combined with an RPC layer for server and client stubs. Its canonical serialization format is binary.
Neither are particularly well suited for human written configuration files. A "JSON without the quotes around keys and allowing Javascript comments and commas at the end of objects or lists and a mechanism to escape multi-line strings" would probably cover most of the required cases.
YAML is an attempt at that, but also attempts to solve a bunch of other problems in a complicated and fault-inducing way (eg, relying on indentation for hierarchy).
One thought that occurred to me several times in the last year or so is that roughly the level of abstraction offered by NestedText might make sense as part of a hierarchy of abstractions that could be built on.
We already have that with text files. Because end-of-line character combinations are special, text files in a given encoding are already more structured than streams of characters.
So, assuming UTF-8 character encoding:
CharacterStream
Text (EOL character combinations)
NestedText
MyNestedTextFormat (domain specific semantics)
With non-text files, this has already happened more than once. For example, both Zip files and Sqlite files are used as base formats for specifying other formats.They are even UTF-8 transparent and can be easily converted to/from CSV, TSV, PSV, and, with a definition of the equivalent of a "close brace" could allow for a multi-level hierarchy.
[1] https://en.wikipedia.org/wiki/Control_character#Data_structu...
Another variant of the idea that I was mulling over was to base a roughly NestedText level of abstraction on UTF-8 transparent characters, and then combine that with what Animats was talking about as a standardized GUI for trees, dictionaries etc.
A recent trend is that programming languages have more than one bijectively equivalent syntax. For example, ReasonML [0] and OCaml are two bijectively text based syntaxes. That idea could be extended to a NestedText like syntax being bijectively equivalent to a text based syntax. Editors like Visual Studio Code infer that sort of information continually on the fly, but it sort of gets lost in the toolchain. Compilers could operate at a higher-level of abstraction than lexing/scanning. Git merge might also work better if could operate at a NestedText like level of abstraction.
[0] https://en.wikipedia.org/wiki/Reason_(syntax_extension_for_O...
We were recently storing tokens in a database, and I chose to use SOH for the metadata and SOX for the text.
One byte width, no collisions with printable text, and that's what they're there for.
I'd love to see a CSV replacement that used SOH ... SOX for headers, RS as "commas", and GS as "newlines". You'd be able to cleanly concatenate multiple files, since the first line is no longer special, and you'd be able to have commas, newlines, and in fact any printable text whatsoever inside the data.
And the semantics are perfectly clean. Again, that's what they're there for. Some small challenges for hand-editing that a competent text editor can easily rise to.
Combined with what rswail wrote about encoding hierarchies, with careful design, those CSV sections could be embedded as tables.
If that was used as a base format for other formats, then objection that encoding for boolean and numerics wasn't standardized might go away.
Like if the column was phone numbers and occasionally there's more than one, that sort of thing. Thinking of each cell as a "record" and allowing it to have more than one "unit" makes sense to me.
But anything would be better than ever having to whip up a script to fix a CSV with comma-separated dollar values in it, ever again.
Very thought inducing... I think the main impediment is that these characters are not visible and not so easy to type. If they were, we might not have got the number of CSV variants that have evolved.
What I'd want is an emacs special mode, that displays RS as a red* comma, GS as a simple newline, US as a red semicolon, and regular newline as a red "\n".
Comma, newline, and semicolon insert the control characters, while M-, etc insert the literal characters. Not sure exactly how to handle header lines but this is the general premise.
*red as in "whatever method of visually distinguishing them as special works for you"
XML might not be the best format for everything, and I for one am glad to use other formats for simple structured data. But when it comes to representing complex content, there is no other format that even comes close to being as useful and usable.
* All digital publishing of ebooks uses XML inside ZIP files.
* All contemporary mainstream word processors (Word, LibreOffice) use XML inside of ZIP as the basic file format.
* Automatic customized conversion processes from Word to InDesign or from InDesign to EPUB use XML at the heart.
* Let’s not forget the web itself, which is still mostly SGML in the form of HTML. Not XML per se, but only different in the details.
Not only is XML the only practical serialization format for working with publication content, but the presence of mature schema tooling is intrinsic to making publication automation robust in a given context.
I’m very glad for JSON and JSON schema in the domain of APIs. But in the domain of content data, it’s all XML.
Every serialization format has a domain for which it is most appropriate (whether or not it is the best choice in that domain.)
I’m really liking the shape of nested text for the domains in which I would have used YAML.
No doubts that there are use cases for this, but calling something that casts everything to strings an alternative to the above formats seems like a bit of a stretch.
1. I like that comments are part of the standard. I wrote my own C++ JSON parser that allows for comments too.
2. Is it strict about indentation? One thing you can never get programmers to do on significantly sized teams is consistent indentation. Is that tab+space? Or spacex5? Is it going to break if a tab sneaks into Git? (Setting up Git push rules just annoys and confuses people.)
3. "without the syntatic clutter of JSON" - I happen to like it. I can compact it quite far if I need to. I also like the fact I can spit it out over a debug server and JS will just magically start reading it.
4. Something really cool would have been the introduction of typed data. One way we achieve this via JSON is to create a template file which would declare something like (in a file named 'template.json' or something):
{
"data" : { "type": "float", "default": 0.0, "min": -1.0, "max": 1.0 }
}
Obviously this requires checking in the code, but it does build up some kind of format checking and sanity. It can also warn you that it's using a default rather than a config defined value.It would be nice if there was then the ability to define type syntax... But I fear this might be going too far.
5. Another thing I do with JSON is inheritance. So you define a 'parent' property at the top of a file, the values are loaded from the parent and then the child loads theirs over the top. Why have this? We usually need some per-application configuration but mostly it stays the same. It saves having to write it multiple times. You can even break it down into sections to keep each configuration file smaller.
(NOTE: For inheritance, a top tip is to implement "maximum depth", encase you get into a loop.)
If everything is a string, and you need to parse values yourself, you can tread "prawda" and "fałsz" as booleans, instead of "true" and "false".
An end user with no programming experience could reasonably understand a file written in this format. This can't be said about json or yaml, not when they don't know what "true" and "false" mean.
They mention that their key/value pairs are ordered. The downside here is that not all languages (eg. javascript), support them.
I also prefer them to be unordered. The downside with ordered dictionaries, is that you need to always be asking "does sequence matter here?". So it adds an additional thing to think about, more tests need to be written, and certain optimisations can't always be made.
1) how does one define a key with a ": "
2) as others said, significant whitespace, specifically trailing, seems to lead problems with keys that have nested key/val, array, multiline string.
3) the github repo associated with the parser has an isue that questions how to verify if the file was truncated.
I think Deco: https://github.com/Enhex/Deco better solves the problem the OP has. It is provably delimiter collision free, unlike this (see my point 1 Unless I'm mistaken).
Au contraire mon frere. What's the point of this data format if it's only intended to be for humans and not computers, because for computers the data type is critically important. For example, I've more than once seen the extremely ill-advised idea to treat zip code as a numeric, which completely screws your data model once you want to support zip+4 or international postal codes that contain letters.
Taking the main example, what do you do if you have more than one vice president?
This:
list of dicts:
-
key 1: Hi
key 2: there!
-
key 1: I'm a list
key 2: of dicts!
parses as {'list of dicts': [{'key 1': 'Hi', 'key 2': 'there!'}, {'key 1': "I'm a list", 'key 2': 'of dicts!'}]}Hidden features effectively do not exist. The parent may have gone overboard with judgement, but this really is the fault of the project maintainers. I'll admit I mentally filed it in the "useless toy" category without this feature.
I went looking for it specifically because it's one of the ugliest and most confusing parts of YAML. To be pathological, what if you have a list of lists?
list of lists:
-
- first sublist
- goes here
-
- second sublist
- is here
produces {'list of lists': [['first sublist', 'goes here'], ['second sublist', 'is here']]}
I don't think either of these are hidden. Yes, explicit examples would be nice, but assuming things not shown are impossible seems strange to me. Would you assume Python can't do lists of dicts, or lists of dicts of dicts? A quick search fails to find any examples on python.org (I did find lists of lists though).Edit: forgot to mention. I personally don’t like most general data serialization formats for configuration, the one I can probably tolerate is XML, but even that I only use when it’s part of requirement. The way I usually implement program behavior configuration is through runcommands(runcom, .rc)
BTW, I happen to like YAML as a configuration format. It's much more readable than JSON. It's not that suitable for serialization, and probably shouldn't be used to create huge config files, but for the rest, it's as good as it gets.
"The format holds dictionaries (ordered collections of name/value pairs), lists (ordered collections of values) and strings (text) organized hierarchically to any depth."
1. Ordered dictionaries are not conveniently supported in all languages I tend to use.
2. The only element type is string which means that parsing of common types has to be separately and differently by each implementation that uses such a file with potential for unspecified differences.
Ordered dictionaries are a fundamental, extremely basic data type present in every language I've used in three decades.
a) What modern, widely used language doesn't have them?
b) Why?
c) Okay, so you've picked a bad language that has only hashtables. You can still implement ordered keys using an array of the keys & values in sorted order, and a hashtable of keys to the array indices.
I don't think this will ever happen, but having a human readable plain text output whenever you select any data from a UI is a very powerful idea. EVE Online did this back in the early 2000s and let you copy virtually anything in UI. This led to awesome third party tools that could work with the game with a minimal learning curve from the user (everything was "paste what you copied from the game"). It did use a custom serialization format... But a plain text one nonetheless.
IO via copy paste between a wide variety of apps in a known, and very simple, structured format would be awesome in my books.
> Unicode characters without encoding them
But JSON doesn’t generally require special encoding for unicode characters.
Technically, JSON is UTF-8 encoded, but that’s true of NestedText too so that can’t be what they mean.
Also, of JSON it says:
> in JSON 32 is an integer, 32.0 is the real version of 32...
But JSON doesn’t distinguish integer from real. It only has number.
On the format itself, I wonder if it needs some testing and hardening. The definition seems ambiguous, but maybe just needs a formal grammar. Just from reading it, tab handling seems like an issue. I think you could have documents that look right but have invisible issues due to tabs. E.g., it sounds like tab can be a character in a dictionary key name, which looks like indentation.
What is the way to represent a map/dictionary where the keys are not strings? In json I also don’t have a good way except I guess a list of objects with a key and value field but it feels heavy. In Lisp I would probably use an alist (whereas a plist might be used for the “string” (or symbol, rather) keys). In fact, how does one even write a list of lists with this?
It just shifts problems to other areas, which some may find even worse.
In my indian music notation publication tool Patantara, I use a text preamble format that's limited and far simpler for people to use than any of these "data structure capable" formats - including NestedText/YAML/JSON. Here is a quick description -
1. You specify key value pairs with lines of the form -
key with allowed spaces = any textual value
2. The first " = " on a line separates the key from the value. The key is normalized to lower case and spaces in the key are normalized to single space, and the key is LR-trimmed. The textual value is mostly kept as is (except for LR-trimming), but can accumulate more content.3. Lines which don't have a " = " are just strings that get appended to the value of the immediately preceding key. So
key = value1
value2
value3
will result in the key "key" having the single string value "value1\nvalue2\nvalue3".4. If the same key is given multiple times, the values are concatenated with line breaks. So
key = value1
key = value2
key = value3
gets you, for the key "key", the value "value1\nvalue2\nvalue3".5. The reader doesn't care about data types. However, the application can decide how to parse the string values. One convenience I use for lists of items is "comma or whitespace separated values", where a list of values is given as "value1, value2, value3" or "value1\nvalue2\nvalue3" or "value1 value2 value3" or any combination including "value1, value2,\nvalue3 value4".
Nested structure is not possible with this format (unless you impose more structure on the value in the application), but it serves very well as a metadata format for text files in my application.
Ref: https://blog.patantara.com/blog/2017/04/16/controlling-who-c...
In Emacs, org files are often used for configuration similarly.
While I'm a fan of EDN most of the time, which doesn't suffer from the issues they mention JSON having, not having to wrap things in quotes for text, and using indent for nesting definitely would make it nicer for non technical people. So I can see a use case for this as a simple text interface for forms that is targetted at non-technical users.
One might call this 'argument by gotcha' interesting to imagine the kind of lifeworld that makes one think this is a worthwhile form of persuasion.
1. comments aren't part of the parse tree, so I can quickly strip all comments by a quick parse print cycle. This is really helpful if you want to do automate changes to the files.
2. I can freely truncate the file at (almost) any point, and still have a valid config file. It's not obvious up front, but if your file gets snipped at some point, it's won't cause parse errors (usually) consumers of the file.
Wonder what the best start of file and EOF-equivalent sentinel would be to use for YAML-esque files.
Something like a double colon?
::
item1:
- etc..
::2. isn't too bad in yaml, since json is valid yaml, a broken file can be reliably picked up with leading '{' and trailing '}'
With regard to 1 - both comments and nonstandard formatting like the above {} hack, are currently hard.
Anyway, got to spend some time outside, and away from screens. Apologies for the snark.
Pydantic would work great in conjunction with this to coerce types.
These felicitations were brought to you by the .ini gang
I think they made some good decisions. All types are strings and left to the application to parse and document. Excellent.
And, I think the desicion for the angle-bracket-multiline syntax was great because I think one could create a parser with awk/grep. Simple.
>
> >I don't see this thing getting any traction. JSON is pretty much the standard at this point and I don't see anything changing that anytime soon.
1. Doesn't org-mode already solve this problem nicely, in a more widely-standardized way? (Maybe not; I'm just learning org-mode ATM.)
2.There is no "i" in Topeka.
2. The addresses are of course fake; the zip codes don't match either. Same with the phone numbers. "KateMcD@aol.com" may be real, but "margaret.hodge@uk.edu" isn't.
I would also like to see YAML converted to NT and back to YAML just to see if there is any information loss.
; Comments!!!
; An Array with no commas!
greatDocumentaries: [
'earthlings.com
'forksoverknives.com
'cowspiracy.com
]
importantFacts: [
; Multi-Line Strings! Without Quote Escaping!
emissions: {Livestock and their byproducts account for at least 32,000 million tons of carbon
Goodland, R Anhang, J. “Livestock and Climate Change: What if the key actors in climate change we
WorldWatch, November/December 2009. Worldwatch Institute, Washington, DC, USA. Pp. 10–19.
http://www.worldwatch.org/node/6294}
landuse: {Livestock covers 45% of the earth’s total land.
Thornton, Phillip, Mario Herrero, and Polly Ericksen. “Livestock and Climate Change.” Livestock E
https://cgspace.cgiar.org/bitstream/handle/10568/10601/IssueBrief3.pdf}
burger: {One hamburger requires 660 gallons of water to produce – the equivalent of 2 months’
Catanese, Christina. “Virtual Water, Real Impacts.” Greenversations: Official Blog of the U.S. EP
http://blog.epa.gov/healthywaters/2012/03/virtual-water-real-impacts-world-water-day-2012/
“50 Ways to Save Your River.” Friends of the River.
http://www.friendsoftheriver.org/site/PageServer?pagename=50ways}
milk: {1,000 gallons of water are required to produce 1 gallon of milk.
“Water trivia facts.” United States Environmental Protection Agency.
http://water.epa.gov/learn/kids/drinkingwater/water_trivia_facts.cfm#_edn11}
more: http://cowspiracy.com/facts
]
Some references: https://en.wikipedia.org/wiki/Rebol | http://www.rebol.com/article/0108.htmlI can't figure out what's happening here. Is everything just a string? They complain about how YAML handles booleans (fair criticism), however I can't see how NestedText fixes this.
This feels far inferior to TOML (everything NextText does, and much more) or JSON (very unambiguous, works natively in most languages).
I see this a lot in startups and open source projects. They exist solely as a criticism of their competition (which is fair!), but don't work to understand the real problems and fix them in any meaningful way. You can spot them when they talk a lot about what's wrong with the competition ("AWS is too expensive", "YAML is too ambiguous") but never explain their solution.
Honestly after a few years using YAML you remember to always quote anything you know you need to be a string, and the different multi line types.
An argument against this is "why not just design it to be simpler" — well, then it's less useful in as wide a variety of applications, and stands less chance of actually becoming a standard like JSON or YAML has. And you end up with the XKCD problem.
YAML came in and solved a few of JSON's most glaring problems (multi line and comments) with a usable approach.
This new format seems like it's "YAML, but a little different."
If I wanted to usurp YAML, I'd focus on the greatest pain points for most, whitespace and schema support.
AFAIK they were both independently created in 2001 and YAML was not created in response to JSON.
One of the biggest issue is just validating the actual data
Edit: A lot of people seem to be accidentally clicking the downvote icon.
if (something == somethingElse)
doSomething();
--counter; // added afterwards
when whitespace has actual meaning.Your editor adds the whitespace to match the current scope, so you don't have to type it manually.
What's your issue with significant whitespace?
I dislike significant indentation for configuration because the nesting tends to get deeper than for programming languages and it can be harder to see what's going on.
Modern languages repair this with mandatory block braces: {C style} (Rust, Go, Swift...) or Modula/Ada style (if x then foo else bar end). C style is easier for editors.
> What's your issue with significant whitespace?
1. Many cases when leading whitespace is still spoiled, including editors (which use can't be avoided due to corporate tool limiting), web formatters, etc. Last decade spreading of mobile-aware tools made this worse.
2. With grouping by indentation, it's hard to determine block start if multiple blocks are ended at the same line, like:
a:
b
c
d:
e
f
p
and let's imagine the most nested block occupies 2-3 screens (not rare for configs or code, even if it is against good coding rules). You have no means to decide what block start is to be matched when you ask editor to find match.I code in Python continuously since 2004 and I deem grouping by indentation is bad idea. Not fatally bad, of course, but with explicit grouping it would be a bit better.