YAML: probably not so great after all (2017)
arp242.net
arp242.net
A second thing to consider... YAML was created before it was common that tech companies actively contributed to open source development. There are lots of things we could have done differently if we had more than a few hours per week... even a tiny bit of financial support would have helped.
Finally, YAML isn't just a spec, it has multiple implementations. Getting consensus among the excellent contributors is a team effort, and particularly challenging when no one is getting paid for the work. Once you have a few implementations and dependent applications, you're kinda stuck in time.
It was an special pleasure for me to have had the opportunity to work with such amazing collaborators.
We did it gratis. We are so glad that so many have found it useful.
* I mean, it uses an ill defined subset of YAML. The definition is "whatever the Symfony YAML parser supports".
eh?
{ "firstName": "John", "lastName": "Smith", "comment": "foo", }
I know it isn't the same as #comments, but who cares really.
the person who came up with HOCON, probably
{ # comment with a note about the value of foo "foo": "bar", # comment with a note about the value of baz "baz": "qux" }
Without driving myself and future readers insane with fooComments and bazComments?
What if I need a multiline comment explaining a yak-shaving story for why a key is set to a certain value?
What if the object in question is a set of keyword arguments, and adding new fields changes the behavior of whaever is parsing the document?
{
"#": "A foo variable",
"foo": true,
"#": "A bar variable",
"bar": false
}
Alternatively. {
"# A foo variable": "",
"foo": true,
"# A multiline..": "",
"# .. bar variable": "",
"bar": false
}
Presto!Also...
print (json.dumps(json.loads(js_data), indent=2))
{
"bar": false,
"foo": true,
"# A multiline..": "",
"# .. bar variable": "",
"# A foo variable": ""
}
Presto! ;-)Drupal 8 file format discussion was in 2011, predating it by two years. https://groups.drupal.org/node/159044
I don't like it because it uses the = symbol which seems imperative rather than declarative. (Same with HCL, it might be a nitpick but these are languages I'm going to be using all the time.)
HOCON is interesting but at first glance it seems it might be too ambiguous for my tastes, because like YAML, because it supports both js-style ("//") and shell-style ("#") comments.
JSON plus comments is beautiful because it adds minimally to an unambiguous language which lends itself to automatic formatting (stringificiation).
Technically, this is not JSON. You won't be able to use a standard JSON parser without stripping comments first. But you can use a simple, JSON-like language with comments for config.
[1] https://code.visualstudio.com/docs/languages/json#_json-with...
YAML can be employed as a simple JSON-like language with comments.
YAML is much, much more complicated than JSON.
> simple
YAML is much, much more complicated than JSON.
Quoting a single word from the parent’s sentence is misleading. The sentence "YAML can be employed as a simple JSON-like language with comments." is true because JSON is YAML, so you can parse a JSON file with #-comments using a YAML parser.So, this unnecessary parser complexity is a usability issue. You should use a parser for the config language you actually intend to support.
Supports comments, trailing commas, single quotes, multi-line strings, and more number formats.
[
1
2
3
]
New lines used by humans, computers should do a good job as well.
required: [
'firstName'
'lastName'
]which is the same as:
{
"required": [
"firstName",
"lastName"
]
}in JSON :)
PHP.
I understand why some languages rely on common configuration file formats.
I don't understand why the popular dynamic script-y languages don't more commonly use the natively-expressable associative/list data structures that they're famous for making convenient.
You picked the wrong language... PHP comes with its own JSON parser. And INI and XML and even CSV.
But, the reason is that, generally, you want config files to describe data or state only. Yes, you could just make your config native code, but then the temptation to add functions and methods and logic to that becomes irresistible and soon your config is an application that needs its own config.
Config formats need to be simple, and preferably not Turing complete.
See The Configuration Complexity Clock. https://mikehadlow.blogspot.com/2012/05/configuration-comple...
INI is still simple, and JSON doesn't support logic, so the madness can be held at bay at least for a time.
XML and s-expressions are lost causes, though.
There's no language that I'm aware of that can natively generate PHP syntax, and there's no common multi-language-platform library for generating PHP syntax. I think that's most of the reason.
To contradict myself, though: Ruby encodes Gemfiles and Rakefiles as Ruby syntax. And Elixir encodes Mixfiles, Mix.Config files, Distillery release-config files, and a bunch of other common data formats as Elixir syntax.
And, of course, pretty much every Lisp just serializes the sexpr representation of the live config for its config format (which means that, frequently, a lot of Lisp runs code at VM-bootstrap time, because people write Turing-complete config files.)
This is a solid argument against using PHP (or any such language) as a cross-language data interchange format. There are others :) And I totally agree you want a language independent format for anything you might have to feed across an ecosystem of tools.
For a PHP-system generating/altering its own config files... PHP's `var_export` generates a PHP-parseable string representation of a variable (though it sadly doesn't use the short array syntax).
Turing-complete config files probably have some hazards, like Lisp itself does. YMMV regarding whether those hazards can be avoided by circumspect developers or need to be fenced off.
Django's settings.py sucks. I've used Django since the 0.9 days. It's extremely impractical and needs to be worked around constantly.
And you won't have to generate your config files (parsing, maaaaaybe), because those needs are covered by the fact that the files are programs. They are _already_ generating a configuration.
Yes, theoretically, if settings.py was a "generator" format that you ran as a pre-step (like you do to get parser-generators like Bison to spit out source files for you to work with), and this generator actually spat out something like a settings.json, and all the rest of the infrastructure actually dealt with the settings.json rather than the generator, then, yes, it wouldn't matter. Tools in other languages could just generate the settings.json directly.
As it stands, none of those things are true, so tools in other languages actually need to do something that outputs settings.py files.
That means that if I wanted to configure Vagrant with JSON, there is no force in the universe that could stop me.
If the config file is actually a normal program, then it can do normal program things, then any benefit from using JSON instead is nullified by the fact that you can still use JSON. In turn, if your tools primary configuration is via a more limited settings, you're stuck with it. Not even "generators in other languages" allow comparable runtime flexibility.
Actually, I've had to use PHP to output a PHP configuration array for a project that required config in PHP.
`var_export($foo)` will output valid PHP code for creating the array $foo. In my case I was doing horrible things to create the array in my pseudo-makefile, then using `var_export()` to output the result. Note that you can run php from the Bash CLI with the `-r` flag, which helps.
$CFG = random() > 0.5 ? "yes" : "no";
...is likely "too powerful". It'd be nice if there were ways in certain programming languages to do something like "drop privileges" to avoid loops, function calls, external access, etc.
Using code-as-data works really well in Lisp-like languages. Reading a Clojure project's project.clj file or a Lisp project's project.asdf file is pretty pleasant. A programming language's choice in how it decides to handle library config info for building and specifying dependencies (XML, makefiles, JSON, YAML, INI, nothing, etc...) will be a good indicator for the culture of the language around config files in general. Composer for PHP only came out in 2012.
John Ousterhout explained in one of his early TCL papers that, as a "Tool Command Language" like the shell but unlike Lisp, arguments were treated as quoted literals by default (presuming that to be the common case), so you don't have to put quotes around most strings, and you have to use punctuation like ${}[] to evaluate expressions.
TCL's syntax is optimized for calling functions with literal parameters to create and configure objects, like a declarative configuration file. And it's often used that way with Tk to create and configure a bunch of user interface widgets.
Oliver Steel has written some interesting stuff about "Instance-First Development" and how it applies to the XML/JavaScript based OpenLaszlo programming language, and other prototype based languages.
Instance-First Development: https://blog.osteele.com/2004/03/classes-and-prototypes/
>The equivalence between the two programs above supports a development strategy I call instance-first development. In instance-first development, one implements functionality for a single instance, and then refactors the instance into a class that supports multiple instances.
>[...] In defining the semantics of LZX class definitions, I found the following principle useful:
>Instance substitution principal: An instance of a class can be replaced by the definition of the instance, without changing the program semantics.
In OpenLaszlo, you can create trees of nested instances with XML tags, and when you define a class, its name becomes an XML tag you can use to create instances of that class.
That lets you create your own domain specific declarative XML languages for creating and configuring objects (using constraint expressions and XML data binding, which makes it very powerful).
The syntax for creating a bunch of objects is parallel to the syntax of declaring a class that creates the same objects.
So you can start by just creating a bunch of stuff in "instance space", then later on as you see the need, easily and incrementally convert only the parts of it you want to reuse and abstract into classes.
What is OpenLaszlo, and what's it good for? http://www.donhopkins.com/drupal/node/124
Constraints and Prototypes in Garnet and Laszlo: http://www.donhopkins.com/drupal/node/69
All configuration files were Tcl data structures that were sourced on server start.
[1] An example: https://github.com/spc476/mod_blog/blob/master/journal/blog....
Lua is great.
There's also an argument about whether making configuration files able to execute arbitrary code is a good idea. You get straight into the JavaScript 'eval' problems which we've spent a decade escaping.
Your configuration file is one of your program interface. It's something that must be well define. If your configuration file is a programing language this interface is not that well defined.
Also you expose yourself to all kind of weird bugs because some (too smart for their own good) people will monkey patch your software using it.
It adds a lot of unnecessary stuff in the configuration file, things like ';' or '$' are not really useful.
Lastly, common configuration file format are good because there are... common. You can have 2 pieces of software in 2 different languages accessing the same configuration file. A common example of that is configuration management, There are a lot of modules/formula in salt/ansible/puppet/chef doing fine parsing of the configuration files and permits fine grain settings, and I'm not mentioning augeas. If your configuration is a php/python/perl/ruby file good luck with that.
I know it's really common for php applications to do configuration files in php, but frankly, it's a bit annoying.
While I do agree with the rest of your comment I don't think they were advocating using the full language for configuration, just the maps/arrays/etc. (e.g. Python's `literal_eval`).
Something like:
config = {'key1': 'value1', 'key2': 'value2'}
could be written as:
config = {}
config['key1'] = 'value1'
config['key2'] = 'value2'
With large chunk possible between the 3.
It basically transforms the configuration file into an API like any library, which is not really what you want for an end user program.
Consider this: Design and optimize for the common case.
Why do we have config files? Because developers actually want a place dedicated to simple or structured application configuration data, for which PHP assignments with arrays + primitives can function at least as effectively as JSON. Most developers would prefer that config data get loaded quickly so the application can get on to doing actual app-y things. Using the language for this means you're parsing at least as fast as you can interpret and you can also take advantage of any code caching that's part of your deployment (especially nice in the PHP-likely event that config settings would be reloaded with every request).
Abuse isn't likely to be the common case. The end users you invoked certainly aren't going to be the ones looking for opportunities to insert code over data. Developers have other places to put code and, as mentioned, probably actually want a place dedicated to data. You're still right that of course someone will do it, just like someone will inevitably create astronaut architecture hierarchy monstrosities in any language with classical inheritance or make potentially hidden/scary changes to language function using metaprogramming facilities.
But potential for abuse doesn't automatically mean a feature should be disallowed.
A lot of the time it's better to let people who can be circumspect have the benefits of a potential approach, and if somebody thinks they need to solve a problem by using a technique that's arguably abuse, well, let them either find out why it's a bad idea or enjoy having solved their problem in an unusual way. Not the end of the world. Possibly even legit.
But to play the devil's advocate, how would JSON be able to support round-tripping comments like XML can, since <!-- comments --> are part of the DOM model that you can read and write, while JSON // and /* comments */ are invisible to JavaScript programs. There's nowhere to store the comments in the JSON object model, which you would need to be able to write them back out later!
On important feature of JSON is being able to read and write JSON files with full fidelity and not lose any information like comments. XML can do that, but JSON can't. To fix that you'd have to go back and redesign (and vastly complicate) fundamental JavaScript objects and arrays and values, to be as complex and byzantine as the DOM API.
The less-than-ideal situation we're in isn't JSON's fault or JavaScript's fault, because JSON is just a post-hoc formalization of something that was designed for a different purpose. But JSON is rightly more popular than XML, because it's extremely simple, and nicely impedance matched with many popular languages.
YAML suffers from the same problem as JSON that it can't round-trip comments like XML can, but it fails to be as simple as JSON, is almost as complex as XML, and doesn't even map directly to many popular languages (as the article points out, you can't use a list as a dict key in Python, PHP, JavaScript, or Go, etc).
You can sidestep some of JSON's problems by representing JSON as outlines and tables in spreadsheets, without any need for syntax and sigils like brackets, braces, commas, no commas, quoting, escaping, tabs, spaces, etc, but in a way that supports rich formatted comments and content (you can even paste pictures and live charts into most spreadsheets if you like), and even dynamic transformations with spreadsheet expressions and JavaScript.
See my comments about that in this and another article: https://news.ycombinator.com/item?id=17360071 https://news.ycombinator.com/item?id=17309132
At the end of the day I'm sure the reason we don't have JSON comments is somewhere listed in this page: xkcd.com/927/
> I removed comments from JSON because I saw people were using them to hold parsing directives, a practice which would have destroyed interoperability. I know that the lack of comments makes some people sad, but it shouldn't.
https://plus.google.com/+DouglasCrockfordEsq/posts/RK8qyGVaG...
And why to model JSON syntax closely after JavaScript literal object syntax (which is actually more convenient, by the way) which, being taken from mainstream programming languages, naturally evolved to be written by humans in small amounts not by computers in large dumps? :)
But you really only need comments in JSON if you're doing stuff like storing configuration in JSON, and JSON's too fiddly in general to be a great config file format (too easy to do something like forget a comma; no support for types beyond object, array, (floating point) number, and string). Something more like YAML without the wonky type inference would be better, IMO.
It's still in its early stages, so if anyone's got any comments I'm interested in hearing them :-)
It doesn't support it for whitespace in general (if you deserialize into JS object model or equivalent), so why would it be any different for comments specifically? It's just not a design goal of the format.
Although, of course, it's quite possible to have a JSON parser that preserves representation. It'll just have a non-obvious mapping to the host language because of all the comment and whitespace nodes etc.
While not mandated by the YAML specification, it doesn't prevent creation of a parser that round-trips comments.
In fact, the ruamel.yaml project for Python provides one.
In fact YAML is probably more complex than XML; the specification of YAML, when I print it into PDF, is about three times as long as that of XML 1.0. (And XML 1.0 also describes DTD, which is kind of a simple type validation for XML and thus includes much more than just serialization syntax.)
Jsonnet is a relief. Kubernetes should have been a dumb json config from the get go. JSON is ridiculously simple to parse and emit. It has huge interoperability as well with lots of programming languages.
P.S. You Romanian by any chance?
Interesting. How did it happen then that, quoting the YAML 1.2 spec, that "every JSON file is also a valid YAML file"? Although the previous spec documents don't mention JSON.
Was that an intentional design decision for 1.2 or was it some kind of convergent design due to Javascript?
> The primary objective of this revision is to bring YAML into compliance with JSON as an official subset.
When I say "JSON didn't exist", what I mean is that it wasn't popular or known to us when we were working on YAML. So, please excuse my sloppy wording. For me, the work on what would become YAML started with a few of us in 1999 (from SML-DEV list). In January of 2001 we picked the name and had early releases. It took a few years of iteration before we had a specification the collaborators (Perl, Python, and Ruby) could all bless.
Anyway, with regard to Crawford's excellent work, JSON. It is a coincidence that YAML's in-line format happened to align. Although, it's probably because of a "C" ancestor, not JavaScript. The main influence on the YAML syntax was RFC0822 (e-mail), only that from my perspective, it needed to be a typed graph. In fact, we documented where we stole ideas from, to the best we could recall at that time: http://yaml.org/spec/1.0/#id2488920.
Out of curiosity, did you see the parser linked to at the end of the article? ( https://github.com/crdoconnor/strictyaml )
That was my attempt at giving YAML a haircut. I'd be curious to know what you thought.
Thank you for creating YAML, by the way. Even though part of that rant was quoted from me, I'm not negative on it like the author - I think the core was brilliantly designed. If you put two hierarchical documents side by side - one in TOML and another in YAML the YAML one is much, much clearer and cleaner.
If on the other hand you are using a YAML library... I've had pretty good success using YAML compatibly across Python, Ruby, C# and Go projects. Do you have a particular issue in mind that the existing Ruby implementation doesn't address?
That said, StrictYAML seems to be a tad bit more of a hair cut than I'd imagine. I'd keep nodes/anchors, since I think a graph storage model is underrated; I think that data processing techniques just haven't caught up with graph structures.
Further, I'm not sure everything can be easily typed based upon a schema. Hence, I'm not sure about completely dropping implicit types, perhaps you may want to provide a way for applications to resolve them if they wish. For example, an application may want to attempt to treat anything starting with "[" or "{" as JSON sub-tree. Perhaps keeping "!tag" but handing it off to the application to resolve might also be a good idea in this regard. Even so, typing should be done at the application level and default to something very boring.
Thanks, that's very flattering.
> I'd keep nodes/anchors, since I think a graph model is underrated
Well, you can create graph models without it (and I do) - you can just use string identifiers to identify nodes and let the application decide what that means.
I always thought the intent behind nodes/anchors was not so much graph models but rather to take repetitive YAML and make it DRY. That appears to be how it is used, e.g. in gitlab's ci YAML.
>I'm not sure about completely dropping implicit types, perhaps you may want to provide a way for applications to resolve them if they wish. For example, an application may want to attempt to treat anything starting with [ or { as JSON.
I think that would cause surprise type conversions. There will be plenty of times when you want something to start with a [ or { and you won't want it parsed as JSON.
I embed snippets of JSON in YAML multiline strings sometimes and I usually just parse it directly as a string. Then I run that string through a JSON parser elsewhere in the code.
>You might wish to give Ingy a ring.
I would like that.
YAML has traditionally been used as the basis of higher-level configuration files for particular applications. What I'm saying is that implicit typing should be permitted, but delegated to those applications.
Conversely, I'm not saying that StrictYAML should do anything by default with unquoted values, except reporting them to the application as being an unquoted value. This way the application could choose to process the value differently from those that are quoted.
I think a reason this won't necessarily fix the problem with unmet expectations is that identical constructs in different but analogous yaml files would be likely to end up with very different semantics and users effectively have to remember which particular idiosyncratic YAML dialect choices various apps make. Say
version: 1.3
means the string "1.3" in app a), the float 1.3 in app b) and a version number in app c) one. Furthermore let's assume that app c) required a version number, whereas a) and b) required strings.Another, more subtle problem, is that such a scheme would make it more likely that applications would end up parsing raw string representations themselves (with ensuing subtle differences even for things which are nominally meant to be identical, say dates or numbers and possibly security problems as well).
That's how I use it too. When I read about competing formats, that's the first feature I check for. It's really key for readability and usability in some use cases.
Other use cases such as dumping any in-memory data structure from memory, perhaps out of a sense that we needed full completeness, actually didn't have any end-user usability testing. Round-tripping data seems in retrospect to be a diversion from the primary value that YAML provided.
YAML is an invented serialization format, JSON is a discovered one. As CrOCKford points out, JSON existed as long as JS existed, he just called it out and put a name on it.
Anyway, XML is a strong anti-pattern (too much security, even if you get it right on your end, the other party likely screwed something up). YAML seems to be going down that path too.
TOML seems to be "the JSON of *.ini" (ie: discovering old conventions, rather than inventing new ones), and I'm glad to have been exposed to it.
If you define JSON as the underlying practice that Crawford later named and documented, then sure, what I wrote reads completely wrong headed. However, when we were working on YAML, JSON was not yet called out and given a name.
I believe the most important convention that YAML and JSON shared was a recognition of the typed map/list/scalar model used by modern languages. Further, as far as conventions go, I think there's quite a bit to be said about languages that use light-weight structural markers such as: indentation, colon and dash.
But then again I was already used to using Perl data structures as dumped by Data::Dumper for config, because I was taught a lot about Perl by a Lisp programmer who had used Lisp data structures for the same purpose since the 1980s. So using JSON didn't feel original or clever. It seemed like I was simply using a well-known technique in yet another dynamic language.
Then again our reaction to XML was the stupid thing other people were doing that you had to do to interact with the rest of the world. I got used to holding my tongue until I went to Google a decade later and found that my attitude was common wisdom there...
At any rate, it's worth mentioning that in the conclusion I wrote:
> Don’t get me wrong, it’s not like YAML is absolutely terrible but it’s not exactly great either.
I still use YAML myself even when I have the freedom to use something else simply because – for better or worse – it's very widespread, and for many tasks it's "good enough". For other tasks, I prefer to avoid it.
I think that a stricter version of YAML (such as StrictYAML) would make a lot of people's lives easier though.
Do you have a link to the document you were pointed to when you got penalized? If it was Google who penalized you, they must have pointed you to a URL with documentation on why you got penalized and how to resolve it.
I ask this because I run a few websites with even lesser markup than your site but I have never got penalized. I once got penalized due to excessive number of spam comments on one of my websites and they pointed me to https://support.google.com/websearch/answer/190597 ("Remove this message from your site") to resolve the issue. This issue did not affect the search ranking much though (dropped by only about 2 or 3 places in the list of results). But never had an issue with abnormally low markup.
The markup in your website looks pretty reasonable to me, so I am surprised you could get penalized for that when I have had no issues with even lesser markup and they still appear at the top of the list of results for relevant search terms.
> I have had no issues with even lesser markup and they still appear at the top of the list of results for relevant search terms.
It seems people are finding my site, whether or not it's being penalized. I mean, someone other than me posted it here, right?
Using it a lot more lately as I'm diving into Ansible, so I'll be interested to see if I run into problems.
Ansible does make some efforts to limit jinja templating to variable substitution, but it's sill not that great, you have all kinds of weird stuff that can happen specially with colons.
The worst one is saltstack, the resulting syntax is just atrocious and border line unreadable, I not a big fan of map.jinja files[0] and on the yaml side, things can get ugly quite fast [1].
I know it's not a popular opinion, but I would rather use the puppet DSL, even with its step learning curve.
[0] https://github.com/saltstack-formulas/salt-formula/blob/mast...
[1] https://github.com/saltstack-formulas/mysql-formula/blob/mas...
I also have come to agree that a DSL is the best solution, though Puppet's particular DSL is not a great example. Projects that re-implement the same thing from scratch like mgmt[1] are on the right track, but probably won't gain enough traction.
[1] https://github.com/purpleidea/mgmt/blob/master/docs/language...
(While constructive criticism is fine, those rare people who trash it are... nonsensical to me. I'd like to see them do one-tenth as good under the same conditions!)
"Anyone who uses YAML long enough will eventually get burned when attempting to abbreviate Norway."
Example:
NI: Nicaragua
NL: Netherlands
NO: Norway # boom!
`NO` is parsed as a boolean type, which with the YAML 1.1 spec, there are 22 options to write "true" or "false."[1] For that example, you have wrap "NO" in quotes to get the expected result.This, along with many of the design decisions in YAML strike me as a simple vs. easy[2] tradeoff, where the authors opted for "easy," at the expense of simplicity. I (and I assume others) mostly use YAML for configuration. I need my config files to be dead simple, explicit, and predictable. Easy can take a back seat.
[1]: http://yaml.org/type/bool.html [2]: https://www.infoq.com/presentations/Simple-Made-Easy
As the saying goes, "there are only two kinds of languages: the ones people complain about and the ones nobody uses." YAML, like any language, isn't perfect, but it's withheld the test of time and is used by software around the world—many have found it incredibly useful. Sincere thanks for your contribution and work.
It's[1] just so blatantly unnecessary to support any file encoding other than UTF-8, supporting "extensible data types" which sometimes end up being attack vectors into a language runtime's serialization mechanism, autodetecting the types of values... the list goes on and on. Aside from the ergonomic issues of reading/writing YAML files, it's also absurdly complex to support all of YAML's features... which are used in <1% of YAML files.
A well-designed replacement for certain uses might be Dhall, but I'm not holding my breath for that to gain any widespread acceptance.
[1] Present tense. Things looked massively different at the time, so it's pretty unfair to second-guess the designers of YAML.
That doesn't help you, of course, when using a multitude of existing systems whose yaml parsers are based on 1.1...
I'd still love for a better means to resolve ambiguities like this, but I've found always quoting to be a fairly reliable approach.
Ansible extends yaml so that:
cmd: a b c
is actually but not quite identical to:
cmd: ["a", "b", "c"]
It also embeds JINJA2 templating part-way (!) through the YAML parsing process.
The gotchas that these and other bastardizations cause is only partially documented at the bottom of this page: https://docs.ansible.com/ansible/latest/reference_appendices...
I like ansible, but its decision to use a bastardized YAML is a major pet peeve of mine.
Your first example is an Ansible convenience feature, it's not extending or changing the YAML syntax in any way. You can simply specify `cmd` values as lists or strings, since working with one or the other may be easier depending on the use case.
The templating is unfortunate in some areas, especially where the jinja2 syntax conflicts with what YAML expects (for example starting an object with '{'). That's due to a combination of templating engine choice and YAML, though, and not some custom implementation of YAML. Unless I'm misunderstanding?
I do think going with YAML was a trade-off for Ansible, but it's hard to see Ansible getting to where it is today if it had gone with a custom DSL (or JSON, thank god). I'd take Ansible's YAML over Chef's Ruby or CloudFormation's JSON any day.
The most recent offenders for bastardizing YAML I have seen are the different CI services:
* Circle CI using moustache-like templating and interpolation with things like {{ .Branch }} available in certain steps [1]
* GitLab CI adding an "include" type directive to declare YAML dependencies [2]
I've also experienced this professionally. At my last company, somebody decided to add a feature to enable interpolation in some parts of the YAML deployment data. It ended up being used by a handful of people who were confused why interpolation worked in some places and not others. The weird trend of "extending YAML" seems to be going against any sort of benefits you might have by trying to use it.
[1]: https://circleci.com/docs/2.0/configuration-reference/#save_...
YAML is for convenience for hand-editing configuration/task files; if you're doing anything that doesn't require hand editing/readability, use JSON.
In contrast, JSON is super intuitive and basically self documenting. The only real quirks are that you need to use double quotes, and objects can't have a trailing comma.
The only good thing I can see about YAML is that it's super easy to convert and re-export to JSON.
Personally I've found the exact opposite when dealing with 'normal' people. Most people can get basic YAML, but unless they're a programmers (or at least know how to program) most people fail miserably at writing JSON by hand.
Also detects repeated hash keys. This is a very good compromise between user-friendliness and machine-friendly specification and serialization.
Thank you very much for YAML! It is a critical user-interceptable interchange mechanism at several companies that I have worked at.
Intuitive usually means "close to what I'm used to".
Especially since leaving the trailing comma is considered the best practice in every other language.
The real schism here, IMO, is Programmer Intuitive vs. Natural Language Intuitive.
You misrepresented yourself by saying "object" when you meant "dict" and saying "practically identical" when you meant "vaguely syntactically similar". I wasn't trying to nitpick your semantics; I just had no idea that "Python objects are practically identical to JSON" meant "Python dicts are to Python what JSON objects are to JS, oh and also Python dicts have some syntactic similarlities to JSON" or whatever.
I'd expand the list of quirks... JSON lacks comments (both line-level and block level). Fine for data transport but super super bad for configuration files.
{ "ConfigKeyComment": "This is for blah blah blah", "ConfigKey": "Foo" }
Obviously this wouldn't work in all cases (you're putting more work on your parser to interpret unused keys basically), but if we're talking config files specifically, I see this as an acceptable approach since there's little chance you'll be parsing such files more than once each (plus, writing a simple tool to strip the comments out would be very trivial).
# Never enable this config, because if you do the space-time
# continuum will collapse into itself and the cloud servers
# will disappear in a puff of steam. However, if you really
# must enable it, remember that it's boolean and go read
# TICKET-8675309 for the extensive list of side effects.
TurboFactorRenoberation = false
...so people just end up writing stuff like: {
"ConfigKeyComment": "TICKET-8675309",
"ConfigKey": "TurboFactorRenoberation",
"ConfigVal": false
}
[edit]: formattingMaybe I'm lazy but avoiding increasing cost to commenting is one of the few absolutes I abide by. Often I find myself tired after a long stretch of code, trying to convince myself that's it's understandable on it's own.
This is one of those systematic rules I have to enforce to shutdown my lazy lizard brain.
edit: But I can see how highly structured comments could actually come in handy as well for viewing configs in a gui
-Douglas Crockford, creator of JSON
There is no issue using JSON with comments for a config file.
/\/\/.*/gm
My point is only that this isn't a big issue. I don't understand why so many see it as a large issue. Instead, projects use non-standard YAML or other problematic solutions only because "JSON doesn't have comments".
JSON5 is a good compromise.
For example,
JSON.parse(`{"foo":"bar",}`)
throws a syntax error.This is a typo, I meant "can't" not "can". Off course browsers don't support JSON5 or my message makes no sense whatsoever.
But I'm curious what you mean by "in the wild"? If you're using (producing) it, something needs to consume it, and you would probably have control over both in whatever project you were using it for.
{..., "birthday": "2018-03-25"}
If my server is located in New York City, and the user is in Sydney, then my server isn't going to wish them happy birthday in time.So maybe we could do:
{..., "birthday": "2018-03-25", "location": "Sydney/AU"}
But at this point we might as well use a standardized time format (UTC) with a timezone offset. Maybe I'm thinking too far into it? Year:
YYYY (eg 1997)
Year and month:
YYYY-MM (eg 1997-07)
Complete date:
YYYY-MM-DD (eg 1997-07-16)
Complete date plus hours and minutes:
YYYY-MM-DDThh:mmTZD (eg 1997-07-16T19:20+01:00)
Complete date plus hours, minutes and seconds:
YYYY-MM-DDThh:mm:ssTZD (eg 1997-07-16T19:20:30+01:00)
Complete date plus hours, minutes, seconds and a decimal fraction of a second
YYYY-MM-DDThh:mm:ss.sTZD (eg 1997-07-16T19:20:30.45+01:00)
where:
YYYY = four-digit year
MM = two-digit month (01=January, etc.)
DD = two-digit day of month (01 through 31)
hh = two digits of hour (00 through 23) (am/pm NOT allowed)
mm = two digits of minute (00 through 59)
ss = two digits of second (00 through 59)
s = one or more digits representing a decimal fraction of a second
TZD = time zone designator (Z or +hh:mm or -hh:mm) “_comment”: “blah blah blah”,
SimplesNot that that matters when applications take it upon themselves to re-save the config file in some kind of normalisation effort. Bye-bye comments, hope they're checked in somewhere.
I'm lookin' at you, kubernetes...
This has caused me so much misery in the past, especially since none of the tools will tell you which line the offending comma is on. Great, somewhere in my thousand-plus-line JSON file is a tiny syntax error but you won't tell me where.
Ended up having to regex for them. Didn't do wonders for my trust in JS tooling.
python -m json.tool < somefile.json
This will tell you where your file is messed up.There is also https://github.com/zaach/jsonlint
First of all, don't ever try to edit a YAML file by hand. You will introduce whitespace or other characters that will break the file, and you will not know until you run it and it breaks something.
The reason you will not know? Not all YAML parsers are the same. Some will interpret it correctly, and some will break. You'll have to get reference implementations of every "supported" YAML parser and run every config you have through them all, and diff them all, before you can trust them.
YAML may be easier to read than JSON, but its added complexity (the parser is significantly more complicated) and obtuse "features" are just not worth the effort. Not to mention, have you ever tried to maintain a very large indented YAML file by hand? Pain in the ass. Just shove everything into JSON files. The fact that it's so limiting is freeing, and everything can parse it. But don't edit it by hand.
And IMNSHO, you shouldn't use either YAML or JSON as a configuration language. They are for data structures, not configuration. If you want a configuration language, go get something designed as a configuration language.
- Allows code reuse.
- Allows configuration to be as dynamic as you want.
- Can use environment variables.
I suppose there are some cases where you can't trust the user in this way (running configuration code), but I think in a lot of cases you can, and it's generally more convenient.
For a sensible language, I would model something after Apache's. Simple, direct, easy to read, easy to write, easy to extend. It's like a server admin who barely knew HTML 1.0 wrote a config format. Perfect for the things it should actually be doing.
Another option is to take a simple format and extend it with another format or language. For example, you could add SQL to a simple file format, and suddenly tons of people can extend the config with some complex logic. But I also think templating and macro languages should generally die in a fire.
INI files aren't bad. They aren't a language, but they are good for simple use cases and a flat structure. Yes, you can have hierarchical section names, but it's a pain. If you want to use INI, you should probably use TOML. But there's very little incentive to add a TOML parser to a simple app when they could just suck in a JSON file. (Personally, I use JSON files, but only because I'm lazy, not because it's a good idea)
The biggest problem with things like Ansible is they'll give you enough rope to hang yourself. First you get defeated by whitespace. Then you get defeated by the stupid YAML rules. Then you get defeated by complexity like inheritance, namespace conflicts, and the shittiest debugging output ever. Then you get Jinja madness embedded inside Ansible madness inside YAML madness, and nobody knows how it works and can even touch it for fear of breaking everything. And of course, there is nothing that can parse it other than Ansible.
I think if Ansible had been TOML+Jinja it would have worked. It would have been ugly and clunky, but it would have worked. (The engine itself and their stupid rules about structuring your project should also die in a fire, but that's a different subject)
An example :
```
// This is a node with a single string value
title "Hello, World"
// Multiple values are supported, too
bookmarks 12 15 188 1234
// Nodes can have attributes
author "Peter Parker" email="peter@example.org" active=true
// Nodes can be arbitrarily nested
contents {
section "First section" {
paragraph "This is the first paragraph"
paragraph "This is the second paragraph"
}
}
// Anonymous nodes are supported
"This text is the value of an anonymous node!"
// This makes things like matrix definitions very convenient
matrix {
1 0 0
0 1 0
0 0 1
}
``` author "Peter Parker" email="peter@example.org" active=true
This is like XML attributes, which I've always found annoying to deal with in programs. It doesn't really map to any native data structure in most (all?) programming languages, so you need a special class/struct which supports it.Simply using something that maps directly to a hash map/object/associative array would be much better, IMHO.
Other than that, it looks like an interesting project.
SDL documents are made up of Tags. A Tag contains
* a name (if not present, the name "content" is used)
* a namespace (optional)
* 0 or more values (optional)
* 0 or more attributes (optional)
* 0 or more children (optional)
So it's like an XML node, but the `0 or more values` means it has a list/array for a "body".For a general-purpose human-writable structured data format, I guess the ugly nonstandard hack that is "JSON with comments" is probably good. It's certainly faster to parse than YAML.
As much as there is a lot to not like about YAML, it is the easiest one for humans to consistently write in my experience.
We don't judge Excel and Librecalc by how easy it is to open their files and produce valid spreadsheets without good tooling.
If they're working with structured data, why can't they use tools/editors which work with the structure, reveal it, and enforce it?
So when we have text with markup the text part is meant to be there for humans and the markup part is solely for computers. Now let's remove all text; now there's no content for humans at all, only for computers. How is this different from general-purpose data serialization?
(Some of the samples you give, like SVG, may not have any text content at all; it's basically a drawing language.)
Given that XML ecosystem has quite a few tools (e.g. several type description languages or a declarative data transformation language just to name a few) it's a very good general-purpose data serialization format.
XML's incredible verbosity is a problem for computers too. I've spent time performance-tuning message parsing code that had no good reason to be slow except that our use of XML bloated the data and decoding time by an order of magnitude or more compared to a binary protocol with a schema.
I've gotten incredible speedups just by switching to SAX parsing in those cases.
I've been slowly ripping out YAML support and converting configurations to TOML.
Then there is floats without a leading zero. Missing colon after the key. And yea, naked keys. The need to wrap the entire file in { } or [ ] is just icing.
Honestly I feel the most bare simple conf format of [first-word] [rest-of-line] is enough for many programs that end up using but never taking advantage of more powerful formats.
JSON is often minified though - you're going to need something to use as a delimiter
Actually, I think you could leave out the delimiter altogether and still be syntactically unambiguous, since quotes are required around keys:
"key1":"value1""key2":2.75"key3":true"key4":"whatever"
though this looks terrible and there's probably some edge case I've forgotten. (Also it misses the point of JSON in that it's no longer valid JS. I don't know whether that's important anymore since you should be calling JSON.parse() not eval() anyway.)The one major reason I could see to use "JSON" as a conf file is in trusted node.js apps because you can then easily embed functions and logic in them if you need more advanced/customizable configurations. And you can do it with full syntax highlighting in your editor. And comments, and trailing commas, and naked keys.
Of course this is no longer JSON, it's straight up Javascript config files. But it has come in handy a few times when I want to override standard behavior on a per config basis, and most of the file is still just plain key: val
it's a shame json doesn't support them though. oh well. would be nice to restart the universe and get all this right next time. :-)
JSON is great in terms of flexibility, but .INI files are really easy to read because everything is on the left side of the screen/window at all times.
<parameter>
<name>ApplicationName</name>
<value>WhizBang</value>
</parameter>
...
So that they could pass schema validation and still have some hope of extensibility.<parameter name="ApplicationName" value="WhizBang"/> ?
Your option is better, but XML is very (maybe too) flexible and is bound to be made a mess of.
ApplicationName=WhizBangIn the same way, if we receive a data transfer in XML and there is a schema, simple validation catches a lot of problems quickly. You'd be surprised how many times a company gives you a schema and then sends you xml which doesn't validate. In JSON, you have to write a program to get even basic validation.
Don't get me wrong: XML has problems, some inherited from html/sgml (entities!), and even more after serious abuse by consultants, archicture astronauts and enterprise vendors (SOAP! namespace overuse! 10 XML parsers in 1 app!). But it was also miles better than what came before and I feel the XML hatred pendulum has swung too far.
Today, JSON is in vogue, and I've seen enough IT to not swim against the tide. It is a reasonable solution for problems caused by XML abuse. Besides, there is value in going with the majority,even if it only fixes 80% of your problem. But I can only weep for the miserable date, numeric and comment support, and their endless stream of incompatible workarounds.
For your parameter example: You can't both strictly validate and have full freedom at the same time. Something has to give a bit. Some less horrible alternatives I've seen:
<parameter name="X" value="Y"/>
<subsystem name1="value1" name2=value2 ... /> , add newline for each attribute
<name>value</name>Edit: XML is just a proper subset of SGML by definition, hence it didn't introduce a single thing that wasn't there before. It only introduced XML-style empty elements and DTD-less markup, and SGML was extended in lockstep with XML to support these as well
XML is more than the part inherited from SGML, it's also the XML culture surrounding it. Namespaces are an example of something that created an XML dialect. And of course SOAP, which actually needs the WS-I standard to explain what parts of the WS-* standards to use or ignore, and how to interprete them. And even then 2 WS-I stacks will rarely interop without trouble. Lets not blame SGML for that monstrosity
<a:log /><b:log /><c:log />
can mean a math function, a text file that records what's happening, and a cut-off trunk of a tree and there will be no confusion whatsoever.(I really don't get how SOAP is relevant here.)
<element name="addressBook">
<zeroOrMore>
<element name="card">
<attribute name="name">
<text/>
</attribute>
<attribute name="email">
<text/>
</attribute>
</element>
</zeroOrMore>
</element>
http://www.relaxng.org/tutorial-20011203.htmlPlus there's a non-XML "compressed" version as well:
element addressBook {
element card {
element name { text },
element email { text }
}+
}
But I agree that XML like you posted is nasty, it's no harder to write <parameter name="ApplicationName">WhizBang</parameter>
instead if you need generic parameters, or <ApplicationName>WhizBang</ApplicationName>
if not.https://www.owasp.org/index.php/Top_10-2017_A4-XML_External_...
Not so much. Sexps don't provide a place hang "extra" information. It's been a pain point. While some lisps allowed decorating runtime things (eg objects with attributes, and symbols with property lists), their printed/readable representations were implementation dependent.
There's also a widespread misconception that Scheme is easy to parse. Numbers and all. It's actually very hard to get right. Real scheme parsers are quite large and hairy.
> XML was a complete and utter waste of time.
While XML was ghastly, there was an unmet need. There still is.
I think JSON is more efficient to write, but XML often ends up being more efficient to read due to comments and the fact that XML tags often give you better context. I think most programmers (myself included) tend to heavily optimize towards writability when we should think about readability a little more.
An example of this is ElasticSearch, where your queries are in JSON and often end up tons of levels deep - it is super easy to get lost in a sea of closing brackets, whereas XML would let you add comments in and the fact that closing tags have names in them would give you better context about what you were doing.
If your configuration file is so long it's unreadable in YAML, then maybe you need to break it up into more than one file? I can't imagine any syntax would be easy to read once you reach more than 100 or so lines.
Do any configuration file languages support type hinting? Adding (int) in front of a YAML key would be easy enough to read, and would keep some of the confusion at bay.
So, would this work?
ports:
https:
enabled: yes
!!int port: 443
Then if someone is copy/pasting and tries to use "blah" as the value, the !!int tag would cause the yaml parser to throw an error. Right? ports:
https:
port: !!str 443
to parse the same as {
"ports": {
"https": {
"port": "443"
}
}
}
in JSONIt just makes me wonder what the hell they thought they were doing all that time...
It's like designing a tool called YACC, and ending up with Yet Another Interpreter Interpreter!
It's like a standard for storing all your pornography in a folder called "Definitely Not Pornography".
https://en.wikipedia.org/wiki/YAML
>Originally YAML was said to mean Yet Another Markup Language, referencing its purpose as a markup language with the yet another construct, but it was then repurposed as YAML Ain't Markup Language, a recursive acronym, to distinguish its purpose as data-oriented, rather than document markup.
The YAML project was a convergence of several different efforts at information representation including people from Perl, Python and Ruby, each with our own ideas. I happened to be involved in the outer ring of the XML community, in particular a group SML-DEV where we were looking for a better information model more suitable to data serialization that would use an XML compatible syntax.
At that time, especially since serializing data with XML was all the rage, "ML" or "Markup Language" was commonly associated with data serialization. In fact XML is very inconvenient for actual markup, even though it derives from SGML (of which HTML is an example).
The "YA" part did deliberately come from YACC, the reason why is that XML (and we hoped YAML) would be the basis of domain specific serialization languages. Hence, you can think of it as a meta-language for building sub-languages, like the application I was working on, a serialization for accounting data.
Hence, that's the origin of YAML acronym, which happened to have a domain name open. In fact, you can look at archive.org starting in 2001 and you'll see the 1st public pass of YAML more closely followed XML like bracket syntax. I disliked it, but, it was what came out of the SML-DEV collaborations. Even so, the important thing was the information model, which was a simple typed graph and not an element tree.
The syntax evolved with many revisions after the model and goal were set, with lots of feedback to address concerns and usability tests with domain experts. The syntax got more and more lightweight, inspired from RFC0822 (email) while adding dashes for list items. Testing the syntax with domain experts, e.g. accountants and the like, was exceptionally important part of the process. Users liked this serialization syntax since it made their data "pop".
So. A year or so passes while we focus on getting things to work and helping people. Then, the product differentiation question comes up. Since XML had the dominant position in the data serialization mind share, how does YAML compare? Well, the YAML model was a typed graph, while XML was an element tree. One required no special libraries to manipulate, the other require a DOM to translate. But perhaps more importantly, it's because XML borrowed its model and tags from SGML and SGML was a true "markup language". So, XML was a "markup language" where it was impractical to do actual markup. Then, it dawned on us, well, of course a data serialization language isn't markup problem. I'm not sure any of this was obvious at the time.
Anyway, in a very fun chat, Oren pointed this out and then Ingy said: well then, YAML Ain't Markup Language! So, the new name actually represented what we had set out to do in the first place. Further, I would suggest that the industry understanding of how serialization languages are poorly supported by markup approaches (XML) is at least somewhat due to our name change and fun filled articulation at conferences.
I've only got one FOSS project using HCL but I think of its' bundled HCL config file is an attractive part of its UX: https://github.com/LukeB42/psyrcd
JSON is lacking in some respects but it's still really close to perfect for its use case.
Being a super-set of JSON is YAML's best feature.
I would never consider it for untrusted input though.
Still, I'm optimistic, especially now that Prettier is getting YAML support. [2]
Dhall has schemas, types, imports, and even functions (though it doesn’t let you write code that would e.g. loop infinitely).
You can even use an executable[0] to compile Dhall to YAML and JSON, excluding unsupported features such as functions.
[0] https://github.com/dhall-lang/dhall-lang/blob/master/README....
Octal 13 is decimal 11.
JSON should actually have the same issue. When I enter { 013: "11" } in the web console I get '{11: "11"}'. And YAML is backwards compatible to JSON.
That's IMO the actual problem of YAML. It could have supported a reasonable subset of JSON and not the whole nine yards.
Also JSON doesn't treat a 0 prefix on a number as special. There are no octal (or hex) literals in JSON. In fact, a JSON number literal cannot even start with 0 unless that's the only digit (before the period), e.g. 013 is not a valid numeric literal in JSON.
JSON doesn't have this issue. `{ 013: "11" }` isn't valid according to the JSON spec for a couple reasons, the important one being that multi-digit numbers cannot start with zero[1]. Try this in the console: `JSON.parse('{"11": 013}')`.
JSON.parse('{ 013: "11" }')
which should produce a syntax error (as already mentioned by siblings to this comment).If there was some deficiency with JSON5, just simply use JSON with comments. It's that simple.
JSON is one of the best things to ever come out of the CS disciplines.
For those that whine about comments in JSON, Douglas Crockford, the creator of JSON, himself said to do it.
>Suppose you are using JSON to keep configuration files, which you would like to annotate. Go ahead and insert all the comments you like. Then pipe it through JSMin before handing it to your JSON parser.
I keep configuration in EDN, which avoids all of the problems described in the article and has other advantages, too.
It's not as a big of a deal in Python because as others have mentioned, you usually don't end up writing large functions to begin with, and tab identation is supported. It still means that if you try to send a code snippet to someone over email or Slack, IM, etc it may not work because whitepsace may be trimmed etc. With YAML, it's way worse for reasons the author outlined.
Something may look good on paper (or in screenshots) but practical considerations need to factored in when designing a grammar, and there are myriads pretty-formatting utilities that could be used to that end if one cares for that sort of thing (see also: clang-format).
I think a big part of the unpleasantness comes from how non-obvious some things are. For example, why does the first automation example require a "-" in front of "platform" (under "trigger"), but the second doesn't? The comments explain when it's required, but not why? Shouldn't the parser be able to figure it out from the significant whitespace?
I found it so unpleasant that I'm using Node-RED[2] to do anything even vaguely automation related and have relegated Home Assistant to being the UI and communications abstraction layer.
In contrast, XML, while overly verbose, has a more reasonable structure. I haven't played with JSON much, but it also seems pretty reasonable.
[0] https://www.home-assistant.io/
Then there is the significant whitespace, difficult to remember quoting rules, etc.
I think the only reason it had gained any traction is because it is relatively easy to write multi-line strings in, so it is good for shell scripts, e.g. in CI configuration.
Someone should really come up with a sane alternative that works well for that use case though.
Everything that you can do in YAML + more, you can do in Home Assistant using Python. Have a look at my PyCon talk [1] from 2 years ago or check the available functions in the docs [2].
About the dashes. In Home Assistant when an option takes a list, you can omit the list if you are just passing in a single entry.
[1]: https://youtu.be/Cfasc9EgbMU?t=1038 [2]: https://dev-docs.home-assistant.io/en/master/api/event.html
Definitely not. I'd expect 0x42 to be 66. (Not kidding, if 0x42 means hex notation!).
Point taken.
XML -> doesn't record what programs think of as 'data'. So a program needs to convert it's data representation to XML and back again. Usually there is no formal spec. This is the exact same issue you have with databases. But at least a database has a formal representation and data types.
Part of SOAP is Microsoft trying to bolt schema's onto XML. SOAP seems to generate a lot of unhappy programmer noises. But my ex roommate said, 'well when you get it working it works'
JSON and YAML do but they don't have schema's to parse against. So programs need to do their own validation.
Why would you ever use YAML for user-provided input? At that point, it's better to just use JSON.
> Many other languages (including Ruby and PHP1) are also unsafe by default. Searching for yaml.load on GitHub gives a whopping 2.8 million results. yaml.safe_load only gives 26,000 results.
Maybe that's everyone's using JSON where it would be unsafe to use YAML.
> YAML files can be hard to edit, and this difficulty grows fast as the file gets larger.
And... this isn't the case for XML or JSON?
Ok, so reindenting a section might be a pain, but if your YAML is containing large amounts of data, maybe that data doesn't belong in that format if you're manually editing the YAML.
> especially since 2-space indentation is the norm and tab indentation is forbidden
Good. ;)
> And accidentally getting the indentation wrong often isn’t an error; it will often just deserialize to something you didn’t intend. Happy debugging!
Which is unlikely to happen if you're using a YAML library or only editing small-ish config files by hand.
---
As noted, YAML has a lot of quirks. As a configuration language, I love it and am used to the little edge cases. Could it be better? Definitely. But I would still consider YAML to be great in the domain where it excels: human-readable configuration. Using it to store and transmit large amounts of data, especially in ways where a human is manually editing the YAML, is a terrible idea.
1. If you're writing Lisp, just read S-expressions in.
2. Use Python or Skylark and have a step that executes it into a configuration format. Obviously this is not something you would want to use for a data interchange format, but no one thinks that they can blindly run a random hunk of untrusted Python. Right? ...Right?
I once did about 98% of the work to support Schemas properly in Psych (https://github.com/ruby/psych) but the maintainer said he didn't want to "maintain it".
So, there you go. What else can one do? You can't blame the spec for decisions of implementors.
(That's not to say the YAML spec couldn't use some improvements, but it's far from "not so great".)
But the author is mostly right, when adding support for YAML to my code I spend a lot of time disabling all of its nifty misfeatures. Wish it was simply an indented JSON with comments and fewer quotes.
[1] Shameless plug: https://www.frontiersin.org/articles/10.3389/fninf.2018.0001...
Well at least it's the least worst then, do that make it the best?
Frankly I hope the future will be make with indented languages. Curly braces languages often allows too much liberty, and it's annoying. The fact that go enforce the curly brace style is really the tipping point of curly brace languages.
Granted there need to be some good compromise for ambiguous details when parsing an indented syntax, but readability matters much more than anything to me.
if foo
...
else
...
end
It solves some of the issues that you can have with Python, while at the same time also avoiding the whole nonsense with the braces. I find it's a good trade-off between the strengths of both approaches.https://news.ycombinator.com/item?id=17309132
>Recently I've been working on kind of the converse of this problem with JSON and spreadsheets, and I'll briefly describe it here (and I'll be glad to share the code), in the hopes of getting some feedback and criticism:
>How can you conveniently and compactly represent, view and edit JSON in spreadsheets, using the grid instead of so much punctuation?
>The goal is to be able to easily edit JSON data in any spreadsheet, copy and paste grids of JSON around as TSV files (the format that Google Sheets puts on your clipboard), and efficiently import and export those spreadsheets as JSON.
>[...]
Since I wrote that post, I've cleaned up and refactored the code into a portable little library that will run in the browser, or inside of Google Sheets:
https://github.com/SimHacker/UnityJS/blob/master/UnityJS/Ass...
Here's an example spreadsheet (also check out the examples in the other sheet tabs):
https://docs.google.com/spreadsheets/d/1nh8tlnanRaTmY8amABgg...
I haven't come up with a trendy marketing name for it yet (except for the source file name sheet.js), because I think it's important to first discover what it is by using and refining it for a while, writing some documentation to explain it, and getting feedback from other people (work in progress), before trying to name it -- otherwise you might end up calling it "Yet Another <something it's not>".
If I don't enjoy a movie, I'm under no obligation to suggest another. It's an opinion. It can stand alone.
YAML can actually be pretty useful. For example, recently I wrote a tool to generate OpenAPI/Swagger files from Go source code, and outputting that as YAML works pretty well, as YAML is quite easy to read for that.
YAML can also be a good choice to serialize some things to disk, like program state. JSON can also be a good choice for that, as can TOML.
But for other things ... it's not so great. I should probably write an article detailing this at some point, but ...
$ ls -1 /data/code/arp242.net/_drafts | wc -l
64
So much stuff I need to finish :-( I've added your suggestion there though!Furthermore, all of the "better YAMLs" that exist solve a different subset of issues based on the whims of the author. I like indentation-based syntax for config files (though not for programming languages, go figure), so half the alternatives look worse rather than better to me, and reasonable people can also disagree on things like when strings should require quotes and what should be a valid hash key and so on. There are so many bikes to shed that I don't see us ever settling on an alternative without major buy-in from one of the big players in tech.
Until then, I'm happy to let YAML win. It's just not broken enough for me to get worked up about.
How does this entire section not also apply to JSON?
https://docs.bazel.build/versions/master/skylark/language.ht...
0 - https://github.com/lightbend/config/blob/master/HOCON.md
The syntax resembles JSON, but simpler, and keys don't need to be quoted. Trailing commas are allowed, and so are multiline strings.
You wouldn't, but the language is lightweight enough (according to Wikipedia, the interpreter is ~180kb compiled) that including it as a dependency probably won't matter. It's less overhead than an interpreter for JSON or YAML at least.
>Is a parser library available for mainstream languages?
It's used commonly in game development so yes for C and C++, and the site[0] mentions "Java, C#, Smalltalk, Fortran, Ada, Erlang, and even in other scripting languages, such as Perl and Ruby" as well.
The first criticism, where he's embedding exectutable code in YAML I kind of have to agree with - that seems crazy. I don't know why YAML would support this.
The remaining criticisms seem to relate to (a) the spec overreaching in terms of complexity and (b) differences in implementation, which I guess is some kind of an extension of (a).
I maintain however that YAML, JSON and XML are different.
If you want to make me feel cross and insulted give me a JSON file to edit. I think JSON is probably the best commonly used format for M2M and storage serialisation.
I wouldn't want you to use YAML for that though. There's two many different ways to do it and any kind of ambiguity never makes for good M2M.
For configuration-files it's great though, as long as you stay away from some of the more exotic features I suppose. It effectively provides a "user interface" of sorts by which your users can specify non-trivial configuration details.
The complaints about overlong and overcomplex yaml files could be extended to other commonly used formats.
With regards to XML, I'd say that YAML provides all the features you'd want to use from that format in a format that's easier to hand-edit as text. XML is "okay" for M2M but probably better for document storage where you have some kind of custom editor.
It's horses for courses. I wouldn't want to use YAML where I'd want to use JSON, or use XML where I'd use either of them, and neither of these have the semantic richness of XML either and so wouldn't be appropriate in whatever spaces XML should be used (which is far more limited I suspect than its current span of applications).
Ultimately when each is used to its strengths they're not interchangeable formats.
> YAML is insecure by default. Loading a user-provided (untrusted) YAML string needs careful consideration.
That is trivial. Everything you execute that come from outside is potentially dangerous.
> It’s pretty complex
I would say, that YAML is more expressive and has more features. On the other hand is true that TOML is a valid alternative in many use cases.
> 3.5.3 gets recognized as as string, but 9.3 gets recognized as a number instead of a string
That is correct, in my opinion, 9.3 is a float while 3.5.3 is a version. If you want both to be strings, use quotes.
It is. At API design time, it would have been trivial to replace them with `load()` (which does the same as `load_safe()` now) and `unsafe_load()` (which does the same as `load()` now) and probably avoid this pitfall altogether. Now? Much more difficult to solve.
I think maybe YAML just went a bit too far: because everyone hates defining key names with quotes when it's unnecessary 99% of the time, but it would have been enough to relax that rather than go all the way to suggesting relaxing all quotes and then attempting to infer value types.
I have not yet taken a look at strictyaml, but after years of use the spec definitely needs YAML, The Good Parts Treatment.
One thing the author did not mention was how slow the out of box the Python YAML parser can be. This can be sped up with a call to libyaml, but then you lose the safe_load method.
I used the basic strictyaml.load function, without any schemas.
README page said speed is not a current priority, and that appears to be true.
I've done informal polls on it and every time it's an even split between .yaml and .yml
- .jpg / .jpeg
- .tif / .tiff
- .htm / .html
- .cpp / .cxx
It's frustrating at times, but not all file formats have One True Extension.
As for C++, I'd also blame Microsoft. The plus character (+) is a reserved character for MS-DOS, so the obvious extension ".c++" couldn't be used (nor could the case-sensitive ".C" extension). So people either toppled their plus signs (".c++" becomes ".cxx"), or replaced them by the first letter of "plus" (".c++" becomes ".cpp"), or treated them as a repetition sign (".c++" becomes ".cc").
Eventually I tracked it down to an ID inside a yml file. Turns out the live environment was running in 32 bit mode which interpenetrated the number as a string.
I love this bit explaining why schemas are essential to having a human-friendly config file:
[1] https://github.com/servo/webrender/
[2] https://github.com/ron-rs/ronPersonally, I don't understand why all these projects have defaulted to using such an esoteric markup language.
But wow do I wish JSON supported comments in some form.
https://github.com/EamonnMR/Zond/blob/master/RedShift/src/co...
It ends up looking like this: https://github.com/EamonnMR/Zond/tree/master/RedShift/assets...
It may seem that way, but it's not. Re: cryptographic functions
- when loaded safely
- when used by humans to edit obviously structured data of hashes and arrays
- and the humans are trained to double-quote potentially problematic entries.
This wins over xml or Json hands down.
I know things like TOML and JSON, while they have good DX, for me YAML has the better UX.
> python: 3.5.3
> postgres: 9.3
> {'python': '3.5.3', 'postgres': 9.3}
Surely that's reasonable?
If I had to write OpenAPI files manually I'd probably choose to jump out of the window.
That actually made me laugh out loud.
Seems like YAML tried too hard to be predictive of intent. I never got into YAML myself simply because it seemed like "JSON, but less ubiquitous and more hassle to find libraries that support it"
also what was wrong with ini files?
http://codesolvent.com/config-node/
it is however difficult to pull off and requires productization, in other words not low-level tooling in a text file.
The key is that you have to have a way to create such graphs beyond code (ie new Resource({"attr":value}..)).
Text file formats are always going to be an issue the moment your requirements are complex.
Here's another example of what I am referring to:
Edit: I actually did download whatever this thing is. What is this thing? The README is Jetty's README. There are dozens of dirs with crap^W code in them.
It really is an answer to YAML/JSON, surely.
the "So Much More" button is a link to the docs:)
Here's the ConfigNode doc: http://codesolvent.com/doc/config-node/
The platform doc: http://codesolvent.com/doc/webapps/
If you're interested, send me an email (in my profile).
Yeah... No. It looks like a link to yet more marketing-speak, and the page it links to answers zero answers as to why I would ever want it.
Hiding the product README in some nested folder in the source code is also a brilliant idea.
So, nope. I’m sticking to my YMLs and EDNs.
The documentation link LITERALLY shows how to use the product.
Nothing is hidden, if you read the platform docs (http://codesolvent.com/doc/webapps/) you see it says:
"Solvent is an integrated platform that combines an application container (jetty), a middle-ware and a developer environment to provide a complete solution for delivering web applications."
ConfigNode is something you'll put on a server as a config management environment, the output can be json/yaml/xml..etc
I think you might find it quite unique if you just give it a chance (no marketing) :)
Bingo.
I’ll never put it on a server for very obvious reasons
Is this supposed to be a feature? One of the great things about simple config files is that you can use standard GNU tools to view edit, and diff them, you can put them in source control, you can be sure that you can edit them on a remote server no matter what's installed, etc.
Eliminating all those benefits would require an extraordinary jump in functionality as a tradeoff, a jump in functionality that most things frankly don't need.
Here's another example of ConfigNode used to manage Akamai configurations:
Akamai configurations can be very complex and need to be maintained and managed, you can't really do that effectively using text files.
With a binary file, you of course have to be using something specific to work with them. And if lots of people were doing that, that would in turn drive a lot of specific, good options to choose from. The tools would be better at their specific task. They could provide actual _user interfaces._
But there's not enough standardization in user interfaces for to reduce the cost of relearning each tool, and the tools would need both an interface and an API to automate them (we don't tend to get both for free), and we don't have anything great for chaining together APIs. So text files it is, which kind of provide these things, but they don't extend too well.
The trend in computing tools is to slowly invent what you could get easily with more specific formats... using text files. Automatic formatters, so you can pretend your project files are really the AST you care about. Smart IDEs with autocompletion, because you're not really typing arbitrary characters. IDEs that will collapse a lot of unnecessary information for you, like showing only the first snippet of JSDoc. Type systems that show you what's available and what you can plug together in a sane way. Version control that pretends it knows how to solve the problem of diffing/merging.
However the output comes after evaluating the object graph. In other words it doesn't just reassemble a bunch of static values but rather actually executes objects (think POJOs) to produce the fields that make up each object.