XML is better than YAML – Hear me out
changelog.com
changelog.com
I get that XML is about as sexy as mainframes, and that a lot of folks here probably have PTSD from working with Java/Spring web apps, but YAML is about the worst of all worlds.
Though I think the real problem is that real-world configuration files are way too complicated for a simple/dumb/logic-less representation like a .ini/.conf file, so someone thinks to add some logic to is - which is just config-as-code. In a terrible programming language.
If you want config-as-code (and you want to!), just do it properly and use a proper programming language for it. Don't care which one, be it JavaScript, Python, Go, PDP-11 Assembly, or Rust. But please stop with these half-measure DSLs that just don't cut it.
https://bitbucket.org/snakeyaml/snakeyaml/issues/561/cve-202...
He claims with a straight face that every software using his library only loads YAMLs from sources 100% trusted to execute code. He is given example after example to the contrary, which he ignores instead opting to constantly blame "low quality tooling" for generating "false reports" about his perfect software. People can be really weird sometimes.
There are two categories of constructors. One is for data that should not be executed, the other is for trusted data that should be executed.
There are two libraries. One has default constructors that can execute data, the other has default constructors that don't execute data.
He's saying to rtfm and choose the library with the correct defaults, choose the correct constructor from that library, and stop trying to take away the choice.
Nobody was trying to "take away the choice".
The problem is that you have to explicitly opt-in to be safe. If you followed the code snippets from the README, your application would be vulnerable to RCE without you realizing it; as people pointed out, it would be more secure to have Constructor (safe by default) + DangerousConstructor rather than Constructor (unsafe by default) + SafeConstructor.
His argument was that "100% of applications using SnakeYaml do not accept untrusted data".
As he explained, this library is, by design, convenient by default. Those seeking safe by default should consider using the other library.
"take away the choice" is my summary of several comments that would have the feature removed. One was about how its existence is a vulnerability if file access is compromised. Another was about how code execution is not in the spec. And so on.
He was speaking literally; he even rejected several of the provided examples because "users have to login first, therefore the data is trusted" which isn't an argument that any security-conscious person would make.
> As he explained, this library is, by design, convenient by default. Those seeking safe by default should consider using the other library.
That is a negligent mindset to have. Log4J added the ability to execute arbitrary dns and ldap calls for the sake of convenience, which resulted in one of the most consequential vulnerabilities of the past decade.
Opt-in security is dangerous and should never be the default — especially when the feature in question is executing arbitrary input.
"Should" statements are always relative to what you value. Clearly he thinks this trade-off is fine for him. His other library accommodates your security needs but this one accommodates his convenience needs. Can the man not make something for himself?
I assume it would be costly for him to make and propagate the changes. Maybe money could persuade him.
That doesn't change the fact that it's a poorly designed API that's insecure by default. There are countless situations where people are inadvertently exposed to risk via transitive dependencies, at no fault of their own.
> Can the man not make something for himself?
He did not make it for himself, he made it to he consumed by others. SnakeYAML is a widely used package.
Making this change would cost him something he values without giving him something else he values in return. According to him, Snake Engine already provides a default safe solution, so he's not leaving anyone without a remedy. It would cost you to switch, but you would get something in return. That seems fair to me.
It's like blaming the browser vendor for an XSS vulnerability rather than the backend implementation not using HtmlEncode.
Here's just a couple examples off the top of my head:
- `$variables` in bash are subject to arbitrary code execution via word splitting without escaping
- PHP register_globals
- PHP, express, and some others parse `?a[b]="foo"` in a query string as an object, allowing for prototype pollution or other exploits
- string concatenation for SQL + escape_string being the default for years
- perl array expansion in function calls
- XML entity inclusion on by default allowing you to read arbitrary files
- log4j executing arbitrary code inside its logs
- passing a variable to printf's first arg
- no difference between escaped and unescaped tags in php
- xargs splitting on whitespace
- yaml allowing arbitrary code execution (it got rails good!)
and there's probably loads more.
Starlark was designed for this purpose.
People eventually want logic inside their config objects. We can argue how much logic but at a certain point, just using a popular programming language just makes sense because it’s familiar. I’d like my coding expertise translate into this part of the ecosystem.
The only downside is that if the objects in question happen to be properly documented in typescript, you will never want to go back.
If only there had been a formalized side by side from the start between JSON, the clean serialization format, and the JSON superset for human authoring (comments and quotes anarchy) that people have been informally reinventing hundreds of times...
Instead, you need to describe the structure of the XML, have preknowledge of prefixes and meanings for namespaces, have to deal with CDATA crap, have directives, config-in-comments, and hosts of other annoyances.
XML sucks. I programmed from 1995 to the present. XML sucks. YAML is far far far far far superior.
The ONLY good thing about XML is XPath. That's it! XSLT? awful. schemas and other validations? horrid.
XStream (java library) was the only thing that made XML usable, and the second JSON (and later YAML) came out, I dropped it immediately.
The problem with this thinking is that you, personally, are then forbidden from arguing for the use of a strictly typed language for development because it’s the opposite position to the one you’re holding here. The exact reasons we use languages like those are the same reasons we should be explicit with our schemas. It’s unfortunate that many people try to argue both sides due to the convenience, as you say, of a single line parse, when years of experience has taught that duck anything is a bug fountain. (Not saying you are arguing both, by the way, it’s just common.)
Try reading back your gripe with the following in mind: do I have a stronger complaint than “it’s difficult” here? I think you’ll find that you don’t convey one effectively.
Robbing the future with deceptive over-simplicity - by creating a bunch of future difficult debugging scenarios and possibly footguns - is the far worse evil, than missing out on the maximally-convenient onboarding (which can be foolishly optimized for, for the sake of short-term popularity). All such crap-tastic solutions will eventually need to be replaced again anyway, creating an endless, hellish, slow churn.
Computers that understand the shape of your data are very helpful friends when you’re pursuing goals like data locality.
"Don't want to have to create a schema? Normalization is confusing? No problem, just chuck JSON in here instead!"
There's this kind of stuff but it's niche and more convention than actual schema.
https://www.xml.com/pub/a/2006/05/31/converting-between-xml-...
You know <key name="blah">key value</key>. It just highlights the extra verbiage, and pushes people towards "just use json/yaml".
You want to access this data for any of dozens of reasons. Graphs, logs, data points, transformations, data feeds, whatever.
I can do that task in json and yaml 10000% faster than with XML. With XML, you may have a schema (hope the document matches the putative schema!). Oh the schema is an http reference? Hope that still exists out there, the internet never breaks links. If you don't, well shit, is this tag beginning a list or a "subdocument"? Am I REALLY using the DOM api to step through nodes and attributes and CDATA? Guess I have to. There goes a day of coding.
Oh, in JSON and YAML, it's ONE LINE OF CODE to get it into something I can easily read, manipulate, analyze?
- say you have an upgrade program. it just needs to read in the old config file, rename some keys, add some new default values, etc. JSON/YAML? I can do that in stupid-simple code. XML? Well, I better hope there exists a library that loads this shit for me in my preferred language, or otherwise lots of fun with DOM. I forget, can I use regex to parse XML? (that is a joke)
- say I want to serialize an object graph pretty quickly for over the wire between languages. Do I want to write a complete XML mapping in two languages, or just do the one-line serialize, one-line deserialize? Yeah.
- say I want my config files to be somewhat extension friendly for plugins / extensions. XML parsing code? Yeah, that will be a ton of custom code. YAML/JSON deserializing to a map/dictionary? Oh, look at that, extension friendly code. Allow them to specify whatever json/yaml struct in their plugin section and pass it to the extension.
This stuff happens with such frequency that I never, ever think "man I wish this was XML".
Do YAML/JSON have some issues? Do I wish XPath and some XML features have json equipvalents? Sure ... very occaisionally. Actually, never.
I'm pretty much consistently on the strict/type-safe side of the "should we have a schema" debate, but there are better options out there to maintain a consistent schema for either data interchange.
JSON is simpler to map, faster to parse, simpler, more lightweight, and less dangerous to add to an online app[1]. You can also use a schema like JSON Schema for inter-app compatibility. It has replaced XML as the standard data interchange format for a reason. It's not great for configuration files, and it's definitely being overused nowadays, but it is a solid data interchange format.
Then you've got binary formats like Protocol Buffers which are even more lightweight and faster to parse and (generally) have schemas that map better to typed languages.
I think the OP has it right: XML is very well-suited as a generic document format. I wouldn't compare it to YAML, because in a perfect world they shouldn't compete in the same categories: Nobody should use XML for configuration or YAML for documents. And I also agree that there are better formats than YAML. I like the ease of writing indented multiline strings in YAML, but the fuzzy typing is pretty terrible. At least YAML 1.2 fixed the Norway problem.
JSON is no simpler (nor more complicated) than XML, if you're using a library. It certainly isn't faster - a SAX parser is faster than a JSON DOM parser (and JSON streaming parsers equivalent to SAX is rare).
> It has replaced XML as the standard data interchange format for a reason.
the reason isn't technical. It's competency (or lack thereof). Most interchange formats are for websites in browsers, where JSON performs well, since there's no native way to parse XML in the browser. So that mindshare from the web has leaked out to other arenas.
Examples:
1. When taking over a project, developers glanced at the code and decided it would be better to spend 6 months rewriting from scratch. The end result was not more readable than the original solution and introduced a new set of issues.
2. Many put too much emphasis on the worst case scenario and do not consider the average case. I worked a lot with many different XML formats and most of them were OK. Not "fun", but simply OK. I have to admit that I did struggle with some complex files, but there were plenty of times where the XML was simple, readable and easy to work with
3. When comparing programming languages they often focus on a few features and don't think about productivity in general. Languages like Java can actually be very productive, even if your favorite language can reduce null checks.
Not even tagged unions can polish that turd.
If you want full imperative just use the AWS sdk or Ansible but I think people have realized how not maintainable that pattern is.
Or generate it, or use something like nix
I designed a complex data acquisition system and after a lot of research, I settled on XML as the only viable option for complex user configuration, that is both readable and rich in content.
I then built a UI system that works with the XML and generates cofig docs with ease.
Sure, XML has been used in SOAP like systems, and it rightfully gets a bad rap, but that is more on the user than the technology.
A hierarchical document with values and attributes and custom tags? That is an almost DSL in itself.
Especially editor integration works great as opposed to YAML which is so ambiguous that it IntelliJ IDEA constantly breaks indentation during copy and paste.
I lost you there. Not that I'm criticizing you since I've gone the same route and built a complex UI tool to manage said config (complete with XSD schema validation and config schema migration using XSLT for version upgrades).
But now I realize that modern developers don't want a GUI to manage their config. We want to store it git, review changes and perhaps even write our own automation and templating around it.
YAML is certainly a flawed format for most of these purposes, but so is XML. It is unnecessarily verbose and it carries a lot of complexity which was designed for a highly-extensible generic document format, but not for configuration files. XSD, Namespaces, Entities, Embedded DTD, CDATA blocks... You can't just ignore all of these, and there are very few parsers out there which work on a well-defined subset of XML. And even there, the whole attribute-vs-child-element choice is a giant distraction and constant source for unnecessary bikeshedding.
YAML has serious ambiguity issues, but there are better alternatives that have great library support like TOML. We don't have to go back to the excesses of the early 2000s and use XML as a configuration format.
Ill tell ya hwhat... I made some simple functions; one to load any yaml/toml/json from file, by just looking at the extension (and providing an override for not normal file extensions). Another to output any data as yaml/toml/json.
I defaulted to using toml for my project, but have provided ways for everyone to be happy, with zero mucking around.
Aaaaaand for the not so subtle promotion, https://pypi.org/project/atckit
Located in the UtilFuncs class.
In any case, the library i linked has the functionality to load any of the three configuration types, so while i prefer toml, if you grabbed my (under development) project, and hate toml, you can use yaml or json, if you wish.
And, on that note, I think that if I could pick as single thing that irritates me the most about YAML, it's that it isn't actually a ML.
The problem is that XML is ill-suited for configuration files.
Why?
I don't understand your comment.
Compare ansible vs helm vs github actions vs literally anything with nontrivial config.
Note that if you asked me for real, I'd be a dhall proponent, but most normal people look at me funny the first time they hear it and then the second time they see the syntax.
If we're limiting ourselves to just JSON, INI, XML and YAML as potential choices, I get why people cling onto one of these suboptimal choices and then fiercely defend it, but there are other options. There's libconfig, JSON5, Dhall, various interpreted languages...
Sadly these alternatives are far from mainstream and have no support in standard libraries, so I think most developers will continue to pick whichever common option's issues they can deal with.
But at the end of the day that's not possible because reality intrudes. You give up a lot for that lack of verbosity, not everyone will agree that it's a worthy tradeoff.
“YAML ain't markup language” is kinda the whole point.
First draft: https://yaml.org/spec/history/2001-03-30.html
Last draft calling it that: https://yaml.org/spec/history/2001-12-10.html
First draft with the new acronym: https://yaml.org/spec/history/2002-04-07.html
That is, if the value is something you would expect to be able to show to a user, then it probably shouldn't be at attribute. If it is a value that changes how you would show it, then attribute makes a lot more sense.
I also think a strong adherence to any preference between attribute and markup is a touch too dogmatic. The only real distinction at the language level is if you allow children, or not. If it can have children, it pretty much has to be a child and not an attribute.
<foo>
<bar>1</bar>
<bar>2</bar>
<baz>3</baz>
<baz>4</baz>
</foo>XML is just to verbose, making it harder to parse in my brain.
YAML is whitespace sensitive which I hate in general.
I think at some point 24 out of 28 Gradle projects I had access to at a certain customer had variations in either kotlin/Groovy style or the way they did or didn't use variables, how they did or didn't do loops or maps and what not.
With Maven you (or someone who know Maven) can immediately look at a rather small, very standardized file and start making educated guesses and so can an IDE.
With Gradle you sometimes have to run it to actually know what it will do.
On the other side I fell in the trap of trying to overcome the limitations of purely declarative config formats by using jinja templates, which also ended up being a very bad idea and a maintenance nightmare.
For most projects, my approach is now to try to be as standard as possible compared to the particular community in the tech at end, and resist the urge to be smart or cute (hard!). Configuration always sucks, and I now prefer to just suck it up and get done with the config part, rather than loosing time reinventing the wheel, ending with a config that still sucks _and_ no one understands.
(More seriously: with Maven shorter and more boring is a sign that everything is correctly configured. Maven works by the convention over configuration principle so if you don't configure something it means it follows the standard. Which again means if you see someone has configured for example a folder or something that usually isn't configured it means they have put something in a non standard location.)
yaml's feature set is also too broad, forcing safe loaders
> yaml's feature set is also too broad, forcing safe loaders
Could you elaborate?
I think you can also make circular references in yaml effectively making a zip bomb
This is why you have yaml.safe_load in Python and SafeConstructor in Java's snakeyaml package
The end result is that you should not use yaml to handle untrusted data unless you are also explicitly handling it safely
It's possible someone could come along and write a json library that would support this, but somehow we have made it this far without it and that's a good thing
The point is that yaml and xml both have side effects in the form of require and eval that json won't, and frequently people are unaware of this
https://owasp.org/search/?searchString=json
Perhaps yaml and xml have _more_ ways to inject behavior into an application, but I would still not consider JSON safe in any way. Why would JSON.parse() even exist if `require()` and `eval()` were safe to use?
not all configurations are maintained by developers, or at least theoretically they shouldn't be.
I think, like almost all software, it depends on personality and other factors like how frequently someone has worked with a language/tool/etc.
If someone has the traditional Unix "read all the man pages in their entirety" kind of brain, they'll probably never use a GUI.
If someone has more of a "learn by example, then refer to the docs to cover cases the examples didn't handle" brain, a GUI can be a much easier way to clearly expose all the potential functionality of a tool. It's much easier to understand (IMO) than e.g. a template configuration file with every possible option present but commented out.
No. You need to define what configuration files are valid or not, but a subset of XML on the basis of syntactic and structural alternatives is lame and nonstandard.
Get a library that parses actual XML, allow anything reasonable as input (for example, no external entities, DTDs, PSVI etc. to keep the file self contained) and "flatten" CDATA sections, entities, namespace prefixes, idrefs etc.
I don't feel the same ease with JSON or YAML though.
I imagine the same is true for people that don’t like XML. There’s just many more people that find it easy to read JSON, even though there are a few that find it easy to read XML.
It looks like you're experience is mainly with XML as an underlying format, with human only dealing with it either at the coding level or through a tool generating the needed confs. In that kind of scenario I'd wager any coherent file format would probably work, even if the configs where encoded in brainfuck in the final step.
XML gets hated because we also had to read it and hand edit it as humans, when dealing with system configuration files (the source that will generate the rest of the configuration) and other upstream documents that are the base input for the system to read downstream when the GUI aren't available for that.
I've worked with a Symphony code base that used XML for all the routes and DI declarations, and yes it was an utter pain to write an edit tags for such simple and repetitive configurations when even an ini file would have been good enough. And god have mercy of the guys that put CDATA sections in the middle of that just to be sure CJK chars wouldn't accidentaly trip the syntax.
The nice thing about JSON, TOML, and YAML is they have implicit structures for arrays and key/values encoded within them.
XML has a lot of different ways of representing data that way, which is what makes it a challenging configuration file format.
That's as unambiguous as you could possibly make it IMO.
i think you're mixing representation of xml data vs the representation of them in a programming language.
XML does have arrays. They care called child elements.
["string1", "string2"]
to <list>
<e>string1</e>
<e>string2</e>
</list>
then each element has about four bytes overhead (<e> instead of " and </e> instead of ",) plus some overhead for the list itself that may be offset by putting the name of the list itself into the element.However, the issue is that you have to write a custom parser. There is no direct mapping between your data structure and the XML file. This developer ergonomics is a big win for JSON and consequently YAML.
i think that's by design tbh.
it's only a big win for JSON (and YAML) because the default case works OK - but every time someone has a problem parsing numbers in JSON (because the value is bigger than Integer.MAX in the host language), this is the cause.
Take any random REST API for example. If it returns JSON, you can integrate it more easily than if it returned XML. If you need special cases like large numbers (or date-times), you handle only those.
With JSON, you can mostly do the same. Such that I don't necessarily see this as a huge advantage of XML, mind. Having a schema does have some advantages, though.
<foo>hello</foo>
<foo>world</foo>
the <foo> and </foo> serve the same purpose as the double quotes in "hello",
"world"
with the added benefit that the type system can be much richer (i.e. not everything is just a nondescript string value).And you don’t even need a comma to separate the values! ;)
XML claims to solve the problem of attributes vs children but then falls short at the first hurdle by not discerning between a single complex object as an attribute and an array of complex objects as children.
JSON and YAML do not have this problem as they are explicit in their representation.
YAML example:
parent:
child: name
vs parent:
- child: name
Try converting each of these to JSON. The former will give you an object property called child, the latter will give you an array property called child with one elementYou're either providing complex objects as properties or you're providing a list of complex objects. Worse, you can have a combination of both. Without a schema it is not possible to infer whether either or both is happening.
This seems like an absurd thing to say. XML went through an extreme bout of popularity. Maybe you could plausibly say that people have been soured on XML by ill-conceived uses for it that don't demonstrate its strengths... but you think most people have never worked with it? Come on.
https://truelist.co/blog/software-development-statistics/
So, according to the link - The average software developer age is between 25 and 34 years.
I think we should also define - are we talking about people who have worked with XML because they made a google site settings xml file OR people who have done serious work with XML and know what they're talking about?
First type - pretty much everyone.
Second type - not very many. I'm pretty much the only person who knows anything about XML wherever I go.
if you are 25 you have probably not done anything with XML or at least not anything important.
If you are 34 you might have, but I mean the last time I did anything really important with XML was 2013, I did a few other things since then because I knew XML was the best solution but that was me or because there was a very niche thing I was doing and the company was providing an XML api.
I bet most of the age 34s have not done anything meaningful with XML either, even though if you are 34 I suppose you probably had some ticket at some point that took you a week and you thought wow, my extensive experience with XML now gives me the right to grouse about how bad it is! If only everybody knew as much as I the world would be a better place!
on edit: my example of google site settings file is an example of some trivial usage, not meaning that pretty much everyone has done that exact trivial usage.
How was it more readable or richer than JSON? There's more stuff in XML, but in my experience that stuff doesn't actually help you any. The schema validation has a lot of detail, but since it can't access your actual system you can't really validate in that detail (like, maybe you can validate that an ID is between 6 and 8 digits long, but really you just want to validate that it's an ID for something that's present in your database). The distinction between attributes and nested tags feels like it should let you express more, but in practice it usually just gives you two equally reasonable ways to write the same thing and causes more confusion. Comments and non-tag text nodes feel nice, but complicate your parsing more than they're worth.
> Sure, XML has been used in SOAP like systems, and it rightfully gets a bad rap, but that is more on the user than the technology.
If one person uses the technology wrong, it's a problem with that person, but if most people use the technology wrong, it's a problem with the technology.
But yaml has the exact same problem. People use it where they shouldn't.
This is just an example.
Besides Microsoft frameworks that use XML as configuration, and Java frameworks and servers using XML for configuration (Tomcat comes to mind), no one else uses it. Maybe traditional software whose programmers don't know any better.
So yeah, I think the world doesn't like XML for really good reasons.
So pretty much all enterprises use it. Not bad for tech not in fashion.
There is a subset of XML that a decent language for some use cases. In particular it is good for documents whete you want a _markup_ language.
But I've seen it used for a lot of things where it wasn't a great fit.
And xml has too many features, which leads to implementations that are inconsistent with each, often slow, and have security problems such as xxe.
I wish that there was a standardized simplified xml format, that would avoid many of xmls problems and meet the needs of 90% of applications where xml is a good fit.
this plus, like GP said, java ee xml bloat culture made it too much of a pain
And I haven't changed my opinion since then, having worked extensively with XML: it is a plague that was brought on into the world and it needs to be killed with fire. I understand that when it was released there were no other alternatives so "it was better than nothing", but it should have died right after JSON was invented. But no, why have cleanly formatted file when you can have an XML one...
https://rimworldwiki.com/wiki/Modding_Tutorials/PatchOperati...
XML has its place, but I wouldn't want to replace my yamls with XML.
* JSON doesn't handle multiline strings.
* JSON is not especially readable (no structure enforced, braces and double quote mandatory).
* JSON can't load YAML but YAML can load JSON.
(lots more)
It's annoying though because it's so mature compared to everything similar (even YAML, which I obviously quite like); like if you make an XSD for your config files everyone gets free editing features (not just validation, completion as well, even completion of attributes)
Those are used in JavaScript to give you the option to provide an expression in place of a single term. But that's not even allowed in JSON, so the commas aren't serving any purpose at all.
The problem comes when you try to use it as a programming language in and of itself (which MS encourages because it's a route to vendor lock in). It's a shitty, shitty programming language with shitty debugging tools.
To your second point: yes, yes, and very much yes.
Creating complex release workflows within GHA is hell given the debugging situation and that's what we're using because 'it's there, use it, everyone's using it'. Send help
Templating and significant whitespace is the worst. 10 commits in a row of "error on line 1" until helm finally parses what you want.
I am a software developer with extensive distributed systems experience. I like thinking about systems, not guess-and-check on config templates that become a bespoke DSL with no proper way to debug.
Our config structure requires similar but individualized configs for on-prem and the cloud. These can deviate for the region, the data center, the logical cluster, and at the individual node level for A/B or canary reasons. And of course local dev. You need to have a way to know where the deviations are, like different regional hostnames for services you connect with, different node sizes or replica counts, and different labels, any of which may or may not be correct and will cause an incident if wrong.
Search you yaml PRs, and I bet you will find a commit along the lines of "because yaml."
YAML is only simple on the surface.
One bad apple can ruin everything though:
from config import School, Teacher, Course, MRS
MNH = Teacher(MRS, “Marissa”, “Neve”, “Harman”)
BIO = Course(“Biology”, MNH)
ST_SIMONS = School([BIO])
Oh what a nice tidy config you have there! Let’s ruin it! courses = [BIO]
if str(today()) < “2023-04”:
courses.pop(BIO)
if today() > DIVORCE:
for c in courses:
if c == BIO:
c.teacher.name = (“Marissa”, “Cox”)
c.teacher.title = “Ms”
# gibberish ad nauseam
I suppose it’s possible to write nonsense code in any language including YAML, so I don’t know if this is a very good point or not. To put it diplomatically though: your cadre of YAML editors, should they be moved over to using Python, are probably the ones who need the most help writing clear code.But on that topic, I feel whitespace sensitivity is a bad idea even for generic-audience configs. People learn how to use parentheses in primary school. Explicit grouping and nesting isn't black magic.
But the whitespace sensitivity of YAML is particularly bad because of e.g. Helm. Templating a whitespace sensitive language with string interpolation is such a terrible idea.
This is exactly what I fear will happen. If you look at the YAML some people produce and extrapolate that to one of the most dynamic languages available, it is going to get ugly.
Programming languages for config are better, and I will throw up if I ever have to see “list comprehensions” and null checks in Terraform ever again, but it requires people who can code. If you simply replace YAML with Python and that’s folks’ first time using a normal language, it won’t work. Hence I’m happy to stay with YAML mostly, in such a scenario.
Unit tests?
Not to speak of tooling support. I can write a Python application with the strictest type settings and have mypy do a lot of heavy-lifting for me, before even running the app once. A bit like Rust. Check out the typestate pattern for what I am on about. It's invariants enforced in the type system, by the compiler. Impossible to misuse: your code simply won't compile. All of that is impossible to have if your types are strings only, with the odd bool and float inbetween.
I will accept that we cannot have ops people be at least medium-grade developers, which would be needed to apply these topics (I consider myself an in-between, leaning dev). That's simply infeasible. It's two different worlds. I will not accept the premise that these things aren't objectively better though! They're practically inachievable, sadly (or at least a decade away).
So no, in that advanced sense, YAML is not code. And if you can do YAML, yes you will get Python syntactically correct with a little practice. That is only 5% of the way though. Note I'm also not advertising for Enterprise Java-level of code... but more than "YAML but in Python".
It is not trivial to tell what code is doing, and as soon as your configuration is code then really you're developing a new application which implements another configuration language.
Most scenarios nowadays are _not_ like that anymore. There isn't 5 servers next door, there's 300 serverless whatevers half a globe away. Are you going to have 300 list entries, 230 of which nigh identical but 70 subtly different? Trivial in code, almost impossible to express statically.
If your configuration is getting more complex then that though, again, it's not really configuration anymore - it's a management application which needs to be developed and treated like that. And why it's that complicated should be re-evaluated - i.e. how come this is "configuration" and not something the application detects for itself? Why is it being surfaced to the user (operator) at all?
I think the creation and take-up of yaml over something like s-expressions is an indicator that the influential practionors in the cloud industry are all young and inexperienced.
If the creators of all these cloud tools display such poor judgment, it makes the tool itself suspect.
All the languages/tools that use it as a config/data format were started during that period, before people realized that all they’d seen of yaml were toy examples.
The idea is that there's a single config.yml file in the project root. It's mostly empty but could grow to ~100 lines if all config options are are changed, which realistically will be almost never. Most times it'll be 0-20 lines long.
In that case xml would be needlessly confusing. It's not easy for non technical users to read.
really?! XSLT is an abomination. XSD is bearable since it gives you the validation and autocompletion etc.
I wish more tooling used Nix, it's a great lisp that isn't ugly.
Then again I'd love to be shown that this partially-considered opinion is wrong.
(Also, isn't Nix's inadequacy as a lisp one of the reasons for the existence of Guix?)
The only difference between Nix and a hypothetical JSON++ is that Nix is lazy by default.
I.e., it is nothing at all like any Lisp.
In the `actions` directory is all the Nix files. There is some glue code in `lib/flakes` to generate the YAML files from Nix.
I really like your sense of humor, much appreciated. Question is whether the PTSD came out of the XML or the Java usage.
Also indeed we're eager to see finally some reasonably good Assembly configs, it'd definitely be a blow in the face of Rust magicians.
I would vote for Prolog as a config language, though. If memory serves right - some 10 years ago Matt Sergeant of Perl community did a very good talk about this approach, as result of his explorations of this area. Interestingly he has some definite experience with XML also https://www.xml.com/pub/au/22.
This is a mistake every generation has to repeat over and over again until the end of times. So it is written.
what is the state of the industry then?
No, you don't. Any program that generates a config will itself want to be configurable. Now what?
I will never understand your problem with it. JSON or YAML should be enough for 90% use cases. And INI or TOML the rest. Let the XML relic rest in peace please.
The state of the industry is... perfectly fine? I really don't get why people hate YAML.
To me the "programming language as config" idea is just like Lisp-like macro. Yeah it's powerful, but once you have more than 2 people working on it you get DSLs (s for plural).
People, at least Lisp people, like to claim the line between code and data is very blurry. In my experience in real world it's blurry 1% of the time. In most cases code is code and data is data.
I remember years ago I was writing my own parser for YAML and when I came across this problem I posted on a few mailing lists and got confirmation that the behavior is per-spec. I've never touched YAML since and still won't.
But I miss the days of XML, XSD, and dare I say it ... XSLT. XSLT is stupidly good at what it does as long as you don't abuse it.
Likewise, CUE Lang is built for config (esp merging docs with shared refs) and is highly under-appreciated. You can express powerful computations if you puzzle over the logical inferencing for a bit.
[servers]
[servers.alpha]
ip = "10.0.0.1"
role = "frontend"
[servers.beta]
ip = "10.0.0.2"
role = "backend"
is ugly to my eyes. And I'd rather swallow my own tongue than have to hand-edit XML in the most common case where there's not a dedicated editor for that specific doctype.In fact, I'd echo the linked article's argument back: I don't know of a case where XML is the best option. For human-edited files, pick almost literally anything else. For serialization, JSON handles the common cases and protobufs-and-friends are better when JSON isn't enough. There's not a situation I can imagine where I'd use XML for a greenfield project today.
I'd still take XML over JSON for human-edited files. At least XML supports comments.
> For serialization, JSON handles the common cases
Counter-point, JSON sucks and is way overused. The types are too fuzzy, the syntax too quirky, and validators/schemas are almost never present. You can bolt that all on, but it wasn't designed for it and it shows. It was designed to be eval()'d, which you should also never do because it's a terrible idea. It's flawed at the foundations.
But I will say that the first time I used a JSON API that had replaced an XML one, I almost wept with relief. Perhaps because JSON is so simple, it pushed APIs toward having simpler (IMO) semantics that were far easier to reason about. Concretely, I'll take an actual REST API (that is, not just JSON-over-HTTP) over the SOAP debacle any day of the week. I know you can serve XML without using SOAP, but to me they're both emblematic of the same mindset.
Still, at least neither of them are EDI.
Do you mean non-json types? Because the supported types seem pretty straightforward. (besides perhaps supporting null bytes in strings in things like postgresql)
> the syntax too quirky
Care to explain? This has always seemed like one of jsons strengths. The syntax for what is valid is pretty straightforward.
What is a "number"? Is it a float? int? short? double? BigDecimal?
What about time values? Or dates? Oh, you have to just shove those into strings and hope both sides agree? That's fun.
> Care to explain? This has always seemed like one of jsons strengths. The syntax for what is valid is pretty straightforward.
One example is the json "spec" on json.org does not allow trailing commas yet many parsers do
XML's validation, schemas and typing are far more complicated and equally useless - the impedance mismatch is too big, all they do is give you a whole bunch of extra ways to shoot yourself in the foot, particularly in the presence of namespaces. If you want something fully structured, protobuf or equivalent is the way to go (and converting back and forth between protobuf and JSON is relatively painless).
Also the use of attributes vs nested tags seems pretty arbitrary and in my experience attributes are hardly used at all.
So. Many. Conflicting. Namespaces. Often for the same “kind” of document, at least as far as the user is concerned.
So much flexibility. So many ways to paint yourself into a corner. Such a nightmare.
Also, the reason we have jobs.
Of alternatives, I think EDN is the closest to being a satisfactory replacement because it supports namespaces.
I don't like XML, YAML, or JSON.
If I am doing server-side PHP, my config files are executable PHP. It's my responsibility to make sure they can't break the system too much.
If I am doing host-side stuff (like Swift Package Manager), I prefer executable Swift files.
But I cut my teeth on BNF[0], and X.500[1]. They worked well, but were painful. They make YAML look good.
Comments that result in:
#Frontend server is called Alpha, and runs HTTPS
frontend:"beta":http
cause me to die.I tend to preface my config stuff with fairly substantial comment blocks that discuss the reasoning behind the configuration.
Here’s an example: https://github.com/LittleGreenViper/LGV_MeetingServer/blob/m...
# Diff for interactive merges.
# %s output file
# %s old file
# %s new file
merge="sdiff --suppress-common-lines --output='%s' '%s' '%s'"
it's useful the first time you dive in, not having to read the man page. But over time, the comments can get out of sync especially if you don't carefully merge in the package-maintainer's version every update.They have a parse.y that sort of gets traded around and joins new projects. Nothing super formal but it does mean most openbsd service configuration feels like each other. but each config is tailored to it's application. and being properly parsed the error messages can be better.
In fact that is my biggest beef about yaml, I mainly use it in the context of ansible, and the parsor usually has no clue where in the file the error actually is. You have to depend on remembering where you last edited to actually find the error. My other big problem with yaml is that the ansible context is trying very hard to make it a programing language... And while it is an okish config language it is a terrible programing language.
In fact this is a common problem with many complex environments. They want to try and push this complicated setup into a config file and claim "look it is easy, no programing required" when really what they have done is to push a programing situation into the worlds worst programing language. see also: xslt
Not without a trauma therapist on speed-dial.
Anything more complex and it becomes one of the worst choices due to the confusing/unintuitive structure (especially nesting), on top of having less/worse library support.
YAML's structure is straightforward and readable by default, even for fairly complex files, and the major caveats are things like anchors or yes/no being booleans rather than the whitespace structure. I'd also argue some of the hate for YAML stems from things like helm that use the worst possible form of templating (raw string replacement).
I think Python's pyproject.toml is a great use of TOML. The format is simple with very little nesting. It's often hand-edited, and the simple syntax lends itself nicely to that. Cargo.toml's in that same category for me. However, that's about as complex of a file as I'd want to use TOML for. Darned if I'd want to configure Ansible with it.
A few months ago I made a "mini ansible / cookie cutter" ( https://github.com/linsomniac/uplaybook ), and it uses YAML syntax. I made a few modifications to Ansible syntax, largely around conditionals and loops. For YAML, I guess I like the syntax, but I've been feeling like there's got to be a better way.
I kind of want a shell syntax, but with the ansible command semantics (declarative, --check / --diff, notify) and the templating and encryption of arguments / files.
It's not hard to infer that they're referring to nesting as a footgun: make it harder and you lose some power but you keep your feet.
Config files are a poor place to complex and deeply nested relationships. If it's not ergonomic to reach for nesting people tend to be forced to rethink their approach.
And of course to minimize impedance mismatch, the structure should be similar to the domain.
So yes I want a "config file" to handle at least a dozen levels of nesting without getting obnoxious.
And I don't disagree. The problems of nesting objects "at least 12 levels deep" aren't going to be solved by the right format. The tooling itself needs to expose ways to capture logical dependencies other than arbitrary deep K-V pairs.
There is no escape, you can't win. If you want the nesting, and assuming you can't remove it from the problem itself (as you often can't, or at least shouldn't), there's only one thing you can do: move inner things out, and put pointers in their place. This is what we do when we create constants, variables, and functions in our code: move some of this stuff up the scope, so it can be used (and re-used) through a shorthand. It loses you the ability to see the nesting all at once, but is necessary (among other reasons) when the nesting is too large to fit in your head.
Of course once you do that, once you introduce indirection into your config format, people will cry bloody murder. It's complex and invites (gasp) abstraction and reuse, which are (they believe) too difficult for normies.
The solution is, of course, to ignore the whining. Nesting is a special case of indirection. Both are part of the problem domain, both are part of reality. Normies can handle this just fine, if you don't scare them first. You need nesting and you need means of indirection; might as well make them readable, too. Conditionals and loops, those we can argue about, because together they give a language Turing-complete powers, and give security people seizures. And we have to be nice to our security people.
If you need 12 levels of nesting, add indirection, or live with the fact no one is designing formats to enable your oddball mess of a use case.
12 levels of nested braces in a single function is already a crappy idea: it's an even more crappy idea in a config file because of the generally inferior tooling, and now there's a downstream component that needs to change to support a cleanup (meaning it almost never gets fixed and the format just gets worse over time)
Agree. I've recently inherited a python project, and I'm already getting tired of [mentally.parsing.ridiculously.long.character.section.headers] in pyproject.toml.
Seriously, structure is good. I shouldn't have to build the damn tree structure in my head when all we really needed was a strict mode for YAML.
> I'd also argue some of the hate for YAML stems from things like helm that use the worst possible form of templating (raw string replacement).
I was literally speechless when I saw helm templates doing stuff like "{{ toYaml .Values.api.resources | indent 12 }}", where the author has to hardcode the indentation level for each generated bit of text like a fucking caveman.
This is by far the worst use of YAML I've seen, but I'd nominate kustomize for an honorable mention. Most of the various was to "patch" template YAML files are just plain bad: https://kubectl.docs.kubernetes.io/references/kustomize/kust...
The tiny examples might look kinda okay, but when someone has stacked 10 different patch operations in a single file, it gets a lot harder to keep track of what's going on.
You’re discouraged to nest your configs.
is it uglier? sure. but it's peace of mind...
in terms of config i'm using myself, say Kubernetes stuff, i really love YAML...because i know exactly what it's doing and i generally keep things simple. it's nice for that, it just does way too much IMHO...
If only Kubernetes hadn't gone and used it as the markup language of choice to represent... everything.
{
"servers": {
"alpha": {"ip": "10.0.0.1", "role": "frontend"},
"beta": {"ip": "10.0.0.2", "role": "backend"},
}
}
YAML also isn't great: {
"servers": {
# Frontend server is called alpha
"alpha": {"ip": "10.0.0.1", "role": "frontend"},
# Backend server is called beta
"beta": {"ip": "10.0.0.2", "role": "backend"},
}
}
Or, depending on your preference: servers:
# Frontend server is called alpha
alpha:
ip: "10.0.0.1"
role: "frontend"
# Backend server is called beta
beta:
ip: "10.0.0.2"
role: "backend"
I don't think XML is much better or worse: <?xml version="1.0" encoding="UTF-8"?>
<servers>
<server id="alpha" ip="10.0.0.1" role="frontend" />
<server id="beta" ip="10.0.0.2" role="backend" />
</servers>
All of them suck in their own way. All of them work fine with autocomplete, type analysis, and autoformatting. The more things change, the more they stay the same.JSON lacks comments, that's the biggest differentiator in my opinion.
It also allows additional trailing , in lists. Something I hate but is a great feature if you use a macro system on top of it. (Like jinja2).
If you're sticking to certain variants, you may as well use YAML, which supports JSON notation, as well as comments and various other improvements.
I am not sure I may as well be using yaml. I don't like it for the multiple reasons in the OP and this thread.
If you are using python, I have found it to be quite easy to support both json5 and yaml, as well as converting between them for people who feel strongly about yaml. Not trivial but low effort.
- introduce the least idiomatic form of YAML as "depending on preference"
- add an optional preamble to the XML example
- add comments when there were none, but also not to all of them
If you didn't artificially stretch the different examples to match, there'd be a much clearer difference between them all, especially considering the fact one of your three examples is a superset of the other.
The odd one out can't even capture an integer vs a string without a schema.
But I think the point of adding comments was just to show that comments are possible in some formats, and not others. Omitting possible comments from the XML example might have just been a sign of fatigue over this topic ;)
At any rate, I find the (idiomatic) YAML example to be -- by far -- the most readable of all, including the GGP's TOML example.
- Most XML files I encounter come in this format. You can skip the preamble but it wouldn't match my real life experience.
- All readable config files I encounter have comments. I forgot to add comments to the XML representation, but I can't edit my comment anymore. I think everyone who ever encountered XML knows how to add comments, though. JSON simply doesn't support comments unless you use a niche JSON derivative.
As for the string versus integer problem: you always need a schema, or you'll run into very funny problems down the line. None of these formats intrinsically know what keys refer to an object and what keys refer to a string, that's all based on your schema anyway.
"10" is a string. 10 is a number. [10] is a single number in an array. {"number":10} is an object.
You're conflating advice for databases with advice for data serialization formats: XML captures less information about the data it contains intrinsically.
_
Also please don't use YAML as "JSON with comments", you're just asking to run into some obscure bug/corner case
If you're willing to do weird things there's always JSON5
This creates a serious footgun in Ubuntu netplan, leaving a server totally unbootable, but simultaneously not triggering "netplan try" as any sort of parsing problem:
% echo '{"address": "2::0"}' | yq -P -
address: 2::0
yq seems to be happy with unquoted colons.String with ": " inside have to be quoted, that's correct, however you won't have space in your IPv6 address unless I miss something.
Getting in the habit of not doing so will lead to schema violations, like the Netplan problem you linked, which can crash the program trying to read your config. If it bails out at an unfortunate time, like most networking tools seem to do, you'll need to use a recovery boot image or serial console to fix your config.
Been there, done that. A good config file or linter should’ve complained and not allowed me to commit such misstake.
Lack of comments for JSON isn't a huge issue considering you can make the keys fairly verbose. And it would be actually pretty easy to add this into the spec, and parsers would still be backwards compatible.
I present jsonc: https://www.npmjs.com/package/jsonc
Making keys verbose tells you what a thing is, but comments are there to tell you why a thing is.
The only real way around this seems to be to create some kind of a "_comment" key.
Json but with comments and allowed commas on the end.
https://github.com/json5/json5
It's my preferred configuration file format, it fixes all the problems I have with JSON (trailing commas, comments) without turning it into a mess full of gotchas like YAML.
It’s without a spec
A theory - they don’t want a third party code in a hot path (they do care about performance in vscode), they already have a very performant parser and they don’t want to add a complexity there.
But that’s just a theory.
Lack of comments in JSON is a huge problem for config files. It's not an issue if you're just exchanging data between APIs, but for config files, comments are essentials.
There are some JSON specs that will do comments, but they rarely specify what dialects of JSON their parser accepts. There are also workarounds that abuse the fact duplicate key handling isn't part of the spec by specifying each key twice, once with a comment and once with data as most parsers only make the second key stick; those are even worse.
You can't add backwards compatible comments to JSON, there's no space in the JSON spec to retroactively insert comments somewhere. The closest you can do is the duplicate key trick, but as the spec doesn't state which of the keys to read as a value, that trick only works with specific parser implementations.
You can easily add comments to JSON spec, by just writing every parser going forward with the added comment parsing. It would read old JSON non commented files just fine.
(servers
;; lispy comment
((name alpha) (ip 10.0.0.0) (role backend)
(notes "very
large multiline
string")
((name beta) (ip 10.0.0.1 10.0.0.2) (role frontend))
)None of them handles it well. YAML has 7 different modes to do it so you will inevitable mix up and use the wrong one, otherwise it's actually the only option that supports it. Json requires inlining \n. Xml only does it with whitespace indentation.
I just don't like TOML. Just rubs me the wrong way for some reason. Don't know why.
in short: all these have their use-cases. i use YAML where it fits and TOML where it fits better. I never use JSON because JSON is just for machines, I generate JSONs.
From the article:
> “I’m making a new kind of book, and I need to annotate all of the verses in the Bible, and have the chapter headings and stuff.”
XML is a markup language, and works great for marking up text. It came out of SGML and attempts to make machine-usable documentation.
I'd favour its use in something like a datasheet, which is a combination of human-readable information, nested objects, and lots of stuff that needs to be machine parseable in fairly precise ways.
IMO it also works "fine" for other structured document formats that aren't text, like SVG, but that's not a strong opinion. JSON and other formats compete more sensibly here, but I'd never want a protobuf-based format for... writing an essay, for example.
The solution was to uninstall netplan.io Debian package where it didn't belong - get that YAML out of there. My hosting provider, OVH, figured it would be a good idea to shoehorn netplan - with its accursed YAML - into Debian for Network configuration. Bad move.
+1 for TOML. I love it's usage in innernet config files to set up new clients with a single generated "invitation" file.
allowed-ips: [0.0.0.0/0, "2001:fe:ad:de:ad:be:ef:1/24"]
Note that the IPv4 addy didn't need the double quotes, but the IPv6 addy did.The parser should have picked this mistake up, and the server shouldn't have been crippled to the extent of needing a rescue.
More docs seen here: https://netplan.readthedocs.io/en/latest/netplan-yaml/
Perhaps the trap is the complacency that YAML induces by not requiring quotes around keys/values, and so text risks being interpreted in unexpected ways. The infamous Norway Problem has the same root cause.
This is incorrect. The colon needs to be followed by whitespace for it to indicate a key-value pair. You can check this with the reference parser (and a bunch of others!) online: https://play.yaml.io/main/parser?input=YWxsb3dlZC1pcHM6IFswL...
I bet the parser did actually fail, because you can't parse a misformed dictionary into a string. Your problem is that the tool probanlt took down the interface before trying to parse for format, and then failed to bring the interface back up.
Similar to bogus data in /etc/network/interfaces, resetting the network interfaces with bogus data will end up with your server having no or limited connectivity capabilities.
At least with YAML there are command line parsers available to check your work. Plaintext config files often end up being a game or chance to see if you've got the format right.
I'm not convinced they'd define a sane schema for either.
Like you've noticed... the trying logic is naive. IIRC if this passes the mild sniff test it will go ahead and apply.
It can't realize changes on specific/named interfaces -- it insists on all or nothing.
As above, you can’t ask netplan to sanity check a config.
You can’t create a draft configuration, apply it, and save it if it works well. Every self-respecting network config system since at least Cisco IOS can do this (and does it by default!).
Interface renaming can’t filter by being a physical interface, which means that the system tries, and fails, to rename VLANs, because their MAC matches something that should be renamed. (networkd can handle this, but the networkd config written by netplan is wrong.)
Deleting virtual interfaces (e.g. VLANs) seems to be essentially unsupported, at least on 20.04. I think it’s slightly, but only slightly, better in newer releases.
Not impressed.
I've grown to enjoy NetworkManager. I know, like all things, that is probably controversial to some.
Two things I really appreciate about it:
- You can 'up' a connection/interface in an idempotent way; only changing whatever is needed.
- it's *very* scriptable. Values can be given with +/- operators
I was surprised/frustrated with networkd initially, but enjoy it now.It getting involved with packet forwarding was an unwanted surprise during a modernization effort.
As soon as Windows/MacOS/Android/iOS clients want in on the fun, alas, you'll need something more complicated than innernet to accommodate these other clients.
That said, you only need to look at the abomination that is an OpenAPI spec that has been annotated to work with AWS to see that the pitfalls are still mostly the same today. (I can also complain about the breaking changes from the old specs.)
Mostly unused.
> YAML/JSON templating to replace XSLT
Generally much more user-friendly.
> jq to replace XPath
And much better at it, if only due to not shooting yourself in the foot by default with namespaces.
The problem with XML isn't that it supports schemas and namespaces. The problem is that it forces everyone to pay the costs of those things up-front, even if they don't use or care about them.
I will happily cede that data is a huge area where "duck typing" is far and away the correct choice. Taxonomies of data fail all the time. With odd rules that are largely defined by their exceptions.
For example, either your format is a programing language, or the data evaluation must never need an internet connection. XML isn't one, yet still requires it.
A lot of parsers do not implement the standard, and they are better for it, because this can create huge security issues. But it's something you have to be always aware of, and could change on any minor version update.
So, agreed it is a bit of a footgun.
Maybe we need a new configuration format...
ducks
Actually I think something in the space of JSON5, Jsonnet, GCL, HCL, CUE, etc, will be the one to win out in the long run. JSON-but-fix-most-of-the-warts.
Then there's things which almost sit between fully declarative and turing-complete. Maybe Dhall, if it picks up a few more language implementations. Or CEL.
IMO Jsonnet, CUE, etc are just in a different category of complexity since they include code and require a full-blown interpreter to read.
- No footguns (compared to the alternatives)
- Basic text-templating
- Can do structural manipulations of data. As opposed to many other config-generators that only
do string-interpolation of text. Helm and j2.... kill me please, who ever thought this was a good idea
- Basic looping, transformations and function calls. Without falling into the full imperative
programming language trap. Configuration is still deterministic and directed.
- Supports loading data from other sources
- Supports commentsOk you win. What a joke of a configuration language.
https://www.bram.us/2022/01/11/yaml-the-norway-problem/
Norway abbreviates to "NO" which YAML parses as "false"
When it comes to config formats this is pretty reasonable. People are just primed to expect true/false to be reserved.
---
enable_frobulator: yes
use_thopojog: no
This was also the issue presented in the OP where they were mad that their version which is a string was interpreted as a number because it wasn't in quotes. Like I don't know how you expected YAML to fix that for you. If you don't quote a number in TOML it'll also be wrong.Similarly in xml it’s a curse that
<Foo><Bar>1</Bar></Foo>
and <Foo Bar=”1” />
are semantically the same but considered different in most systems.Just avoid ambiguity and make it impossible to face a choice. The same semantic meaning should have one expression.
>>“I need to configure this server and the server needs to know if this value is true or false.” >No, that’s bad. Don’t do that. That’s not a good use for XML.
First, the primary case that people use YAML for is not appropriate for XML? I don't agree with this but, taking it on face value, it creates a YAML strawman to attack.
Second, the rest of the argument boils down to "I don't like how floats are parsed in my language of choice". Guess what? You're storing your version strings in the wrong data type. Stop storing semver and it's variants in floating point types, that's not what they're for.
Finally, this is ultimately a "considered harmful" article, which is always a red flag.
Two years ago I started a small data quality checker software where users could define their alerts, frequencies,.. all in config files instead of modifying code.
I initially chose JSON as config format, but then realised comments are necessary to guide users in defining alerts. I moved to YAML, but after some "indentation incidents" started using HOCON conf [0] and never looked back. I don't see any reason for choosing YAML over one of JSON or HOCON, except being forced to because of some dependency. Features such as inheritance and text block support which were essential for me are nicely supported in HOCON.
Anyways sure. We can talk.
Humans can screw up anything. And more text often allows it to hide for longer.
XML has them sorta built in (basically all (notable) libraries support them), but it's not like it's required or somehow innately protected because of that. It's just a bit easier to adopt.
JSONSchema? It's pretty standardized and well accepted.
All JSON evaluates to some value in YAML. But it's not always the value you want it to.
Oh hey, I had some while making https://hn-blogs.kronis.dev/
I wrote more about it on my blog: https://blog.kronis.dev/articles/ever-wanted-to-read-thousan... but the gist of it is that I had to parse thousands of blog feeds and some article from 2009 included a SOH control sequence inside of otherwise valid XML and this broke everything, until I added additional error handling.
Ever use the earliest IDE incarnations of Eclipse on a PC still measured in hundreds of Mhz, or when RAM was still in the dozens of MB? That was the world XML found itself in during its inception.
But, I agree that this is one of the worst parts of XML
Nah, I have an easier time editing XML in a dumb text editor than YAML with the best tooling I've ever found for it.
In fact, I have an easier time with any of its competitors people post around than with YAML. What is distressing, because the language has clearly the goal of being easy to write.
“test this against Go 1.20”
> It interprets that as Go 1.2.
Well, understand your data format before complaining. YAML, like JSON, understands data types. If you were to write { value: 1.20 } in JSON, it would similarly be interpreted as the numeric value 6/5. The only reason this works "magically" in XML is because XML itself doesn't have data types, it only has text, and the interpretation is left to the user rather than done by the parser.To add, the XML specification on decimal data types [0] explicitly says: Precision is not reflected in this value space; the number 2.0 is not distinct from the number 2.00 -- so a decimal data type in an XML document would have the exact same problem as the YAML example in TFA; the only difference is that with XML, the authors of the tool would have to actively shoot themselves in the foot by annotating that element as a decimal type rather than text.
this is like saying Java is untyped because the source files are just text
My full sentence is more like saying Java is untyped if it were possible to run Java source files from the AST while skipping the type validation step, which seems pretty much a truism to me.
The fact that versions often contain numbers separated by decimal points or that often times, versions only have two components or that minor versions may rarely exceed 9 for a particular product are merely coincidences.
Contrast that with JSON, which provides booleans, numbers, and the list and object composite types. (side rant: as a standard JSON does not define whether it's numbers are integers, floats, or decimals! a conforming implementation can use whatever type it wants).
Yes, XML has (multiple) schema definition languages that can be used to enforce that the strings can be coerced into specific types, but XML itself conveys no type information in-band about the values. I think this is one of the reasons it is difficult read and write by hand.
node @attr="bob saget"
subnode @foo="john"
subnode @ham="bar baz"
html:h1
p $"this is literal text"
It's ok.I am searching for this, any clue ?
However, before you start, there were bugs that I found, and Augeas is a bit hard to work with.
Xcode's plist editing ability is a mild improvement over manipulating XML text directly, but could use more obvious shortcuts/hotkeys. Even writing XML in an IDE like IntelliJ isn't great, autocomplete should do a lot more.
- Syntax for the configuration is straightforward.
- Syntax for the definition is straightforward.
- It supports comments (end of line and full line).
- It's not NP complete and there is no weird parser meta-language or includes.
- There are parsers for any major language I've used in the last decade.
- It gives me a strongly typed definition so I don't have to infer or parse values.
- It's easy to verify the validity during lint tests by using protoc to encode to a binary proto.
It'd be nice if it supported unicode a bit better, but it's configuration, so that's not a common concern.
servers {
alpha {
ip: "10.0.0.1"
role: "frontend"
}
beta {
ip: "10.0.0.2"
role: "backend"
}
} grocery_list {
fruits: "apple"
fruits: "banana"
fruits: "pear"
} grocery_list {
fruits: [ "apple", "banana", "pear" ]
} (grocery_list
(fruit "apple")
(fruit "banana")
(fruit "pear"))
That's pretty much exactly XML. Alternatively: (grocery_list
(fruits "apple" "banana" "pear"))
but the semantics are different.My motivation is expressing docker-compose.yaml in the "best" way I can imagine, as a design exercise. In that case, I'd rather have a bunch of (service ...) forms, instead of a "services" object. Not sure why. Maybe it's a pun on interpreter-driven formats like that used by [guix][1].
[1]: https://guix.gnu.org/manual/en/html_node/Shepherd-Services.h...
Haha, what?
30% implicit defaults that are waiting to break when you upgrade something
25% boilerplate
20% derivable from some other configuration but you just gotta set them all explicitly
10% cargo culted in through copy-pasting from the last project
5% can only ever be set to one particular value or nothing works
5% silently ignored due to typos - (some of these are causing subtle bugs and some are preventing them)
3% critical to actually doing the thing
2% used to have an effect but that was 3 versions ago
Of these flags and values only 40% have meaningful and correct documentation and 66% are different in staging and production but we are 85% sure that's ok.
Yes XML, YAML, and JSON have a ton of warts, but a good part of the problem is how we layer configurability into software systems in the first place. The serialization format can only be blamed so much.
For example, the identifier blah is interpreted as a string; but 01 is interpreted as a number, which makes it equivalent to 1. Similarly the identifier nope is a string while no is the Boolean literal false, and there’s no clear indicator that the former is a string while the latter is a Boolean.
I think editing xml is just as bad as editing anything else and is only made better by having some kind of plugin which understands xml + a specific schema.
{:a 1 2 :bar [1 2 3] :baz}[0] https://github.com/bbatsov/clojure-style-guide#opt-commas-in...
This is the case with any hash-map in any language, modulo some of them making certain keys illegal.
But "string":"string" can be reversed in pretty much any language's hash-map.
>here's no way to tell at a glance what e.g. the value associated with:bar is.
You've certainly constructed and formatted a hash-map that is a little tricky to understand, but 1> this combination of key and value types is going to be pretty rare in the wild and 2> to the extent it exists, people would format it to be easier to parse, either by adding commas or newlines.
That's not true. What you're seeing here is a map where [1 2 3] is the key for the value :baz
Per that style guide, the above map should be formatted like this:
{:a 1
2 :bar
[1 2 3] :baz}
Most maps written in EDN have keys of a consistent type. A map whose keys are consistently keywords would look like this when formatted: {:a 1
:bar "https://example.com/"
:baz [1 2 3]}
Or like this, when condensed to one line: {:a 1, :bar "https://example.com/", :baz [1 2 3]}
The first map had keys of three different types: keyword `:a`, integer `2`, and vector `[1 2 3]`. Why would one want a format that supports maps with mixed-type keys? As a contrived example, mixed-type keys let you define a sparse 2D tile-based map for a game that is indexed by x and y coordinates: {:default {:type :grass}
[0 1] {:type :npc-rival}
[2 3] {:type :wall}
[2 4] {:type :wall}
[100 -5] {:type :treasure, :contents :sword-of-slaying}}
HN comment formatting help: https://news.ycombinator.com/formatdoc. I indented the above code by two spaces.Some differences: EDN has characters, keywords, lists, and sets Ion has dates, decimals, timestamps, s-expressions, and blobs
Ion has a binary serialization, hashing standard, and schema language.
YAML is easy to read like TOML or INI, has comments unlike JSON, and has dictionaries unlike TOML or INI. It's not bad.
It's not YAML's fault they didn't quote their numerical strings.
I recall (perhaps inaccurately) seeing that notion somewhere in http://www.catb.org/~esr/writings/taoup/ and it was at once a shock yet seemed so obvious. Might have been in something else esr wrote, but I've always considered that one to be his masterpiece.
Yes, it is. If you design a format in such a way that type parsing is ambiguous, or in this case "keep trying to parse it over and over, starting with the most restrictive option and working your way down", people are going to commit errors. That's just life.
I would absolutely love it if YAML parsers supported a mode where quoting your strings was required. Or, while I'm wishing: a new YAML version that requires that (even though that's impossible to do backward-compatibly without yucky things like having to specify the YAML version in the document itself).
But I do still use YAML for config files, because there's enough about the other options that I don't like even more than YAML.
I don't know what I'd prefer, exactly. Probably JSON, or better yet JSON5.
Then I used k8s for the first time. And after that, I understood, because k8s manifests are an abomination. Super error prone and hard to read. And I think that this has a whole lot more to do with the data model than the markup format. YAML isn't perfect (in particular, the way you specify an array of dicts is hot garbage and confusing), but it's not the main problem. The way the data is laid out in k8s would be awful to work with in any format.
But XML is not even a comparable standard. XML addresses an entirely different problem space. It cannot be directly mapped into baseline common data structures in most other languages like JSON/YAML/TOML.
Saying XML is “better” than YAML requires defining your problem space. Otherwise it just looks like you don’t really understand the difference between the two.
> “test this against Go 1.20”
> It interprets that as Go 1.2.
Without actual YAML (“test this against Go 1.20” YAML would interpret as the string “test this against Go 1.20”), its hard to know the actual complaint here, but it sounds like there was a YAML file with something like:
testTarget:
- language: Go
- version: 1.20
Where the tool interpreting expected a version string in “version”, but also accepted a number and implicitly converted the number to a version string. This is not a YAML problem, this is a “code that accepts and implicitly converts invalid data problem”. Its true that a schema and validating parser would help with this, and YAML doesn’t have a broadly supported standard schema language, but I bet the real problem here was that the underlying tool was written in JavaScript, as its the main popular language where even the most naive attempt to parse what you expect to be a string value would not fail if the value was actually a number.The issue is not isolated to JS either, the same could have happened in python, php, or any other untyped language.
Versions of the YAML standard from the last 14 years don't do the later, and supporting numbers as a basic data type is, honestly, a weird thing to harp on as “too far”.
> The issue is not isolated to JS either, the same could have happened in python, php, or any other untyped language.
Indexing an associative array of runtimes that is keyed by a string, or doing almost anything else that expects a string, when you get a number instead, will fail with an error in most dynamically typed languages including Python and PHP; much fewer string-expcting operating operations (and particularly not object indexing) will fail in JS with a number.
No, this is a YAML problem, because YAML is specified to allow for unquoted strings, and then it uses heuristics to decide if you meant a string or a number or (god forbid) a boolean.
So the "code that accepts..." that you're talking about is literally every conforming YAML parser out there. And they do that because the spec tells them to. So yes, it is a YAML problem.
And to top that off, most of the examples and tutorials you'll find on how to use YAML don't quote their strings. I get the idea: fewer characters to type, more human readable. But god it's a minefield.
As an aside has anyone tried making Python object notation? For instance you'd get tuples, sets, complex, hex, etc. along with dicts, lists, and all the other stuff json has. I know you can use literal eval but it has some security issues json doesn't have.
Good points re: XML and its misuse as anything other than a markup language (its in its name, afterall). After using things like HAML and whatnot for a few years I went back to plain HTML. I like it much better.
YAML, meh, I choose to use it in Hugo because that's what I'm used to and I'd rather not learn a new config language until I'm forced to. I prefer config to just be in the language I'm working in, though many people disagree for various reasons but, you know, like, whatever, man.
If think about it, thats not a problem of the format but of the editors people typically use. Simple text editors used in a command line use a line paradigm rather than a tree paradigm. Thats not well suited and its not the fault of the format.
Given that trees is about as fundamemtal as it gets as a data structure maybe what is really needed to end these format wars is not a new format but a new editor or editor plugin.
I also had to laugh at this statement:
the YAML specification has all these features that nobody ever uses, because they’re really confusing, and hard, and you can include documents inside of other documents, with references and stuff
That's a pretty funky argument to use in favour of XML.
I still like GNUstep's extended version of NeXT property lists.[1]
XML and JSON seemed like a step backward on macOS at least.
I think the problem may be that yaml systems I've used haven't done this well.
XML also has these?!? In fact I'd guess YAML was developed to mirror the feature set of XML while having nicer syntax for humans.
This one bit me recently. Parse it and see what happens.
- yaml:
- 01
- 08
This one works as expected: <list>
<value>01</value>
<value>08</value>
</list>
And I can qualify it <xs:schema attributeFormDefault="unqualified" elementFormDefault="qualified" xmlns:xs="http://www.w3.org/2001/XMLSchema">
<xs:element name="list">
<xs:complexType>
<xs:sequence>
<xs:element type="xs:string" name="value" />
</xs:sequence>
</xs:complexType>
</xs:element>
</xs:schema> - yaml:
- '01'
- '08'
Complaining about this is like complaining that { json: [01, 08] }
will be interpreted as [1, 8] in json... It's not a problem with the tool, it's a user error.The JSON format avoids this problem because it requires that all strings be quoted. If you start with `["09", "08"]` and change one of the digits inside the quotes, the string will stay a string. In JSON, strings becoming not-strings – removal of quotes or addition of backslash escaping – is more obvious than in YAML.
“I need to configure this server and the server needs to know if this value is true or false.”
No, that’s bad. Don’t do that. That’s not a good use for XML.
But on the other hand, if you need to mark something bold, then XML is a great choice.I see both cases as being quite equivalent: whether the server needs to know that a value is true/false or that it needs to display something as bold feels quite the same thing, isn't it?
Or does he mean that you cannot easily retrieve the value of an arbitrary XML element? Whereas to display a document you just process it sequentially and do not need to 'retrieve' arbitrary data. Is this it?
Well, life isn't black and white, but somewhat greyscale. From my experience, none is better and they don't share the same use-cases.
This makes sense:
<element>Some text here. <bold>Some more text here.</bold> And yet further text here.</element>
This does not: <value>true</value>
That should be: value: true <setting name="foo" value="true"/>It's born out of being a mark-up language, but it's true strength is as a data exchange format that can be self-documenting and human-readable at the same time.
Yes, it has a tendency to make people's eyes bleed with how verbose it can be. But if you were to open an XML file without knowing its format, or what each entry is supposed to be... it is likely you will understand (or be able to suss out) the purpose of the data being exchanged.
Schemas are very powerful, well-documented, and can be used to validate that an XML document is well-formed. People who tend to hate XML and Schemas tend not to understand how or why this is needed.
Contrary to the article's one point about config: I believe one of XML's strength is as configuration documents. It's better suited to configuration that is meant to be shared between different projects or platforms. If you only need a configuration file for a few keys and values - you're probably going to be fine with anything else.
But if you need to dump a configuration file into a document to be read by another program or exchange, then you're better off using XML because that format is more tolerant and can be well-documented by using Schemas. You can version those schemas, and use XSLT to convert the configuration document into another format that is far more readable.
That said, XML is not an be-all and end-all format for everyone. I believe, more importantly than anything, in using the appropriate tool for the given task. XML is simply not the best tool for every task.
<?xml version="1.0" encoding="UTF-8"?>
But more often than not, it also has a url.But so does every json file I get from Azure. The other day I exported a power app and each app parameter/variable has its own folder with an xml file and a json file all, which each reference versioned schema and all sorts of stuff all to essentially just say “key: value”
Who does this and why?
Version Information: It specifies the version of the XML standard being used. In this case, it's version 1.0, which is the most common version of XML.
Character Encoding: It specifies the character encoding being used in the document. In this case, it's UTF-8, which is a widely used encoding that can represent a vast range of characters from different languages and character sets.
Standards Compliance: It signals that the document adheres to the XML standard. This declaration helps parsers and software that process XML documents to interpret and handle the document correctly.
Interoperability: By including this declaration, creators of XML documents ensure that their documents can be correctly interpreted and processed by a wide range of XML tools, libraries, and parsers.
In essence, the XML declaration helps ensure that XML documents are self-describing and can be processed consistently by different software and systems. It's a crucial part of the XML standard and is included at the beginning of most XML documents to set the context for how the document should be handled.XML is a notational tool. A notational tool is a tool for a human to write something by hand and then process with a computer. The important part here is “by hand.” E.g. I’m processing a corpus of someone’s letters and see a phrase like “on Monday I saw him the last time” and I know that I need to mark “on Monday,” “I,” and “him” with references to a date and people I have inferred from elsewhere, and the only way I can do this is by hand, there is no automation. Yet once I place those references, I can mechanically index the corpus, which is my goal.
But writing a configuration file is the same task. Here again I am doing a thing that in the general case has to be done by a human, but the goal is to process the result mechanically.
All other uses of XML are a misuse. Yes, you can use it for data interchange, but it is similar to programmatically calling a third-party tool via a command line. Possible, but involves much overhead and is way less convenient than using that tool via a library. If the data are not generally composed by hand, then they should not be in XML. (But I would argue they should not be in JSON or YAML either; we do this mostly because we have no suitable tools.)
As a notational tool XML is actually rather good. Yes, it is verbose. You know what is not verbose? A special language you create for your special case. It will outperform any generic notation out there. If you decide to do that, then you will have to write a parser, process the text and get an abstract syntax tree. But note that XML is an abstract syntax tree. It is the intermediate result you get if you decide to solve your case in a perfect way. So maybe you could start with XML, get the syntax tree right, write the code to process it, and then see if you still want a parser.
On XML verbosity: what if we compared not the overall length, but the number of syntactic symbols?
<server id="a" ip="1.2.3.3" role="back" /> -- < "" "" ""/>
{"id": "a", "ip": "1.2.3.4", "role": "b"} -- {"":"","":"","":""}
I’ve extracted syntactic symbols on the right and you can see XML is much quieter.On XML not being mappable to objects and arrays: it is straightforward if you remember that the data are supposed to be a syntax tree. They are indeed far from the final representation in the same way an arithmetic expression is far from the code that will evaluate it.
<server id=a ip=1.2.3.3 role=back /> -- < />
why do you need quotes when you have nothing to quote (no whitespace etc)?
The author also mentions that JSON is better:
"So if you want to write YAML, you can. But it’ll just take that YAML and turn it into JSON behind the scenes. Then they also have a specific Caddy language. So you can give it the Caddy language and then it turns that into JSON behind the scenes. And you can give it an NGINX config and it’ll turn that into JSON behind the scenes. If you have the cycles and time to spare, that’s probably the best solution for most people…"
When the schema says so — you even get auto-complete for that in any half-decent editor.
rfc8259 is more detailed and it mentions that most parsers will have limited precision and references ieee754 to explain what the expected precision after deserialisation should be for most parsers.
Its spec is not "complete and unambiguous"!
Non ambiguous structure, parsable, nested values, lists, maps, and so on.
If JSON had comments it would be ideal - and for configuration you can support it.
Just check VS Code for what is possible to do directly regarding configuration with JSON, and also to build on top.
country: no
YAML will interpret that "no" as a boolean "false". Which... c'mon.So I started thinking... maybe I should use something else instead.
TOML? Nah, my config file requires around 4 levels of nesting, and nesting in TOML isn't great, at least when done the more idiomatic way.
JSON? No, I want comments, and I hate having to double-quote all the keys.
XML? No, way too verbose; the author of the article goes into more detail why XML is bad for a configuration language.
HOCON? I've used it in some Scala projects, but I'm a little worried it's not mainstream enough and users might be confused at the syntax (or annoyed that they need to learn a new format, even if it's simple).
CUE? I'd never heard of it before it was mentioned in this article. "Validate, define, and use dynamic and text-based data" -- um, that sounds scary. I don't want a language, I want static, declarative configuration.
Sigh. I guess I'll stick with YAML for now.
I feel like we need that in giant bold red letters at the top of yaml.org. Nobody gets it.
You may disagree with my reasons for not liking the others I mentioned, but that's just your opinion, and when I'm writing my own software, my opinion holds more weight.
Besides, JSON is already human-readable when well formatted. Just lacking in hand-typing ergonomics.
I don't get what makes YAML a serialization format. And if it was intended to be such, then it sucks even more than most people argue it to.
Agree that it's not a great serialization format, though...
What do you mean by originally a Ruby thing? I am curious.
Ignorance is only bliss until it gets you in trouble.
But how am I supposes to create my docker-compose configs?
xmlstarlet is a great tool if you want to query xml on the commandline using xpath or do some xslt transformations.
Similar to jq's utility for processing json.
In any case, my little experience with it had made me hate YAML. Generally speaking, I have come to dislike any language with significant whitespace other than Haskell.
Sorry, but this seems like a very silly, petty argument. There are other reasons to not like today's pervasive use of YAML, but this is not a very compelling one imo.
The same is true in XML, practically speaking, there just isn't another option - everything must be quoted or wrapped in tags.
So just ... use quotes in YAML. You can force it through a linter [0] if the optionality is what you're hung up on.
[0] https://yamllint.readthedocs.io/en/stable/rules.html#module-...
Because NONONONONONO . .
OK seriously. Seriously. These people all chillin' here in the future and talking about how great XML is need to jump in my DeLorean back to 2002. Or maybe try using XML for a few years. Or decades.
* Schemas break XML *all the time*[1],
* "XML-aware" diff/merge[4],
* No Such Thing As Line Breaks[2],
* Sneaky proprietary entities[3],
* NAMESPACES,
* "1NF? Is that a sex thing?",
* Computability[4],
* FRICKIN CHARSETS,
* asemantic but pretends to have semantics,
* Hierarchy Fetishism
And so so so much more. The combined effect of this is that it reduces the volume of the tool ecosystem for a given XML spec. Don't believe me? Run the metrics on gitlab/github/npm/pip/DaSEA[5].So the tool ecosystem is - very often - only as big as a singular project. It's one of the reasons there's so many XML editor vendors. In S1000D, it's pretty common to have a special vendor for each project.
XML completely nukes, by its essential nature, any possibility of using standard tools. There is an entire category of emergent technology - Lightweight Markup Languages - that were invented, by individuals, working for free, for no other reason than to get out of XML.
Using XML at scale is something that should never happen, for any reason, ever. It's this horrifying perfect storm of non-technical academics steering a crew of malcontents still angry at how GML went. Ol' Linus was a bit of a butthole, but he was right on the money when it came to XML. ALL of the problems YML has that are mentioned - they can ALL be found in XML, but they're magnified times fifty bazillion because of the inherent lack of support.
[1] Leading whitespace in attributes? HOW CHARMING. Yeah, that's in a schema, a very popular one. So each schema is its own language, and both DITA and S1000D allow for virtually any level of customization on top of that, and on top of THAT, in S1000D, you have to contend with each of the Issues. Seriously, it's a flashback to the pre-ATA100/JASC 1930s "shop manual" systems.
[2] No such thing as "normalized" when it comes to XML whitespace, which means no lines, no tabs, no spaces. Everything is elements. Oh, ha ha, unless there's dual-mode DTD/XSD validation . . which should REALLY have its own bullet. Do you realize, in any way, how incredibly radioactive external entities in a internet-facing parser are?
[3] REVBARS! Oh, and FRICKIN CGMs. Good luck processing those, because they were golden tickets handed out to ISO-favored software vendors.
[4] Infinite arbitrary nesting combined with whitespace agnostic means it's REALLY hard to make any sort of compute optimization unless you load the WHOLE thing into memory. An xml-aware git repository has performance several orders of magnitude worse than a normal one, and if used in quantity with goofball schemas, you can actually choke a Bitbucket CLI.
In YAML it's kinda more baked in. You have anchors and alias, which are part of the spec itself.
Yeah, in retrospect Billion Laughs was a bit of a cheap shot. It is, however, hilarious. And no one ever put forward any sort of mitigation or fix, for decades[1]. Meanwhile, in the YAML dev world open issues . .
if (refDepth > maxRefCount && node.kind === Yaml.Kind.ANCHOR_REF) {
I don't really have a dog in the YAML fight - apart from Asciidoctor-pdf template files[1] - but the YML people are patching, and the XML people didn't, for a very long time.Why is that? I'm going to go back to the basic notion of XML as the Everything for Everything, which was encouraged by its design pattern insistence on fake semantics. YML has, no doubt, a big ol' dose of the same sickness, but with a lot less overhead, and it makes maintenance easier.
Keep in mind, we're now debating "How YML is perhaps just as bad as XML"
[1] This has resulted in a lot of software and IETM files (even whole devices) getting pulled from USN vessels in theatre; there's more than a few vulnerabilities that ride on the SGML/DTD Billion Laughs. Bunch of other ancient file formats getting the same treatment, something we in the industry saw coming since 2007. Just a ticking bomb until you fight a peer.
Then enjoy your complimentary security vulnerabilities.
Jokes aside what is it used for that XML Schema or other XML technologies can't do better?
USAF hasn't yet gotten nailed with the DTD attacks the way the USN was[0]. And that was a complete musterfluck. First they pulled all the handheld maintenance devices, then they basically mandated that all the stuff getting stuffed into entities could instead get shoved into a black box XML element stuffed full of Base64 or reference to an external binary or - hell - whatever you want. That's the current solution: the //multimedia element.
You'd think USAAF and USAA[1] would have learned something from this . .
[0] That's changing as we speak; DIA has a bunch of hardass new IT policies rolling out. God be praised.
[1] Although the USAA spec has more flex in it when it comes to geometry and other extremely specific rendering behaviors. It's much easier to optimize because it's not insistent that a frickin PDF parts catalog have draftsman-perfect line art.
Here's what the entities (specifically, CGM, the 800 lb gorilla of external entity references) do that can't be done in XML+SVG: ISO/IEC CGM:1999 line types (your dashed lines are exactly right); ISO/IEC CGM:1999 nurbs (so that the curves are just right). I have a bunch of counterarguments to these things and more, but the easiest one is : how much is a perfect dashed line worth? Is it worth twenty two million dollars? Because that's what it cost the Navy. That's assuming the PLAAF/PLAN doesn't hop inside your maintenance network off the east coast of Taiwan. Then you can buy your dashes at the reasonable cost of a few hundred dead sailors.
They need to swap out //multimedia for a standardized, text-based format yesterday, though. Either that or release an ISO profile for SVG, which honestly would be, like, a week's worth of work at most . . if you wanted to see it done, of course. Oh ISO Technical Steering, you and your loveable scamps made up almost entirely of stoneage software industry reps.
When I’m reading the `maxItems` property from it, I know it’s an integer. Why do I need the document author to also tell me it’s an integer?
For statically typed languages, you have to tell it what type you’re expecting, so it can interpret the value at that point. `config.getInt("maxItems")`.
Even for dynamic language, most of the time you still want to validate the types up front to avoid an incorrect type blowing up in a random place in your code. So write a schema, use the schema to drive how the values are interpreted.
By replacing all the non-structural types with string, you can have the clean quote less format they’re after but without any of the “Norwegian problem” issues.
Because this allows well-designed client libraries to detect conflicts between the document-author intent and programmer intent, enabling mismatches to fail as errors rather than being read other-than-as-intended.
> For statically typed languages, you have to tell it what type you’re expecting, so it can interpret the value at that point.
Yes and a client in such a language should fail (or have a mode in which it fails) if the value is not actually, as read in YAML semantics, a type compatible with what you are asking for.
> Even for dynamic language, most of the time you still want to validate the types up front to avoid an incorrect type blowing up in a random place in your code. So write a schema, use the schema to drive how the values are interpreted.
Schemas are for validation, not interpretation. If you use them for interpretation, then you get JavaScript-esque weak typing, and the errors that come with it (basically, magnifying the kind of YAML 1.1 problems that YAML 1.2 tamped down.)
> By replacing all the non-structural types with string, you can have the clean quote less format they’re after but without any of the “Norwegian problem” issues.
What you propose is an explosion of Norway-problem-style potential for values to be interpret other than as intended by the document author, not a mitigation.
For a similar spec that offers no types outside of string, list, and dict, there's NestedText.
We learnt that XML was a bit too verbose for serialization and moved to JSON. Now we need a good configuration format, especially for the advanced usecases, and YAML ain't it. The known ambiguities are a minor thing - the real problem is that it's not typed enough and significant whitespace. We need a language optimized for custom datatypes. That's why properly used XML is actually better here (it's better typed), but it's far from optimal. We could really use a different option.
Nothing like a good old type safe compiled language to cut down on the verbosity, copy paste usage, silly syntax errors, weird undocumented you just have to know the magical incantations, etc. Kotlin or similar languages are the way to go. Much safer, more compact, easier to cut down on the copy paste reuse (which is just miserable drudgery), easy to introduce some sane abstractions where that makes sense. You get auto completion. And if it compiles, it's likely to just work.
People keep on moving around the deck chairs on the proverbial Titanic when it comes to configuration languages. Substituting yaml for json or toml just moves the problems. And substituting those with XML just introduces other issues and only marginally improves things. Well formed xml is nice. But so is well formed json. Schemas help, if the urls don't 404 and you have tools that can actually do something with them. Which, as it turns out is mostly not a thing in practice. And without that, it's just repetitive bloat. XML with schemas becomes very hard to read quickly.
There's a reason, people started ignoring XML once json became popular: json does most of the essential stuff well enough that XML just isn't worth the effort. And if you have something where you'd actually need the complexity of XML, it's likely to be some really ugly bloated kind of thing where the last thing you'd want to do is edit it manually.
I've dealt with cloudformation in XML form at some point in my life. It sucks. Not just a little bit. It's an absolute piss poor format for a thing like that. Since such a thing was lacking at the time, we ended up actually building our own little tools to generate that xml. Hand editing it was just too painful. One mistake could corrupt your entire stack. And it takes ages to find out if you actually got it right. In Json form it's hardly any better. It's just one of those convoluted over-engineered things. Anyway, Json support for cloudformation was not there at the time and the difference is like asking whether you'd preferred to be shot or stabbed. It's going to suck either way.
YAML is great for human-readable config, and there are footguns.
XML is terrible for human-readable config. JSON is not great for human-readable config. TOML is okay for human-readable config.
The problem is there's no clean way to abstract strings, bools, lists, objects, trees, etc. into a human-readable configuration syntax that does not have a footgun.
My favorite take on this by far is that any types beyond string, list, and dict don't belong in the format at all, and should be left to the ingesting code. I started on the path thanks to StrictYAML and found a home with NestedText.
I considered generating 2 separate reports, in JSON for machines and in PDF for humans. The PDF part turned out to be difficult.
In the end I settled with XML. Its machine readable by nature and with XSLT it becomes human readable in the browser. The programming language provided XML encoder in the standard library which made the task very easy.
I still find XSLT a pain to deal with though!
For wire/data formats I don’t care at all. Perf/size and library support is the only factor there but xml or json will do up to the point when I can use protobuf or something similar. I don’t like the conventions/opinionated-ness of obit though. The absence of a set of things is usually not the same as an empty set of things, for example.
Most of the time I don't need a format with support for some esoteric charset or namespaces or any of the other baggage that comes with XML, I just need a quick way to exchange regular, structured information, possibly with comments, in UTF-8, that is easy to read and write.
Show me a better format that doesn't make my eyes bleed from the angle brackets (or the curly braces from JSON) that is commonly supported across lots of programming environments and I'll happily switch to it.
Using a car engine as a pizza oven would also result in problems if you tried to cook a pie with it. But having a toxic half cooked pizza doesn't mean the engine has footguns. The engine works fine if you use it to drive to a pizza place to pick up a pizza.
How? It's even worse as a serialization format. Why would you want to serialize into human readable blob anyways? https://github.com/cblp/yaml-sucks
> It interprets that as Go 1.2.
In these types of configurations, _everything_ should be a string and data types are parsed with an additional helper function. It was a mistake to have per-platform, per-implementation parsing.
> Now, I understand why people do YAML, but there are better choices.
JSON. XML is great, but a little too verbose and can be difficult to parse. JSON is simple, well structured, easily human readable (if a sane structure is used), agnostic to indentation.
The example give about go 1.2 vs 1.20 is a bad one. Go should have gone with 1.02 instead of 1.2. Also you can just convert that into a string.
The reason why yaml is great is because I’m tired of learning “better” solutions that only fix the tiny percentage of issues that yaml has, but has a much more tremendous learning curve. XML is overengineered to the point that it’s a confusing mess.
Give me yaml any day of the week.
One is a document markup language. The other is a data notation.
YAML is better understood as a family of notations at this point. For every gripe about the standard YAML implementation, there is a "safe" implementation that does not have that problem. You don't have the throw the readability baby out with the footgun bathwater.
For another example: YAML 1.1 would treat "yes and "no" as boolean true and false. I've heard this called the "Norway Problem":
country: no
Whoops. I had never actually considered the issue in the article, and I wonder if I've ever made a similar mistake with numbers and didn't realize it.Honestly I think I would be fine with YAML if strings were required to be quoted. Yes, that would make things a tiny bit more verbose, but it would remove a big footgun.
OTOH, I’ve seen xml schemas that would make a rattlesnake cry.
But it's too hard for cowboy programmers to write valid XML.
Hi I'm a singer.
My name is:
Paul my name is
My last name is:
McCartney my last name is
My children are:
My child is:
Her name is:
Heather her name is
Her age is:
15 her age is
...my child is
...my children are
...I am singerUnfortunately YAML was even worse in that regard, as it allowed arbitrary code execution as seen in recent CVEs...
We got there. We got the extensible markup tooling we needed. Gosh, it took a long time though.
> And I think the fact that now everybody is writing React… and they’re writing it with JSX… and JSX is basically just inline XML… I think that shows that there are cases where actually XML is pretty good
No, popularity does not imply quality
http://move.rupy.se/file/logic.html
I used it for DB schemas and game pixel art.
But for now there is only a go implementation.
Someone will eventually implement it in rust, and then all languages may adopt it.
XML is a whole other beast.
So it really depends on your usecase. Do you need to be able to import several independently developed vocab and use them, possibly namespaced, in a single document... Seriously, go XML.
Then we could have endless debates about the relative merits of our preferred styles/layouts of our favorite .porn files.
I do agree that TOML is a better option. Easy to write and easy to read.
If you keep your comparison between those two, you will lose every time.
YAML spaghetti is so much worse than dealing with XML tags, no tooling, no schema validation,...
No... it doesn't...
What the heck is this misinformation?
num: 1.20
str: "1.20"
str2: Go 1.20
str3: "Go 1.20"
These all parse as strings except the first one.TOML also parses the following as number:
num = 1.20