How I learned to stop worrying and love the YAML
leebriggs.co.uk
leebriggs.co.uk
The idea that we can replace skillful software engineering with the right abstraction just hasn't panned out for me.
Every framework I've worked with had some issues, and because it's a framework it becomes a herculean effort to switch to something else. If it's only a single functionality that's much easier.
What would be the way to orchestrate those libraries? Some sort of orchestration library?
Say you use a web server library to help you handle requests on the wire, you’ll probably end up writing something generic where you set up a server, and pass in just the application code to receive the request and return the response which the library will send back.
You just wrote code that surrounds your application code — you made a framework. Then you realize that handling things like logging shouldn’t be in every handler so you invent middleware, you realize that your big ole block of app code is actually lots of small ones in if statements about the path so you factor that out and invent routes, etc. etc.
I hear this argument a lot, but I don't think it really ends up working that way. You build an application with some abstractions, not a framework. The difference is that the code you write does what you need, and doesn't need to support a bunch of different use cases. When you need to change it, you can, without worrying that you're breaking someone else's use of it.
And because you only need to support the one application, you can make it as simple as it can be to support that application.
I currently maintain a minimalistic website, and using a static site generator has been golden.
I consider that a framework. The moment I need to do anything the framework author didn't think of, I'll end up spending hours. But because I make so few changes, the bulk of the website content is Markdown, and the very limited styling is done with the HTML/CSS templating system. Basically, because this website is a low priority, anything that proves complicated is a bad investment. A framework is great here.
I agree with this but it does have tradeoffs as well. Libraries, especially if they are many and small, can get out of sync a lot. I feel like I spend too much time on dependency maintenance.
[0] https://www.carexpert.com.au/car-news/platform-sharing-the-m...
The Kurtosis hypothesis is that we can abstract away most of these things with automation (e.g. force devs to update the changelog, force versioning of the API, handle all the automated client generation) while still allowing users to pop the hood and do the low-level if they really need to (but popping the hood shouldn't be easy).
If you’re not building a tool for junior developers only, you might want to change the tone of your pitch to more of a “we make this super slick” instead of “we built black box #93784583”
Any sufficiently powerful language has the same hazard. It's natural to want to know everything about the technology you use and want to use everything you know. It's how we expand our skills, and the sheer pleasure of it is why a lot of us got into programming in the first place. But not everybody considers the interests of their team and their coworkers when they choose outlets for that impulse. A single selfish person can run a C++ or Scala codebase clean off the rails.
It takes strong management to make sure the team has good collective boundaries and to crack down on people who go too far. Since senior developers are now largely left to manage themselves, you're at the mercy of the savant who knows 98% of the YAML standard to understand why they're being asked to only use it 70% of it, and to pay attention to the opinions of people who know less (about YAML) than them. Or you can just use JSON or TOML.
Between all those links there doesn't even seem to exist a design document, much less a full specification for the language.
That is really not enough. If it solves your specific use case, great, use it. But it's not a solution to the YAML problems.
"No fancy YAML bullshit allowed" aaaaaaaand done.
Because they are programmers though, they might actually try to make the “code” better:
First they refactor everything using ‘extends:’ to use inheritance for infrastructure “code” sharing.
Then they started sharing “code” between blocks by using YAML references. This removes duplication, except for the edge cases that need to be different. Hmmm.
Then the YAML file became a Jinja / Jsonnet / M4 / C-preprocessed template. That means there’s now another programming language involved so on the upside there’s some hope of finally being able to express computation, but with the downside that it somehow plays second fiddle to the YAML despite the latter having nothing to do with computation.
The downsides of trying to prematurely capture your ideas as config are real. If you don’t focus on all of the key variation points you’re going to see people start hacking and templating at some point. It’s inevitable. I like how Ansible didn’t even try to avoid templating and ran with it straight off the bat.
Config DSLs are hard. OpenSMTPD’s is beautiful but they didn’t get it right until version 6. Even then I bet some people are generating the config. Caddy’s dual approach of letting you use their DSL or give it generated JSON is nice. They have a config DSL for the use cases they could think of, but recognise that there will probably be other cases they didn’t think of, and so support you calling their API with machine built configs instead.
Configuring a system with a programming language is difficult too. If you had a Python module to set up Apache, and your config was particularly lengthy, you can guarantee someone on your team is going to come up with a “clever” abstraction at some point via a pull request that links to the DRY Wikipedia page.
It’s just that with Python etc — vs YAML “programming” in “code” — they are more likely to actually get it right.
> It’s just that with Python etc — vs YAML “programming” in “code” — they are more likely to actually get it right.
Are they, though? My impression from various projects involving a good amount of YAML files (for CI pipelines, k8s deployment and such) has been that countless things that take me seconds in python (declare and pass around a variable, find out what a line of code does, find out where else a variable is used) take me orders of magnitude longer in declarative DSLs (like anything YAML-based) than in imperative languages like Python. In 99 of 100 cases I have to either consult the docs (because I don't understand what a given declaration means) or resort to string search (yes, string search) across one or – god forbid – multiple repositories. Meanwhile, imperative code at least tells me what it does and tooling support (usage search etc.) also tends to be very good.
YAML hell is real.
I'm not decided whether I'd want something vim-like for infra as code. On one hand, in VimL, simple things remain simple and complex things are possible, alas you could say perfect design. On the other hand, everyone who has programmed something non-trivial with vim knows the pain.
To me the badness is something like:
- just want a static unconditional config -> YAML
- ah but I want to pass around some variables -> YAML with some custom syntax for vars
- ah but I want conditionals based on system properties -> ...
And you end up with an awkward programming language with YAML syntax. I don’t much care if it’s XML or JSON at this stage, I just don’t want to write programs this way.
Yes I do use environment variables for secrets/passwords but that is about it.
If you need weird complex stuff, use a real scripting language like Lua or something. It should be faster than a YAML parser, at least.
Source: empirical study from me
Thanks for not being that person.
* for new projects, use a well-established, reputable standard
* for existing projects, use what is already in place if it is consistent
* in case of inconsistency, fix to the closest industry standard
* Use an auto-formatter and enforce it in CI
Github even has public beta support for it, though they require a specific name for the ignore revs file: https://docs.github.com/en/repositories/working-with-files/u...
But if this project made a weird decision I'm ok with it, as long as it's enforced consistently.
I've had one project - consumer investing web application - that has had a number of iterations. a Flash / Flex app, then when the ipad came by and Flash died, a BackboneJS + Bootstrap app, then rewritten to AngularJS, then some ex googler came by and decided it should all be rewritten again in Polymer for some reason. Last I heard is there was "rebellion" and they were building things in React.
I mean I kinda knew it was all futile anyway, but that experience, where years of my own labor was just discarded in favor of an experimental and unfinished technology, made me real jaded about things.
If you just need a configuration format for your application then none of this applies, there is however a whole school of tech job that has to deal with this every single day
script:
- true # oh no*
No* pun intended.*Pun also intended.
(To the uninitiated: YAML automatically casts the bare string “true” into a boolean for you. Fairly reasonable. But it does the same with “yes” and, per the parent comment, “no” too. This is problematic if you have a list of ISO language codes including Norwegian.)
The application knows exactly the schema of its data, let it do the parsing.
But yeah, the YAML one is crazy.
I do like "strict yaml."
Shameless plug if you wanna read a longer analysis: https://beepb00p.xyz/configs-suck.html
A great example is pyinfra https://github.com/Fizzadar/pyinfra#readme Think Ansible but instead of YAML you write Python. It provides a set of primitives/DSL and some rules you need to adhere to, but otherwise you just write regular python code.
There are a few keys:
1. Don't reinvent constructs. If your XML looks like <element name="foo"><attribute key="bar" value="biff" /> </element>, you've got a problem. It should be <foo bar="biff"/>
2. Provide reasonable defaults. Everything should be optional.
3. Break up files into manageable chunks
I wish everything in /etc/ used a common format, like YAML, or a small set of common formats.
I also wish for good configuration libraries, which handled:
- Parsing / validating YAML config files, command line parameters, environment variables, database overrides, defaults, etc.
- Printing pretty error messages
- Worked in a few languages (Python, JavaScript, and perhaps C)
- Guided developers towards sane designs.
Also, if you're putting objects in XML, documents in YAML, or strictly tabular data in either, you're probably doing something wrong.
- YAML: Configuration files, objects
- XML: Documents
- CSV/TSV/TSVX: Tabular data
This is what makes the Windows registry such a success. I think too many Linux people hated Windows for too long that they ignored this problem. Microsoft already had the solution in Windows 95. Early iterations maybe weren’t perfect, but they stuck with a standard API and everyone got onboard. It completely removes an entire set of problems that Linux people keep reinventing the wheel for.
I am glad it's standardized. I'm not glad it's binary, huge, and opaque.
Intuitive behavior - unification by design is unordered (ie. it doesn't matter in which order input is processed, it'll produce the same result), you get what you see.
Verification - ie. you can do one-liner policies easily to have some structure with desired types/values etc.
It's just well designed (based on good theory), practical language for configuration.
We use it in critical project to describe services on different envs, it works very well.
> YAML may seem ‘simple’ and ‘obvious’ when glancing at a basic example, but turns out it’s not. The YAML spec is 23,449 words; for comparison, TOML is 3,339 words, JSON is 1,969 words, and XML is 20,603 words.
Grafana Labs, Datadog's Vector, the Istio Service Mesh project, Dagger, sigstore, the SLSA framework all use CUE in some fashion. Dagger wrote a great blog post on CUE [0]
You can't make me love YAML!
(but I'll still use it and put my hatred to the side!)
Well. This makes a whole lot of sense.
- YAML is fine when you can use it for last mile configuration and don't have to use it for everything else
Well, I agree. Take Helm, for example. values.yaml is a perfectly nice configuration file that you can even enjoy editing. The chart itself is an abomination of endless boilerplate (not Helm's fault, but inherited from k8s) and the worst possible templating mechanism bolted onto it (totally Helm's fault).
Other than the slow pace to bake and test golden images I still prefer this approach. It is simple and a proven design and easy to build-out on every single cloud provider.
I might be biased tho, this is how we manage the host's for our several thousand Kubernetes clusters. I'm just kind of used to it.
- Local variables
- Pure functions with basic math and string manipulation capabilities
Possibly with an extended version that also has:
- Input variables (that must be provided by the code reading the file)
- Exportable functions that are callable by the host code
This extended version would be perfect for use cases like defining CI pipelines.
A language that did this but was otherwise simple and allowed no access to the host system, and had widely available library support across most common languages would be just awesome.
From what I hear people saying I'd guess its main usage is being a general purpose language that creates those shitty YAML based formats as outputs.
Anyway, I'm in complete agreement with the article. CI pipelines and infrastructure as code should both be defined in code, real code. Configuration files are the place you set the values for your variables, not the place you do calculations with them.
I think creating an entirely new Turing-complete language just for configuration is a waste of brain cells for everyone involved.
That being said, if you're only evaluating the file one time, just write a script to spit out your config.
Also, YAML has more features, and YAML parsers have to be a whole lot more complicated as a result. This is important if you use that parser at a security boundary!
This all sounded like a good idea until I read this. Now I too shudder with fear.
Significant whitespace is also a problem but I think that war has been lost.
A couple benefits: 1) all fields are documented and strongly typed, with backward and forward compatibility (old binary + new config, new binary+ old config works). 2) you can use any language to generate the config, simply specify your proto-generator binary as dependency. 3) if your config is simple enough and willing to lose some backward/forward compatibility, a text proto file is perfectly human readable/writable.
change my mind.
disclaimer: hate both
--- --- - --- --- ---
\ \ / / / \ | \/ | | |
\ / / ^ \ | \ / | | |
/ / / / \ \ | |\/| | | ---|
---- --- --- --- --- ------
(At least http://www.yamllint.com/ says so.)no action required
But YAML, as it is "intended" to be written? I think I'd rather have XML as well...
If you're only reading and writing the files via a library, then XML is probably better.
Isn't the problem more the tooling than the configuration language specification? Can you imagine authoring office documents by hand? Or maintaining them?
My text editor, KeenWrite, allows users to reference YAML variables as paths while editing documents. The editor includes a collapsible tree view of the variables, which makes navigation and maintenance a breeze.
https://i.ibb.co/qp0vj1Q/yaml-editor.png
Also, when inserting a variable into the document, the YAML tree expands automatically to the location the variable is defined. From the end user perspective, the data format is irrelevant.
Here's an example that produces a user-friendly UI for a JSON schema:
Another thing that would help in these template languages is if they'd pervasively track the sources of things, and let you view the final output of the templating process with some mechanism for viewing those annotations. Then when you look at the final result of some templating operation and can't figure out why "AWS_ENV_VAR" is "lol don't deploy this", you have some chance of figuring out where it came from.
This is yet another entry in the class of things that aren't too hard to add from the beginning, but good luck retrofitting it on after the fact.
Just compare Ansible/Salt YAML+python (because you can't do anything YAML-only) to Guix System, just see k*s YAML or Home Assistant ones: they are horrible to manage by a human. They are not much effective respect of anything else for a machine.
Long story short IMVHO it's about time to recognize that the old "data and language together" is the way to go, erasing common DSLs/formats/parsers etc.
[1]: http://ix.io/3XTk
What if the way you wrote the component does not allow some sort of dynamic config you need through YAML. Now you are working in two separate files and two different languages trying to make a change. Things almost get worse when the writer of the component and the consumer of the component are different people. Either massive horrible cludges happen to the YAML, or a weird back and forth needs to happen.
I usually run into it with webserver config files where (in both apache and nginx) I have to do something moderately contorted to get the end result I had in mind, but in the end the general case is just another trade-off.
There is a tree of defaults from the repo and a tree of config which overlays it with any changes.
Most of the settings are 0/1 or a one-line string, but a few are multi-line.
This also allows me to store templates under the same tree, which is where most of the code lives.
Because I use one call for getting and setting a config setting, I can probably rearchitect it in the future to e.g. a key-value store.
But for now, it just makes it so easy to maintain from the shell, and to create overlay themes, and so many other possibilities that I'd never dreamed of before.
TOML is less than ideal with (as was pointed out) nested tables – the problem I ran into is that some valid TOML cannot be serialized with serde.
You are limited to very few types, important ones like dates are missing, JSON is a pain to write and format without a "smart" editor. It’s also a pain to read as soon as you start listing complex things.
Not saying here that YAML or TOML are better (though I personally prefer TOML) but more that nothing is really nice.
It’s just that if we have to give up on a config format to be easily editable with basic tools, maybe we should have stuck with XLM. At least it had solid format definition and enforcement. It’s a semi-troll but maybe we could have something in between XLM and YAML.
ie. YAML. Write your JSON with comments and parse it with a YAML parser. Most decent libraries spit out readable YAML or even JSON if you tweak it to always quote strings.
if only people could get past the terror and see the light...
Of course if you can stay away from somewhat complex hierarchies then great! I can't imagine TOML ever working for something like a CI system or terraform config though.
Having recently migrated from cloudformation to terraform on a project the use of HCL is wonderful in comparison and when I'm all ready to commit I just type terraform fmt and everything is aligned properly. Considering it's just as trivial to do this in any JSON prettifier...I don't understand the need for yaml besides pure asthetics which is why you either love it or hate it, because it doesn't help or hurt it just annoys you or it doesn't.
I don't have a better solution but it just feels needlessly complex for text.
What you say the point of the article is? Submarine story is an acceptable answer.
We really need a language and not serialization format to describe configuration.