Why are we templating YAML?
leebriggs.co.uk
leebriggs.co.uk
I think this approach is a byproduct of thinking about infrastructure and configuration -- and the cloud generally -- as an "afterthought," not a core part of an application's infrastructure. Containers, Kubernetes, serverless, and more hosted services all change this, and Chef, Puppet, and others laid the groundwork to think differently about what the future looks like. More developers today than ever before need to think about how to build and configure cloud software.
We started the Pulumi project to solve this very problem, so I'm admittedly biased, and I hope you forgive the plug -- I only mention it here because I think it contributes to the discussion. Our approach is to simply use general purpose languages like TypeScript, Python, and Go, while still having infrastructure as code. An important thing to realize is that infrastructure as code is based on the idea of a goal state. Using a full blown language to generate that goal state generally doesn't threaten the repeatability, determinism, or robustness of the solution, provided you've got an engine handling state management, diffing, resource CRUD, and so on. We've been able to apply this universally across AWS, Azure, GCP, and Kubernetes, often mixing their configuration in the same program.
Again, I'm biased and want to admit that, however if you're sick of YAML, it's definitely worth checking out. We'd love your feedback:
- Project website: https://pulumi.io/
- All open source on GitHub: https://github.com/pulumi/pulumi
- Example of abstractions: https://blog.pulumi.com/the-fastest-path-to-deploying-kubern...
- Example of serverless as event handlers: https://blog.pulumi.com/lambdas-as-lambdas-the-magic-of-simp...
Pulumi may not be the solution for everyone, but I'm fairly optimistic that this is where we're all heading.
Joe
Because your build then becomes an actual program (i.e. Turing complete) and you have to refactor and maintain it! This is the common problem of using a "programming language as configuration" (e.g. gulp?)
Dhall solves exactly this problem: https://dhall-lang.org
It has the same premises of Pulumi, but without the Turing completeness (I don't know if/how Pulumi avoids that, but if it does it should be part of the pitch), so you cannot shoot yourself in the foot by building an abstraction castle in your build system/infrastructure config.
We use it at work to generate all the Infra-as-Code configurations from a single Dhall config: Terraform, Kubernetes, SQL, etc.
And there is already an integration with Kubernetes: https://github.com/dhall-lang/dhall-kubernetes
This is the key bit and not something which is pitched well enough from the Dhall landing pages: using straight YAML forces you to repeat yourself in multiple areas for each Individual tool being used, and these repetitions have to stay consistent across multiple tools. What Dhall does is allow you to write a single config and use it to derive the correct configurations for each tool that you use. So you can write a single configuration file from which, eventually, every single part of your system is derived - Terraform infrastructure, Kubernetes objects, application config, everything. When you pull it off, it's simply magical.
You can think of it like this: JavaScript is a horrible, no-good, very bad language, and yet all browser programming is done in JavaScript because every browser supports it - so too, are JSON and YAML horrible configuration languages. But JavaScript gave rise to abstractions like TypeScript which are much better languages which compile down to JavaScript for compatibility. TypeScript is to JavaScript what Dhall is to JSON and YAML - the fact is, pretty much everything is configured with JSON and YAML, and Dhall makes it much, much easier to live in that world, with no need for the systems being configured to support it.
Considering the relative obscurity of Dhall, it's basically the best-kept secret in the DevOps world right now, and it's a shame more people don't know about it.
Writing Dhall code look exactly like programming to me, and the programmer must possess the necessary programming skills to produce good Dhall code. A random guy with a text editor will make an equal mess in Dhall as they would with a “real” programming language.
I don't see how the restrictions in Dhall really help much in this regard. Turing completeness feels like a red herring to me.
For TC languages, comparing if two programs (original and refactored) do the same thing is not solvable in general. If the language is not TC then it is more feasible.
Sure, a TC program may not finish to produce output you can compare, but in my experience that's only a theoretical problem.
Give me a real language any day over dhall or jsonnet.
https://github.com/dhall-lang/dhall-lang/wiki/Safety-guarant...
The real reason for giving up Turing equivalence was probably to get dependent types. This gives very powerful static guarantees, including the presence/absence of fields under non-trivial record operations such as merge. In using dependent types, they have also had to give up significantly on type inference, which is really going to annoy the average JavaScript/Ruby programmer.
You already have to do that, so why not do it in a reasonably powerful language?
Also you might be familiar with the Rule of Least Power: https://en.wikipedia.org/wiki/Rule_of_least_power
I think that you’re right, and I think it’s great, because we have a programming model in which code is data and data is code: Lisp & S-expressions.
It’d be downright awesome to have a Lisp-based system which used dynamic scoping to meld geographical & environmental (e.g. production/development) configuration items. But then, it’d be downright awesome if the world had seriously picked up Lisp in the 80s & 90s, and had spent the last twenty years innovating, rather than reïnventing the wheel, only this time square-shaped. But then, the same thing could be said about Plan 9 …
I’ve not yet had the time to take a look at Pulumi, but I hope to have time soon.
"Any sufficiently complicated C or Fortran program contains an ad-hoc, informally-specified, bug-ridden, slow implementation of half of CommonLisp."
http://wiki.c2.com/?GreenspunsTenthRuleOfProgramming
Seriously, this has happened again and again and again. You have software, so you configure it via a clean and simple text syntax, then the configuration needs to be generated and the syntax becomes more complicated, then the next system you do has an "API" instead so you can configure it via programming, which is too complicated so the next time you Do it Right and go with a simple text file, which is then outgrown when the configuration it stores becomes too complicated...
It's like a circle of life thing.
Compare with: strongly vs weakly typed languages
The only requirement is a commitment to doing things imperatively in a real programming language. It’s hard to resist the temptation to do things declaratively (because it’s easier to imagine a declarative interface that describes your problem than an abstraction of the procedure which will solve it) but you are never forced to.
It has become yet another community that's fighting a struggle that everyone else ended years ago, like the few Japanese in jungles who refused to surrender. I'm not entirely sure why it's not been adopted, but I suspect it's because most people strongly prefer (a) visually semantically different scope delimiters and (b) function-outside-brackets syntax ie f(a, b) rather than (f a b).
Or you could go the other way and say that JSON is s-exps with curly brackets so it should be made executable as such, and build that language.
That's probably true, but I think it's useful to fight the good fight regardless. Even if Lisp & s-expressions don't, in fact, take over the world (and I think they will), arguing in their favour might help increase the chance that whatever inferior technology does end up getting adopted is better than it could have been.
> Or you could go the other way and say that JSON is s-exps with curly brackets so it should be made executable as such, and build that language.
The problem is that without symbols, that ends up being hideously ugly. This:
["if",
["<", 1, 2],
"less than",
"greater than or equal to"
]
is appreciably worse than: (if (< 1 2)
"less than"
"greater than or equal too")
And alternatives like: {"if": [[1, "<", 2], "less than", "greater than or equal to"]}
are so much worse that I don't think anyone could seriously expect to use them.Nice imagery, but the wrong point.
Except for the syntax, everybody else joined Lisp.
"We were not out to win over the Lisp programmers; we were after the C++ programmers. We managed to drag a lot of them about halfway to Lisp." --Guy Steele
Flash back to the mid-1980's (when the mainstream was C, Pascal, BASIC, FORTRAN, COBOL, etc.) and it's Lisp/Scheme (and Smalltalk) that have features like Garbage Collection, interactive development, lexical closures, decent built-in data structures, dynamic typing.
The fact that all of this is commonplace today, both justifies a lot what Lisp did in the first half of its existence and undermines its (technical) competitive advantages now.
> but I suspect it's because most people strongly prefer (a) visually semantically different scope delimiters and (b) function-outside-brackets syntax ie f(a, b) rather than (f a b).
It's not technical. I don't think it ever was. So much of it is around social concerns: a performance stigma dating back to the 1970's, fear of being able to hire people to do the work, fear of what VC's will think, worries that the language will still be available... And then at the end of the day, the problems whatever language will solve are a tiny fraction of the overall problem of doing something relevant and lasting and useful to others.
> As the kids say: stop trying to make Lisp happen, it's not going to happen.
Life is too short and the world is too big to try to confine other people's ideas of how they should think or work.
The point of the market economy and of the scientific process is that people get to try what they think is going to be useful and then let the world decide. The fact that Lisp is still in the conversation at all, when its contemporaries (Autocoder, Fortran) either aren't or are highly specialized, says a lot that we can learn from.
Mean girls came out in 2004, no kid knows that movie
I also feel this is where JSX got it right. Instead of creating yet-another-templating-language (looking at you Angular!), they used JavaScript and did a great job of outlining how interpolation works. Any new templating language is always going to be missing some key feature you expect out of a general programming language and your customers will continue to ask for more features.
Take for example Terraform and HCL, they're continually adding more and more [templating features](https://github.com/hashicorp/terraform/blob/master/website/d...) and [functions](https://www.terraform.io/docs/configuration/interpolation.ht...) because there's so many different ways to skin configuration/infrastructure as code. What if TF just expect a "computed" JSON object and it's left up to the developer to figure out how to put it together?
I'm gonna keep an eye on Pulumi and hope to be able to use in a real project soon.
<AutoScalingGroup name='Main cluster'> <LaunchConfig imageId='ami-xx'>...</LaunchConfig> </AutoScalingGroup>
Paired with Typescript, we would have the clearness of a declarative language, with the power and flexibility of a real language that is also easy to extend and navigate.
As a bonus, most tooling already exists.
https://docs.microsoft.com/en-us/aspnet/core/mvc/views/razor...
In ROS2 the launchfile can now just be a Python script. Very much learned all this the hard way and the solution was to just support Python. I think it's brilliant.
Configuration files for each component were a DSL made of Tcl functions. Each module just sourced the respective file on load.
- the django like situation: the configuration is pure code, and it's a mistake. It was not necessary, it brought plenty of problems. I wish they went with a templated toml file.
- the ansible like situation: the configuration is templated static text. But with something as complex as deployment, they ended up adding more and more constructs, until they created a monstrous DSL on top of a their implementation language, with zero benefits compared to it and plenty of pitfalls. In that case, they should have made a library, with an API and documentation making an emphasis on best practices.
- and of course a big spectrum between those
The thing is, we see configuration as one big problem, but it's not. Not every configuration scenario has the same constraints and goals. Maybe you need to accept several sources of data. Maybe you need validation. Maybe you need generation. Maybe you to be able to change settings live. Maybe you need to enforce immutable settings. Maybe you need to pub sub your settings. Maybe you need to share them in a central place. Maybe they are just for you. Maybe you want them to be distributed. Maybe you need logic. Maybe you want to be protected from logic. Maybe the user can input settings. Maybe you just read conf. Maybe you generate it.
So many possibilities. And that's why there is not a single configuration tool.
What you would need, is a configuration framework, dealing with things like marging conf, parsing file, getting conf from the network, expressing constraints, etc.
But if you recreate a DSL for your config, it's probably wrong.
It may have its problems (I don't have many issues with it) but it doesn't seem to have this problem of attracting ever more layers of abstraction on top of it. It works.
There should have a schema checking the setting file. There should have a better way to extend settings, and make different settings according to context, such as prod, staging or dev.
There should be a linter avoiding stupid mistakes like missing a coma in a tuple, resulting in string concatenation.
There should be variables giving you basic stuff like current dir, log dir, var dir, etc. We all make them anyway.
And there should be a better to debug the import settings problem.
But all in all, it's quick and easy to edit, and very powerful.
The different context stuff can be handled by using env vars, and a nice python wrapper, like python-decouple.
Does anybody here personally suffered those problems that the Turing complete Django configuration creates? (I mean, not the ones caused by lack of a completness checks, or good library support, but the ones caused by too much power.)
If so, how do those problems look like?
I never had an untrusted party editing my config, nor did I use data from any.
Also, you can make the same mistakes in the setting file that in any code file, but it's not more or less important.
In fact, all the problem I had could have been solved by better integration: solving the import problem, making composition easy, adding checks, allow loading data from several sources and merge them, presenting them in a unify interface.
If I'm being honest, problem with settings.py may have not been that it's Python, but that it's a flat file with no strong conventions, tooling or best practices.
I could raise the issue that you can't read the config from another language, but I never had to, and good tooling would allow a synced export or an API to consume the settings.
Same for writing, or live settings.
Unfortunately, what described here is good in many level, but not excellent in any.
If you are OK to describe the complexity of your infrastructure in a programming language similar to the general purpose language, then a well abstracted API built on original APIs from cloud providers are more familiar to devs. And it will be more reliable performance and flexible.
If you want a config experience, something like kustomize is leaner and more compatible with the text config model.
I also cannot see how this interoperate with other tools, which will seriously limit it's appeals to people using other tools.
This has long been a problem in the python/pip community, as its basically impossible for the build tools to determine the dependencies of a package without fully downloading and running the setup.py file.
Static config files are static for a reason!
So, what's my point? My point is configuration languages help in that they push "the one true way" and help to enforce it. Sure there are times you end up having to work around the one true way but given very powerful tools of a full language for configuration leads to chaos or at least that's my experience. Instead of being able to glance at the configuration and understand what's happening because it follows the one true way you instead end up with configuration language per programmer since every programmer will code up stuff a different way.
I appreciate that their revenue model doesn't require making the open-source version frustrating or stupid and I appreciate that they're incredibly responsive. And some of the stuff you'll see around cloud functions/Lambdas and the deployment thereof will fucking blow your mind.
It's good. You should strongly consider it.
Joe Beda will be doing a deep dive on Pulumi on the TGIK videocast tomorrow, so it's a timely opportunity to check it out: https://twitter.com/jbeda/status/1092963296565587969
https://leebriggs.co.uk/blog/2018/09/20/using-pulumi-for-k8s...
I do think this is more like what we should be doing, but as dismayed to see Pulumi’s free tier get sunsetted
This is the why I prefer to use a JS file for configuration instead of native JSON or YAML file if those options are available.
That being said, the concept of defining a function in, essentially, a config file seems like a step in the right direction. I don't think I'd trust that functionality outside of builds or infra-as-code, though.
It probably only seems like magic because you didn't build a fundamental understanding of how it works before using it. I use some massive webpack configurations and I understand them all quite thoroughly thanks to well-written, modularized configuration files.
Webpack now does simple config as well with the 'mode: "production"' and 'mode: "development"' presets.
DSLs ought to be type safe and type checked since getting things wrong means all kinds of trouble. E.g. with cloudformation I've wasted countless hours googling for all sort of arcane weirdness that amazon people managed to come up with in terms of property names and their values. Getting that wrong means having to dig through tons of obscure errors and output. Debugging broken cloudformation templates is a great argument against how that particular system was designed. It basically requires you know everything listed ever in the vastness of its documentation hell and somehow be able to produce thousands of lines of json/yaml without making a single mistake, which is about as likely as it sounds. Don't get me started on puppet. Very pleased to not have that in my life anymore.
On a positive note, kotlin recently became a supported language for defining gradle build files in. Awesome stuff. Those used to be written in Groovy. The main difference: kotlin is statically compiled and tools like intellij can now tell you when your build file is obviously wrong and autocomplete both the standard stuff as well as any custom things you hooked up. Makes the whole thing much easier to customise and it just removes a whole lot of uncertainty around the "why doesn't this work" kind of stuff that I regularly experience with groovy based gradle files.
Not that I'm arguing using Kotlin in place of Json/yaml. But typescript seems like a sane choice. Json is actually valid javascript, which in turn is valid typescript. Add some interfaces and boom you suddenly have type safety. Now using a number instead of a boolean or string is obviously wrong. Also typescript can do multi line strings, comments, etc. and it supports embedding expressions in strings. No need to reinvent all of that and template JSON when you could just be writing type script.
I recently moved a yaml based localization file to typescript. Only took a few minutes. This resulted in zero extra verbosity (all the types are inferred) but I gained type safety. Any missing language strings are now errors that vs code will tell me about and I can now autocomplete language strings all over the code base which saves me from having to look them up and copy paste them around. So no pain, plenty of gain.
And yes, people are ahead of me and there are actually several projects out there offering typescript support for cloudformation as well.
Thanks for plugging!
We are actively working on https://github.com/pulumi/pulumi/issues/2430, which will make it easier for our small team to manage multiple languages. Once that lands, I would expect this to be high priority.
Some of our amazing community members have been prototyping this, and it's looking pretty promising: https://twitter.com/MikhailShilkov/status/109278757393889689....
Joe
Powershell would be great, it has nice support for building DSLs.
I don't like it in Python either, but for some reason, when I write Python, it's a lot easier. Maybe YAML is just a bit more complex (and Python has better IDE support..?)
Okay, I'm gonna be the asshole in the room, but how hard is it to just use consistent indentation? I can't count how many times I've heard people complain about significant whitespace in languages.
Not only is it not difficult to begin with, but every code editor and IDE will show you where there's a syntax error in your YAML. People are free to dislike YAML, even for its significant whitespace, but how does it "kill you"?
Look at this example from the article:
```
something: nothing
hello: goodbye
```This is pure sloppiness, and anyone who has trouble carelessly adding pointless bytes to code, no matter the language, is sloppy. I don't understand why people criticize YAML and Python because "whitespace is hard".
P.S.: There's a similar configuration language called ArchieML, which is similar to YAML but doesn't have significant whitespace.
- "cut and paste and edit" is broken. You can't autoformat the pasted code into the right place, you have to go back and fix the whitespace. Since whitespace is semantically significant, this can introduce bugs.
- visually identical whitespace may not be textually identical whitespace. Unless you go around breaking the tab key off your colleague's keyboards you'll trip over this. Especially (again) if you paste. Occasionally seen in merges too.
- editors can no longer give you 100% correct indentation.
Depends on how your editor is configured / it's feature set. Which makes me wonder how editorconfig would handle this when enabled. It seems like a insignificant issue to me, you can auto-PEP8 the code before pasting it. You should probably be following PEP8 anyway (as far as spacing is concerned at least).
> - visually identical whitespace may not be textually identical whitespace. Unless you go around breaking the tab key off your colleague's keyboards you'll trip over this. Especially (again) if you paste. Occasionally seen in merges too.
I turn on show all whitespace on my editors regardless of programming language. I've been burned by Sublime Text not just figuring out the already defined whitespace ruleset for a file by what it's using and just shoving in it's own defaults. I wish all editors would base whitespace on what the file's structure looks like, if there's mixed spaces, give me a warning.
> - editors can no longer give you 100% correct indentation.
I don't understand this, it sounds like you've got your editor configured poorly or something? But it goes back to how unintuitive the nice editors can be. You can use editorconfig to define the indentation project wide, then any editor should pick it up, of course if you define PEP8 at a minimum it guarantees spacing settings.
I'm not sure if PyCharm covers a few of those cases, since I use it so seamlessly I don't usually have complaints.
I ended up converting it to yaml, making the edits and converting it back to JSON.
Before anyone asks the obvious, how do I handle deeply nested code in brackets. Simple, I don’t. When things start getting nested deeply, I use my IDE to Extract Method.
In the Python case it's much better, because people less often casually edit .py files without editor support, and because Python has good diagnostics and it's much much harder to produce syntactically correct but semantically wrong Python by whitespace mixups.
However, missing/extra whitespace is not "hard". You would be docked points in an English paper and you should be docked points as a programmer.
So, whitespace aside... Tell me what is easier to edit without built-in syntax support: JSON, or YAML?
If we define "easy" as "how long it takes to complete a task" or "how quickly you can grok the structure of a given block of code", then YAML beats out JSON every time.
1) YAML is a configuration file format, and it's targeting user groups and environments where people use ad hoc terminal based or os-bundled editors, such tools being nano or Notepad, and such users being sysadmins for example. 2) YAML implementations (=parsers) have poor diagnostics compared to Python, separate from the editor issue, and 3) YAML syntax is more prone than Python to parsing correctly but producing unwanted semantics when you make a mistake.
I think there is value in your English paper analogy: many/most people editing YAML files don't know YAML syntax very well compared to this scenario. If their knowledge of English was at the same level, misplaced whitespace would not be chief of their problems in a graded English paper.
It is of course a structurally valid (philosophically consistent) argument that people should not make mistakes and they should suffer when they do, but this goes generally against the consensus of configuration language usability thinking.
How hard is it to use HN formatting? I can’t count how many times people screw it up.
It’s not difficult to begin with, the documentation is free, yet here I am reading your comment with broken formatting.
something: nothing
hello: goodbye
Anyone who has trouble with this is just being sloppy. No useless backticks! You might think you’re doing it right, but unless you check, maybe you’re not. …
[4]mode:[1]mode-specification
[4]acls:
[6]-[1]acls-specification
[4]context:
[6]user:[1]user-specification
…
"AWS CodeDeploy will raise an error that might be difficult to debug if the locations and number of spaces in an AppSpec file are not correct."Great. There couldn't possibly be an easier format to use, could there?
[1]: https://docs.aws.amazon.com/codedeploy/latest/userguide/refe...
Another take, perhaps: Assigning deep semantic significance to invisible symbols is simply stupid. It is stupid to a much greater degree than wanting to be free from having to care about the amount of invisible symbols is “sloppy”.
It's pretty annoying when you don't have access to an IDE or decent editor.
{
a: 42,
# But you can have comments!
b: "hello world",
c: "and
multi-line
strings!", # and trailing commas!
}https://github.com/crdoconnor/strictyaml
Worse than JSON though, is the Norway problem.
If you remove this stuff and start validating it properly it becomes much easier to maintain.
Huh:
https://hitchdev.com/strictyaml/why/implicit-typing-removed/
I've always liked YAML, it's always seemed pretty intuitive to me coming from Python, and I like human-readable resource files, but those are some pretty damning counterexamples.
And depending on how exactly it was accessed, this can go two ways. The best case is that the app just gets the value via the untyped API, casts it to string, and blows up with an invalid cast - best because you actually know what went wrong.
The worst case is when the app specifically tells JSON.NET that it wants a string value (via generic type parameters), at which point it will helpfully implicitly convert the actual date value back to a string... except it can reformat it, and even helpfully adjust it from one timezone to another. Semantically it's the same date, of course, but it's not at all the same string, and sometimes that matters a lot. So this is the worst case because it's just silent data corruption.
For some mysterious reason, the author believes that this is acceptable default behavior - i.e. "it's a feature, not a bug". It's especially ironic to look at all the mentions in GitHub ticket, as various projects that rely on the library run into this issue (one of them is mine):
This is insane. I already knew writing a YAML unmashaller was a needlessly complicated affair - and that was before I realised this could happen.
It's a step up from XML though.
Also, I'm guessing I'm in a tiny minority who loves YAML but hates Python's semantic indentation...
In programming languages, it makes me twitch. I don't have any problem with "accepted" formatting styles (i.e. linux kernel c style), but for the language itself to enforce that for some reason feels like it's adding perpetual cognitive overhead (like whenever I use python). I don't know why; it shouldn't be any different than using a particular formatting style voluntarily in a more flexible language, but somehow, it feels different.
None of these problems are hard, why are all the solutions so awful?
foo: bar # {"foo":"bar"}
foo: "bar" # {"foo":"bar"}
foo: 42 # {"foo":42}
foo: "42" # {"foo":"42"}
Other than that, no major complaints. My editor understands YAML and shows the indentation level in the background (highlight-indentation-mode) and auto-formats files so they all have consistent indentation (prettier-mode). As a result, it is not much of a nightmare to edit, despite the fact that semantic whitespace COULD cause you a lot of problems. foo: yes # "foo": true
bar: YEs # "bar": "YEs"
baz: YES # "baz": trueI wasn’t aware anyone liked it
I'm sure it will be just like XML where it's the trendy thing for a while in the early days, then everyone stops and hates it for a while. Except XML at least has a handful of applications where it's the right tool for the job (it has a nice streaming mode), YAML doesn't even have that.
Opting for YAML over XML/JSON/whatever doesn't make me a tool. It made life much easier for myself & my colleagues.
I guess YMMV, but after you've used both YAML and JSON for a while, you might appreciate YAML a little bit more.
See also ~5 years ago: https://news.ycombinator.com/item?id=7325735
Then it's not valid JSON, and if you try to treat it as such bad things will happen.
> I guess YMMV, but after you've used both YAML and JSON for a while, you might appreciate YAML a little bit more.
I've used JSON a lot, and XML and s-expressions and MessagePack and ini and YAML and a whole bunch of other formats.
I usually have to fire up Google to read YAML. YAML is the only one where I routinely have to Google for a syntax cheatsheet and wade through tables of redundancy and edge-cases.
YAML made sense before JSON became a thing. Why people persist with it in new projects is baffling to me.
With YAML it's difficult to even know the structure of what I'm looking at, due to anchors and extensions. It's also hard to discern structure from skimming, since strings can appear unquoted, and may contain unescaped lexical tokens (depending on which particular symbols it started with); hence we must carefully consider each and every character, rather than just skimming for the next token.
If I know I'm looking at a perfect YAML file, than I should be able to guess the gist of what it says, since I can make assumptions about what the syntax means. If I want to be sure, I'd be Googling for cheatsheets. Yet as a programmer, I mostly look at files when they're buggy, meaning I can't just assume that, say, an unescaped quotation mark won't terminate the string; or that a certain piece of text is allowed to run across multiple lines; or that the indentation corresponds to the nesting; etc.
My life improved a lot since I got an YAML mode for emacs. Now things would be just perfect if haskell's cabal migrated to dhall...
Looks like he doesn’t believe the code approach is viable as much as other people are claiming in this thread.
It also has very nice bindings with haskell and nix
Since the author of Dhall comes from the Haskell community, he's kept this style.
The advantages are:
- All of the separators are in the same column, along with the opening and closing characters. This makes it trivial to check if we've missed a separator.
- Appending new lines to the end will not affect previous lines (i.e. we don't need to go and add a comma). This avoids making mistakes and polluting diffs.
Unfortunately the error-prone diff pollution we avoid at the last line instead occurs at the first line. It's still less error-prone than trailing commas, since we can look in the separator column and either spot that it's empty, or that it contains two opening braces (depending on whether we inserted or copy/pasted).
Allowing trailing commas, like Python does, would be really great. Unfortunately trailing commas already mean something: (a,b,) is a function that still takes 1 argument to make a triple. It's called "TupleSections".
I think there are three very different kinds of tools that people use for this:
1. Interpolation/preprocessor languages: This is what the author is talking about. There are delimiters/tags/sigils to distinguish "the templated parts" from "the rest" and the primary operation done by the template engine is substitution. "The rest" is literal content that's already in $FORMAT and it remains mostly/entirely unchanged during template rendering. Languages of this type are basically glorified `sed`. This can be nice because they're agnostic as to their embedding (any string will do) so they're very portable/flexible (you don't have to create "handlebars for YAML", "handlebars for HTML", "handlebars for CSV", etc; one implementation does it all). Languages of this kind can work in the small but don't scale well for all the reasons mentioned in the article/comments. The language doesn't know anything about the semantics of $FORMAT and that can cause all kinds of pain. Examples include golang templates, PHP, ERB, handlebars, the C preprocessor, Jinja, etc.
2. Compilers/code generators: These are "complete" languages that compile to $FORMAT. The difference between these and and interpolation/preprocessor languages is that the entire input is the language, not just specific chunks/tags. This kind of language can be nice because you have complete control and can therefore guarantee valid output and do tricks like supporting multiple different output formats for the same input, but the downside is that you're working with an entirely new language so there's a learning curve, you need specialized syntax highlighters and other tools to work with templates, etc. Examples include HAML, Jsonnet, Dhall, etc.
3. Embedded DSLs: Templates of this kind are valid $FORMAT from the beginning, but have embedded ways to specify transformations to be applied to the parsed AST. These languages are homoiconic with respect to $FORMAT. First $FORMAT is parsed, then the template engine iterates through the AST to perform evaluations, then the result can either be used as-is in memory or serialized back to (a possibly different) $FORMAT. This is sort of like an interpolation/preprocessor language with the evaluation order swapped: preprocessing is "run the template engine, then parse $FORMAT" while this is "parse $FORMAT, then run the template engine". A downside of this approach is that it is less general, e.g. it only really makes sense when $FORMAT has a well-defined structure (you probably can't template plain english sentences with this approach), but these days most "data languages" have converged towards being semantically equivalent to JSON (lists, dictionaries, and primitives) and this approach works well for any of them. An upside is that like compilers/code generators you can guarantee that the output will be valid $FORMAT no matter what the template looks like. Examples include JSON-e, Lisp macros, CloudFormation templates, etc.
It's unfortunate that all of these get called "templating languages" because they're very different beasts from one another, and usually when I see conversations about this stuff these distinctions get blurred and you end up with apples-vs-oranges comparisons. If I had my druthers we'd reserve the word "templating" for the first one and use different terminology for the others, but that ship has sailed.
Even though YAML is not optimal, it is a human friendly compromise between too verbose XML and machine only JSON. It lacks native templating, leading to funny constructs e.g. with Ansible files. However human kind has made progress and will make progress further, so it is just a matter of time until someone comes up with sane "native templated YAML" and all projects will adopt it.
I actually really like the idea behind XSLT: machine-friendly, human-tolerable, structured data + declarative rules for turning that data into a display, or a report, or whatever else.
The execution was horrible though: incredibly verbose, lots of overcomplication due to XML weirdness/asymmetries (e.g. attributes vs elements vs text, namespaces, ...); mixtures of different languages hidden inside each other (e.g. XPath hidden in attributes); etc.
I would really like to see what this could look like if done in a more minimalist, lispy fashion (normal code-is-data stuff in Lisp is similar, but I think term-rewriting is a more appropriate evaluation mechanism for such rules)
jq occupies the same role as XSLT, but for JSON. It can be used for templating but it's not quite as declarative as XSLT (you must pipe things through).
Yes, I didn't mean to imply that XPath itself is bad (although it also has to handle XML quirks like element/attribute/text, etc.).
Rather that the reason to write XSLT as XML in the first place is that it's machine readable, we can mix and match elements from different vocabularies, etc. yet most of the heavy lifting ends up as opaque string attributes :(
PS: I've done a few projects which make heavy use of jq; it's really nice, but as you say it's more of a pipeline.
(parrot (@ (type "African Grey")) (name "Alfie"))
<parrot>
<type>African Grey</type>
<name>Alfie</name>
</parrot>
It's more the one-variable-per-line pretty-printing than the syntax as such, but still.In most if not all of these cases an existing and well designed turing complete programming language would likely have better served them.
Your point is valid though, the power seems to end up being needed, sometimes, in some parts, in some cases. Escaping to a full language when needed seems to retain the benefits of both worlds.
The idea behind pure function is nice, but in practice you ended up with hundreds of mini plugins in Java or Python.
I also rather liked thorough extensibility. Namespaces were the right idea, despite clunky syntax. Today you can see Clojure doing something similar in Spec.
And while we're on the subject of XML, XSLT and Clojure; I feel like this is the best solution for readable serialization of tree-like data, and an associated ecosystem of tools (to validate, transform etc). Note some nice features for humans, like the ability to comment out a specific node, in addition to the usual line-oriented comments.
Then we have built a whole new domain on things on the top of UTF-8.
I export custom queries for data to xml, Json, CSV and HTML.
Where HTML is XML + XSLT. It works great. And clients can even theme it
Horrible syntax is kind of a forgivable offense, lots of things have horrible syntax and work fine.
Generating YAML with go templates though.. it is just horrible on so many levels.
I definitely feel the 90s can take the higher ground over 2010s on this one.
kustomize is much more sane for your own stuff: https://github.com/kubernetes-sigs/kustomize
It is actually a little bit too magical for my taste, but I continue to use it because it hasn't done anything stupid. I have one file that maps logical names to images in a container repository. If I create a service called "foo" pointing to selector.app.label="foo" in the base, then in production it's called foo-prd and the label magically updates to foo-prd for the selector. It actually understands what it's generating, and while they might have taken it a little bit too far, it's far better than just dumb text replacement.
Here is why everyone should use Helm:
Helm 2.0 introduced package as a first-class concept for Kubernetes and created the standard to distribute applications, thanks to Helm thousands of people could discover and collaborate on cloud-native deployments of the open source software https://github.com/helm/charts/tree/master/stable published and managed by organizations and contributors all over the world.
Helm 3.0 keeps innovating, it adopts the most forward-thinking approach to package management and Kubernetes config management by using higher level domain specific language based on Lua to create expressive package management system:
https://sweetcode.io/a-first-look-at-the-helm-3-plan/
Helm is also backed by CNCF[1] and is the best choice so far for organizations to create a reproducible CI/CD pipeline in a Kubernetes cluster.
I don't really trust Helm to do anything that's actually useful in the long term. It will get something running very quickly, but whether or not it's maintainable, I am yet to be sure of. For example, very early on, I installed the helm chart for prometheus. Now I want it to live in the kube-system namespace because I am tired of seeing its resources in the default namespace. For some reason, I highly doubt that changing values.yaml to change the namespace is going to do anything other than give me a fresh instance of prometheus running in another namespace. It's not going to use the already allocated storage volume to satisfy the persistent volume claim in the new namespace. It's not going to update the other stuff in my cluster to refer to prometheus-pushgateway.kube-system.svc.cluster.local. It's not going to update my Grafana dashboards to refer to the new namespace, even though I installed Grafana with Helm! So what did I really gain? Helm isn't giving me the ability to manage the long-term lifecycle of third-party software. It just explodes some API objects all over my cluster and lets me delete most of them automatically. That's all it does.
I get why Helm is popular. You can get some piece of software running in Kubernetes with minimal effort. I would have never successfully made some random complex piece of software work correctly in Kubernetes on day 1, especially using something that assumes you deeply understand the core API objects like kustomize does. What that boils down to is that Helm doesn't go far enough, and in its current state, just encourages people to make mistakes early.
I've been waiting for it for over a year now...
YAML is insanely over complicated; it's as bad or worse than XML for config files, and it doesn't even have the nice streaming mode. Not to mention that it's a bit of a security nightmare (seriously, who put pointers into the YAML spec?).
And, on a more subjective note, YAML is just confusing: between all the significant whitespace and the random single character symbols that no one ever remembers what they do, I never get a YAML document right on the first try.
Templating it really does add a whole new level of headache too.
XML works very well for config files. It's schema-optional (but is there), well-specified, human-readable, has plethora of supporting technologies (making things like templating easy), and is well supported by every language.
At the very least it is way better than JSON.
If your config file requires more tooling than that you fucked up.
And for complicated stuff, you're going to spend a lot more time reading the manual than you will actually typing those closing tags. In fact, in most cases you'd be copy/pasting bits from the manual as well.
technically yes, practically maybe not so much, especially with e.g. CDATA sections
Machine generated XML can be noisy but the target for those are other machines, and the extra context is there for a reason.
You can certainly make XML as obtuse and complex as you want.
Codelite's xml based project files can be easily read and modified by hand. Diffing them yields useful information about files added and moved, config values changed, etc.
Eclipses project files also written in xml are an Eldritch Horror.
I think the failing of xml is also it's strength. It doesn't do typing and schemas, doesn't even try. Which means that can be sane. Or not.
<?xml version="1.0" encoding="UTF-8"?>
<fof>
<!-- Common settings -->
<common>
<!-- Container configuration -->
<container>
<option name="componentNamespace"><![CDATA[MyCompany\MyApplication]]></option>
</container>
<!-- Dispatcher configuration -->
<dispatcher>
<option name="defaultView">items</option>
</dispatcher>
<!-- Transparent authentication configuration -->
<authentication>
<option name="totpKey">ABCD123456</option>
<option name="authenticationMethods">HTTPBasicAuth_TOTP,QueryString_TOTP</option>
</authentication>
<!-- Model configuration. One tag for each Model. -->
<model name="orders">
<!-- Model configuration -->
<config>
<option name="tbl"><![CDATA[#__fakeapp_orders]]></option>
</config>
<!-- Field aliasing. One tag per aliased field -->
<field name="enabled">published</field>
<!-- Relation setup. One tag per relation -->
<relation type="hasMany" name="items" />
<relation type="belongsToMany"
name="transactions"
localKey="foobar_order_id"
foreignKey="foobar_transaction_id"
pivotLocalKey="foobar_order_id"
pivotForeignKey="foobar_transaction_id"
pivotTable="#__foobar_orders_transactions" />
<relation type="belongsTo" name="client" foreignModelClass="Users@com_fakeapp" />
<!-- Behaviour setup. Use merge="1" to merge with already defined behaviours. -->
<behaviors merge="1">foo,bar,baz</behaviors>
</model>
<!-- Controller, View and Toolbar setup. One tag per view. -->
<view name="item">
<!-- Controller task aliasing -->
<taskmap>
<task name="list">browse</task>
</taskmap>
<!-- Controller ACL mapping -->
<acl>
<task name="dosomething" />
<task name="somethingelse">core.manage</task>
</acl>
<!-- Controller and View options -->
<config>
<option name="autoRouting">3</option>
</config>
<!-- Toolbar configuration -->
<toolbar title="COM_FOOBAR_TOOLBAR_ITEM" task="edit">
<button type="save" />
<button type="saveclose" />
<button type="savenew" />
<button type="cancel" />
</toolbar>
</view>
</common>
<!-- Component backend options -->
<backend>
<!-- The same options as Common Settings apply here, too -->
</backend>
<!-- Component frontend options -->
<frontend>
<!-- The same options as Common Settings apply here, too -->
</frontend>
</fof>
is at all preferable to this: (fof
(common
(container (component-namespace "MyCompany\\MyApplication"))
(dispatcher (default-view items))
(authentication
(totp-key ABCD123456)
(authentication-methods (http-basic-auth-totp query-string-totp)))
(model orders
(config (tbl "#__fakeapp_orders"))
;; Field aliasing. One tag per aliased field
(field enabled published)
;; Relation setup. One tag per relation
(relation items (type has-many))
(relation transaction (type belongs-to-many)
(local-key "foobar_order_id")
(foreign-key "foobar_transaction_id")
(pivot-local-key "foobar_order_id")
(pivot-foreign-key "foobar_transaction_id")
pivot-table "#__foobar_orders_transactions")
(relation client (type belongs-to)
(foreign-model-class "Users@com_fakeapp"))
;; Behaviour setup. Use merge="1" to merge with already defined behaviours.
(behaviors (merge 1) (foo bar baz)))
;; Controller, View and Toolbar setup. One tag per view.
(view item
(taskmap (list browse))
;; Controller ACL mapping
(acl
(task dosomething)
(task somethingelse core.manage))
;; Controller and View options
(config (auto-routing 3))
;; Toolbar configuration
(toolbar "COM_FOOBAR_TOOLBAR_ITEM"
(task edit)
(button save)
(button saveclose)
(button savenew)
(button cancel))))
(backend)
(frontend))
There's just no way.Anyway, to each his own, but I think XML holds up very well and I do find it more readable and easier to work with that your lisp example.
I also never said XML is the best configuration format. For simple configurations a simple property file is by far the best option. For anything complicated (as in your example) XML does a great job. To contrast, JSON would fall flat on its face with this. Not to mention the fact XML parsing is typically part of the standard library of most programming language and most people are familiar with it.
That was a conscious decision, because the verbosity of XML prevents clear understanding of a data model, while the cleanness of S-expressions enables a clarity of vision which enables prudent judgement when laying out a data structure.
> You also removed some comments.
Yes, because they were akin to:
// Add 1 & 2, assign to X
x = 1 + 2
If you really want an S-expression version of the XML in that example, here is SXML[0]: (*top*
(fof
(*comment* " Common settings ")
(common
(*comment* " Container configuration ")
(container
(option (@ (name "componentNamespace")) "MyCompany\\MyApplication"))
(*comment* " Dispatcher configuration ")
(dispatcher
(option (@ (name "defaultView")) "items"))
(*comment* " Transparent authentication configuration ")
(authentication
(option (@ (name "totpKey")) "ABCD123456")
(option (@ (name "authenticationMethods"))
"HTTPBasicAuth_TOTP,QueryString_TOTP"))
(*comment* " Model configuration. One tag for each Model. ")
(model (@ (name "orders"))
(*comment* " Model configuration ")
(config
(option (@ (name "tbl")) "#__fakeapp_orders"))
(*comment* " Field aliasing. One tag per aliased field ")
(field (@ (name "enabled")) "published")
(*comment* " Relation setup. One tag per relation ")
(relation (@ (type "hasMany") (name "items")))
(relation
(@ (type "belongsToMany") (name "transactions")
(localKey "foobar_order_id") (foreignKey "foobar_transaction_id")
(pivotLocalKey "foobar_order_id")
(pivotForeignKey "foobar_transaction_id")
(pivotTable "#__foobar_orders_transactions")))
(relation
(@ (type "belongsTo") (name "client")
(foreignModelClass "Users@com_fakeapp")))
(*comment*
" Behaviour setup. Use merge=\"1\" to merge with already defined behaviours. ")
(behaviors (@ (merge "1")) "foo,bar,baz"))
(*comment* " Controller, View and Toolbar setup. One tag per view. ")
(view (@ (name "item"))
(*comment* " Controller task aliasing ")
(taskmap
(task (@ (name "list")) "browse"))
(*comment* " Controller ACL mapping ")
(acl
(task (@ (name "dosomething")))
(task (@ (name "somethingelse")) "core.manage"))
(*comment* " Controller and View options ")
(config
(option (@ (name "autoRouting")) "3"))
(*comment* " Toolbar configuration ")
(toolbar (@ (title "COM_FOOBAR_TOOLBAR_ITEM") (task "edit"))
(button (@ (type "save")))
(button (@ (type "saveclose")))
(button (@ (type "savenew")))
(button (@ (type "cancel"))))))
(*comment* " Component backend options ")
(backend
(*comment* " The same options as Common Settings apply here, too "))
(*comment* " Component frontend options ")
(frontend
(*comment* " The same options as Common Settings apply here, too "))))
Which I think is still indubitably and inarguably clearer & cleaner than the XML version.Technically, the XML spec requires whitespace preservation, so really it’s this:
(*top*
(fof "
"
(*comment* " Common settings ") "
"
(common "
"
(*comment* " Container configuration ") "
"
(container "
"
(option (@ (name "componentNamespace")) "MyCompany\\MyApplication") "
")
"
"
(*comment* " Dispatcher configuration ") "
"
(dispatcher "
"
(option (@ (name "defaultView")) "items") "
")
"
"
(*comment* " Transparent authentication configuration ") "
"
(authentication "
"
(option (@ (name "totpKey")) "ABCD123456") "
"
(option (@ (name "authenticationMethods"))
"HTTPBasicAuth_TOTP,QueryString_TOTP")
"
")
"
"
(*comment* " Model configuration. One tag for each Model. ") "
"
(model (@ (name "orders")) "
"
(*comment* " Model configuration ") "
"
(config "
"
(option (@ (name "tbl")) "#__fakeapp_orders") "
")
"
"
(*comment* " Field aliasing. One tag per aliased field ") "
"
(field (@ (name "enabled")) "published") "
"
(*comment* " Relation setup. One tag per relation ") "
"
(relation (@ (type "hasMany") (name "items"))) "
"
(relation
(@ (type "belongsToMany") (name "transactions")
(localKey "foobar_order_id") (foreignKey "foobar_transaction_id")
(pivotLocalKey "foobar_order_id")
(pivotForeignKey "foobar_transaction_id")
(pivotTable "#__foobar_orders_transactions")))
"
"
(relation
(@ (type "belongsTo") (name "client")
(foreignModelClass "Users@com_fakeapp")))
"
"
(*comment*
" Behaviour setup. Use merge=\"1\" to merge with already defined behaviours. ")
"
"
(behaviors (@ (merge "1")) "foo,bar,baz") "
")
"
"
(*comment* " Controller, View and Toolbar setup. One tag per view. ") "
"
(view (@ (name "item")) "
"
(*comment* " Controller task aliasing ") "
"
(taskmap "
"
(task (@ (name "list")) "browse") "
")
"
"
(*comment* " Controller ACL mapping ") "
"
(acl "
"
(task (@ (name "dosomething"))) "
"
(task (@ (name "somethingelse")) "core.manage") "
")
"
"
(*comment* " Controller and View options ") "
"
(config "
"
(option (@ (name "autoRouting")) "3") "
")
"
"
(*comment* " Toolbar configuration ") "
"
(toolbar (@ (title "COM_FOOBAR_TOOLBAR_ITEM") (task "edit")) "
"
(button (@ (type "save"))) "
"
(button (@ (type "saveclose"))) "
"
(button (@ (type "savenew"))) "
"
(button (@ (type "cancel"))) "
")
"
")
"
")
"
"
(*comment* " Component backend options ") "
"
(backend "
"
(*comment* " The same options as Common Settings apply here, too ") "
")
"
"
(*comment* " Component frontend options ") "
"
(frontend "
"
(*comment* " The same options as Common Settings apply here, too ") "
")
"
"))
But I think that rather proves my point: XML obscures that which should be obvious.(and apologies for these terribly vertical posts — I think that they go a long way towards demonstrating the need for a compact information representation).
We get better and better tools each year but it still seems unavoidable. We're ultimately building incredibly complex systems with each layer using multiple development approaches, style choices, language choices, degrees of quality/time investment by the creator, etc.
TLDR: you can't help bang your head against the wall in any real-world day-to-day programming
Let's agree to disagree here. No human should ever write XML. No human should ever be forced to read it.
YAML is very readable and writable if you stay away from the corners. Templating allows you to stay clear of the corners (the 1 char operators that concatenate stuff, b64 stuff and so on).
But they're super useful. Some examples from Ansible.
- name: Do something annoying.
command: >
./my-really-annoying-command
-a yep
-c it
--really-does=take-a-lot
--of-options
subcommand
--target=blah
--timeout=60
--callback=/usr/local/bin/callback-x245.sh
https://example.com/this-is-the-worst.aspx
varable_that_I_need_to_preserve_whitespace: |
#BEGIN_LICENSE_KEY
...Templates try to bandage over that by drilling down the abstraction to key-value pairs themselves. And imperative constructs that sneak into templating languages are an artifact of wanting to gain expressiveness without losing the benefits of declarative form -- but really, the two are at odds.
YAML is a red herring -- we had the same headaches with XML a decade prior. The problem is always that there's relationships among the data (or even multiple instances of the config) that we care about, but that the structure of a single config file at rest cannot model.
Databases -- let's say, an SQL one -- are actually among the better solutions, because they allow the universe of config items to live in structured places without overspecifying the exact form the data must take when serialized into a file. Then, data can be normalized where it makes sense to avoid repetition and introduce propagation. An SQL database gives all the tools needed to accomplish this, using mostly declarative code.
Databases in a KV sense are often used for configuration, and SQLite's rise has increased richly structured configs that are specified at a higher level than what's typically done with other serialization formats, but the full approach has not caught on outside big enterprise systems and complex applications. Which is a shame, because it's hardly more complex than the current awkward pairing of a full serializer and a templating engine.
I feel this article is missing the bigger problem - one that for some reason just cannot die.
The problem is that of gluing strings together. YAML is not an unstructured text file, it's a tree notation. Whatever "templating" or "generation" mechanism you want to use, it needs to respect the tree nature of the language it operates on. It needs to respect semantics.
Gluing strings together is literally what causes SQL Injection to exist. It caused countless of defacements on the web, and countless of broken websites. I would think we've learned our lessons, but for some reason, I see these template languages still alive and kicking.
Here's an example (adapted from some real-world code) where I specify the k8s cpu limit in one place, and then look up that info in several other places to avoid needing to change multiple values later:
{
local container = self,
requests: {cpu: 5.5, memory: "2G"},
limits: container.requests + {memory: "4G"},
environment: [
{
name: "NUM_THREADS",
value: std.toString(std.ceil(container.requests.cpu)),
},
],
}
Note how I can patch the container.requests object with an alternate memory limit, and how I can calculate an expression for the NUM_THREADS value in order to automatically set it to ceil() of the requested cpu.(edited for nicer formatting of the code)
I also don’t think we have a workable definition of “configuration”. 70% of the config at my work is hard coded service discovery. If we moved the service discovery anywhere else (say, consul, kubernetes, hell - Docker swarm), we’d need far fewer sets of config than we have deployment environments. When there are only two or three you don’t need templating.
How often do you really have the same service deployed twice in prod and legitimately want it to work differently? I can count the scenarios I know on one hand and none of them have occurred for me in almost ten years, except read replicas and that shouldn’t be more than a few lines of config.
So far the only things close to this are
- Azure pipeline's syntax: https://docs.microsoft.com/en-us/azure/devops/pipelines/proc... - Something called Jasonette: https://docs.jasonette.com/templates/ - Something called Jsonnet: https://jsonnet.org/
Azure Pipeline's approach I think is closest to what I've been looking for.
Anything else in this space?
Has the advantage/disadvantage that it's still valid json/yaml
I wrote a scary command line wrapper for it:
https://wryun.github.io/rjsone/
Has libs for Python, Go, and JS, and there's a bazel interface.
JSON requires constant quoting, can't support multiline strings, has no comments, has no/little typing (e.g., no datetime type). It's not good if a human needs to encode data.
For configs, I think TOML beats YAML hands down; I think YAML's spot is at encoding data structures that humans need to read/write.
I do agree that YAML, the spec, is fairly complicated. But YAML, as used in most projects, by most people, is not, and can be picked up fairly quickly. It is easier to visually read as it removes much of the clutter that would exist in the comparable JSON. It isn't typically necessary to know the entirety of the YAML spec to be useful with YAML, and most of the parts you won't know will get introduced by an obvious-looking sigil, which can be used to figure out what you're dealing with.
When I've actually sat with folks struggling with YAML, it's almost always in configuration tools, and it's also always around the templating bits. Ansible, in particular, has a bizarre templating: it happens after YAML parsing, which is not the mental model most people use when approaching it. I've also found that most of the people I've spoken to intertwine YAML and Ansible's templating functions, thinking they're one in the same.
I do not think Ansible makes good use of YAML: I would rather write task files in an actual programming language, since they are — at their core — a program. (The tasks do have some metadata attached to them, but the core task itself is a program. A function in some real language can get metadata attached to it in a number of ways, and that would be a better solution.)
If you compare a medium sized, 3 level deep TOML document with an equivalent YAML document you'll find that the TOML document is up to 50% longer.
None of that additional verbosity adds meaning - or readability.
Part of this is the extra syntactic noise but most of it is because all of the key names have to be defined explicitly above the key values for each value whereas in YAML you just need to put the values below an indent.
TOML is tolerable for small, very lightly nested config files but even medium sized TOML configurations get ugly very fast.
So while Ansible may not be great, other solutions also have drawbacks, and it isn't quite as easy as do <x>.
That being said, I find it funny that the author is crying afoul, saying that YAML isn't as well suited to be templated as JSON is, when XML had schemas sorted out ages ago.
People ran away from XML, saying it was this verbose, ugly, bandwidth hungry (back when we were sending XML over AJAX) behemoth (which it totally is)... but I think when your usecase is complicated enough to worry about templating, you should take a hard look at XML and ask whether it might suit your better.
Write JSON and use your editor tools to format it with nice indentation, and you are sweet!
That said Yaml makes an excellent format for reading, but not for writing.
To me yaml (especially with templating) seems like a very contrived way of avoiding having to program, while still effectively programming. I much prefer json, or if more intelligence is required, actual javascript objects.
The only big downside of json where yaml does shine is the support for comments within your files.
a:
- b
- c
has a list inside an object but the list is not further indented. There's in fact a hierarchy relationship but absolutely no indentation.YAML gives you options, or varies necessity somewhat arbitrarily on the structures, which is (marginally) good for reading, but a lot of headache in writing
That's fine if you're primarily concerned with computers exchanging data (JSON) but where readability and writability matters, extra syntactic weight is a headache.
I can't imagine writing JSON by hand but I write plenty of YAML by hand every day without issues.
1. there is no language where whitespace does not matter
2. the very purpose of indentation is to ease cognitive load.
Now, you may not want it to be inflexible or mandatory but your argument needs elucidating.
Non-semantic white space can be deceptive. Semantic whitespace isn't.
So we're left with the cases where you'd like compact code / one-liners but are forced to use multiple lines for language syntax reasons.
I'd say that that's a pretty small set of cases and long way from the phrase "I don't like whitespace sensitive languages" which you (and many others!) use.
I conceded it can be annoying in Python and I do wish there was an escape route sometimes.
Writing YAML is easy with a good editor like VSCode. Install the YAML code outline extension for sugar; works great for OpenAPI specs. YAML flow style offers some good options for keeping the file compact.
The problem for me is that the multiple ways to do the same thing result in something that's pretty opaque to clear specifications that aren't just very simple examples / structures.
YAML hides most of the scary brackets and quotes so that new users can focus on copy-and-pasting semantic tidbits.
That said... I don't like YAML...
Also, it is true YAML has too many features, though I find they are typically ignored or disabled.
And partly it's because the syntax for multiline string literals is very minimal, which is kind of a nice feature for CI since you tend to have a lot of them.
It's still insane though. TOML is much more reasonable.
That said I agree with all the criticisms of templating yaml. We have to do the same (with helm and other tools), and I have pushed hard to adopt conventions that we only use flow-style and not block-style to avoid all the white space problems when splicing together chucks of yaml. And on the plus side we get trailing commas and other such niceties which don't exist in json and make it harder template.
>One of the clearest signals I’ve gotten from users is that Dhall is “the YAML killer”, for the following reasons:
>Dhall solves many of the problems that pervade enterprise YAML configuration, including excessive repetition and templating errors
>Dhall still provides many of the good parts of YAML, such as multi-line strings and comments, except with a sane standard
>Dhall can be converted to YAML using a tiny statically linked executable, which provides a smooth migration path for “brownfield” deployment.
Source: http://www.haskellforall.com/2019/01/dhall-year-in-review-20...
We have a couple ways that we template this out, but mostly we literally just do this in bash:
sed -e "s/\$CI_COMMIT_SHA/$CI_COMMIT_SHA/" kube-deploy.template.yaml | kubectl -n $ENV apply -f -
(Where CI_COMMIT_SHA comes from gitlab)
ENV comes from our gitlab CI file.That all being said, the extent of our k8s integration is lots of stuff like that. We could write a JS file that creates a JSON k8s template, but honestly, that would be more work and more learning than we had to for what we're doing. Why would we do more just because we want to avoid templating in a YAML file?
That means you write it, send it, store it, operate on it, etc. with little or no modification.
The author says "converting between the two is trivial" which may be true, but the developer overhead is less trivial. And it will always be JSON in the client - JS doesn't support YAML objects.
Is that really the selling point? I interact with numerous services that speak JSON...and most of them aren't written in JS. YAML and JSON both have to be parsed to be used, even by JS. Otherwise it's just a string.
To me the real selling point in JSON is the dead-simplicity for humans to read and CHANGE. YAML is human-READABLE, but frankly I often screw up changing it because the formatting is a little too magical and I'm a little too unfamiliar. JSON is downright picky and obnoxious...but that makes it really easy to make a change. Screw up the quotes? It'll complain about the quote. Dangling comma? It will complain. Did I cut-and-paste from an HTML display that screwed up my whitespace by compressing everything down to a single space? Nothing cares.
I remember mongodb started with json, but switched to in-house bson pretty early because of the json limitations.
Not making this up! https://reactjs.org/docs/introducing-jsx.html
- <template /> (Templated HTML)
- <script /> (JS OR Typescript)
- <style /> (CSS OR SASS/etc)
Whereas in old PHP files you could mix it in anywhere and the files were a big mess. Including inline SQL into your view templates which is hardly a good separation of concerns. While a Vue component can be separated into separate files, at least as one it all represents one isolated piece of the interface.We have done the classical memcached+database custom solution, but I was wondering if there is any accepted library/tool to change application run-time behavior. We have tried consul KV store [1], but does not quite fit in our environment.
My ideal solution would be a webapp with some text editor (think codemirror). Changes in this text file would push the configuration data to a running application.
It makes me throw up a little in my mouth every time I see hundreds of lines of YAML to configure something like Traefik with Kubernetes. The worst is when people say they prefer that because "I don't have to write a config file for my backend". That's true but instead now you have extremely verbose configuration mixed in with other verbose configuration.
But in YAML's defense I think it's more of a problem with the tools that use it more so than YAML itself. Ansible is a great example of how amazing YAML can be to manage complex configuration in a concise way.
Your data does not need to be "expressive", it just needs to provide input to a program. If your data files need to be complex, you need a program to generate them for you.
I've danced the dance of ini -> json -> yaml -> weird hybrid -> embedded logic, and it ends with "program that asks for what the thing you want looks like and generates data files". Industrial software design figured this out ages ago.
People keep increasing complexity unintentionally precisely because they don't realize that code = data. There's no real distinction. Code is data is code.
You will end up having Turing completeness somewhere, it's just a matter of choosing (or blindly selecting, like most people do) where. For a popular product, it eventually gets embedded in the configuration language, turning it into half-assed programming language (see most web-related templating). For less popular products / more enterprise'y settings, you can probably get away with embedding the Turing-complete part in your bureaucracy. That is, I can't code my config to make it do what I want, but I can pay you to get developers to write some code and export it to the config language as a keyword. There's a spectrum to this, and tradeoffs galore.
But ultimately, YAML is nothing but a tree notation. Tree notation is enough to represent high-level programming languages. Lisp without parenthesis, if you might, or Python, if you squint your eyes.
If you embed "code" in "data", you made your thing way more complex and subject to software design patterns. But in software operation, we already have to contend with highly complex systems, so we want to remove as much time and effort and complexity as possible from the instrumentation.
To put it another way: if you had to run a nuclear reactor, do you want to instrument it by constantly writing new code, or turning a dial? I'd rather turn a dial. That means I have to develop the code for that dial ahead of time, but in the end, actually using it will be safer.
* all configuration, without exception, is XML.
* all configuration may be generated from any other format imaginable, but it's sure as fuck going into the Big Main Godlike Application as XML.
Separation, interfaces, etc. Disclaimer: I work in .NET almost exclusively. The .NET configuration APIs generally work, as long as you only ever use them for reading; treating config as something the application itself can fiddle with is a fast route to madness.
I find it pushes me to write plain YAML files for variables and defaults (Ansible), while allowing strong templating of generated files (Jinja) and letting the result be readable (YAML). By readable I mean minimum programming bloat (code spread on many lines just to write a for loop) and minimum extra syntax that clutters the screen (brackets and quotes). It also lets me write very little custom code (aside from variables, obviously).
If I had to use a more "powerful" YAML-like replacement, it would mix all these into files, written differently by people with different styles, and it would have bloat all over the place.
The main issue I have with helm is that values.yml is not templatable by default so you have to generate it if you want reusability.
YAML does one thing and does it well, it's readable and bloat-free. Maybe we need more tools like "kubectl explain" to know the syntax though.
I'm not proud of this (and like to think I could come up with something better these days), but this code was a bit of a nightmare for that reason:
https://github.com/mschaef/vcsh/blob/master/vm/number.c#L256
Lisp macros do better, but they have the problem that the macros (and their potentially unusual evaluation rules) can easily just blend in with ordinary function calls.
You can extend it, convert it to JSON if necessary, and it is easy to read.
On the other hand, Ansible uses yaml and there it works great. I feel like Ansible uses yaml in a way that's easier to understand and the way it was meant to be written. With Ansible, you're writing configuration, not templates of configuration. I don't think a layer on top of Ansible like Helm is for K8s wouldn't make sense.
I and @akx wrote a templating tool called Emrichen that's specifically designed for producing YAML and JSON from YAML templates:
https://github.com/con2/emrichen
In contrast to other template systems, Emrichen templates are not just "based on YAML", they _are_ YAML. YAML tags like "!Var varname" are used to perform things like variable substitution, loops etc. Variables can be of any JSON type, not just strings, and the template is evaluated top–down.
Why template yaml when you could just generate it? Or generate json, or toml, or xml, or environment variables...
This is actually a solved problem, and you shouldn't be doing it in your YAML/JSON templates. You should be using an external parameter store to do this, and using a single template for everything.
See https://aws.amazon.com/blogs/compute/query-for-the-latest-am...
This is a simple use case: I want to deploy the latest AMI (Amazon Machine Image) in any region, so I always get the latest patched Linux base image to run my application on. I don't (and shouldn't) want to update my YAML/JSON every time a new image is published.
So, why are people having to go to these crazy templating macro lengths? Just store the changing bits in an external config/parameter store like etcd and let your infrastructure as code templates remain unchanged.
like why have:
{"foo": "<%= bar %>"}
when you could just have
{"foo": bar}
the problem with text based templating is the templating language has to make a decision about escaping and it is sometimes the wrong one. for example rebar used an erlang haml [at some point... maybe they fixed it :)] which meant it escaped html special characters by default. but this makes almost zero sense when generating erlang configuration files.
i guess the reason that it is like this is because it is just easy and for most YAML/JSON/etc configuration files it is not a problem because you are basically doing static substitution or 'dynamic' substitution but with a safe range of characters. so the reason things are 'bad' is because the current solution works for 99% of the use cases and no-one wants to spend time fixing it when they could spend that time fixing a real problem. heh :/
You do need some way of saying "this is the base configuration and this is what we change for staging and production". Helm is a way to do that, and a popular one, but it's pretty ugly. Hence this article.
Parameters: AMI: Type: AWS::SSM::Parameter::Value<String> Default: /aws/service/ecs/optimized-ami/amazon-linux/recommended/image_id
That gives you the latest, fully patched base OS image no matter which of 19 regions you launch it in.
Even hardcoding that in a K/V store is going to get outdated unless you manually update it. Parameters like this are great because you can simply write your code once and never have to update unless you're adding new functionality. All base parameters and external systems (APIs, etc) are parameterized and never need to get updated, except by your SaaS partners that update them for you.
For who, exactly?
“This is not a problem for my specific set of use cases around AWS” is not a “solved problem”
Adv:
- It's still quite declarative.
- supports expressions.
- supports merging objects using the spread `...` operator. This enables breaking up large configs files into smaller files.
- supports type constraints.
- serialises to JSON easily.
the only reason config files in another language might be required is when you require configuring the application after it is compiled to native binary code.
Obviously this doesn't apply to all these applications written in scripting languages.
I do wonder how you create type-safe config files though. I currently have a `development.ts` and a `production.ts`, where the `production.ts` is only loaded when `node_env === production`. `production.ts` contains blank/default values and the file is overwritten on the server with production secrets.
The types help prevent erroneous config declarations.
e.g
```
enum LogLevel {
Info = "info",
Debug = "debug"
}const MyConfig = {
logOptions: {
level: LogLevel.Debug
}
};```
And since we serialise `MyConfig`, the config has to be transpiled by Typescript.
[1] or any of that ilk