Typed Config Languages
kevincox.ca
kevincox.ca
[foo]
bar = baz
instead of [foo]
bar = baz
Do not indent TOML!Also, I don't think it's typed in the way the article wished for.
Unless you're already consuming some other API that publishes Cue validation, in which case dhall is just the templating language.
Whether it's going all the way in the direction of supporting general purpose languages like Pulumi does, or with a niche but still Turing-complete language like Nix's expression language, or even Dhall, which other commentators have mentioned. That isn't to say there is no place for simple, human readable schema. I think these tools need a fallback to something simpler.
Nothing these tools do couldn't be emulated by manual processes of hand-writing complex makefiles or YAML or whatever, but what is striking to me is the use of general purpose languages, usually with tooling to do type-checking and IDE assistance to writing these things, which lowers the barrier to entry and empowers someone who "just knows (Python|JavaScript|Go|...)" to contribute in a familiar environment.
A couple more examples:
Envoy: I have had the displeasure, recently, of writing by-hand a configuration for the Envoy load balancer. Envoy uses a "typed config language", which is often represented as YAML or JSON, but it's very painful to write by hand. On the other hand, if you have the protocol buffers/gRPC schema available in code, it's vastly less painful to use any programming language to build the typed objects and then export to plain text. The xDS protocol is designed for being interacted with via programs, not plaintext.
GitLab CI: GitLab supports writing a program, which runs as part of the CI/CD job, to generate the configuration for a subsequent pipeline. This makes writing complex jobs, or repetitive monorepo tasks much simpler. A 10 line Python program that effectively does "for each folder in `python`, emit this block of YAML" is incredibly powerful.
That last example is salient to me: markup languages are really easy to parse, but challenging to read when they're dynamic. Wouldn't it be nice to be able to mix and match? YAML/JSON where it makes sense and the intent and meaning is self-evident, and to write code where you need dynamism?
Full disclosure: I work for Pulumi, opinions are my own, etc.
This was generally considered to be a mistake. It opens up a huge threat surface for your program and few people ended up using the more advanced capabilities it created. If your config format supports simple globbing that covers the 95% use case for otherwise needing a turing compete language for your configuration files.
Pity. Tcl allows you to create safe interpreters within which you can disable any commands you want in order to have a trustworthy environment for running configuration scripts.
Tcl itself uses them to build its internal list of available package modules.
Guix uses Guile Scheme end-to-end, but you really don't have to know any lisp as a user. The DSL just reads in such an obvious way. However, since it's Just Lisp, you also have a full programming language and ecosystem available for the times you need it.
In fact, this has been how I have begun to absorb lisp! I started off writing simple package definitions and slowly started tackling ones that require more hand-holding. Now I find myself even able to quickly sketch up scheme scripts for common system tasks and the like.
EDSLs are great!
For anyone who knows more about Cue: right now you can go from Cue<->yaml (in fact, their docs on yaml also use the "no" case as an example: https://cuelang.org/docs/integrations/yaml/) to integrate with existing systems, but I suppose eventually the goal would be to have direct support in libraries like Serde?
Cue is a very cool language, but it is quite different than the "typed config language" that I have described here. Maybe I picked a poor title but in the post I am talking about using the type information to "improve" parsing. IIUC Cue does not due this, it parses in a "dynamically typed" manor, then uses the type system to evaluate the turing complete (or close to it) expression language.
> Cue does not due this, it parses in a "dynamically typed" manor, then uses the type system to evaluate the turing complete (or close to it) expression language.
As I understand it Cue would help in two ways currently. 1. It would be able to type-check existing yaml files to catch things like the "no" case. 2. if you write your config in Cue, it would output properly-typed yaml to avoid things like "no".
An example (also see it in the online editor[0]):
users {
andy
beth {
admin true
}
carl
}
The author is right that you gain syntax benefits when you define a schema. For those who say this adds cognitive overhead, it actually doesn't; the schema and compiler are able to reduce that overhead, because if you make a mistake, you get a nice, accurate error message.[0] https://gilbert.github.io/zaml/editor.html#s=N4IgzgxgFgpgtgQ...
It's albeit clunkier and less freeform than YAML. And if you ever only plan on using rust the proposed solution here is probably cleaner.
Having portability over multiple languages maintained by large organizations can be useful in some cases though.
What you really need is Cuelang. Cuelang does graph unification over a type-value lattice. This allows the user to do progressive type -> value refinement (e.g. type->range->value).
For configuration, this is both better than regular type systems, and better than inheritance.
https://darrenrush.medium.com/ruby-is-the-ultimate-config-fi...
If you now add an invisible type layer on top of them, it doesn't get better, it actually gets worse. Now the interpretation of a value in the file not only depends on the literal interpretation by the human brain but on some type definition one needs to be aware of, adding a semantic interpretation that can be non-obvious. That is why we have micro-formats: they are obvious (mostly). http:// 03-13-1980 123e4567-e89b-12d3-a456-426614174000 "I'm a string"
There's a delicate balance between readability and correctness and safety. YAML et. al. are not hitting that balance.
And if you really, constantly, need to edit config files and can't handle a format that prioritizes correctness over readability: Build a UI.
TOML is an excellent contemporary choice that is not unpopular. HCL might be another, but unfortunately not very popular outside Hashicorp tools.
Your IDE can help you write only acceptable config files (functions) this way.
In case you do not want to recompile for conf file changes, many languages come with some kind of interpreter. You may even make it hot-reloadable for some properties.
You need more, you probably need a config service, which you can build in a type safe fashion as well.
I think it's better in general but in any case it gets away from the "everything has to be a keyed map" silliness and you get sums/products.
Most of the time, what's missing an accurate JSON Schema for a configuration. I usually encourage owners of those configurations to sit down and write it for everyone to benefit from
Edit: Should be fixed (might need a hard reload). It turns out inverting something twice doesn't really do much.
Other than that it was a nice read, and the grey background is nice and easy on the eyes. I really like the breadcrumbs too, they act as a minimalist menu and easily clickable URL.
Typo fixed.
Thanks for the flow feedback, I thought the full-width code was cool and it can help for wide code without making the page text too wide (it allows the code to grow to the right) but maybe it is more confusing than it is worth. And thanks for the feedback on the background and breadcrumbs, it is good to hear both positive and negative thoughts.
A minor note since the author seems to be around (and noboby has mentioned it that I could see). There's a typo in the example:
allowed-countires
should of course spell "countries" more like I just did. The same error occurs twice.Muphry's law is an adage that states: "If you write anything criticizing editing or proofreading, there will be a fault of some kind in what you have written." The name is a deliberate misspelling of "Murphy's law".
Yaml is just horrible; I've never had a good experience having to use it.
Your application then becomes a library that is invoked by the config file.
This also enables advanced features such as configuring call-back functions.
Because I simply don't want to expend the same amount of cognitive load to read config files as I do for code.
Yes, yaml has some minor ambiguities. These are easily solved. To use the example from the article:
countries:
- ca
- "no"
- us
There, done. The problem was solved with 2 extra characters and remembering the fact that `no` is special in yaml. Comparing that to the amount of typing I have to do to define a scheme, the syntax of which I have to learn, which I also have to read or remember and keep in mind every time I read the config, I take the 2 extra double-quotes.And, speaking of statically typed languages: This problem would be caught immediately anyway if the config is read into static types.
From my use case of config files the code that is reading them knows the type anyways. So for a setup like Rust+serde there is no overhead to set this up.
> This problem would be caught immediately anyway if the config is read into static types.
That is true, but it still breaks you out of your flow. You get a confusing error, it probably doesn't tell you the exact line number and you need to look over your changes. If you changed a lot of places in the file it may be easy to miss that adding `no` to a list was the mistake. Because problems like that are easy to understand in retrospect, but if you keep reading "no" as "Norway" it is easy to look straight at this mistake and think it is fine before hunting elsewhere in the file.
I think you are right. It is still unclear if the cognitive overhead when writing the file is worth it, but from my point of view the upsides are much more valuable then you make them appear to be.
Except, we're not done.
YAML has multiple footguns like this, which I have to remember, forever. And anyone who works with YAML. It's unintuitive and confusing, and costs space in my brain that I really should be using for more important things.
Not to mention that if you -don't- know about these ahead of time, debugging them can be confusing.
A type system is marginally more work for decreased cognitive load and eliminating stupid, idiotic bugs that nobody should have to waste their time tracking down.
With IDE integration, the cost is pretty much negligible other than learning the syntax, which, c'mon, is not difficult and we're being paid to do it.
There are even tools like Dhall [0] that auto-generate yaml for us.
"countries":
- "ca"
- "no"
- "us" enable_feature: yes
Is totally natural. Nobody reads the spec though. If you’re outputting YAML documents with string builders you’re headed for ruin no matter what. You don’t need Dhall, you need yaml.dump which handles the types too.The strength of types, in my opinion, is composability. Most config files I've seen have ultimately pulled in input from another source and used that to create their output. Types would allow the configuration to be checked for correctness even in the face of unknowns.