A reasonable configuration language
ruudvanasseldonk.com
ruudvanasseldonk.com
What the author understands that so many don't is that a language can be a home-cooked meal. A programming language is nothing more or less than a tool for use by a programmer. Compilers have a lot of mystique about them because of all the crazy optimizations that production-grade compilers put in, but fundamentally a compiler is just a pipeline that transforms a data structure from a format that the human would prefer to interact with into a format that a specific machine can work with.
rcl is a home-cooked meal kind of configuration language. It was never intended to serve a wide audience, it was intended to solve a specific human's pain points and help that specific human with their task. Its value lies in that it doesn't need to try to be anything else.
[0] An app can be a home-cooked meal (1051 points, 288 comments) https://news.ycombinator.com/item?id=38877423
My intention is to make it something that can update existing configuration files just as well as -be- a configuration file, such that if somebody finds it useful, they can use it without their colleagues needing to be aware and with the same commit diffs as if they'd done the same work by hand.
How well that will work in practice is still an open question, though a bunch of my experiments in that direction seem to have worked out nicely so far.
But so long as it helps -me- make changes to a repository that are good, so long as I can ensure nobody else working on that repository needs to care that I used it to help, I figure it's worth the attempt.
(also, hey, I'm having fun :)
You're neglecting the second function that a language serves, which is as a communication device to communicate a program from one human being to another, and in that second function, a language is utterly useless if it lacks critical mass in terms of how many people "speak" it. This runs completely counter to the "home-cooked meal" model.
You run into problems, even if you only have one human being trying to communicate the program to their future self, because there's a tendency to forget a language you don't use regularly, so you might cook up a very brilliant language to solve a problem that you have right now, but that you don't have sufficiently frequently. So you don't practice the language regularly, and then you come back 10 years later and want to change something about the program and you might find yourself in trouble.
What is a config language really? It's a language that doesn't allow side effects and evaluates to json.
Once you start getting dependencies from others, it seems like a configuration language could drift into package management.
Downloading dependencies isn't usually done by config languages. Config languages generate a config, JSON or YAML, from source. The config file can be used by engine that applies the config including downloading dependencies.
If you step outside then sure, it has side effects, but so does simply running the program (any program), in 100% of the cases.
You start with an INI/yaml/json
Then you might have several overlaid INIs
Eventually: https://docs.spring.io/spring-boot/docs/2.1.13.RELEASE/refer...
Go ahead and make a java crack or something similar. Every line of that hierarchy is bathed in the blood of IT workers in the real world.
Wait! Not even close to being done.
Templating, expansions, "classes", embedded expressions: now we have computation folks!
I don't actually know the computational power of HCL, but eventually you end up with a Turing Complete configuration input that the poster wants.
And I'm actually leaving a lot of things out.
So many systems go through this, picking their own path. Kind of like workflow engines, almost all major systems will end up with a workflow engine in it (and the builds are workflow anyway). But there aren't well-conserved implementations of these overarching meta-patterns, so everyone bespokes it or picks from a mind numbingly vast array of pseudo-matching options.
Good luck!
As you say, I'd start out with INI files, then move on to JSON, after that it would come a hack to allow the JSON to be generated by executing a shell/ruby/perl script. (i.e. "--config=!xx" would execute XX and parse the output, instead of reading a file).
Later still we started embedding lua, or similar, to allow things to be templated and "dynamic" on a per-host basis.
I was always fond of the Apache-style configuration system, and I guess HCL is close to that, but there aren't any great universal solutions unless you go all-in with scripting, and then you end up with emacs!
Level 1 is just values in a file. The Linux kernel uses that.
Level 2 is a list of values, e.g. ini files.
Level 3 allows nesting. JSON, XML, and YAML are here.
Level 4 allows computation but limited. Dhall and Starlark are here.
Level 5 is a Turing-complete language. Python, Javascript, etc.
RCL seems to be level 5, so I'm not sure if there is really an advantage compared to Python.
Given how many configuration languages eventually end up with a half-assed level 5 (Greenspun comes to mind) it strikes me as, at least, an experiment worth performing.
Existing JSON/YAML/etc could be compiled to WASM that simply builds the corresponding DOM.
Funny joke?
I think we already have an almost perfect language i.e. Lua Everyone knows it (or can learn it easily). tiny runtime to embed. sandboxed by default. garbage collected. Only feature missing is static typing.
It's worth mentioning that in the opposite direction, if you wanted to have an interpreted config option with fewer compile steps, there's also the option to use lua — a language for which the compiler itself is quite small and portable
See,
[1] : https://zellij.dev
For the ones that do go down the hole, you are almost always better on separating them into components, to minimize the ones that need complex configuration, and accelerating the journey of those small pieces. (But yeah, it would be nice to settle on a good workflow description language that isn't a makefile.)
Now, it seems that everybody that is pushing those high-complexity languages wants an infrastructure description language, and just keeps calling it by "configuration language". Since those two problems are completely different, insisting on using the wrong name is quite harmful to the goal.
As one of the authors working in a 2000+ module Python codebase, nothing gives me greater pleasure than to drive one part of the codebase by creating another module. Nothing gives me greater sadness than to be forced to interact with something through a YAML config file, command line flags, or launching GitLab pipelines. All three of those boundaries interrupt the powerful electromagnetic force fields that pervade the system: type checking, linting, and symbol finding (IDE integration). In time, I vow to destroy every last one of these non-code boundaries. Another unwanted-boundary demon on the exorcism list is polyrepos.
There once was a time when the world worked like this but instead of Python it was C. It wasn’t as rich as the dynamic world we have today — I’m certain I don’t want to go back to configuring my softest by recompiling it — but it did have a lot of the advantages of working in one language environment across all the things.
For me, the panacea is to be able to easily do both. I want to be able to use an application as a library from a repl or script by "configuring" it via dataclasses. And I want to be able to run the same application from a cli in different configurations in different production environments by checking in sets of config files that I know are more constrained in what they can do than a generic python runtime.
There's no fundamental reason it is hard to expose both of these interfaces without repeating yourself a bunch!
In Ruud's case, Jsonnet might have been worth looking at as Hashicorp tools can be configured with json in addition to HCL. But that would have been less fun I guess ;-)
I hope for Ruud it finds its niche, there's quite some competition in this field!
3: referal link: https://www.udemy.com/course/jsonnet-from-scratch/?referralC...
> I never properly evaluated Jsonnet, but probably I should. Superficially it looks like one of the more mature formats, and in many ways it looks similar to RCL. Its has a page comparing itself against other configuration languages.
[^1]: https://ruudvanasseldonk.com/2024/a-reasonable-configuration...
I have been working on a couple of video tutorials that could be helpful to get started
Happy to help if you need anything
I understand why data formats like JSON and YAML are valuable, and I understand why it would be valuable to use a programming language to automate the generation of such formatted data. But the niche of the configuration language remains mysterious to me.
Not exhaustive list but generally:
Usually constrained to reduce complexity and try to eliminate the need for testing config.
Interopable between many different programming languages.
Readable by programmers working in different languages.
I think this usually makes config languages favour declarative over imperative which usually eliminates most general purpose languages.
Another topic is why do we use configuration at all and what is the difference between code and config.
That is at odds with using a generic Turing complete language, as figuring out what the configuration is requires running code that may do all kinds of weird things (download data, delete files, etc) (that is similar to the postscript/PDF case. To figure out how many pages a postscript file has, you have to run code)
Now, some will say they trust those writing config files to not do stupid things and only use that power when it is absolutely required, but that’s not an opinion shared by all.
Embedding your configuration in your application language only works well as long as you have one application language. (Or perhaps you use Lua or TCL which are both almost trivially embedded into every other language.)
A bigger issue imo is packaging. General purpose languages aren’t typed well-designed for producing single, understandable standalone files like config languages are. Like I think a simple config DSL in TypeScript could potentially be a perfect way to solve this problem, except that no one wants to lug around a package.json, package-lock.json, and node_modules directory just to write a bit of config. Yet the ability to bring in npm modules—especially those containing relevant types—is where a lot of the attraction of using TypeScript comes from.
Config is dynamically parsed at runtime. Do you want to bundle a c++ compiler in your code, for example? Config shouldn't be difficult to write, or to parse and process.
Perhaps zig and rust can be used as configuration languages as they both do some interesting things that other languages don't.
> Config shouldn't be difficult to write
Couldn’t agree more! In fact, it doesn’t matter if it’s hard to parse. Because that’s what computers are good for. What a human does is write and read configs. Whatever we can do to make that easier we should do.
Using a known programming language, rather than a dedicated confit language, means the user is already familiar with the language and uses it day to day. That’s easy. What’s hard is learning a totally separate language with its own rules, syntax, tool chain, and capabilities.
Sure, or you could turn around and ask:
Why do we need languages at all? Surely we can express everything with assembly?
Or, shouldn’t one language work for all cases? Why not just use C everywhere?
And the answer becomes obvious: some languages are better at certain tasks than others! Then, a “configuration language” is just going to be a programming language that is best suited to generating configuration.
The same way that C or Rust have different use cases than C# and Java.
But, like, the answer is because "people had an itch that they couldn't scratch".
When you look at it from the perspective of someone who already knows a computer language well it doesn't seem useful, but the point is to address the users who don't know any computer languages particularly well.
And 24 years ago I thought we needed to just use general purpose computer languages and turn system administrators into programmers (or fire the ones that wouldn't learn) but that has clearly never happened, and we're retaining the less-technical-than-SWE roles who manage infrastructure, although its now DevOps people managing k8s with YAML. If just using a general purpose language met the social needs that we have then it would have happened already. The existence of templated YAML and its overwhelming success proves that there's a need that has to be met. I'm still skeptical that any of these configuration management languages is thinking clearly about what need is driving templated YAML though, but its nice that there's an explosion of them so that hopefully sooner or later one of them will really stick and become popular.
A programming language that allows you to enable language constructs individually might be interesting (imagine being able to easily turn on list comprehensions but not allow imports or to disable mutation).
S-expressions beg to differ.
FWIW I agree with you, but driving adoption would be hard.
I have the opposite opinion here. Commas should be like semi-colons: only required if you want multiple statements on a line.
Eg
fruit = [
“apples”
“oranges”
“bananas”
]
No commas yet still extremely explicitCommas as separator is pretty much orthogonal, and… well, better overall.
In Murex (my own language) both commas and semi-colons can be dropped if the statement or expression is terminated by a line feed.
list = %[
chair
table
“bed-side table”
]
YAML, for all of its warts, also behaves similarly too.It’s a much better way to handle lists and objects because the comma only exists for the same reason the semi-colon does in C-like syntax: as a parser hint. But if you’ve got a new line then that hint is completely redundant.
Yeah, I did say "several" not "all".
There is very little benefit in new languages making terminator tokens necessary when followed by LF. We are long past the point of needing to optimise language syntax for parser efficiency.
list = %[
chair, table
]
And this? list = %[
chair, table,
] » %[ chair, table ]
[
"chair",
"table"
]
» %[ chair, table, ]
[
"chair",
"table"
]
Here you can see that internally it understands you're writing a JSON object however it relaxes the semantics a little (without going full on YAML) because chasing that erroneous comma is probably the least fun part of programming.However having a comma terminator when followed by a LF doesn’t. The whitespace is larger and more visible than the comma. In that scenario the comma doesn’t add enough visual distinction to add value to the syntax.
1fruit = ["apples","oranges","bananas"]
2fruit = [apples oranges bananas]
was only talking about commas being required like semicolons since multiple statements use space within, so need something else unlike lists, so it's not comparableotherwise you need parenthesis around expressions with operators: [1 (2 + 3) 4]
unary operators don't create a need for commas, nor the f(x) function call syntax with a single parameter: [1 f(x) ~2 ~f(x)]
now, mild redundancy in syntax is a good thing. it enables detecting whether some code is mistaken (perhaps because a botched copy and paste, for example), giving syntax errors
[1 2+3 4]
doesn't need a comma, and the rule "if you value space in values, quote/bracket those" is a fine simple ruleRedundancy is also a bad thing, it's a tradeoff, so the proper approach is to allow enabling it for those and in those case when the value of the good part is higher
I’m not someone that subscribes to the notion that saving keystrokes increases productivity (eg function key words being “fn”). I only think it’s worthwhile removing tokens if they add to the cognitive overhead. In the case of trailing semicolons and commas, it’s something you need to remember to add. So why not make it optional?
However using whitespace as a delimiter and having expressions require zero spaces between tokens is adding to a developer’s cognitive overhead. As well as inviting hard to find bugs.
Expressions don't require zero spaces, use (), this was just an illustration that there is no need for a comma.
That’s exactly why I’m advocating that trailing commas should be optional ;)
> Expressions don't require zero spaces, use (), this was just an illustration that there is no need for a comma.
You’re example specifically stated using zero spaces for expressions though. So that’s what I’m replying to.
I have no issues with parentheses being used to denote an expression. In fact that largely mirrors how my own language works.
I understand this is just an example, but FYI the modern solution is to use CDKTF rather than HCL for Terraform.
That allows you to choose your favorite general purpose lang: Python, TypeScript, Go, Java, C#.
> Aside from Nix, Python, and HCL, which I’ve already discussed extensively,
He doesn't seem to have much more to say about the latter than that.
'My reply', bluntly, is something like 'maybe it has shortcomings, you haven't addressed them; its most fundamental key advantage is not addressed at all, about it or any other language, and is not a feature of your new one'.
CDKTF is not accessing the Terraform state directly, it's allowing you to express your configuration that you would normally write in Terraform in a programatic way.
That's an odd take. Are you saying that because it's newer? I would push something like Crossplane as more "modern" in that it solves the critical issue of Terraform not having any sort of reconciliation loop.
There are Terraform alternatives...Cloudformation, Pulumi, Crossplane...but that's a separate matter.
https://blog.boltops.com/2020/10/08/terraform-hcl-for-in-and...
Correct, but that has nothing to do with HCL and everything to do with stored state.
https://blog.boltops.com/2020/10/06/terraform-hcl-nested-loo...
Quote:
Consideration: Updates with Removal
There’s a subtle but important consideration with the current code. It happens when the code gets updated, particularly when previously added elements are removed.
For example, let’s say we first use the code above and run a terraform apply. That creates security groups with rules. Then we delete the rules from the code. Running terraform apply again will not remove the rules.
This is because when there’s an empty List, the for_each loop never iterates. If you wish for the security group rules to maintain its current state set outside of Terraform, you may want this behavior. However, this is probably unexpected and undesirable behavior.
If you want to have Terraform remove all the security group rules, then ingress needs to be assigned directly with a List. We’ll cover how to do that shortly.
I agree CDK for TF can be better than HCL, but it is still more like a template preprocessing utility for Terraform and thus still carries the limitations of Terraform.
Worth a read, and worth trying out sometime, if nothing else because I too don't use jq often enough to remember more complex usages.
At its basis, a valid Lua file can consist of one table. Then the data description will be very similar to how people operate with JSON files. At any point the user is able to write code or use other standard library functions... which is pretty limited, but at the same time it's easy to sandbox plain Lua away from "dangerous" functions such as os.exec or the debug library.
There are only a few caveats:
- Lua uses a 64-bit IEEE float format with some truncated precision
- The standard interpreter doesn't handle very huge tables for parsing
- Treating .lua as code, not data, you'd likely prepend "return " to your table to just receive the table object from the config file
- Unkeyed arrays start from 1, though storing at index 0 is possible
- Cascading of data definitions is possible, like fall-through from one semi-filled table to the default table. Use the __index metamethod
Personally I find Lua's table syntax vastly more readable than JSON. Although the language naturally has comments, the parser ignores them, if that's something you want. But the same workaround works: just store the comment in some keyname_comment value.
Apart from sandboxing you'd only need to limit the script execution time to stop attempts like "while true; end"
PS: Lua started as a configuration language. Wiki:
> Lua's predecessors were the data-description/configuration languages SOL (Simple Object Language) and DEL (data-entry language).
I make no claim that anybody who does look will like either of them, but I do claim they're worth a look even if it turns out that you don't.
You noted at the bottom that RCL type system may take the path of TypeScript. This is brilliant idea because structural type systems match perfectly with a configuration language.
yaml-get -p '.tags[has_child(amd)][parent()].name' machines.json
That can replace rcl query --output=raw machines.json '[
for m in input:
if m.get("tags", []).contains("amd"):
m.name
]'I think at the end of the day, k8s YAML is focused on the datatype, i.e., what would be the result of your `rcl evaluate`. I do think this would be a much saner path than what Helm provides, though, and Helm's textual templating, as opposed to understanding values and building the actual datastructures, and then converting that to a format like JSON or YAML … is the wrong path. (I.e., I can see RCL being a possible replacement for Helm.)
> The language is a superset of json.
I'm going to introduce what I think is pretty much a universal law: languages claiming to be supersets of other languages are not supersets.
In the case of JSON, there's basically one counter-example that breaks all alleged supersets:
"\ud83d\udca9"
And if we try it: » printf '"\\ud83d\\udca9"' | cargo r -- evaluate
Finished dev [unoptimized + debuginfo] target(s) in 0.01s
Running `target/debug/rcl evaluate`
stdin:1:2
╷
1 │ "\ud83d\udca9"
╵ ^~~~~~
Error: Invalid escape sequence: not a Unicode scalar value.
Help: For code points beyond U+FFFF, use '\u{...}' instead of a surrogate pair.
Now! This to me is not a bug in RCL: this particular facet of JSON is utter crazy town, and I would strongly encourage you to not adopt it. (Since down this road lies madness, like unpaired surrogates. JSON's grammar & standard is sloppy here, and JSON/JS's syntax of using the UTF-16 encoding, and not just the scalar value … it's the part of JavaScript that is just what JS is, but is the part that we shouldn't be copying.)It is much saner to just have a flag/way to indicate "my input is JSON" and then to just parse it via an actual JSON parser. Then let RCL evolve on its own merits. If there's a lot of happy overlap, and most JSON documents are blissfully polyglots with RCL, that's fine too / a happy little accident. (But the option is important if you want to apply it somewhere programmatically, where the inputs are JSON.)
YAML breaks the same way, too: it too is a "superset" of JSON that isn't.
I somewhat agree with your point on having an explicit way to indicate JSON input. OTOH, in practice I'd rather prefer having a section about limitations in the docs.
And that's the right way for infrastructure and build configs. You do want over-verbose, no-logic code in the configs. Unlike app code, you rarely need to change it, and readability is even more important. 6 duplicates is not too many yet for infra code.
In terms of copy-paste tolerance, I'd rank from least to most: app code, tests code, infra configs, and build pipeline configs.
Clearly I'm not the only one. Gitlab even sees the use enough to support syntax that is invalid yaml: https://docs.gitlab.com/ee/ci/yaml/yaml_optimization.html#re...
This, of course, causes yaml-language-server to complain. What I wouldn't give to have a language where I could, say "goto definition" for things like this rather than get LSP errors.
Yaml, JSON, xml, text files... all work great for configuration.
But the assumption is that your configuring a piece of software on an already existing system.
A config file as means of setting up a system, installing software, and establishing how its going to run is exceedingly stupid.
Write your app so it can run on bare metal, install with apt, yum or your tool(s) of choice for your org. Build it so it scales diagonally, and works against spot instances where you can (because sometimes more or less cores are cheaper). Dont even get me started on the nonsense of everyone and their ideas for "secrets management", it makes me miss LDAP.
It sucked.
Config should be easy and flat and simple. IF it isnt that source it from code, or from a service... Most of the nonsense of config madness is a byproduct of, or hidden in, containers. If we were writing installable software much of that would go away...
If it's running on bare metal, why are you using apt or yum? Those are parts of an OS you wouldn't have if you're running your program without an OS.