The Dhall Configuration Language
dhall-lang.org
dhall-lang.org
Jsonnet is far from perfect, but it gets the job done pretty well and has been relatively easy for engineers across the company to pick up.
That said, after a 3+ year journey, the shortfalls in our original design have become more noticeable and we're giving more thought to writing a more robust tool using a more traditional language like Go to solve some of our configuration/deployment and data templating problems.
Cue^2 is something I'm keeping an eye on these days as well.
[1]: https://jsonnet.org/
[2]: https://cuelang.org/
BCL code is the 3rd largest human written code in Google internal code base. In 2019 Aprial, there is 180M lines of BCL, while C++/Java sits at ~300M.
BCL configuration's large scale use probably is beyond any other infra as code use cases known to human.
And the learning and ideas over more than a decade, is manifested in CUE.
Personally, this is enough to convince me to comfortably ignore anything else on the market.
It's... not obvious to me that that's a good thing? The ratio of configuration code to code in the things being configured being that close makes me think that BCL is something that's ill-suited to what it's now being asked to do but there's too much of it to realistically replace.
Maybe it's amazing and the problem it's solving is so complicated that even a language designed very well for solving that problem leaves you needing a lot of it, but I don't really consider "a staggeringly large amount of code has been written in this language" to necessarily be an endorsement of that language's quality. Citing Google's 300M lines of Java would also be a silly reason to pick Java in a project where you'll never interact with the Google ecosystem.
But the point is that this large scale application offered the space for exploring the space of configuration as code, infra as code, and many relevant technical problem space for designing and deploying configuration.
Can it be used to successfully build and maintain configuration at extremely large volume and complexity scales? Yes.
Also we're not talking about using this language, but its spiritual successor.
I have personally written several thousand lines GCL (the generic version of BCL used at Google) and I can say that it can be pretty frustrating.
The difficulty and complexity of defining configurations using it really depends on the system you are configuring since you are (generally) just defining a set of static fields that are packaged into a protobuf and fed into whatever system you are working with.
Outside of syntax issues, it's up to the system you are configuring to provide concise config semantics and helpful error messages
That's just typical of almost all human involved affairs. Some dudes hard support but get 20% of what they deserve. Whereas some few get 800%.
As for BCL, it literally had no investment since 2010, yet still reliably March on ward with little support, in which case 99.99999% of software would simply die into oblivion. BCL just shows it's power and strength.
To name the importance of any single software for Google's success, the 120k lines of CPP code in borgcfg stands on a high peak that look down upon all the other dwarfs with ease.
But I sure can tell you a few easy changes for BCL ecosystem can make a huge difference:
* Testing: despite the common knowledge of testing everything, BCL has not testing facility. No man power to support this easy (easy in relative scale) feature.
* Code search: BCL has inheritance like semantic, but no support in code search to navigate the code base. That's ridiculous. If anyone claim Google cpp code can be supported by lack of code searching navigation, then BCL is rightly to be called mysterious and lack sane language design.
* Packaging: BCL is always halg mixed and half separated with executable. That's just plainly stupidity.
* Lack of sane abstraction from Borg: Borg was designed without package concept. Borg's package was entirely an invention by BCL and borgcfg, as an idempotent shell command executed before starting Borg job. Go figure how much a shaky ground BCL is relying on. And how much blame Bcl has to suffer on behalf of Borg's lazy design.
But the issues with it were well known, so it makes sense to me that after learning all that the guy who invented it could have come up with a really good successor.
Agree with the other comment that the amount of it at Goog should have a mixed interpretation. Obviously it was useful, but although Google systems were complicated, it still doesn't seem right that it's on the same magnitude as the main system langs.
CUE was in its infancy when we were evaluating Jsonnet (and I wasn't even familiar with it at the time). If we were picking a data templating language today there's a good chance we would choose it. We very well may end up migrating to it or incorporating it in some way.
You might want to give Pulumi (no affiliation) a look. I've been using it with Typescript, but it supports Go, too.
I heard you liked configuration languages, so I wrote a configuration language for your configuration language.
How long before you stick a Yaml config into your Go configuration library for easier maintenance? And so the cycle continues.
I really wouldn't use it for anything until tooling improves (but then against I'd much rather use something like Starlark).
Although each implementation of jsonnet has some quirks, take a look at sjsonnet^1 (scala-based) or go-jsonnet^2 for improved performance. We generally prefer go-jsonnet.
There's also a Rust version^3 that claims to be the fastest yet^4, but I haven't experimented with it at all.
[1]: https://github.com/databricks/sjsonnet
[2]: https://github.com/google/go-jsonnet/
[3]: https://github.com/CertainLach/jrsonnet
[4]: https://gist.github.com/CertainLach/5770d7ad4836066f8e0bd91e...
ytt lets you embed logic via a python-subset (starlark) and also provides "overlays" as a "replace/insert" mechanism. and all valid ytt files are valid yaml files, so they can be passed-through other yaml parsing stages.
Have you looked at https://github.com/kris-nova/naml ? :)
Configuration is opaque. Undebuggable. Unmaintainable. By its very nature. No matter what "language" you use for it.
You should strive to keep configurable things to an absolute minimum.
Although it's worth making a distinction between "configurations" and "settings": when something is configured wrong, the application basically fails to operate as intended. For settings, there can't be "wrong" settings, all settings are valid, and if the input for a specific setting was invalid, the application can simply ignore it and use the default value. Example of a "setting": font size for a text editor. Example of configuration: the connection information for a database.
Larger team(s) using free form yaml configuration and templating quickly becomes messy.
Setting a boundary at “configuration” and exposing configuration options to dev teams through settings works really well.
It requires you to have a tooling team with a few developers, but it scales!
Much better is default configuration. I.e. all "conventions" are explicitly generated into a default-configuration file which you can then change/overwrite/update in whatever way you want.
But even so I think you're probably right that convention is better because it forces everyone to use the same structure for stuff.
Why do you need configuration? Because you build your application by gluing together a multitude of different applications that need to configured properly to be able to talk to each other.
Once you understand this, the solution is obvious: don't build your application this way.
Build it in one language. If you want to "reuse" existing solutions, reuse them as libraries, not as separate programs.
Teams have agreed to use postgres for sql, this is by agreed convention.
A tooling team implement hashicorp vault and build a pg module/deployment that is packaged and accessible through whatever tooling you build.
This makes a datasource of type pg sql available to all teams in a couple of minutes.
These modules/deployments are governed by a bunch of conventions, but those are of no special concern to consuming dev teams.
Use “something” to keep track of service upstream/downstreams: consul, a database, or something else.
When a service is connected to a database, say using consul upstreams you have the job gen api inject the appropriate env vars according to an agreed upon convention.
Now dev teams have an option to deploy incredibly secure pg databases, and by simply setting a service as downstream it will automatically receive rotating connection strings.
There’s no configuration to be managed by dev teams except keeping a service graph up-to-date.
Lot of conventions above, and a tooling dev team with a lot of infra and tooling expertise.
I believe that for scaling dev teams you should look at tooling like a product domain and treat it as such.
What makes you think so? How do you define configuration?
I mean, in the extreme case you can see Python as a configuration language for the behaviour of the Python interpreter. Does that make Python a configuration language? Where do you draw the line?
If you have ever worked in the web startup industry, their use is endemic in everything. You often cannot deploy any web applciation without properly understanding tons of configuration files.
Because everyone builds applications out of services that must be glued together, and these services often glue themselves together via configuration files. The python side of the application code needs to figure out how to connect to all the databases (postgres, redis, elastic, memcached ... etc). Each of these databases sometimes needs its own configuration before it can meaninfully operate. And because everyone is using Docker, you need docker files for all of your services. The list of things that need text based configuration fiels goes on and on.
Go to any web company and ask around to see who can meaningfully understand and edit or maintain these configuration files, and you will find the majority of the people have no idea what's going on inside them. It's very common that only two people have a decent understanding of all the configs.
Now to answer your core question:
Why are they opaque and undebuggable?
Because to understand the configuration files, you have to understand the programs they are inteded for and the environments they end up in, and the language they are written in.
For example, to understand a three-line docker compose configuration for ha_proxy, you have to understand the language of docker compose, and the way configuration files are aggregated. For example, you can define some VARIABLES in one config files and have them be available for reading by another docker compose config file in a totally different place. But you have to understand under what conditions which files are included together. For instance, you can have 10 files define different values for one VARIABLE and each of them runs in a different environment. You have to understand all of this. And we haven't even gotten to ha_proxy yet. Now you have to understand how ha_proxy is influenced by the configurations you have specified. Let's say it's telling it about another server to connect to, or a directory to serve static assets from. The path of this directory depends entirely on how other docker images are composed together, which in turn is usually influenced by the contents of many other docker compose config files that are - again - spread out across many files. You need to figure out _which_ config file is used to make a certain directory available at a certain location, etc etc.
There's no tool that can debug this mess. There's no substitute for deep and intimiate knowledge about all the details of the system.
It's very possible for the whole system to collapse when you make one tiny small mistake in one configuration file. There's no tool that can help you debug what's going on.
Seems like the right approach. To use something, one has to understand how to use it. Why is this unusual?
But that's not what I said.
What I said is:
> Configuration is opaque. Undebuggable. Unmaintainable. By its very nature. No matter what "language" you use for it.
The underlying point - perhaps not metioned explicitly - is that they are way too complex for what you achieve with them.
Of course, complex systems require a deep level of understanding. But everything being equal, complexity is a thing to avoid - when you can do it without loss of capability.
I manage DNS in Terraform, and since every Terraform provider uses different objects definitions, and every object definition is rather verbose, Dhall would be a way to specify my own DRY types and leave the provider-specific details in one place. Adding new DNS entries and moving several domains between providers would be a matter of changing fewer lines.
Dhall also has Kubernetes bindings:
https://github.com/dhall-lang/dhall-kubernetes
Although I'm tempted to just stick to Helm here: even though it's less type-safe, Dhall's verbosity makes me reconsider.
I'd like to hear if anyone has used dhall-kubernetes if they like it.
Are there any pitfalls you have learned to avoid?
LE: The "garbage collection" feature is very interesting, but I didn't get to experiment with it enough yet: https://tanka.dev/garbage-collection
We have all of these tools that try to propose declarative configuration, then run into the fact that people really do want dynamic systems with some abstraction capabilities, and then have to overlay that into their systems.
They look like Python, and they occasionally act like Python. But occasionally they don't, and the mask of looking like Python obscures those times. Then you never really know which Python facilities you have and which ones you don't.
https://github.com/oilshell/oil/wiki/Survey-of-Config-Langua...
Nope, so now I have no incentive to use your config format because it's established something is wrong and it's completely non-obvious.
Thanks for not wasting my time, I guess.
The "Hello, World" example is nothing more than JSON, showing a string repeated 3 times. On the third time, "bill" is misspelled with "blil". The next tab shows how Dhall uses a variable definition to prevent this type of error.
On small viewports, the “right-hand side” is down page and off screen, and the instructions I’m following reference only tabs visible above.
Would be cool if they had a three panel story:
* using json with the typo
* using dhall with the typo
* using dhall with variables to avoid the typo
Might be easier to understand.
The project looks cool though.
Another ~similar project not mentioned in that thread is https://ucg.marzhillstudios.com
The syntax for objects with dynamic keys seems unnecessarily verbose.
The arithmetic capabilities are very restricive. No subtraction, division, or modulo. Addition and multiplication only work on Natural. No numeric comparison. = And != Only work on booleans. There isn't really much motivation given for why it is so restricted.
Maybe I'm missing something, but it seems like if a type has a lot of optional fields, which is pretty common, you have to explicitly pass None for all of them. Maybe the pattern is to merge a default record with a smaller record that has what you actually want? But I didn't see any examples of that. Also, how do you deal with cases where null, and the absence of a value are treated differently? For example if leaving the value off means use the default and null means to turn a feature off. From what I can tell optinals are either always null or always absent when None.
It also seems like it would be annoying to have to specify types whenever calling functions like List/length.
IDK, maybe in practice these aren't as bad as they seem.
You can specify defaults for you record types via Type/default pattern https://github.com/Gabriella439/dhall-manual/blob/develop/ma...
Full disclosure: I maintain the [dhall](https://pypi.org/project/dhall/) package on PyPI
The only thing is needed to run the config-generation process in a sandbox and restrict how long it can run. If it takes too long then we stop it - the same we should do with a Dhall config as well, even if Dhall theoretically will finish in a finite amount of time.
That way everyone can use their preferred language to describe and maintain the configuration. We just have to agree on the format that is being spit out at the end, such as json, yaml, dhall or even a proprietary format.
That's a very brittle restriction at best, and if implemented strictly on time, it's a very flaky on as well.
We are not talking about applying the config - we are just talking about using some inputs (variables) and generating a string / text-file that we return to the entitiy that then validates and uses the config.
Not sure why that would be flaky.
If you allow that for Dhall, then you also have the problem of a config potentially taking extremely long to load / to be ready. So you will want to set a timeout anyways and I don't see a big difference here from a practical point of view.
The major advantage of a language that isn't Turing complete is not having the major risk inherent to Turing complete languages: asking if any non-trivial program will produce any given result or behavior is undecidable[1].
> write a program that runs until the heat death of the universe even in a Turing-incomplete language.
The Halting Problem is just a simple example of program behavior. The undecidability extends to any other behavior. Asking if a given program will behave maliciously is still undecidable even if we only consider the set of programs that do halt in a reasonable amount of time.
When you are using a regular language or deterministic pushdown automata, questions about the behavior or even asking if two implementations are equivalent is decidable. It is at lest possible8 to create software/tools to help answer the question "is this input safe." When you use a non-deterministic pushdown automata or stronger, you problem becomes provably undecidable*,
I highly recommend the talk "The Science of Insecurity"[2].
[1] https://en.wikipedia.org/wiki/Rice%27s_theorem
[2] video: https://archive.org/details/The_Science_of_Insecurity_ slides: [pdf] https://langsec.org/insecurity-theory-28c3.pdf
It's not obvious to me that I should care about this property in a configuration language. For a given configuration use case, I probably have a good idea about the extreme upper-bound for a correct program--say, 5s. If the program runs for 30s, the supervisor kills it.
> The Halting Problem is just a simple example of program behavior. The undecidability extends to any other behavior. Asking if a given program will behave maliciously is still undecidable even if we only consider the set of programs that do halt in a reasonable amount of time.
What's a malicious action in a configuration use case that a Turing complete program could muster but not a non-Turing-complete program?
From Dhall's docs (emphasis mine):
> Note that a “finite amount of time” can still be very long. For example, there are some short pathological programs that take longer than the heat death of the universe to evaluate. The main benefit of evaluation being finite is not to eliminate long-running programs but to make them significantly less probable. In practice, you will discover that you will rarely author a configuration file that takes a long time to evaluate by accident.
> For example, Dhall does not provide language support for recursion. If you try to define a recursive expression or function you will get a type error. Lists are the only recursive data structure and the only way to build or consume lists is through safe primitives guaranteed to terminate, like List/fold. This restrictive programming style keeps code simple and makes expensive code more obvious (both to the code author and reviewer).
(Same for the recursive equivalent.)
It really is Haskell for configuration language and undoubtedly superior than YAML.
But for my grug brain, function currying and the lack for loops made it hard.
Functional languages continue to be a great place to steal from but pure functional requires warping your brain quite a bit.
Maybe Tanka is more my style
Dhall: A Non-Repetitive Alternative to YAML - https://news.ycombinator.com/item?id=20355405 - July 2019 (177 comments)
At the top of the main page.
See also the Hello World and other examples.
But yeah, you'd need to carefully select your subset of the language that is both useful and total, and practical to implement.
It's still a massive improvement, but it could be so much better if the typechecker was smarter.
Dhall is a tool to generate configuration such as json, yams, xml. It’s useful if you have *complex* configurations such as a AWS Cloudformation stack or Kubernetes yaml files.
10 years ago people were running hundreds/thousands of VMs/Servers with full operating systems on them. They would often be long-lived and you needed puppet/ansible to update or run stuff on them.
With kubernetes deployment and updating your app/data is all native. The servers/nodes often run some sort of cut down operating system and are immutable.
Of all my gripes, only two still stick in my mind: • No JavaScript implementation; I want to be able to use it more places, but I always end up having to convert it to JSON to use it • the built-in formatter is way too aggressive; I'm fond of the never-contract-only-expand formatters that don't pressure me to collapse things I don't want and the way it is isn’t to optimize Git diffs where conflicts can arise.
Has the situation improved?
{ home = "/home/bill"
, privateKey = "/home/bill/.ssh/id_ed25519"
, publicKey = "/home/blil/.ssh/id_ed25519.pub"
}
This is the most awkward abuse of formatting I've seen in a long time. Just allow/require a final trailing comma, and this nonsense goes awayEg ("Foo", 2,) is different from ("Foo", 2) in Haskell thanks to TupleSections. For innocent bystanders: in Haskell ("Foo", 2) is the tuple you'd expect it to be. But ("Foo", 2,) is a function that takes another argument and creates a three-tuple. A Python equivalent would be
lambda x: ("Foo", 2, x)
{
, home = "/home/bill"
, privateKey = "/home/bill/.ssh/id_ed25519"
, publicKey = "/home/blil/.ssh/id_ed25519.pub"
}
Or, to make it markdown-ish: {
- home = "/home/bill"
- privateKey = "/home/bill/.ssh/id_ed25519"
- publicKey = "/home/blil/.ssh/id_ed25519.pub"
}
This is just to play with the idea that leading punctuation may be preferable because it all aligns in the same column.This actually works just fine.
{:home "/home/bill"
:private-key "/home/bill/.ssh/id_ed25519"
:public-key "/home/bill/.ssh/id_ed25519.pub"}Try it, I doubt you'll go back (unless someone's stupid parser doesn't let you).
(And that's what the comment you reply to suggests: make Dhall allow a final comma.)
In practice, this style reads just fine once you get used to it. Indentation is mostly there as a human convenience, so just has to work well with human brains (and be understood by a computer, if it's significant), but doesn't have to necessarily follow some abstract unified theory of syntax trees.
"Comparisons between CUE, Jsonnet, Dhall, OPA, etc." https://github.com/cue-lang/cue/discussions/669
I would not want to use a bash script as a config file, for example.
And, have you ever worked with giant configs in json or yaml? It becomes incredibly painful to manage.
Even something as simple as variable substitution becomes very useful in these cases so that configuration can be a single file and substitution delivers the staging or production config. Functions and operators allow more complicated configurations to remain DRY. Checks prevent the stage servers from using the production database connection string, etc.
{- You can optionally add types
`x : T` means that `x` has type `T`
-}
let Config : Type =
Config's type is Type?Clear as mud.
[EDIT] I mean, you can almost "who's on first?" this.
"OK, what type do you want this to be?"
"Type."
"Yes, what type do you want it to be?"
"Type."
"Great, ok, yes, the type, what type is Config?"
"Type."
"WHAT IS THE NAME OF THE TYPE YOU WANT CONFIG TO BE???"
"Type."
flips desk
let ConfigOf : Type -> Type = \(type : Type) ->
{- What happens if you add another field here? -}
{ home : type
, privateKey : type
, publicKey : type
}
let Config : Type = ConfigOf Text
and the rest of the example still works and evaluates the same.Also in a later example it has the expression `generate 10 Config buildUser`, which also works because of first-class types. Instead of needing generics, you just take a type as a regular parameter.
Looks great otherwise
It is a pretty decent "take your existing yaml and make incremental improvements" setup though. Rewriting all of your config from scratch is rarely an enjoyable experience.
They have a great writeup on their safety here: https://docs.dhall-lang.org/discussions/Safety-guarantees.ht...
But sometimes, as you indicate, you want VERY dynamic configuration. But I would argue then that such logic goes in your application itself and is not in fact part of your configuration.
Anecdotally, I've heard a lot of GCL horror stories, and many Xooglers have chosen to create things like Jsonnet or Skycfg (https://github.com/stripe/skycfg) instead.
Despite the minimalism it is not necessarily simple. There are features like inheritance and late binding. They can be quite complicated, but the thing that really stands out about GCL is it's not complicated to use. The simple cases are easy, straightforward, and readable. The complicated cases are made possible.
I've used it extensively to store and transform data that has to be manually edited by humans. It's a really great format because I can define all the translation rules which can get pretty complicated, but the complexity I expose to the humans who are writing the data is minimal. But if they need to do some transformations on the data or even just, like, string substitutions...the language is there and they can use it. That's what makes writing your data in a configuration language nice.
(Honestly though I think jsonnet is at least vastly superior to skycfg and the whole starlark ecosystem. I swear I'm so sick of languages that intentionally look like Python without being actual Python!)
From that link:
> However, the early design of GCL went for something simpler that coincidentally was also incompatible with the notion of graph unification. This simpler approach proved insufficient, but it was already too late to move to the earlier foreseen approach. Instead, an inheritance-based override model was adopted. Its complexity made the earlier foreseen tooling intractable and they never materialized. The same holds for the GCL offsprings that copied its model.
Needless to say I disagree with the Cue author. I think the inheritance-based override model is fantastic and has made for a great and straightforward configuration language.
Damn network effect, ruining everything :-/