What is wrong with TOML? (2019)
hitchdev.com
hitchdev.com
This is the major problem with most comparisons of config file formats: the actual semantic domain of a config file format is extremely limited, which means the main thing left to disagree over is syntax, which is highly subjective and extremely difficult to get people to agree on.
Add too many syntactic features and a lot of people will disavow you for being too complicated. Add too few and you'll be missing someone's pet feature. Make white space significant and you'll frustrate people. Require extra characters to delineate and you'll frustrate another group.
It's worth noting that this article is primarily talking about TOML in the context of the Python ecosystem, and I think that's a healthier way to talk about config file formats: How well suited are their syntactic choices to the community they're targeting?
> How well suited are their syntactic choices to the community they're targeting?
Also, "best" practices. One could reduce the pain of the other, but by no means is the right solution to a deeper problem at hand. For example, if one has very deep and complex nesting for configs, TOML "may be a lot nicer" compared to YAML, but that doesn't mean using TOML will make all the config parsing problems go away. It just mask away code smell. Maybe time to check if they're overcomplicating configurations in general.
So, why not use Scheme ?
Personally I think the right answer for configuration files is to define them in terms of a generic object model. A program could even support multiple formats (TOML+JSON+YAML). If a user dislikes all the supported formats or the file is generated with something like NixOS, it can be handled with straightforward conversion.
I invite you to check out Symfony where you configure your app using yaml, attributes in code, code itself, or a mix of all the above.
You will cry.
(Don't shoot the messenger, I've done my fair share of scheme. I've also done a lot of thinking about why some people are so turned off from the syntax, and it's certainly not that the opening parenthesis is in a slightly different place on the prefix function calls)
Or worse: make whitespace significant and strike a blow in the eternal tabs vs. spaces holy war at the same time, like YAML.
Even the customizability argument makes less sense now, IDEs could just change the width of leading spaces, I'm not sure why they don't.
Like a C++ developer crying foul because inheritance doesn’t exist in YAML.
Not being DRY is a good thing in a config file - it makes it much easier to understand and work with just one section of the file (which is what you most often want to do), because the context information is right there without having to jump around and figure things out.
And whatever the downsides of syntactic typing are, requiring a schema file to go along with your config file is far more of a downside; it's one more point of potential failure, another thing to maintain and sync up and keep in your head, and not worth it for most use cases.
And that's the crux of it: it all depends on what you need from your markup language, what your use case is, today and through the lifetime of your project. "What's wrong with TOML" makes much less sense as a question than "What's wrong with TOML(/JSON/YAML/etc.) for _this project and its needs_".
I love TOML and will continue to use it as my default choice for configuration files, because I think most applications simply do not need the power and flexibility of YAML, even if The outright safety problems are mostly resolved in YAML 1.2. But I do agree that the inability of the syntax to convey nested structure is a limitation and it definitely gets annoying in larger configuration files, such as pyproject.toml files that tend to accumulate in larger Python projects. I have considered just manually indenting nested table blocks, even though that would look pretty ugly and is decidedly non-standard.
I'm sure you'll be happy to know this is getting relaxed in toml 1.1, wherever that comes out (and the implementations adopt it): https://github.com/toml-lang/toml/issues/516
Though the difficulty will then be knowing whether a given piece of software uses 1.0 (and single-line tables) or 1.1 (and more flexible tables).
Most every common GCedlanguage these days supports native JSONlike object, YAML can represent things outside of strings, lists, dicts, bools, numbers, and null.
Lack of nested structure is a positive in some applications. Flat is better than nested. I've seen way too many config files where someone says "add foo=3" to the file and you can't even figure out where in the structure it goes.
And worse, sometimes people reorganize things into options. They'll move all the stuff for one subcomponent into its own nested thing, and you can't configure it without knowing the full architecture.
With flat stuff you get an obvious single way to represent any given config options. Maybe not the nicest way, but it's obvious and unique.
>Not being DRY is a good thing in a config file - it makes it much easier to understand and work with just one section of the file (which is what you most often want to do), because the context information is right there without having to jump around and figure things out.
If the contextual information is relevant that's true. However, syntactic noise of the form of lots of [s, ]s and equal signs isn't necessarily relevant. A YAML file can exhibit identical information with fewer characters and that makes the files easier to read and maintain - especially if they are information dense.
>And whatever the downsides of syntactic typing are, requiring a schema file to go along with your config file is far more of a downside; it's one more point of potential failure
Schemas aren't technically required by strictyaml provided you're happy mapping everything to a dict, list and string, but they're recommended because they make it much easier to prevent something from going wrong and it means you can directly use the types you were expecting.
Schemas in config files are equivalent to static types that generate compiler errors in a program. If you can use them, it's an easy way to get your program to fail fast on invalid input and save on debugging time.
If you don't have a schema and some invalid data gets put into your config file, instead of getting an error that says "didn't expect key "ip addresses" on line 14" you tend to get a really cryptic error a bit later on when your program tries to get a key from a dictionary that doesn't exist.
This is an example of the principle of https://en.wikipedia.org/wiki/Fail-fast design.
I don't know that ': ' is any better than ' = '. I get being opinionated about it but this feels squarely in the realm of the subjective to me. Further, adding an errant '-' and accidentally creating another object is real common in YAML, which is something you can't do with TOML's lists. I think this washes out, tbh.
There is an element of opinion here, but there is no question that equivalent TOML files are longer, and most of that is syntax.
It's much more pronounced when you have more than one or two levels of nesting. With 4 or 5 levels of object nesting TOML files grow huge, whereas YAML is still fine.
>Further, adding an errant '-' and accidentally creating another object is real common in YAML
Yep, this is one of the things that type safety helps with though. Similarly it's quite easy to mess up an indent in YAML, but a schema can catch that stuff.
Yes. That's a clear indication that either you configuration type structure is badly designed, or that you aren't trying to create a configuration type and should use some data encoding or programing language that fits your problem better.
I guarantee you the people using your software will curse you every time they try to edit a YAML file with more than 3 levels of nesting.
But it's declarative, which is really cool! What else would you realistically use for like, AWS infra? The JSON version is so much worse. Using Python or whatever is its own can of worms.
This is the kind of thing that stuff like CUE and Dhall are aiming at, and I welcome it. It feels like they're the way out here.
Programing languages can be declarative too. As you already pointed, CUE and Dhall exist. (And lisp, and prolog...)
(And in fact, a lot of people just use Dhall anyway and derive the actual configuration files from it.)
> ...but a schema can catch that stuff.
100%; it's amazing to me the number of projects that essentially accept arbitrary structures and/or values here. I think no matter whether you're using YAML/TOML/JSON/etc. you oughta be using something to validate.
Especially when I'm scrolling through a file, I encounter myself backtracking to understand again its structure.
XML parsing is also not a pure computation.
That said, a XML-lite would quite probably be the best data encoding language one could create right now. It would still suck as a configuration language, as it's a very different problem.
mynode /-"commented" "not commented" /-key="value" /-{
a
b
}
A problem with a node-based design similar to XML compared to something that parses to nested lists and dictionaries is that you pretty much need a query API.
Implementing it takes nontrivial extra work.
This affects KDL implementations:
KDL specifies a query language,
but I think no Python KDL library has implemented it so far.
It limits how useful KDL is in Python, despite there being multiple working parsers.Edit: Rephrased the second paragraph.
But I do disagree on the work. A node-based language parses into a sequence of (name, node_type, data) tuples. What is just as easy to travel as the (name, data) from the data map ones. The thing is, a query language is much more useful for them than for a data map, so it's worthwhile to implement one. (There's probably some cultural aspect here too.)
yaml? who knows. Whether and how (and the limitations) depends on whoever cooked up that pile.
I don't think this is a particularly good yardstick. Code without comments is shorter than code with comments, but I wouldn't call comment-less code easier to read; the more information dense, the worse, really.
I first encountered strictyaml years ago and have used it, happily. I especially appreciate that you made clear arguments for what ought to be excluded from a configuration format itself, and how proper validation ultimately requires real code anyway.
I was always disappointed however that your project didn't amount to a formal specification (at least at the time, unsure if that's changed).
More recently I came to know and love NestedText, which seems very close to what a strictyaml spec could be/have been. I'm curious if you've engaged with that project/format, and what you think of it.
I would dearly love to do this. I would ideally like to work with somebody who can help me though because it's a lot of work and I struggle to find the time.
If you want a "language" for expressing data (like configuration data), you might be interesting in having a look at EDN. https://github.com/edn-format/edn
According to that logic, binary files would be easiest to maintain and read, which is obviously bogus.
Also unit test code is expected to be a lot less DRY.
[params]
profile = {
name = "Gareth",
tagline = "..",
}
contact = {
enable = true,
list = [
{class = "email", icon = "fa-envelope"},
{class = "phone"},
]
}
The whole business with the syntax typing has no "one correct way" to do it. No matter what you do it will cause problems and headaches for someone at some point somewhere.> Dates and times, as many more experienced programmers are probably aware is an unexpectedly deep rabbit hole of complications and quirky, unexpected, headache and bug inducing edge cases. TOML experiences many of these edge cases because of this.
Eh? In the original text it links to three issues to back this up:
That first issue it links to is "failed to parse long floats like x = 0.1234567891234567891" and the third is a feature request for hex values (v = 0xff, not even a bug report). That has nothing to do with dates? The second issue did relate to dates, but was just a simple bug, not an "unexpectedly deep rabbit hole of complications and quirky, unexpected, headache and bug inducing edge cases".
This just seems repeating a tautology. I maintain a TOML implementation that sees some reasonable use. Dates have not been a huge source of bugs, confusion, or other issues. All you need is to be able to parse RFC 3339 style dates (and some things derived from that), which is usually just calling strftime() or whatever your language has for this.
I do think TOML had some bits I wouldn't have added (not dates though), but the feature sets and complexity of TOML and YAML and not even comparable; it's like comparing Iceland (pop: ~300k) to Ireland (pop: ~5M). Yes, they're both islands and both are small countries, yet the scale of their "smallness" is just completely different.
https://github.com/toml-lang/toml/blob/main/toml.md#inline-t...
Now I just have to hope enough TOML tools support this syntax, lest I end up in the same boat as JSON5.
Out of curiosity, how does the deserializer know which "standard" to use?
AIUI one can pin a version of YAML via its directive: %YAML <https://yaml.org/spec/1.2.2/#681-yaml-directives> with a missing one implying 1.2 although (heh) the 1.1 version says that documents which are missing their yaml directive are implied to be 1.1 <https://yaml.org/spec/1.1/#YAML%20directive/> so ... versioning, it's hard!
Given how many comment remover util functions I have written to make almost JSON config files proper JSON ... why could they not just have included those in the spec.
I can’t be the only one who feels this way; isn’t it this use case the thing that sucks? “plain text but not really” config formats that aren’t “code” but have special syntax but lack a debugger and handy IDE tools and you’re never really sure of what you’re doing … isn’t that the thing that sucks?
This specific example is like somebody saw cucumber and went “I think this should be a lot worse, and in an ecosystem which doesn’t want to come anywhere near that too”.
This would probably be half the size and actually comprehensible if it were just a pytest test file.
With the stories in YAML two new use cases (which are demonstrated in the example projects) are enabled:
* Automatically updated how-to docs. I used to write these types of docs manually and if they existed at all they would ALWAYS get out of sync with the code and were painful to maintain manually. Now I have YAML files and a simple jinja2 template I can push out new markdown how-to docs on each new build - with snippets of JSON, screenshots from the app, whatever.
* Tests that rewrite themselves. E.g. if I have a REST API test of the form "call API x and expect y blob of json", I don't have to actually write that blob of json into the test, I just write the code that produces it and run the test in rewrite mode so it updates the "expected JSON" field with actual json. I can then eyeball it and 20 seconds later it's part of the test and part of the docs.
The productivity improvements from doing both of these things means that writing tests is cheaper so I do more of them. Having how-to docs for all scenarios is way cheaper so I now always have them.
These use cases are impossible with pytest. They are impossible with cucumber.
They would be too painful to maintain with regular (i.e. not type-safe) YAML and those stories have enough indents that they would be an epic unreadable mess of syntactic noise if they were built in something like TOML or JSON.
That sounds like expect tests in OCaml [1]. I've found them quite a joy to work with and I'm surprised more languages don't have something similar.
[1] https://dev.realworldocaml.org/testing.html#expect-tests
It's not that I agree with all your choices. It's that I'll defend to the end your method of making them.
If you need it to be _that_ complex then just write your configs as code...
That said, YAML's syntax is very weird. YAML just sucks. And any such implementation must necessarily be unityped (up to the point where the data is coerced into the configuration structure at your program) and completely preserve the original data.
TOML could be extended to support it. I don't think it's "tasteful", but I see no practical problems with it.
Seems doubtful as that's specifically something TOML was created not to support. If you want unityped ini files you can define that dialect of ini files.
Actually, that, but without sarcasm. YAML is crap. TOML is crap. JSON is crap. INIs are actually fine for flat lists but major crap otherwise. Anything Turing-complete is completely inadequate crap for the purpose.
{:adapter/jetty {:port 8080,
:handler #ig/ref :handler/greet}
:handler/greet {:name "Alice"}}
[1] https://github.com/weavejester/integrantFirst off, it shares the desire for extensibility with yaml (in the form of tags), which is a giant trap. And then it’s quite syntactically noisy.
For user-facing configuration I still favour TOML. I think it’s a bit application dependent, sone applications (eg nginx) have complex configuration needs and for that it makes sense to use something more sophisticated, but for many user-facing config settings, a simpler-TOML would be a great fit. Basically just some basic key-value pairs that can be collected into groups. As the article states, the parsing types should perhaps be enforced by the parser not the written config.
I did like integrant, at least I felt like I understood it, unlike component.
There were two things about edn that made it seem better than json to me, tagged elements (and readers), and symbols. I don't remember exactly my use case but I used symbols in edn to something like namespace-resolution for multimethods. It was something like including a file in a classpath or loading a file, dev vs prod kind of config.
Mainly it’s just personal taste and no deep reasons.
JSON-alikes are opinionated. They don't let you do stuff that requires a plugin for the parser to work with
Is this actually something I want to see in a configuration script? Why not just use a scripting langauge and be done with it? I wonder if the safety features can't be replicated with rigorous testing of say python as config scripts instead of learning yet another programming language?
https://prelude.dhall-lang.org/Text/concatSep.dhall
I think this is the key fulcrum for me: "config is code", sure, but not the same kind of "code".
It -is- compelling to argue for statistically deterministic config code but my practical objection here is 'can we arrive at same safety using testing with a known language?'
Writing this has made consider whether configuration should be conceptually looked at as a database instead of "code". How many people even know how e.g. postgres stores its tables and why would modulo some performance niche would you care anyway?
It seems configuration management is a graph db query and update matter. Standardize on configuration query language (if necessary) and stop worrying how the damn thing is represented by the config management tool.
I run my configs mostly as YAML in Consul and Vault, sometimes in Spring Cloud Config with a git backend. This way I have dynamic config evolution. But I prefer to generate those yaml files from Dhall to avoid unnecessary bugs. After years with Haskell, the syntax is very natural, too.
As for Postgres internals, they do matter if your data set keeps growing.
Xkcd covered standardization. YAML and JSON ASTs are graphs, YAML not necessarily a tree. JSON extensions also support references. As for the ops side, YAML has become a de facto standard, HCL is used with the HashiCorp tools. Nix has its own language.
It’s not about how it’s represented but how you express dependencies across config key nodes. It’s good to avoid repetition and have a syntax linter, a compiler even better. Small static configs are amenable to querying and writing. But you need to separate the writing from the querying the configs. With Dhall you write code to generate the actual config, whether as a Dhall AST or exported to YAML or JSON, with certain correctness guarantees upfront.
I want my configuration to be guaranteed to halt. Turns out it's hard to not make anything useful accidentally Turing-complete!
[0]: https://kdl.dev/
1. the post is previewed with Github's dialect of Markdown and not the SSG's (and definitely with none of the inline configuration applied)
2. the preview is still a Markdown document, so you get no benefits of syntax highlighting or auto-formatting w.r.t. the config header (short of opening in Vim and explicitly declaring the filetype or temporarily changing the extension in Github's editor)
3. you HAVE to put trailing spaces in the YAML/TOML/JSON or the preview pane crams everything into an unbroken paragraph
4. there's not a quick preview of how the configuration will parse, just specific workflows (live update, compile single page) that you can test and then modify as needed. This is either in the rich online editor or your own machine and will require console commands and a browser window
5. I still have to know all of the modifiable attributes as well as defaults, which will be in a separate document and probably not in a Ctrl+Space dropdown
-----
For point 5, it would be nice if configuration formats had completions for common editors and/or their own scripts:
* "generate big config file with all possible keys and default values"
* "condense modified config file so it only contains non-defaults (and hope the schema doesnt change hahaha)"
* "suggest a valid fix for a currently invalid config file"
-----
Sure, this all isn't a direct criticism of TOML, but inlining configuration is a great-to-okay idea that is simply poorly executed. It is extremely unfriendly to non-technical users; I can fully understand why someone would pay a few hundred a year for a WYSIWYG templated website builder to just handle everything.
It also has best-in-class editor support thanks to paredit :)
The only real downside is the hassle of convincing people to use something weird and non-standard, which is really more a problem with people than with the format.
I think my ideal format would be INI, with tagged headings so you could do [md:README] and stick a multi line string section in there, or use something like [comment:Module Attributes] to make an embedded docs section.
Either that, or just some other method of embedding multi line keys. Maybe HTML-like tags, so that inside a heading you can do
<my-key> value </my-key>
JSON is the champion serialization format. But it is hard to edit config files with missing comments and strict quotes and commas. JSON is great for the API, but configs should be converted from something nicer to JSON. I wonder if it would be good to make the conversion explicit so any format could be used.
(But extrapolating, I'm guessing it goes into your crap category.)
I find it particularly useful for configurations that often have repeated boilerplate, like ansible playbooks or deploying a bunch of "similar-but" services to kubernetes (with https://tanka.dev).
Dhall is also quite interesting, with some tradeoffs: https://dhall-lang.org/
A few years ago I did a small comparison by re-implementing one of my simpler ansible playbooks: https://github.com/retzkek/ansible-dhall-jsonnet
The whole point of the parent is that JSON is the Pareto solution, which I agree with.
But it does grind my gear.
It's configuration with a build step. Which is an option, but that's pretty different than JSON, YAML, TOML, etc.
I don't completely disagree with this. However, in most cases TOML is used, it isn't that much of a problem.
And I actually like that the full key is repeated. When you have several layers of nested mappings, it can be hard to determine exactly where the current value is in the hierarchy. Especially if the top level key is above the current screen of text. It can also make it easier to search for a specific key. IMO, this is a case where more verbosity and repetition makes it more readable.
That said, it seems a little arbitrary to me that inline tables don't allow newlines within them. If they did, then if you didn't like repeating the keys, you could use inline tables.
> TOML's hierarchies are difficult to infer from syntax alone
This is a little subjective, and depends on the actual data represented in the config.
But in general, my experience is that when you have several layers of nesting, and the only indication of the hierarchy is indentation, it can be a little hard to follow where a specific value fits in the hierarchy. See above.
And I disagree that meaningful indentation is "generally considered a good idea". I won't enumerate the pros and cons here, as it has been discussed a lot elsewhere, but it is definitely controversial, and subjective.
> Overcomplication: Like YAML, TOML has too many features
This section lists exactly one feature that it thinks TOML shouldn't have. Maybe dates shouldn't have been included, but it isn't anywhere close to the complexity of YAML.
This article has been written in a way that (most likely inadvertently) implies a measure of distance from StrictYAML.
Now I would obviously never ever recommend doing this, but it was certainly an interesting and eyeopening experience.
https://support.microsoft.com/en-gb/office/stop-automaticall...
> "make it easier to enter dates. For example, 12/2 changes to 2-Dec"
12/2 is obviously the 12th of February in my locale. But I need to keep Excel in English as a company policy, so this is not only unhelpful, it's outright wrong.
> Unfortunately there is no way to turn this off.
How does Microsoft justify this choice?
I also love how the support article describes a behavior of their own software as “very frustrating”.
Imagine if you said everything with SQLite instead of Excel, and all of a sudden your just talking about structured config in a database. Not new, not crazy, and generally a decent practice.
SQLite is great for storing application state and config options set from within the app, but it is a pretty terrible format for end users to edit.
For a somewhat trained audience however it can be quite interesting for some specific problem domain ...
I think you are conflating the file with the workflow. A proper UI is the solution to making something not "terrible for end users to edit".
With the original solution, you have an autocontained file that virtually any user knows how to edit, structure, expand, version, email, compare, discuss...
The level of user knowledge that you need for a similar solution based on a standalone SQLite file that you can version is another different world, e.g. to relate two values you would need to perhaps create a view or a trigger. And even with the most knowledgeable user you would still lack functionality such as simply pasting an image as a means of documentation and be able to see it, or WYSIWYG colors.
Gonna have to disagree with you here. Very few people know how to version, compare, let alone if there are multiple collaborators, Excel files. Is there even a decent way of doing that outside of Excel Online / Google Sheets?
Yes, that's horrible. But millions of people can autonomously resort to this, and they would be incapable of doing anything with a SQLite file.
Same with version control: you just add _v17_20230909_final to the name. Yes, it's horrible. Yes, it's buggy. But yes, it also runs the world.
Are you trying to suggest that either you can't do that with sqlite, or that people didn't require training to do this in Excel?
sofixa: version and comparison is not really supported in Excel
harperlee: agree that software support is not there, but in practice average user knows how to do it and is quite comfortable doing it in Excel
btreecat: sqlite also has filenames and you can learn how to use it
harperlee: agree on the filename, Excel files and sqlite files are externally, opaquely versionable in the same way. the point about people learning is moot though, average user already knows Excel because they were forced to in the past, but does not have the time to learn new things.
SQLite is also an auto contained file. There are similar tabular GUI tools that could let you interact with sqlite using a similar workflow. Users knowing that thing is not inherent, they had to be trained on it, and they can be retrained as they will be for other workflows and business tools.
Remember, you still need an external app (Excel) to open it's files, the files themselves are just data, exactly like sqlite. So you could just make an excel plugin to interface with SQLite.
SQLite is as version-able as an excel file, as in not very with standard tools.
Why are you pasting images into excel? Doesn't matter, sqlite handles that fine actually. https://www.sqlite.org/fasterthanfs.html
The main argument against sqlite is that you would need to build an interface or figure out how to train folks on existing tools. That's not a huge argument against it in my experience, it's a strong social/political one in many orgs but rarely a technical issue.
The design of a solution needs to look into way more requirements than just the technical ones, time and money being 2 big ones. I think most of the HN readers would agree that you could end up building an interface with most of the Excel functionality, even more perhaps, on top of sqlite, and have your particular group of users trained on it.
I think ideally you'd want multiple backends, either SQLite, or flat files that are Git/SyncThing friendly.
I was even thinking the file format could have a file UUID and record timestamps, so that if you put a different version of the same file in the same folder(Like with a SyncThing conflict file) it would give you a merged view, with newest-record-wins logic.
Formulas I think would be the easy part. Just write it all in Python and use one of the many Excel compatible formulas implementations, and just make sure all changes went through the app.
Maybe you could even have a REST API to build other things on top of it.
The web frontend wouldn't be too hard, it could just be a Vue3 app with an HTML table element.
Then you could have a cell type who's value was a query, which would embed a DB browser list widget in the cell, and that cell type could have the option to bind its selected row to another cell.
Like VB+Excel+Access+My fork of Freeboard with inputs in one!
I’d assume the issue will be that an sql table is not free form, you can’t randomly decide to write in some other cell.
Early 1990s, college internship. The company did presentations for clients, like many do. They had an unusual way of presenting data that required using actual protractors to draw circles and curves, with pencil, on otherwise computer-generated charts. They read numbers from Excel spreadsheets and plotted them on paper.
I was shocked, to say the least. I proposed writing a program that read Excel spreadsheets and emitted the graphics. They loved the idea, especially from a summer intern.
So I wrote a letter to Microsoft asking for documentation of the Excel file format. A week later I got a thick envelope with a photocopied manual completely describing the format. I remember the word BIFF throughout. I wrote the program, it worked great, and I even negotiated a hefty lump-sum payment to sell it to them at the end of the summer.
It left me with a very positive impression of Microsoft as a developer-friendly company. Makes sense; developers are their platform's customers, and they're good at serving their customers.
What's that? I ddg'd it and I got a bunch of hits for regex \d ...
See https://jmmv.dev/2020/08/config-files-vs-directories.html
Basically every engineer who joined the team thought it was a unforgivable blemish on the system, yet it survived a few years with no major issues, long enough for the team to build an internal backoffice and port the whole sheet structure into a proper CRUD API.
Also, I said “most.” Not, “there exist no exceptions.”
At Google, many internal tools use Sheets as their source of truth for config data, and it works really well.
I agree completely but thats not the point he appears to be making. He never stated this was the use case and he reiterated that it was a bad idea which should never be done.
Regardless, I’m saying those arent really unique advantages to excel. They just look unique compared to json, toml, yaml, etc.
Cross-platform; copious spreadsheet applications; no-one needs training; scales.
(1) A huge dependency in the project for reading Excel files.
(2) Everyone needs training, as usually people in my profession have rarely if ever any need for Excel.
(3) Visuals != contained text, so the config might be different from what you see on screen as the config value.
(4) No proper version control, even a csv file would be better.
(5) scales? lul, have fun trying to solve merge conflicts. Also don't come with any Excel git plugins. It will only make the bad decision worse.
Many software developers usually have no need for something like Excel. Either they use some free/libre alternative, because they know about it existing, or they use an actual programming language, or some might even use something like Emacs org-mode spreadsheets, or they use some library like Pandas for things, where it is reasonable to get out the tools. Software developers are also more likely to be aware of the technical debt incurred by storing anything inside Excel formats and will avoid it, if they are wise.
As such many software developers rarely use Excel, if at all. I personally don't use it at all. All my simple spreadsheet needs are covered by Libreoffice Calc or Emacs org-mode. If I had to use Excel now, I would not know the names of functions (translated perhaps, because Excel does that silly stuff) or how to reference cells (Is it $ and then the number? And : as a separator between col and row?). So yeah, to properly use it, many of us would have to learn at least a little of it.
Many if not most cases of Excel usage are actually due to people not knowing the alternatives, or perhaps knowing they exist, but not having the knowledge to use them (like with programming and quickly dishing out a few Pandas calls or Emacs org mode spreadsheets).
Excel is like Python: the second best tool for a lot of problems. There are problems which take five minutes to solve in the shell or in a text editor (maybe less now when LLMs can straight produce certain solutions with a simple prompt) and they take 10 seconds in a spreadsheet including copy and paste. I really recommend spending some time with Excel just as I recommend reading the table of contents of your primary DBMS’ manual.
I’d actually argue that being able to solve a problem quick and dirty in excel and in then in a more proper way in pandas is a good thing.
I am perennially lost in all but the very simplest spreadsheets.
The problem IMHO is that we're using "configuration" files for things that aren't configuration. The "story" example from the blog post illustrates this. I find it hard to read YAML files that are dozens or hundreds of lines long.
Furthermore, once your "configuration" starts being so long, there's usually going to be enough repetition that you want to extract some duplication. YAML does have some facilities for this (anchors), unlike some other formats, but they're extremely limited.
So what happens is that different tools using YAML all start designing their own mechanism for sharing behaviour. It's all usually very ad-hoc, has edge cases and may not do things in the way you expect them. It also forces you to learn the specific rules for these facilities instead of allowing you to reuse your general programming knowledge.
On top of it all, YAML is essentially just a structureless key/value data structure. You can add schemas, but as far as I know, this isn't really standardised and editor support is... variable. In the worst case, you don't get any indication that you've configured something wrong. This is also part of the reason why I think that significant whitespace is OK for a programming language (still not a fan of it though), but bad for a configuration format, because bad indentation in a program either won't parse or will lead to obvious runtime errors, whereas bad indentation in a YAML file might just mean that a key isn't being set even though you think it should be.
For authors of tools that consume YAML, this means writing a lot of custom validation logic instead of relying on standard techniques like type systems.
I think we're on the wrong track and essentially just repeating XML's mistakes (just slightly less verbose, but also without schemas). We should rather use the programming constructs we know, e.g. by leveraging internal DSLs (I think that's part of the reason why Ruby was popular for tools like Chef for a short period, why Jenkins uses Groovy and Gradle now uses Groovy or Kotlin - these languages make internal DSLs easy). If we're worried about Turing completeness, maybe Dhall or something like it is the answer. But 400 line long YAML files with custom "!reference" tags that my editor doesn't understand doesn't seem like the solution.
I like TOML, I started to look into using Hugo over Jekyll though and the TOML seems weirdly abstract and difficult to follow.
Both are much easier to read, and don't have the footguns of YAML or StrictYAML.
I would generally say JSON5 is more appropriate because it is simpler, but Jsonnet does have some neat features and its IDE support is much better.
JSON is good enough for anything I've done. Not perfect, but no serious flaw that can't be fixed by just adding a simple app-specific post-process step that I will inevitably do for any other format anyway. JSONSchema gives us some typing sanity.
Can we just move on already to more interesting problems? It's not like git fulfills every VCS wish I've ever had either, but I have to move on. Projects and libs that introduce new config formats that continually remake the wheel, whose quirks have to be learned, are not helping my net productivity.
</rant>
I want comments in config files.
{
“”:”// comment here”,
“Entry”:[-1,0,2],
“”:”// next comment”,
“Flag”:true
}But if you use those extensions, all your tooling breaks.
(Aside: I think the real bike-shedding would start when you want to add some syntax for raw string literals, e.g. heredocs; it's one of those features that feels redundant, until the day when you really need it and you can't bear the pain of repeatedly escaping and unescaping.)
If someone wants JSON with extra features like comments and typingz they are better off switching to Ion.
JSON is good enough for anything I've done.
...
</rant>
Except for closing comments, for that, you need XML.I moved the configuration of several of my programs to Hjson. There are still problems but they're less puzzling. Hjson isn't ideal either but might still be the best configuration format we have today.
You've mentioned in the blog that ": it's meant to be written by humans, read and modified by other humans, then read by programs", but is it possible for apps to (roundtrip)-edit those configs preserving all the human syntax intact? It's rather common for apps to e.g. have font size changed, but unfortunately also common to destroy human formatting in the process
I didn't do it in my deserializer because of the big value you have in Rust in being compatible with serde and that wouldn't be. But this would be interesting, probably as an side library.
But I definitely strongly prefer YAML to TOML. It's just makes a lot more sense to me and it's a huge shame that PyPA went with TOML which is so un-Pythonic. I preferred setup.py. StrictYAML is a really good development that I wasn't aware of, though.
I'd argue that's enough for TOML to be more pythonic than YAML.
The main value proposition of TOML is to provide a concrete specification of a INI-type config language. INI is ubiquitous, but it lacked a spec, which led to a lot of wheels being reinvented. TOML fixed that.
If a project needs convoluted config files, I'd argue the project is already broken. If TOML doesn't fit your needs, that's hardly TOML's fault.
I always comment the same thing on these sorts of discussions - JSON5 has been really nice to work with if you can fully control all consumer applications of it (since there aren't great libraries for json5 in every language). Certainly nicer than the hellscape that is YAML.
Which seems fine given TOML was designed as a stricter and more reliable INI, not a half-assed programming language. Can’t say I’m unhappy to know that when I see a toml file, it’s probably going to be pretty simple (exactly the opposite of the dread I feel when I see a yaml extension).
[Lib]
But also
[[bin]]
I wrote a config file format. I took JSON and added comments, strict typing (using one character for any type that needs a marker), the ability to split items with newlines instead of commas, trailing commas, explicit binary data (as base64), and a new type I call "symbols" for things like enums or references. Then I removed the need to surround keys with quotes if they have only a certain set of typical characters.
It looks like [1].
It turns out that JSON was very close. It just needed a few more things.
Edit: one thing about config file formats that I strongly believe is that they must be purely data, no code. If you need code, supplement them with a separate thing, and perhaps use an established language, like Lua.
[1]: https://git.yzena.com/Yzena/Yc/src/branch/master/build.gaml
Lua pivoted to plugins instead.
It is:
- Streamable
- Extensible
- Whitespace-insensitive, but there are formatting conventions for readability
FWIW my best config experience has been with HOCON via typesafe|lightbend/config in Java. The ability to compose environment specific defaults in a reasonable way just felt good. Of course /internal/config to dump the config was a necessity, but trivial, so I tend to be less sympathetic to DRY is not necessarily good arguments.
Have been missing it in Go, unfortunately Java's ability to publish files within packages (that can be imported in config) was a key part of the UX that is missing in any compiled language I've seen.
For those who haven't encountered it before, HOCON is a superset of JSON so all valid JSON is also valid HOCON. Then it starts adding syntax sugar and useful features specifically for configuration files (the "H" stands for Human).
We wrote a tutorial for our product, it has a slider you can move to see how JSON evolves into HOCON along the way:
https://conveyor.hydraulic.dev/11.2/configs/hocon/
It's got some nice features. There's no "syntax typing", programs that use HOCON are thus very forgiving. Conveyor takes that even further, for example, anywhere you would normally need to specify a list of strings you can also specify just one string, it'll be wrapped automatically. There is a formal spec. It supports substitution, inclusions and it defines the semantics of duplicate keys which allows for refactoring of configs out to separate re-usable files. It has a nice clean look that gets out of your way. You can not only include files but also URLs.
On top of that we add a few more features. If you need to express a list of strings then brace expansion is supported, i.e.
foo = "bar-{1,2,3}"
is equivalent to foo = [ bar-1, bar-2, bar-3 ]
But probably the most important feature is hashbang includes. These allow you to include the output of arbitrary external programs: include "#!program --flags"
This lets you get the best of all worlds - the fast loading, simplicity and IDE sympathy of a declarative JSON-based config syntax, but if you hit the limits and need to programmatically generate some config you can do so whilst restricting the imperative logic only to the part of the file where it's needed. The rest remains declarative.All this works pretty well. At some point I want to package this up into a native library using GraalVM so it's available to anything that can load native libraries. Being Java it's accessible to any language that can run on the JVM which is pretty good already, but to use it from Go would require bindings.
There are some downsides. It's not as well known as other syntaxes so syntax highlighting is sometimes missing from things like docsite generators. IntelliJ has a plugin for it but for other editors you might not get good support. It doesn't really have a schema language either, although you could of course just use JSON schema.
What I prefer about TOML for the data I have (dictionaries of dictionaries, no deep nesting) is that the textual representation is flat. I find it easier to read and edit.
For comparison, here is one project's TOML data (formatted using https://github.com/tamasfe/taplo with added comments): https://raw.githubusercontent.com/dbohdan/structured-text-to.... Here is the same data converted to YAML with unlimited line width and formatted using https://github.com/google/yamlfmt: https://paste.dbohdan.com/projects.1694621084.yaml.txt.
EDIT: typos
Significant indentation alone doesn't eliminate this class of bugs. I'm pretty sure I've only seen this once ever, and can't remember the exact combination that caused it, but I have encountered a file that mixed tabs and spaces in a way python didn't barf over, that resulted in visual indentation being different from the programmatic indentation. I can definitely say it was python 2, not 3, so that particular instance may error nowadays.
just few days ago I crafted together some ideas i had couple of years already for a configuration language, syntactically like HCL but without HashiCorp's idiosyncrasies.
Here it goes, BCL (_Basic_ Configuration Language, for a lack of better name yet), Go prototype, I can code Python port and possibly several other as well..
PS. It was a pleasure to code, esp. getting the parser & the reflection right; the latter enables so elegant api. And yes, together with Russ Cox I am in a camp claiming that yacc is alive and well
It also benefits from being easier to copy + paste configs from docs since the entire path is included in each block; users don't have to worry about setting up the correct hierarchy.
Regardless, I strongly agree with points #1-#3. I also wish there was a way to support something similar to inline YAML schemas to catch users' typos in their IDE.
I would love to use HCL or potentially Ion, but IDE support and widespread acceptance isn't strong enough yet.
No semantic whitespace please. Can't tell you how many times I've seen something in GitHub Actions or similar get messed up because someone forgot a space.
I don't think TOML is a perfect format and agree it's not great at hierarchies.
As a library implementor, I wish arrays would hold only one type at a time, but I get that could be useful for users. But as a user, I wish tables were fully defined once (more can't be added up later in the file), especially when using larger files.
The article is great, and so are a few other links. Particularly enjoyed this[0] on the same site - The Norway Problem, which discusses parsing challenges in YAML.
[0] https://hitchdev.com/strictyaml/why/implicit-typing-removed/
If we kept strings double quoted as they were forever it wouldn't have been a problem.
That's the easiest way.
* TRUE / True / true & FALSE / False / false = parse as boolean
* "something quoted" = parse as string
* ...
Same stuff happened with the first version of Angular. For some reason they parsed "no" as false, even though JavaScript coerces it to true.
That's like having an engine problem and replacing engine with Shuttle Thruster. Or Space X rocket engine.
YAML imo will never beat JSON. The spec is so complex and full of edge cases that it will never reach simplicity and speed of JSON.
I implemented a spec compliant parser and it gets around 20 MB/s. In JSON I can get to at least five times that.
As opposed to YAML's double quotes (and single quotes, and unquoted, and folded, and literals) or you mean Toml's double quotes (and single quotes and multi-line quote).
Quotes aren't a problem. Many different types are.
Ironically, all JSON documents are valid YAML. YAML is actually a superset of JSON.
To compound the misery YAML is whitespace indented with some extra bonkers rules. So even your JSON subset in YAML needs to obey it.
Mix of flow and block makes both modes worse. However it also makes a lot of sense.
Don't judge me, as I don't use it anywhere besides personal projects)
I can look at any TOML file and I can see exactly which incredibly nested value am I looking at.
Is my configuration file twice as large as an equivalent YAML? Great, I'll pay 0.000001 more on R2 but I'll retain the ability to easily understand what nightmare configuration I am looking at
I agree with you the "3.14" vs 3.14 difference is not easy on some users, albeit that could be fixed in the business logic. Casting to int is not the end of the world.
I also hate indentation based anything (I hate Python with a passion - especially now when I'm forced to use it because AI people are fond of it)
IDK, having date/time as first class seems very good. It's so common,
foo = "something to use later",
bar = { "a": 1.1, "b": 2.2 },
baz = [ 1n, 999999999999n, ],
/* comments! */
{
[foo]: `\
multiline template string
${foo}`,
...bar, /* spread! */
"baz": [ ...baz, 1234n ],
"more": my_extension("something like !foo in YAML, but just JavaScript function syntax")
}Maybe defining the format as one big comma expression [^1] and disallowing functions+eval would have it's merit... but then, what to do about circular references?
Very creative though :)
[1] So that way, the whole config would evaluate to the expression behind the last top-level comma
Thanks for your elaboration
If I want a key-value file (for some simple preferences), TOML is by far the best out there. For anything more than that, I use YAML (though I should probably reach for a simpler subset of it, as YAML is crazy).
Unless there are deeply nested large maps/dictionaries, TOML is fine.
I prefer UCL, though. IMO, the most sane and usable format for configuration.
When the config file needs to be consumed by a third party, then I’d use TOML.
I basically completely disagree with this, especially namespaces: the ability to combine elements and attributes from multiple document types without ambiguity is extremely powerful and pays dividends when designing a query language like XPath
This is a very interesting hot take if I have ever heard one. I don’t want to agree, I don’t like significant whitespace, but I might have to agree.
I think by the same line of reasoning, spaces for indentation rather than tabs violates DRY as well though? I could be wrong.
Seems to solve a lot of the issues I have with regular JSON, YAML and TOML.
The only thing I wish there was a way to go directly between StrictYAML and dataclasses (or have an easy way to generate a corresponding dataclass from the schema).
Perhaps that was because they didn't consider them or maybe they did consider them but could not think of any reasons not to use them?
Among the things that S-expressions don't address per se are the interpretation of tokens (e.g., numeric vs non-numeric token syntax), how non-numeric tokens may be interpreted as booleans, symbols, timestamps, etc.; whether to use alists vs. plists for associations and the semantics of any duplicate keys within an associative construct; how to specify the schema for a configuration object (required vs. optional elements & the types of each, etc.)
IOW, merely saying "use S-expressions" over-emphasizes syntax while under-emphasizing semantics.
It's all about config files. Are there people using s-exps in config files?
I’d rather write some TOML/YAML/JSON for simple things where there are a handful of key/value pairs. In more complex cases where you want specific typing, or inheritance, or to solve any of the problems in this article, oh how I wish I could use something like a Guile parser to evaluate s-exprs. It would make life so much simpler.
You could also include anything built on Clojure EDN.
And, in the same way as Elisp for Emacs, Guile is an s-expression based language used not only for configuration but also for extension of GnuCash, GIMP, LilyPond and Pidgin.
Tools can parse it as JSON and that’s it.
Environment variables are generally a hack and should be avoided where possible.
I'm assuming StrictYAML is your favourite out of all the different config files you could have currently?
That means it's harder to learn, that it will be supported by fewer tools, and that the tools that do support it will have more bugs.
This has to be one of the worst takes I ever heard.
Comments, trailing commas, multiline strings, and a real date type.
Comments, trailing commas, multiline strings, a real date type, and dropping the quotes on keys.
Comments, trailing commas, ...
JSON really isn't just the "simple tweak" away from being the perfect config language it is often presented as. There's this fun sort of error that only a group of people can make, where someone can stand up and make a statement like "JSON just needs a couple of tweaks to be the perfect config format!" and everyone can individually nod in agreement, and the group thinks it is in agreement. But it turns out every individual actually had a completely different interpretation of the statement and they don't agree at all. When you get a group of developers together to be clear about the changes to JSON you will discover not everybody has the same tweaks in mind.
(& I wrote this before seeing the sibling comment from enriquto who gives another thing necessary for perfection. I don't even 100% know what enriquto means by "standard" floating point numbers versus what is already in the spec, but I bet it's another thing that multiple people could nod in agreement to but in fact mean something quite different about if you nail them down to the spec level!)
IMO, that would make it "good enough" for most use cases.
Having the application parse a date from a string isn't too bad. And dropping quotes on keys is just a nicety. And really, trailing commas aren't strictly necessary either.
But comments are absolutely necessary for a human readable configuration format. And multi-line strings are critical for any use case where you have strings that could be long. Like, say, a "description" field.
As a sidenote, the absence of multi-line strings is a major frustration I have with writing JSON schemas in JSON.
JSON could be much better, and good enough for far more purposes, with the top 3 tweaks there.
I think it's not entirely unlike programming languages. If I sit down and write out all the features I want from a programming language, I can't have it. They straight-up contradict each other. As a simple example, I want simplicity and a rich type system. Nope. Not gonna happen. Intrinsically at odds with each other. And there are many other such places, more subtle than that but no less real.
Config has the same thing going on. I've made my peace with just doing whatever's convenient in the moment and not stressing about it. It certainly isn't worth bringing in a new syntax my local team has never heard about because it's just the perfect little config syntax. It's very hard for any config syntax to be better enough than what came before to justify adding another tech to the stack.
JSON with comments wouldn't be perfect, but it would be unacceptable a lot less often.
I'm not sure who made the decision at the time, but looking back at the hours we have collectively wasted because of this non-feature, I believe they would have reconsidered.
TOML on the other hand is also a clear indication that Tom did not like using (or hadn't learned about) a formal grammar.
And I don't thing you can validate JS with a regex? (Have I forgotten that much of formal language theory?)
There's just a couple inherent problems with the notion of a universal configuration language...
First not all software has the same type of configuration requirements. I don't mean different schemas, I mean some measure of "configurablity". You have things like nginx or varnish that are configurable enough that configs are almost (or actually) programs. You also have programs where you need to store a few KVs that will be inserted into strings at some point and that's it.
That wide range of "amount of possible configuration" is hard to capture easily in a single configuration language.
Next you have a wide range "flatness" - that is if the configuration needs to be a list of kvs (optionally w/ sections) or if you need a tree or graph shape, or something weird.
Finally there's the whole preference and "style" thing. Writing a config language has an appeal to it - it's a fairly straight-forward thing to do (at least seemingly, until you hit the edge cases), and it's infinitely bikeshedable. In a similar vein, it seems to devolve into this sort of thing quite often: https://xkcd.com/927/
This is one of those situations where it's nice to be "old" - I like that I only have to know a handful of config formats these days, rather than keeping track of each program's bespoke config format.