TOML – Tom's Obvious, Minimal Language
toml.io
toml.io
Such file formats are typically used for configuration files. Yet, a substantial amount of effort was expended on data types not commonly seen in configuration files, such as a date-time-offset. Meanwhile, much more common data types typically used in configuration files is missing, such as GUIDs, IP Addresses, and byte arrays.
Most disappointingly, TOML has a "table" type, but unfortunately, just like other "RPC as Config" formats, the column names need to be repeated for every row, which is just crazy. It also makes no sense to talk about "arrays of tables", which are actually just "rows" in standard terminology. An array of rows... is a table.
Here is a terse example of the type of thing that does come up very often in bulk provisioning of the type that needs config files full of data:
@"
Server,Subnet,IP
db01,databases,10.1.2.3
db02,databases,10.1.2.4
user01,filers,10.3.10.1
user02,filers,10.3.10.2
"@ | ConvertFrom-Csv | New-VMBuildScript.ps1
I'm yet to see a config file format that can even approach this in terms of its readability and terseness.The example you give would suck in any configuration file format, and honestly I don't believe it belongs in a configuration file - I'd put it into a data file (e.g. a CSV as you show) and then in the configuration file I'd include a pointer to that data file to be loaded separately.
Plenty of real world applications use static hostnames and rely on DCHP to assign IPs. When you have systems that can fail hard and you need to replace a NIC, it saves having to update either a ton of config files, router configs, or your hosts file(s). Most networking library api calls also take strings as arguments for this reason. Not saying its a better solution, but there times when either an IP or a string hostname should be acceptable.
I've never had to use it, but Qt does have a hostname class, for an example of what such a dtype looks like. The class explicitly handles conversion for you:
After all we got ~16 % IPv6 traffic hitting our repository CDN. (seems to match with an observation of LWN https://lwn.net/Articles/808896/ )
It looks like you have tabular data where you don't need any nesting, every row has the exact same structure, and there's no single column of a row that you'd call "the value." TOML is not well-suited for this, and I don't think it sets out to be. I see how you'd call it "configuration," but it seems like a pretty different category of thing.
(I'd buy the argument that "table" is a poor choice of name for that TOML construct; I'd personally have called it a "multi-value" or something. It's an array of objects, and there's no guarantee that the objects have homogeneous structure, and it's usually useful for them to have different structures.)
In what situation is it not suitable to store these as strings, byte arrays as unbounded hex-encoded integers? 0xabcdef123400000000
As for your csv-example, I’d propose that if you have multi-value entries in a list like that large enough that it becomes unwieldy to repeat the property names every time (like in JSON or yaml), it really doesn’t belong in static configuration but in some kind of data store.
(And in the odd situation where that’s not the case, just use csv like you seem to prefer?)
If the format specifies some stronger types, the decoders can "bubble that up" to the programming language that decode the file. E.g.: a guid turns up as a "System.Guid" type in C# instead of "System.String".
This matters, because all of these are the "same" GUID, but not all GUID parsers can handle all of these formats:
{123e4567-e89b-12d3-a456-426652340000}
(123e4567-e89b-12d3-a456-426652340000)
123e4567-e89b-12d3-a456-426652340000
123e4567e89b12d3a456426652340000
urn:uuid:123e4567-e89b-12d3-a456-426655440000
Similarly, with IPv6 all of the following are the "same" address, but if encoded as a string it's hit-and-miss if the far end can properly handle all of the variants: fe80:0000:0000:0000:01ff:fe23:4567:890a
fe80:0000:0000:0000:1ff:fe23:4567:890a
fe80:0:0:0:1ff:fe23:4567:890a
[fe80::1ff:fe23:4567:890a]
fe80::1ff:fe23:4567:890a
fe80::1ff:fe23:4567:890a%3
Similarly, for your hex-encoded binary example, you'd be shocked at how rare it is for programming languages to provide a "hex-string-to-byte-array" decoding function. If you roll your own, it'll be about 10x less efficient than something done properly using SIMD instructions.Just recently I had to deal with Azure Resource Manager Templates, where even the distinction between numeric and text types has been eroded. There's several number fields that must be a string. There are also numeric fields that weirdly accept a string, but only if it is an ARM template expression that evaluates to a string containing a number at runtime.
Ugh...
C# is a good example. JSON doesn’t have any of these types. It doesn’t even make a distinction between integers and other numbers. Yet it’s common practice to marshal these into the kind of types you mentioned at ingestion.
At the end of the day, if you type it with your keyboard, it’s essentially a string.
I think it’s completely fine that all this functionality is part of neither programming languages core libraries nor configuration file formats - there is no one-size-fits-all here.
There are plenty of configuration frameworks that do what you ask for.
This is the #1 mistake people make when designing, evaluating, or critiquing frameworks, file formats, protocols, and the like.
I don't get to choose the language. I'm not writing 100% of the code that I use.
In fact, I have control over approximately 0.00001% of the code involved in processing a typical JSON file, or TOML file, or YAML file.
OTHER people control the language choice, and it certainly won't be ONE language. It'll be many languages.
If submit an ARM template to Azure, it'll go through at least three languages in the process: C#, Python, and JavaScript. Possibly C++ and F#. Who knows?
Even when the language is reasonably consistent, such as JavaScript, there is very little consistency in the specific parser used.
After the parser-level inconsistencies, there's further inconsistencies in how the stringly-typed data is converted into some more strongly typed format. That's just up to whoever wrote the code that consumes the data. There is exactly zero standardisation of this. None whatsoever. It's almost never documented, and there's just no way to know without experimenting.
NONE of this is in my control, or your control. I can't emphasise this enough.
Stop thinking in terms of a developer sitting down in front of an IDE with a new project, where they type all of the code in, hit compile and ship the binary to some customer.
Start thinking in terms of having to deal with inconsistencies between Terraform, Cloudformation, Amazon, AWS, GCP, and CloudFront.
Starting thinking about having to publish something like a Rust module to three different crate repositories, and cache it locally, and have it work properly with the various IDEs people use.
Start thinking in terms of security vulnerabilities of different parsers at different layers escaping data differently, or smuggling data using encodings accepted by the back-end past web application firewalls that don't understand that particular encoding.
This stuff matters.
JSON famously had a definition so short that it fit on a business card.
Less is more, right?
This "simplicity" lead to fun blog posts titled "Parsing JSON is a minefield": http://seriot.ch/parsing_json.php
The designer of TOML seems to have never read that article, or if he has, he hasn't learned any lessons from it...
I mean .. yes, but that's the exact chaos we're trying to mitigate? The different types may all get serialised as strings, but their semantics are different. And the purpose of strong typing is to make it harder for people to put semantically wrong things in the wrong place.
(This is semantic markup vs <font size=2> all over again, isn't it)
If you want a strongly type configuration that works across many languages, you'll need to also provide the per-language implementation of those types to fill the gaps.
You'll notice that JSON, YAML, INI, and XML also do not have IP address nor GUID types.
What configuration format does meet your approval?
None at the moment.
My problem with all of them is that I can't scan them visually to spot typos or mistakes. They break the content up and move related data too far apart for the inconsistencies to be immediately apparent.
Don't laugh, but in the absence of better options, I prefer to use Excel. With tabular data, things that belong together are visually adjacent. There are no repeated headers, or unnecessary syntax. Formulas are available, but hidden by default.
In the past I liked to use XML, but only with Altova XmlSpy, which had a kind of "hierarchical table view" superficially similar to Excel, but for tree-like data.
Both of the above are easy to use from PowerShell. I even wrote myself an Import-Xlsx module that works without Excel having to be installed. It uses the .NET framework libraries for opening XLSX files. (Under they hood they're essentially just a Zip file containing plain XML data.)
Ideally?
I'd like something like the SQL Server Analysis Services MDX query language, but with a front-end more like Excel for the editing of the data.
It is purpose designed for multi-dimensional data, and in my field that's exactly what I have to do: provision every combination of a bunch of tables.
For example: For each data centre, for each availability zone, for each of PRD/UAT/DEV, deploy each of the following roles, with 'n' instances each.
This is a multi-dimensional cube, also known as a full outer product.
Naively expanding everything is not good either, there needs to be some sort of language for making "exceptions". Such as: The DEV environment is only in this location. UAT has fewer sewers than PRD. Etc...
That's where a nice language would work wonders. You'd want to be able to do things like:
*.*.dev.sql.instances = 0
us-west.*.dev.sql.instances = 2
That would set the dev instances everywhere to zero, except US West. It's selecting a hyperplane through the cube and setting all of the cells to the constant on the right. This is the kind of thing MDX is designed to do, but no "simple" config language ever can.So instead you get monstrosities like the Azure ARM "copy" syntax: https://docs.microsoft.com/en-us/azure/azure-resource-manage...
I'd wager that the major reason no one has taken it that far is that your preferences are very rare.
But who knows, if you build it and make it available maybe others will find it useful too.
You're probably getting downvoted because unless I'm misinterpreting your previous comments, you're mixing up abstraction levels and concerns when you're talking about types.
Personally I think yaml is a terrible and confusing format in general (yes I do understand it quite well), I see your point with JSON, but I wouldn't personally say it's a fundamental issue the way you're describing it.
Sure, if I were provisioning 10K+ more-or-less-but-not-entirely-the same objects, I'd be reaching for a SQL database engine of some sort.
For about 10 things I'd just grit my teeth and click through a GUI manually.
In between 10 items and 10K is where things get interesting. Inefficient tools can waste more time than they save. Full-on programming languages are right out. You have to consider dealing with people who aren't programmers. People that aren't DBAs.
This in-between-land of, say, 50-5000 instances is where configuration data files live. It's where Excel works well enough at the moment, but I feel that something better is just waiting to be invented.
There’s no going back to GUIs when the shell and keyboard become and extension of your mind.
But I think you have a point in that there’s an big user base of somewhat-power-users that prefer visual interfaces and that there’s a better middle ground to be found. The biggest hurdle I see is that tools like that easily become outdated, have selective platform support, aren’t as extendable and customizable, etc. it’s a vastly bigger undertaking to make something that can work everywhere with everything in the way text UIs can.
> Dhall is a programmable configuration language that you can think of as: JSON + functions + types + imports
It's specifically not Turing-complete so it has a number of safeties built in because of this. I've been using it more and more wherever possible as of late.
Is this really an issue when loading a field from a config file? If it was something happening 1m times per second then maybe I could see your point.
The fact that your config will load some tiny amount faster with those optimizations is a side effect of the hot path optimizations, not a requirement for config formats.
I feel like that's something that's more an appropriate concern for the application, and not something that ought to be baked into the configuration format itself. For example:
> E.g.: a guid turns up as a "System.Guid" type in C# instead of "System.String".
Can you not just pass the System.String you get from the decoder into System.Guid's constructor and handle the resulting error should that string not actually be a GUID?
Like, my problem with a lot of config file formats is that they're frequently too clever about assuming types of things when I would much rather they be strings by default for me to interpret later.
> If you roll your own ["hex-string-to-byte-array" decoding function], it'll be about 10x less efficient than something done properly using SIMD instructions.
Unless you roll your own that uses those SIMD instructions, whether in hand-written assembly or in a programming language with a compiler smart enough to use SIMD for this.
This does, however, reek of a premature optimization.
databases:
db01: 10.1.2.3
db02: 10.1.2.4
filers:
user01: 10.3.10.1
user02: 10.3.10.2E.g.: Region, AZ, Network, Subnet, Description.
Let's see it do a tree. (And not an adjacency matrix. That's cheating.)
To connect with the other threads, SQLite would be probably the best candidate proposed here to store a tree.
TOML does a much better job than CSV.
Quoting at least makes it clear what's part of the value.
E.g., having to specify column names again and again is tedious to write, but makes the file easier to read and modify safely.
After normalizing, one could get to a tidy heterogeneous solution.
(Only the finest vaporware spoken here.)
In a tiny yaml file it's pretty easy to understand what everything is; it's huge files where toml's forcing you to full qualify nested tables becomes useful.
Too many huge yaml files where I'm just scrolled somewhere in the middle and have no clue what part of the tree I'm actually looking at.
I think yaml files hold's up well to about 50 lines.
I have only seen toml files up to about 10-20 lines and it looks great at that scale, but I am having a hard time imagining it being nice at 500+ lines.
Because everyone would agree for 80%, but then need different 20%.That's how you end up with 10k pages specs or XML and doc types.
And if you need that just use XML, that is actually what it was designed for.
TOML is for config, config that users edit by hand, not input.
But there is a reason we have a word for config.
It's a specific kind of input: it feeds the boostrapping state of your program, it's not the main data your program process.
The difference is important: structure, complexity and dynamism requirements are not the same.
2 - config sets your program boostrapping state. It's not the main data your program processes. That's your main input. IP addresses are what the script processes. It's not config.
2 - Those IP addresses are probably not the main data the program processes. In my case, they were configuration, describing where an auditing processor scrapes input and how to allocate compute.
opt="
# this is the foo option
a=1
# define the bar
#BAR=ON
#BAR=OFF
BAR=MAYBE
" | egrep -v '^#|^$' | sed 's/^/-D/'
or progs="
one two three
#four five siz
seven ate tree
" | egrep -v '^#|^$' | while read o1 o2 o3
do
program $o1 $o2 $o3
done
sort of like on-the-fly readable configurability - [Server,Subnet,IP]
- [db01,databases,10.1.2.3]
- [db02,databases,10.1.2.4]
- [user01,filers,10.3.10.1]
- [user02,filers,10.3.10.2]
That being said, the CSV does seem more straightforward.I tried Python's INI for configuration, but it was both too simple and too complex at the same time, very frustrating.
YAML is a nightmare, the formatting is way too easy to get wrong and accidentally break your configuration.
JSON is too strict/simple. You can't even have trailing commas in a list!
[info]
name = { first="bob", last="jones" }
and info.name.first="bob"
info.name.first="jones"
This can be seen as a positive (it's a really simple document) or a negative (it's hard to machine-generate nice TOML, and it's easy to think there's more to the structure of a written document than there is).As a comparison, there's not much choice when serializing to JSON - the only variation is around whitespace. When serializing TOML you need more insight to decide what's the best representation for a human. A contributing factor to this issue is that TOML is mostly used for configuration files, so it often matters that the output is readable.
This is why I like JSON!
I have seen the horrors of XSLT and I’m never going back.
In TOML (or YAML) it's just as easy as prepending it with # at the beginning of the line.
Simple, write comments in Morse code using spaces and tabs.
{"comment":"hi mom", ...}* json5
* HCl
* hjson (now unmaintained)
Which is quite sad, as I thought that it was the best one (if memory serves)..
I never liked that anything I pulled from a config file was never checked by typescript, so I just went simple with TS itself.
I understand if I use another language within the same project, it'll be a problem but until then this can't be beaten.
a bit from 5 months ago https://news.ycombinator.com/item?id=22707387
2018 https://news.ycombinator.com/item?id=17513770
2013 https://news.ycombinator.com/item?id=5272634
https://news.ycombinator.com/item?id=5274025
Alternative formats to this alternative format:
* comments
* bareword (not double quoted) object key names
it'd be a much more suitable configuration file format for everything, and I don't think we'd see the motivation to make things like TOML and YAML.But that's a minor quibble. If everyone adopted this thing instead of legacy JSON, we'd be better-off.
Javascript already has this, so allowing single quotes simplifies copypasting object literals, which is a common use case.
On a deeper level, I think they might be frustrated with the ability to develop custom formats (XSDs) which effectively make XML not one format, but a gazillion formats.
There's plenty of room for debate, e.g. element versus attribute and things get unwieldy pretty quickly [1]
The "note"-example illustrates the issue quite well: the order of the elements matters and you end up writing every element identifier twice.
Not to mention the bloat that comes with using an XML library that's actually compliant with the standard and includes all the bells 'n whistles.
with json - you just suck it all into your data structure
with xml - sometimes people put stuff as a tag, sometimes as an attribute. Maybe not you, but someone else will do it.
your comparison to yaml or ini is apt; toml's strength over yaml is syntax simplicity and toml is more-or-less a superset of ini (which itself is poorly-defined).
Modding command and conquer through "rules.ini" was always an adventure.
I honestly find TOML harder to read and more complicated to use than YAML. I tried to look into it but frankly I still haven't found a use case where it made more sense to use TOML.
There's been times I wanted YAML support in a particular app; I've never wanted TOML support. If anything, it's been "ugh, I have to use TOML".
https://hitchdev.com/strictyaml/
It's a smaller project so probably could use some support.
Generally I prefer actual code based configuration though, there's only a small subset of people than can read and edit these config formats but would be unable to configure something like DWM through real code (https://git.suckless.org/dwm/file/config.def.h.html) and this approach gives far more flexibility without code bloat. Once customization through config is removed it's only machine-to-machine configuration that changes and these are easily handled by TOML, INI or even environment variables.
There is a subset of yaml that is almost good. But even then, there are cases where yaml is interpreted differently the average human would expect.
country: noI totally understand that sentiment. But that does seem to contradict the "batteries included" philosophy. Then again pip itself seems like one of the batteries you would expect to be included, so maybe that philosophy is just no longer as relevant to python.
If you are provisioning servers from a TOML files, you're doing it wrong. That's your program main input, not config.
If you have deeply nested values, chances are you are again confusing configuration vs main input.
JSON and CSV are great main input formats.
TOML is a great configuration format.
YAML and XML try to do both, and end up doing average on either, but being abused for all of them (looking at you Ansible and Solr).
So you should use each format for the correct purpose, as usual.
A config file looks like ~/.ssh/config or /etc/nginx/nginx.conf, not like ~/.ssh/authorized_keys or /etc/nginx/site-available/default.conf.
Just because we love to put a lot of data input in config folders (.config and /etc are full of input data) and call that configuration doesn't make it so.
How to know if something is configuration and something is input?
Conf data is usually only manually edited (authorized_keys changes with the life cycle of the program) and doesn't contain logic (default.conf if full of logic). Conf is also how the program is going to behave when performing its main task, the boostraping state, not the data used for it's main task (so not ip for provisioning servers, which is not meta, but the main course, or default.conf, which nginx uses to perform the main task). Conf rarely changes, main input often does. Conf is not piped or redirected to stdin.
`[[foo]]` syntax is not obvious at all. It's a bit weird that top-level declarations use ini-like syntax, but values can use JSON-like syntax, and there are multiple syntaxes to express the same data structure.
It doesn't strike me as very elegant, or obvious. Still, it is less annoyingly-inflexible than JSON, less verbose than XML, and less footgunny than YAML.
If it looks like text, make it as easy as possible for humans to consume and write by hand. Period.
That's a small but very important reason why HTTP rules the world.
Notice I did not mention readability. JSON was always meant to be human-readable; and its popularity is a testament to the rapid turnaround on debugging that the readable aspect of JSON affords. This is, say, in contrast to BSON which is similar but not humanly readable, and unsurprisingly less popular despite its bandwidth advantages.
Config formats are, by raw bulk, generated by humans, and need to be readable (ideally reproducibly across implementations) by computers.
JSON is good for data exchange. Configs? That's crazy talk.
All the benefits of JSON, no new syntax to learn, easy to fit in place in new or existing systems.
All languages should make unnecessary characters optional like trailing semicolons and commas.
There's also the NO/Norway problem (although it seems this instance of the problem may be fixed if you're always using newer versions): https://allan.reyes.sh/programming/2018/06/20/The-YAML-NOrwa...
They all have warts because using the same symbolic notation for data and structure leads to encoding issues that have cognitive overhead.
So instead of getting used to the pratfalls involved in making mistakes in one format/library set, we can do it in a whole bunch of different ones. And of course if anyone uses the new format much much, we'll get multiple, slightly incompatible versions, all with different bugs.
Thanks?
I can't seem to find a date for it, but I think TOML is pretty old, so probably predates a lot of these YACFs.
I think that the problem is that it's such an easy workflow:
1. No existing config format is perfect. Look at this should-be easy thing I want to do that isn't easy!
2. Create a new format that makes the desired should-be easy thing actually easy.
If you have one or two pain points, then it's really easy to address them, and your life gets so much better for a little while. If you're lucky and you're a good YACF designer, then your life gets so much better for quite a while—and, by the time you realise why the design isn't perfect, you're already invested in it.
Not _that_ old. TOML's first release was in 2013.
But as a testament to the config format, it was only about a year old when Rust's cargo adopted it fully.
>"
\uXXXX - unicode (U+XXXX)
\UXXXXXXXX - unicode (U+XXXXXXXX) "