YAML: Probably not so great after all
arp242.net
arp242.net
Surely using an encoder on an object/structure hierarchy (like people do with encoding/json) is the way to go?
On the other hand, the quality of the yaml libraries in Go wasn't great, last time I had to choose a configuration file format.
Yes, we may not have a cleaner way of deploying k8s Deployment configs to different clusters but the desire to templatize YAML is easy for everyone to understand. The decision to abstract or templatize is one rooted in time and cost, not ability to understand data structures.
In my head, serialization is a special case of encoding when B is some type of string. In this case it's YAML.
I could probably have worded my previous comment more precisely.
I just had a shiver recalling a Kubernetes wrapper wrapper wapper wrapper at a former job. I think there were at least two layers of mystical YAML generation hell. I couldn't stop it, and it tanked much joy in my work. It was a factor in me moving on.
oh god why
I'm pretty much thinking I want Go as a pre-config where I can set variables, loops, and conditionals and that my editor can help with auto-complete. Maybe I can "import github.com/$org/helmconfig" and in the end write one or more files for config.
If you use it in anger, you quickly need all language features e.g. importing libraries, namespaces, functions, data-structures, rich string manipulation etc. But you rarely get these.
At run-time, you don’t have a debugger or anything leading to a maddening bug fix experience because config cycle times are really high.
Because it’s a niche language, only one poor soul in a team ends up the expert of all the plentiful traps.
Eventually... you give up and end up generating the config in a proper language and it feels like a breath of fresh air.
Somebody got the silly idea in their head of implementing a templating language in PHP, even though PHP is ALREADY a templating language. So they took out all the useful features of PHP, then stuck a few of them back in with even goofier inconsistent hard-to-learn syntax, in a way that required a code generation step, and made templates absolutely impossible to debug.
So in the end your template programmers need to know something just as difficult as PHP itself, yet even more esoteric and less well documented, and it doesn't even end up saving PHP programmers any time, either.
https://web.archive.org/web/20100226023855/http://lutt.se/bl...
>Bad things you accomplish when using Smarty:
>Adding a second language to program in, and increasing the complexity. And the language is not well spread at all, allthough it is’nt hard to learn.
>Not really making the code more readable for the designer.
>You include a lot of code which, in my eyes, is just overkill (more code to parse means slower sites).
https://web.archive.org/web/20090227001433/http://www.rantin...
>Most people would argue, that Smarty is a good solution for templating. I really can’t see any valid reasons, that that is so. Specially since “Templating” and “Language” should never be in the same statement. Let alone one word after another. People are telling me, that Smarty is “better for designers, since they don’t need to learn PHP!”. Wait. What? You’re not learning one programming language, but you’re learning some other? What’s the point in that, anyway? Do us all a favour, and just think the next time you issue that statement, okay?
http://www.ianbicking.org/php-ghetto.html
>I think the Broken Windows theory applies here. PHP is such a load of crap, right down to the standard library, that it creates a culture where it's acceptable to write horrible code. The bugs and security holes are so common, it doesn't seem so important to keep everything in order and audited. Fixes get applied wholesale, with monstrosities like magic quotes. It's like a shoot-first-ask-questions-later policing policy -- sure some apps get messed up, but maybe you catch a few attacks in the process. It's what happened when the language designers gave up. Maybe with PHP 5 they are trying to clean up the neighborhood, but that doesn't change the fact when you program in PHP you are programming in a dump.
Wow... Yuck!
Lua is at least a "standard" and rather sane language. Not ad hoc insanity.
Also I can't make a head or tail of your statement, no offense intended—I can make out parts, but the whole just has no meaning to me.
I suspect you're front-loading a lot of logic from the app bootup into the config file, and I'm not sure what you stand to gain conflating those two things.
Like—what's the execution order of a config file? Can you refer to a value later in the file? If so, how does it determine which value is executed before the next? If not, how do you setup circular dependencies—can you redefine config values halfway through a file? If not, you're gonna have to fall back on at-boot pre-processing anyway with sufficient complexity, so just treat a config like dumb data to begin with and do all your logic in the bootup and put all the values/whatever in the config file. And god knows, I would be strongly tempted to murder an engineer that introduced a config file capable of rewriting itself—that engineer has clearly never needed to debug another person's shitty code before.
See: https://news.ycombinator.com/item?id=20735231
Also: https://en.wikipedia.org/wiki/Don%27t_repeat_yourself and https://en.wikipedia.org/wiki/Single_source_of_truth
The opposite of DRY (Don't Repeat Yourself) is WET (Write Everything Twice or We Enjoy Typing or Waste Everyone's Time) -- but twice could be ten times or more, there's no limit. Writing it all out again and again by hand as pure literal data, and hoping I didn't make any typos or omissions in each of the ten repetitions, without some sort of logical algorithmic compression and abstraction, would be idiotic.
To what end? I don't get what you could possibly be putting in these files that could consume so much space without refactoring.
I never used helm / kubernetes before 3 months ago.
Not 2 weeks ago I needed to loop in a helm config file in order to basically say "all this same config, libraries, etc., just run this other command instead" ... because someone who makes those decisions had ~100 lines of environment-injected configuration + boilerplate in the yaml that I couldn't get rid of, needed, and would have otherwise needed to copy / paste.
Since then, those environment variables have been pulled out into a different file (refactoring!), and now we replaced a loop over 100 lines of config, with 2x sets of 15-20 lines of config boilerplate. Better, but still a lot of bull. I don't know what the right answer is, because we've got less helm templating bullshit in there, but we still need boilerplate. Because it's not like I can tear down an entire kubernetes + helm infrastructure because I don't like how the config files are written.
Configs / config generation is hard, and generally awful. If you don't see it that way, congratulations; you're either a genius in your field, you've got not enough experience, and/or you're wrong. If you believe it's easy, and we're all missing something - please, by all means, write a book on how / why configurations aren't as hard as the rest of us say they are.
Best of luck to you.
The point I'm trying to make is that you're describing broken frameworks, data flows, and work flows, and blaming it on config generation. If you have a counter example, I'd love to see it. Discussing these things in the abstract is pretty pointless and based in emotional language/semantic quibbling rather than meaningful things people can reason about and discuss, like code comparison or time tradeoffs.
Hell, because no specific GOOD examples of configuration-as-code have been brought up, literally everyone in this thread could be considering a different pet example of theirs. It's OBVIOUSLY a waste of everyone's time without examples. Why bother comment at all—to go out of your way to punch down without contributing to the discourse?
You say this is easy. Seems to me that you're claiming to be elevated above us all with something we don't know, claiming that everyone else is doing it wrong, all the while hiding behind anonymity.
Stop clutching your pearls and faking the victim. No one is punching down; you're claiming knowledge you don't have and are being called out for it.
Look at any one of the references cited in the thread.
Pantomime extensively (and extensibly) uses the simple JSON templating system I described in this other link, which I wrote in C#. Everything you see is a plug-in object, and they're all described and configured in JSON, and implemented in C#:
https://news.ycombinator.com/item?id=20735366
Pantomime Playground - Immersive Virtual Worlds for the Rest of Us
Pantomime Creatures - Real-Time Augmented Reality
Bug Farm – First Minute, First Person 3D
Consumer Augmented Reality Arrives – with Pantomime
If I were to rewrite it from scratch, I'd simply use JavaScript instead of rolling my own JSON templating system, because it would have been much more flexible and powerful.
Oh wait -- I DID rewrite at least some of that stuff from scratch! To illustrate that superior JavaScript-centric approach, here's an example of some other JSON based systems I developed with Unity3D and JavaScript, one for scripting ARKit on iOS, and the other for scripting financial visualization on WebGL, both using UnityJS (an extension I developed for scripting and configuring and debugging Unity3D in JavaScript).
One nice thing about it is that you can debug and live code your Unity3D apps running on the mobile device or in the web browser while it's running, using the standard JavaScript debugging tools!
UnityJS is a plugin for Unity 5 that integrates JavaScript and web browser components into Unity, including a JSON messaging system and a C# bridge, using JSON.net.
https://github.com/SimHacker/UnityJS
WovenAR Tools with ARKit and Pantomime (built with an early version of UnityJS):
A description of UnityJS, and how I use JavaScript and JSON to define, configure and control Unity3D objects:
https://news.ycombinator.com/item?id=19804242
https://news.ycombinator.com/item?id=19748582
https://news.ycombinator.com/item?id=19745034
Here's a demo of another more recent application using UnityJS that reads shitloads of JSON data from spreadsheets, including both financial data, configuration, parameters, object templates, etc.
https://reasonstreet.co/2019/01/31/how-does-amazon-make-mone...
https://reasonstreet.co/2019/08/09/tech-giants-acquisitions/
Here is an article about how the JSON spreadsheet system works, and discusses some ideas about JSON definition, editing and templating with spreadsheets, which is about a year old, but that I've developed it a lot further since writing that article.
https://medium.com/@donhopkins/representing-and-editing-json...
>Representing and Editing JSON with Spreadsheets
>I’ve been developing a convenient way of representing and editing JSON in spreadsheets, that I’m very happy with, and would love to share!
>I‘ve been successfully synergizing JSON with spreadsheets, and have developed a general purpose approach and sample implementation that I’d like to share. So I’ll briefly describe how it works (and share the code and examples), in the hopes of receiving some feedback and criticism. Here is the question I’m trying to answer:
>How can you conveniently and compactly represent, view and edit JSON in spreadsheets, using the grid instead of so much punctuation?
>My goal is to be able to easily edit JSON data in any spreadsheet, conveniently copy and paste grids of JSON around as TSV files (the format that Google Sheets puts on your clipboard), and efficiently export and import those spreadsheets as JSON.
>So I’ve come up with a simple format and convenient conventions for representing and editing JSON in spreadsheets, without any sigils, tabs, quoting, escaping or trailing comma problems, but with comments, rich formatting, formulas, and leveraging the full power of the spreadsheet.
>It’s especially powerful with Google Sheets, since it can run JavaScript code to export, import and validate JSON, provide colorized syntax highlighting, error feedback, interactive wizard dialogs, and integrations with other services. Then other apps and services can easily retrieve those live spreadsheets as TSV files, which are super-easy to parse into 2D arrays of strings to convert to JSON.
[...]
>Philosophy: The goal is to leverage the spreadsheet grid format to reduce syntax and ambiguity, and eliminate problems with brackets, braces, quotes, colons, commas, missing commas, tabs versus spaces, etc.
>Instead, you enjoy important benefits missing from JSON like like comments, rich formatting, formulas, and the ability to leverage the spreadsheet’s power, flexibility, programmability, ubiquity and familiarity.
More info:
https://news.ycombinator.com/item?id=17384078
What this illustrates should be blindingly obvious: that JavaScript is the ideal language for doing this kind of dynamic JSON generation and event handling, so there's no need for a special purpose JSON templating language.
Making a JSON templating language in JavaScript would be as silly as making a HTML templating language in PHP (cough i.e. "Smarty" cough).
https://news.ycombinator.com/item?id=20736574
JavaScript is already a JSON templating language, just as PHP is already an HTML templating language.
UnityJS applications create and configure objects by making lots and lots of parameterized JSON structures and sending them to Unity, to instantiate and parent prefabs, configure and query properties with path expressions, define event handlers that can drill down and cherry pick exactly which parameters are sent back with events using path expressions (the handler functions themselves are filtered out of the JSON and kept and executed on the JavaScript side), etc.
At a higher level, they typically suck in a bunch of application specific JSON data (like company models and financial data), and transform it into a whole bunch of lower level UnityJS JSON object specifications (like balls and springs and special purpose components), or intermediate JSON user interface models like pie menus, to create and configure Unity3D prefabs and wire up their event handlers and user interfaces. Basically you're transforming JSON to JSON, and associating callback functions, and sending it back and forth in messages and events between JavaScript and Unity.
There are also a bunch of standard JSON formats for representing common Unity3D types (colors, vectors, quaternions, animation curves, material updates, etc), and a JSON/C# bridge that converts back and forth.
https://github.com/SimHacker/UnityJS/blob/master/doc/Archite...
This is a straightforward function that creates a bunch of default objects (tweener, light, camera, ground) and sets up some event handlers, by creating and configuring a few Unity3D prefabs, and setting up a pie menu and camera mouse tracking handlers.
Notice how the "interests" for events include both a "query" template that says what parameters to send with the event (and can reach around anywhere to grab any accessible value with path expressions), and also a "handler" function that's kept locally and not sent to Unity, but is passed the result of the query that was executed in Unity just before sending the event. The point is that every "MouseDown" handler doesn't need to see the exact same parameters, it's a waste to send unneeded parameter, and some handlers need to see very specific parameters from elsewhere (shift keys, screen coordinates, 3d raycast hits, camera transform, other application state, etc). So each specific handler gets to declare exactly which if any query parameters are sent with the event, up front in the interest specification, to eliminate round trips and unnecessary parameters.
https://github.com/SimHacker/UnityJS/blob/master/Libraries/U...
The following code is a more complex example that creates the Unity3D PieTracker object, which handles input and pie menu tracking, and sends JSON messages to the JavaScript world.pieTracker object and JSON pie menu specifications, which handle the messages, present and track pie menus (which it can draw with both the JavaScript canvas 2D api and Unity 3D objects), and execute JavaScript callbacks (both for dynamic tracking feedback, and final menu selection).
https://github.com/SimHacker/UnityJS/blob/master/Libraries/U...
Pie menus are also represented by JSON of course. A pie can contain zero or more slices (which are selected by direction), and a slice can contain zero or more items (which are selected or parameterized by cursor distance). They support all kinds of real time tracking callbacks so you can provide custom feedback. And you can make JSON template functions for creating common types of slices and tracking interactions.
This is a JavaScript template function MakeParameterSlice(label, name, calculator, updater), which is a template for creating a parameterized pie menu "pull out" slice that tracks the cursor distance from the center of the pie menu, to control some parameter (i.e. you can pick a selection like a font by moving into a slice, and also "pull out" the font size parameter by moving further away from the menu center, and it can provide feedback showing that font in that size on the overlay, or by updating a 3d object in the world, to preview what you will get in real time. This template simply returns a blob of JSON with handlers (filtered out before being sent to Unity3D, and kept and executed locally) that does all that stuff automatically, so it's very easy to define your own "pull out" pie menu slices that do custom tracking.
https://github.com/SimHacker/UnityJS/blob/master/Libraries/U...
I don't use configs in this way, or if I did, I would not be inclined to call them configs. I can certainly appreciate the problem of processing of JSON objects in many different contexts. I was more referring to a UX concept of providing a configuration interface—short of something like emacs that gives full functionality, simpler and easily debuggable is emphatically better.
The fact that you've never done and can't imagine anything complicated enough to need more than a simple hand-written data-only config file doesn't mean other people don't do that all the time. It's simply a failure of your imagination.
What I can't understand is what you were getting at about "punching down". When you say things like "I would be strongly tempted to murder an engineer", that sounds like punching down to me. And why you were complaining nobody gave any examples, by saying "no specific GOOD examples of configuration-as-code have been brought up". Don't my examples count, or do you consider them bad?
So what was bad about my examples (or did you not read them or follow any of the link that you asked for)? Pantomime had many procedurally generated config files, using the JSON templating engine I described, one for every plug-in object (and everything was a plug-in so there were a lot of them), as well as some for the Unity project and the build deployment configuration itself. It also used dynamically generated JSON for many other purposes, but that doesn't cancel out its extensive use of JSON for config files.
This is exactly what Tcl supports / was designed to do (and in turn is one of my motivations for developing OTPCL). This is also exactly what your average Lisp or Scheme supports.
I think it all depends. Most of the time I would agree that you shouldn't template yaml, but sometimes, it's the lesser of two evils.
Some of the values like DDB table attributes are common across all regions, other values like tags are common across all infra in the same region. Some values are a scalar multiple of others, or interpolated from multiple sources. For example, a DDB capacity alarm for a given region is a conjunction of the table name (defined at the application level), a scalar multiple of the table capacity (defined at the regional deployment level), and severity (owned by those that will be on-call).
To add insult to injury, a stack can only have 60 parameters, which you butt up against quickly if you try to naively parameterize your deployment.
Given all these gripes, auto-generating CFN templates was easiest for me. I used a hierarchical config (global > application > region > resource) so the deployment params could be easily manipulated, maintained, and where “exceptions to the rule” would be obvious instead of hidden in a bunch of CFN yaml. To generate CFN templates I used ERB instead of jinja, but to similar effect.
A side benefit of this is side-stepping additional vendor lock-in in the form of the weird and archaic CFN operators for math, string concatenation, etc. I don’t have a problem learning them, but it’s one of those things that one person learns, then everyone who comes after them has to re-learn. My shop already uses ruby, so templating in the same language is a no-brainer.
I think its opposite, the most lean way to deploy AWS resources. Did you wrote it yourself, in text editor? I was doing it for 5 years now. You can omit values if you're fine with defaults, you only state what needs to be different. Other tip is use Export and ImportValue to link stacks.
I kept on using JSON, even after all my buddies jumped on YAML. JSON is just more reliable, harder to miss syntax errors, and can be made readable by not using linters and keep long lines that belong on one line. Also, the brackets are exactly what they are in Python :)
> wraps CloudFormation with a templating language (jinja2)
Not sure it it is a good idea. Everyone's use case is different, though. A well written CFN template is like a rubber stamp, just change the Parameters. The template itself doesn't need to change.
https://github.com/cloudtools/troposphere
The basic type checking done was quite helpful, and avoided some of the dumb errors that we had run into when we attempted to do everything by hand.
i've collaborated on ytt (https://get-ytt.io) - yaml templating tool. it works directly with yaml structure to bind templating directives. for example setting a value is associated with a specific yaml node so that you dont have to do any manual indenting etc. like you would with plain text templating. defining functions that return yaml structures becomes very easy as well. common problems such as improperly escaped values are gone.
i'm also experimenting with a "strict" mode [1] that raises error for questionable yaml features, for example, using NO to mean false.
i think that yaml is here to stay (at least for some time) and it's worth investing in making tools that make dealing with yaml and its common uses (templating) easier.
We had joy, we had fun, we had seasons in the sun, but as I added more and more features and syntax to cover specific requirements and uncommon edge cases, I realized I was on an inevitable death-march towards my cute little program becoming sufficiently complicated to trigger Greenspun's tenth rule.
https://en.wikipedia.org/wiki/Greenspun%27s_tenth_rule
There is no need for Yet Another JSON Templating Language, because JavaScript is the ultimate JSON templating language. Why, it even supports comments and trailing commas!
Just use the real thing to generate JSON, instead of trying to build yet another ad-hoc, informally-specified, bug-ridden, slow implementation of half of JavaScript.
I originally felt it was overly complex, but after seeing some of the Go text/template and Ansible Jinja examples in the wild, it actually seems like a good idea.
Perhaps we should more strongly distinguish between “basic” data definition formats and ones that need to be templated. JSON5 for the former and Jsonnet for the latter, for example.
It's not just alternate skin for JSON, and yet that's what most people use it for. Some users also want things like map keys that aren't strings, which is actually pretty useful.
I recall there being CoffeeScript Object Notation as well... perhaps that would've been better for many use cases, all things said.
I made it largely because I saw a disconnect with what YAML was, and what people - including me - thought it was (which is what it should be).
Don't agree with non-string map keys though... they're a complication I never saw a use for.
Even if you’re of the opinion that IDs shouldn’t be numeric, there are a lot of cases where you’re stuck with integers—on Linux, user IDs, group IDs, and inodes are just a few examples.
“Hey, I deleted a character in a string and now I am getting this weird schema validation exception”.
Maybe it would be sense if quotation marks were not required for keys with only a restricted character set which are not an empty string, though.
They don't mention mismatched surrogates and otherwise invalid Unicode characters, but they should perhaps be implementation-dependent, like duplicate keys are. (It can either allow them or report an error.) There is also the possibility that some implementations may wish to disallow "\x00" or "\u0000" too, I think.
The thing I disagree is the part about white space. U+FEFF should be allowed only at the beginning of the document (optionally), and other than that only ASCII white space should be allowed. Unquoted keys also should be limited to ASCII characters.
Other than that, I think it is good.
I don't get why TOML is so underrated, it's barely mentioned in the HN discussion
It is outright bad as a human-operated format. It explicitly lacks comments, it does not allow trailing commas, it lacks namespaces, to name a few pain points.
YAML is much more human-friendly, with all its problems.
If JSON was developed more recently on places like GitHub, it would never have ended up like that with that many deficiencies.
This is where comments belong.
If the comments are so critical that it is a problem, then an accompanying file with those comments would be used. Otherwise, it's just a bunch of crocodile tears.
Seriously, our batch jobs for better or worse have configs with a bunch of parameters that are passed around as json, and while most variable names are intuitive and there is documentation on the wiki, and most often the config can be autogenerated by other tools it would still be better if when I manually open it in the config itself I would easily see the difference between n_run_threads vs n_reg_threads, etc...
https://yaml.org/spec/history/2001-08-01.html
XML and HTML are markup languages. JSON and YAML are not markup languages. So when they finally realized their mistake, they had to retroactively do an about-face and rename it "YAML Ain’t Markup Language". That didn't inspire my confidence or look to me like they did their research and learned the lessons (and definitions) of other previous markup and non-markup languages, to avoid repeating old mistakes.
If YAML is defined by what it Ain't, instead of what it Is, then why is it so specifically obsessed with not being a Markup Language, when there are so many other more terrible kinds of languages it could focus on not being, like YATL Ain't Templating Language or YAPL Ain't Programming Language?
https://en.wikipedia.org/wiki/YAML#History_and_name
>YAML (/ˈjæməl/, rhymes with camel) was first proposed by Clark Evans in 2001, who designed it together with Ingy döt Net and Oren Ben-Kiki. Originally YAML was said to mean Yet Another Markup Language, referencing its purpose as a markup language with the yet another construct, but it was then repurposed as YAML Ain't Markup Language, a recursive acronym, to distinguish its purpose as data-oriented, rather than document markup.
https://en.wikipedia.org/wiki/Markup_language
>In computer text processing, a markup language is a system for annotating a document in a way that is syntactically distinguishable from the text. The idea and terminology evolved from the "marking up" of paper manuscripts (i.e., the revision instructions by editors), which is traditionally written with a red or blue pencil on authors' manuscripts. In digital media, this "blue pencil instruction text" was replaced by tags, which indicate what the parts of the document are, rather than details of how they might be shown on some display. This lets authors avoid formatting every instance of the same kind of thing redundantly (and possibly inconsistently). It also avoids the specification of fonts and dimensions which may not apply to many users (such as those with varying-size displays, impaired vision and screen-reading software).
So now no longer a YAML fan...
The big strike against YAML I see there is that it needs a good conformance test suite and implementations need to be tested against it. But that's not a problem with the format but a fairly easy to fix ecosystem problem.
But the syntax was valid, the parsers/validators would've been correct to accept it.
I have found myself using TOML more and more for configuration, though. It helps a lot with keeping things flat and easy to read. I'll still prefer YAML over JSON for human-writable files, but I'm starting to prefer TOML over YAML.
Boxes to check:
1. Self describing format
2. SAX-style parser available to C++
3. Easy for users to understand and ad-hoc parse using command-line tools
4. No document closing necessary, so appending is trivial
YAML looks pretty good:
- cmd: git checkout file.txt
when: 1565133286
pwd: /home/me/dir/
paths:
- file.txt
protobuf is also an option: entry {
cmd: "git checkout file.txt"
when: 1565133286
paths: "file.txt"
}
though I am unsure of how well its text serialization is supported.Any suggestions?
Ps. Thanks for (all the) fish, it's my daily driver shell and keeps me that much more sane c.f. the alternatives.
Having to close contexts is a VERY good 'sanity check' to see if something is malformed or not.
If appending is necessary make the parser handle multiple copies of the namespace and merge them upon output. Unknown keys and sections should also always be copied from input to output (this is how you embed comments).
To clarify the requirement, history could be a JSON array of objects:
[
{"cmd": "git checkout", "when": 1234 },
{"cmd": "vagrant up", "when": 4567 }
]
To append an entry to this file and keep it valid, one must locate the closing square bracket and overwrite it. That work is what I hope to avoid. {"cmd": "git checkout", "when": 1234}
{"cmd": "vagrant up", "when": 4567}
Not sure about other languages and libraries, but Go supports this out of the box[1]. And while we're at it, why not CSV? That can be processed with awk. #cmd,when
"git checkout","1234"
"vagrant up","4567"
[1]: https://play.golang.org/p/sTN9z4Kv3DBThe object stream idea is apparently supported widely and seems pretty strong. Thanks for the suggestion!
Have you considered SQLite? I know it’s not a friendly text format, but it alleviates a lot of the issues with the “append to a text file” approach, such as concurrency. It’s great for this sort of thing.
A file already exists with this content:
[{"cmd": "git checkout", "when": 1234 }]
Another tool wants to add a setting/element/etc, and simply creates a new config object with only the change in question included and APPENDS the existing config.
[{"cmd": "git checkout", "when": 1234 }][{"cmd": "vagrant up", "when": 4567},{"comment": "Comments, notes, etc are kept even if they don't validate to the recognized configuration options."}]
A configuration validator / etc loads the config and merges these in state-machine order, over-writing existing values with the latest ones from the end of the stream of objects and then determining if the result is a valid configuration. (Maybe file references fail to resolve / don't open / there's some combination of settings that's not supported...)
[
{"cmd": "git checkout", "when": 1234 },
{"cmd": "vagrant up", "when": 4567},
{"comment": "... actually kept as above"}
]Better than nothing I guess, but I'd say just use a syntax that supports comments.
Try just using line-delimited JSON objects (http://jsonlines.org/). It ticks all of your boxes, especially 3: "jq -s '.cmd' fish_history | histogram".
Neither YAML or Protobufs are quite as easy as that.
All in all it's ridiculously simple, easy to parse in a variety of languages and each row is a single line that's simple to iteratively parse without loading the whole thing into memory.
Here's a proposal: use a Tree Language.
I created a demo for you called "Fished": https://github.com/breck7/fished.
Took me just a few minutes but already get type check, autocomplete, syntax highlighting, and more.
Tree Notation is early, and there will be kinks until the community is bigger, but I think it may be useful for you.
http://treenotation.org/designer/#grammar%0A%20fishedNode%0A...
Stumbled into the idea. Basically just brute forced it. Tried thousands of things, built a huge database of languages, and tried to keep it simple.
Basically, it’s a narrative discussion of how this solution came to be, the trade offs involved, and perhaps its relationships with prior art.
[0] https://docs.pylonsproject.org/projects/pyramid/en/1.10-bran...
Googling "Design defen{c,s}e" just got me a whole lot of military contractors.
It fulfills all of the requirements. There are several available C++ TOML parsers, including one from Boost.
[1] https://github.com/toml-lang/toml/issues/356
[2] https://github.com/toml-lang/toml#user-content-array-of-tabl...
Personally I really like the array of tables syntax. It is a little unintuitive but it's not difficult to remember. It's useful for fulfilling the OP's "No document closing necessary, so appending is trivial" requirement.
And if you don't want to use it, you can always use inline tables in an array, just like JSON.
I also wonder if you need a text format, or if SQLite or systemd's journal API would work.
CI would be a lot nicer to use if it didn’t rely on a single YAML file to work. And if you want to switch, suddenly you had a build step to convert back to YAML.
Braces also allow for easily copying and pasting blocks of code because the braces delimit the semantics of the copied text. Because your code is already indented, with white space indentation you have to check that a) you pasted the first line at the right indentation level b) every subsequent line is also at the right level relative to the first line. No small feat.
Without braces copy-pasting becomes a context-sensitive headache. That's poor usability.
Quite the opposite.
I like to format my code nicely anyways (or rather, mostly my editor does it for me because I’ve asked it to do so).
I indent with two spaces usually, regardless of language. And have my editors configured to insert two spaces when I press tab.
JavaScript, Rust, Python, C. Same difference, in terms of how I use whitespace.
If you're writing well structured, original code in Python, it's generally cleaner and easier than other languages because the syntax avoids ambiguities that other languages have.
Language becomes popular largely through library ecosystem and resources around it, not just how the language looks. I think Google embracing it had a good role in acquiring mind shares.
I've never had someone that needed extensive help understanding YAML and that's besides reviewing work for people just coming up to speed. Find me an IDE or editor that doesn't have YAML support. Also, YAML supports comments so if you have pitfalls people need to know about you can document them inline.
Your argument is people who don't know things might screw stuff up. Well Yeah! This applies to everything.
You may be surprised to find that there’s significant disagreement on that point.
Will you though?
With learning a tool (almost any tool), its value increases with more use until you find a local maximum. The amount of effort to switch to something with a higher maximum at that point will definitely be considered by most people as part of the cost of that competitor. It is very hard to write off years of your life on anything.
When was the last time you did this?
It could take a few months at a minimum, just to see if you can understand it. Then, perhaps years of building things another way just to get good at it. Would you really give a few years of your life in making configuration suck less forever?
If you're serious, I have a potential solution: We don't have configuration files (or indention) or parsing problems, and users aren't the slightest bit confused on how to configure our applications (although they are often surprised if they've ever had to use a configuration file!). The downside is it's going to require a lot of re-learning on your part, and there's little I can do to make it any easier for you.
The last time I wrote off technology? In 30 years I've been doing this, I would say constantly, especially with tooling. I wouldn't consider the lessons learned using defunct tools to be lost either.
I'm not complaining about serialization formats, so why would I dedicate time to making it better?
Solution? I'm a quick learner so I'm not overly concerned. In a perfect world, you could just point me to your documentation.
Smalltalk environments don't typically have "configuration files" because it's so much easier to just directly manipulate the configuration parameter. Interactive development- the kind impossible with anything but the most specialised of IDE-- simply makes "config files" obsolete. Take a look at the seaside configuration guide[1] to get an idea what this is like.
In k/q[2] we also don't typically use configuration files, but for a different reason: Every type can be serialised over network or onto the disk, and when it's written to the disk it's usually in a format we can mmap() and access transparently. This is how q is also a database -- the data types q supports also includes tables.
k/q also has a built-in event loop that's not dissimilar to what you get when you run nodejs with the debugger port open except it's fast, and it's the regular way k/q processes communicate with each other.
What typically happens is that a table is designed for configuration, and we just expose it to the UI. Production environments are usually locked down so the only UI is to edit existing configuration parameters (and those are permissioned accordingly). These UI are typically quite general for any k/q data type, so they're quite rich and easy for people to use.
Then parts of the application interested in configuration just query the appropriate configuration table - this is only about 1000x faster than you would expect connecting to a remote database, and in many ways it's similar to a python application just storing its config in an sqlite database, except SQLite doesn't let you have a table as a data type so you can't put a table into another table, and you don't have tooling around comments and advice like you do with k/q UIs.
There are other places if you look carefully: Environments people have used (even beyond thirty years ago) that didn't have configuration "files", often had interesting and useful solutions to storing configuration. People tended to build configuration into part of their application, and so the legacy of that has tended to be excellent tooling instead of novel file formats.
[1]: http://www.shaffer-consulting.com/david/Seaside/Configuratio...
[2]: https://kx.com/
"From the point of view of a component, the configuration can be thought of as a Dictionary-like collection"
This is exactly what YAML and JSON provide too. Using a configuration file is a choice, so are parameters, and so are databases. I'm not really understanding what you're trying to get at?
The reason configuration file [formats] exist is because many programs are configurable and programmers are too lazy (or are not specified to, take your pick) to build a configuration tool [that has all their needs]. Configuration files are inferior in every way to an integrated and well-thought-out configuration process except that they may be easier to build and use in less ideal environments.
JSON is a fine format for interchange, and even persistence (i.e. to store configuration) but as a "configuration file" that people are expected to edit in their own way it is lacking, and that's why there are things like YAML and TOML and a million other things.
This is simply wrong. There is nothing in the spec stating that.
Could you elaborate on this? I use Ansible daily and I've never had a problem with YAML once I took some time to understand it. What do you mean by broken parsers? I'm assuming that's something Ansible specific you are referring to.
And ansible's parser is broken, in more ways that I can remember (haven't been writing playbooks and stuff for a couple of months now). If you like pointless pain, try embedding ":" in task names for a demo (or one of other several "meta" characters: the colon is just the one that ends to recur most).
I will give a passing mention to the smug, vague error message "You have an error at position (somewhere in the middle of the file) It seems that you are missing... (something) we may be wrong (they almost always are) but it appears it begins in position (some position close to the first line)" that sets off a hunt for the missing brace/colon/space/whatever and makes me want to do stuff to the person who devised it.
This compounds with the confusion brought on weaving of yaml's and jinjia2 syntaxes and ansible's own flakiness on deciding what is evaluated when - which decides when and if a variable does indeed change, when does yes means "yes" rather than true or 1, but not '1' or "true" (try prompting the user for a boolean variable, and find yourself writing if ( (switch == "true") or (switch == "1")) in short order).
Pity that ansible is so damn convenient, or I would have ditched it long time ago for anything - bash included (OK, maybe not bash).
Is it TOML as the author seems to prefer at the end?
(More recently, I've come around to prefer plain environment variables for configuration, but that only works nicely when the amount of configuration is fairly limited, say 20 values instead of 1000 values.)
For my own use, I do prefer TOML.
More here : https://hitchdev.com/strictyaml/why-not/toml/
I have some projects where I'm frequently writing and midifying content that resembles the example here, and I use YAML there and plan to keep using YAML. For most other things, I'm just doing configuration, so I use TOML. No reason you need to stick to one or the other.
It's overly verbose, and hard to understand XML, it's no comments son, horrors of yaml or some okay format that doesn't have parsers for the languages (plural) you are using on your project.
I've also used the Python port without issues.
TOML, in my opinion, is like a weird mishmash of JSON, ini, and bashisms. Though I have worked with it a lot less than the other formats, so YMMV.
The biggest problem with config formats is they mislead users into thinking they understand the format. The user tries to edit it by hand, and chaos ensues. So only formats that are stupidly simple, or whose warts are already familiar and well documented, are good choices.
Apache had a great configuration format. Nothing else used it (that I knew of) but you could in theory implement "Apache configs" and then people'd just have to look up how to write those, which there's lots of examples of.
JSON and YAML and XML are data formats; they should only be written by machines, and read by humans. Same with protocols like HTTP, Telnet, FTP... You're not supposed to write it yourself, but it's readable to make troubleshooting easier.
Data formats are nice for expressing nested data structures, but then they don't (usually) support logical expressions; at that point you need a template/macro/programming language, and at that point you're writing code, which will need to be tested, and at that point you should just write modules and use a config format to give them arguments. Every complex tool goes through the same evolution.
If you care about your users, write a tool to generate configs based on a wizard. Good CLI tools do this, and it really makes life better. (It's also a great way to document all your config features in code, and test them)
(No, not YAML.)
Aside from lack of comments, the other major thing that can sometime make json a bad config choice is lack of multi-line strings.
You might end up with valid YAML, but you won't know until the YAML consumer barfs.
BTW, all of a sudden XML with DTDs are looking sane again :)
- what draft of JSON Schema are you using 4? 7? Neither?
- what version of Swagger or OpenAPI are you using?
- etc.
Sure, it's great to see ongoing development of schemas, but with each new development we have yet another dialect to consider/support.
In my view, perhaps an even greater problem with structured data formats in general is the void that separates them from programming languages esp. static languages such as Java where static type information is otherwise leveraged. The industry standard solution, code generation, is awful in almost every respect. The Manifold framework looks promising in this regard (http://manifold.systems/).
Sadly, this is utterly wrong.
https://tools.ietf.org/html/rfc8259 https://tools.ietf.org/html/rfc7159 https://tools.ietf.org/html/rfc7158 https://tools.ietf.org/html/rfc4627 http://www.ecma-international.org/publications/files/ECMA-ST... https://www.iso.org/standard/71616.html
XPath is a struggle with namespaces, as well. It’s ... trying.
<author name="pete" />
or <author>
<name>pete</name>
</author>
yet?I disagree. Back in the day we used attributes for everything that was key value and inner tags for anything with structure. We also formatted for clarity:
<lunch
env="outside"
food="sandwiches"
drink="cola"
/>
Compared with what we used to do, I look at attribute-less maven pom.xml with horror. <lunch
env="outside"
envfr="dehors"
food="sandwiches"
foodfr="sandwichs"
drink="cola"
/>
Or <lunch
env="outside"
food="sandwiches"
drink="cola">
<label type="env" language="fr">Dehors</label>
<label type="env" language="de">Außenseite</label>
<label type="env" language="en">Outside</label>
</lunch>
Quite curious about it. <lunch
env="outside"
food="sandwiches"
drink="cola">
<label xml:lang="fr">Dehors</label>
<label xml:lang="de">Draußen</label>
<label xml:lang="en">Outside</label>
</lunch>
says[0]the w3. <user name="alice" group="wheel" />
Oh no: <user name="alice" groups="wheel admin sudoers" />
Or is it one of: <user name="alice" groups="wheel,admin,sudoers" />
<user name="alice" groups="wheel:admin:sudoers" />
<user name="alice" groups="wheel;admin;sudoers" />
Or give up and use elements: <user name="alice">
<group>wheel</group>
<group>admin</group>
<group>sudoers</group>
</user>
Hold on, should there be a container? <user name="alice">
<groups>
<group>wheel</group>
<group>admin</group>
<group>sudoers</group>
</groups>
</user>
XML really needs either richer attributes, or no attributes.It's your job to decide how the data should be structured, in any language.
I would do the following :
<user name="alice">
<groups>
<group uid="wheel"/>
<group uid="admin"/>
<group uid="sudoers"/>
</groups>
</user>When you add verbosity on top of that, working with XML over the long haul is an utter pain.
But in other languages with notoriously irresponsible coders (JS, PHP) I bet to see even more of these problems.
(I coded in all of them)
During my working life using Python I got few meh-moments with YAML. And this is all. Never lost real joy of using it.
Format aside, no matter which you choose, you have to pick library interfaces that don't deserialize to arbitrary, constructed objects.
you can quickly tell if an xml document is malformed (good parsers will tipically point you to the un-closed tag).
Yaml on the other hand would probably load anyway, with the application receiving garbage data, potentially misdirecting the application behavior...
https://yamllint.readthedocs.io/en/stable/rules.html#module-...
JSON was a reaction to the verbosity of XML, but a better reaction would have been to work harder on our text editors so that working with XML would be just as easy as working with JSON in terms of the numbers of keystrokes needed. Better parser interfaces that help you treat the dataformat more like it's part of the language would also help (i.e. SAX and DOMDocument suck to work with, but SimpleXML is almost idiomatic).
Isn't that only solving half the problem? XML is also pretty difficult to read
But why is XML so freaking great? We can’t even tell if whitespace is significant or not. If a schema says it’s insignificant then that’s that!
https://www.oracle.com/technetwork/articles/wang-whitespace-... That alone is TERRIBLE! (Same problem with YML.) Why should I bother with that? JSON can encode strings, hashes, arrays etc. in a way that’s instantly interoperable with JS and is far far more unambiguous.
What exactly is so great about XML that you can’t do with JSON in a better way? Schemas can be stored in JSON. XPATH can specified for JSON. Seriously I never got the appeal of XML except that it was first.
Oh and there's no comments.
[0] https://realprogrammer.wordpress.com/2012/08/17/json-graph-s...
For a lot of data, XML isn't the right form and buries too much data in hierarchy and tag soups - but it's flexible enough to make it into whatever you want, and since XML was buzzworded and XML libs were some of the easiest things to reach for in the 90's, it got pushed into every role imaginable.
If I have data that I need to send somewhere, and I can create the format for it, that's really easy to do.
The problem, every time, is the reverse; receiving some piece of data and trying to figure out what parts of it I care about. Both XML and JSON allow for schema definitions, but in both cases it fundamentally requires me, as a consumer "grokking" what is being sent. And the verbosity of XML simply makes that harder. Working with either is not _that_ hard (though I have run into XML in the wild that is so large a payload, yet so poorly designed, that there is no good way to process it; I can stream with via SAX without writing my own state handling mechanism, and I can't just deserialize it into an object without massive memory issues at scale); the difficulty really is in containing it in my mind, and JSON simply facilitates that better due to it's simplicity and explicitness (yes, explicitness; in XML it's not clear if a child element should only exist once, or multiple. JSON it's obvious)
Per the OP; I cringe every time I see YAML. Pain to write, pain to read; have to have tooling every time or I get whitespace issues.
I’d say this is schema-dependent. If you’re talking about plist files, sure; those are ugly and unintuitive. But on the whole I find XML far easier to read than JSON. With closing tags, what you lose in terseness is made up for with scannability: it’s easier to understand the document hierarchy at a glance, and find your place again after editing. Whereas with JSON I often have to match curly braces in my head, or add comments like `// /foo` which isn’t even possible outside of JS proper or a lax-parser environment.
Attributes are just strings, generally for metadata. I'd probably serialize an object from another language more verbosely. This is where an important distinction needs to be made: the XML format you use for config files or for data exchange from your app to others should not necessarily be just a serialized object from the most convenient form inside your application. If you care about the operators of the app, you'll allow them a more concise format for that kind of thing, and use XSD or your own internal mechanisms to turn that into an object you want to actually work with.
It's the sort of problem people have once, abstract away, and move on from.
I'm so frustrated with our collective attitude of building abstractions upon abstractions upon abstractions, without ever stepping back and realizing that we're using the wrong tool to begin with.
JSON was a reaction to the simplicity of having a data format that JS, ubiquitous on the web, could load via eval (which, once JSON was established, was largely abandoned because it is ludicrously unsafe, but the momentum was already there.)
(Mostly /s. Come at me:))
My one request would be to bring back to SGML-like closing tag abbreviation:
That is, instead of
<foo><bar>qux</bar></foo>
we should be able to write <foo><bar>qux</></>
I think this one change would make XML more "palatable" for the JSON/YAML/TOML crowd. <foo/<bar/qux>>
(Although standard SGML inexplicably specifies a null-end-tag character of '/' rather than '>', so this won't work in stock parsers.)Recently I was writing a custom static site generator for my website. I started with python and yaml using the pyyaml lib. After two months (don't laugh, I wasn't writing this generator all this time; i had a break) I tested if everything I wrote previously was warking. Pyyaml came at me screaming that they deprecated something and I shouldn't use it, otherwise the feds will get me.
Let's go several months earlier still. I was learning Python, using Debian Stretch which has Python 3.5 installed. The book I was learning from used Python 3.7. When I got to a point it became clear that 3.5 lacks a few things which are needed to continue learning Python according to the book. So I compiled Python 3.7, set up virtualenv with it and... The code I had previously written stopped working. With a version bump from 3.5 -> 3.7? That's a minor version change. And now 3.5 was expecting at one point to have a string path passed as an argument, and 3.7 was expecting a pathlib path. That was trivial to fix in case of a small example, but I would dread using such a thing for anything big and then having to debug what exactly broke between different (minor!) versions of Python or its libraries.
These new hip tools seem to have a backwards-compatibility issue.
I eventually settled on using XSLT and a couple really short shell scripts (which all fit on my screen at the same time) and I don't expect them to break in the next two decades.
However, XML is still a pain and I would prefer just using S-expressions and Lisp[1]. It's just that for now my only experience with them is writing things for Emacs and I would like to learn Scheme/CommonLisp to do anything outside of Emacs with Lisp.
[1] https://sites.google.com/site/steveyegge2/the-emacs-problem
> I'd prefer XML because of the stability of the tools available for it.
Python doesn't do semantic versioning. (Just like most language runtimes) You can find deprecations and backwards incompatible changes in the release notes.
Although those breaking changes are not frequent and if https://docs.python.org/3.6/whatsnew/3.6.html#whatsnew36-pep... resulted in broken code, you likely ran into something that's reportable as a bug.
1. Built-time validation (or at least type checking).
2. Built configs can be easy to parse but (potentially) rich enough to avoid confusing templating.
3. Separation of concerns between storing/maintaining configs and applying them. E.g. scoop text configs off a source repo, but send out binary configs over the network.
All this is a fantasy in my head. Right now the closest mainstream thing is protobufs. But they make trade-offs for non-config use cases, and thus don't really cut it in the "... rich enough to avoid confusing templating" department.
Someone mentioned protobufs, maybe with something like those one could have both?
You can use sqldiff[1]. Try adding it to your .gitattributes[2]. If you need TRIGGERs and VIEWs, consider dumping your database[3] instead.
[1]: https://www.sqlite.org/sqldiff.html
[2]: https://git-scm.com/docs/gitattributes
[3]: https://gist.github.com/peteristhegreat/a028bc3b588baaea09ff...
Out of curiosity, is there anyone here who doesn't like TOML for configuration?
BTW. I've handled all of those formats using jackson on Java & Kotlin. It has a flexible parser framework originally intended for json. But it has lots of plugins for different tree like configuration files. Look for jackson-dataformat-yaml and jackson-dataformat-toml on github. There are loads more formats that you can support with jackson. Nice if you need to translate from one to the other or need to support multiple formats.
IMHO Json with some tweaks would be really nice. E.g. just supporting comments and multi line strings would make it a lot nicer. A lot of json becomes unreadable due to the need to escape strings. I've come across Hocon a couple of times (jackson-dataformat-hocon) and it's a strict superset of json, which means that if you accept hocon as input, you implicitly also accept json.
Not to talk about the attribute/content duality and all the ill-defined parsers it leads to.
<t>Some <b>text</b> is here </t>
And I found some things weird, like entities and (external?) DTD references.But it's great for building your own formats. For data interchange it has it's problems, like no types. Everything is kind of a string. With encoding problems and XML in XML everything ends up in CDATA...
I think that's why XML makes for heavy parsers. I remember it was better in Java, because of the strong libraries.
I recall the first time I saw YAML and all I could think to myself was that I have to learn yet another syntax. I find it far less readable than JSON or XML and made me pine for the latter.
Templating YAML is even worse. Templating is an ad-hoc abstraction, and very easy to run into issues. A minimal JavaScript runtime with JSON would be much better, JSON is JavaScript Object Notation after all.
And what is assembler, may I ask? Is it for YAML or Markdown?
* YAML can have several 'documents' in the same file,separated by ---
* there are anchors and references
* easy to read multi line texts
* it's also a superset of JSON
I can see, how choosing YAML when you just wanted readable JSON might give you more headaches than expected.And like someone else said, putting another template engine (or two) on top of YAML is when the real problems start.
https://arp242.net/yaml-config.html#can-be-hard-to-edit-espe...
In the heading, how was sp ligatures in the `espcially` written, is there a name for this? How do you connect the beginning of a `s` to the beginning of a `p`?
They're turned on in html with the following two css lines (though on firefox, either one is enough to have them happen):
font-variant-ligatures: common-ligatures discretionary-ligatures;
font-feature-settings: 'liga' on, 'dlig' on;
[0]: https://www.fonts.com/content/learning/fontology/level-3/sig...2. Format is declared to have x and y problem, new format is invented that is "simpler and better"
Time passes
3. People slowly discover format in #2 has the same issues that led to creating #1.
Repeat.
(Just like attempting to super-generalize anything else)
But here is a big question.
Imagine you can influence the switch of the configuration formats for projects like Ansible, Kubernetes, Docker Compose, AWS CloudFormation, Google Cloud Deployment Manager, et al.
You can take any project with huge user base and all of those project will have one thing in common: JSON-based configuration with an option to write this configuration in YAML.
Since basically anyone talking YAML in the context of JSON is talking about a JSON superset.
So here's the task: propose a JSON-compatible alternative to YAML.
Things to keep in mind:
- backwards compatibility
- easy migration from YAML to a new format
- full JSON compatibility
- relatively cheap to get supported by the project of interest.
Function application was disabled some time ago, and `yaml.load()` logs a noisy deprecation warning telling users to use `yaml.safe_load()` instead [1].
[1]: https://github.com/yaml/pyyaml/wiki/PyYAML-yaml.load(input)-...
I like the _idea_ of yaml, but: - it's overly complicated in the wrong ways - common/simple use cases aren't supported and require post processing (i.e. Merging block maps/arrays, string interpolation, etc)
It's easier to write than JSON (no need to quote keys, allows trailing commas), has reusability through functions and objects, and can output JSON which is much easier to parse than YAML.
Downside: you need another build step for the config.
I wrote a lovely library for parsing them in Go (along with a full lisp interpreter if you like) https://github.com/glycerine/zygomys
Provides comments, multiline strings, and automatic translation into Go structs using reflection.
However at the time, a tree data format that deserialized into native types was quite useful. The alternative was writing event based SAX parsers, or incredibly verbose XML object apis.
For instance, the recent iMessage bugs that project zero announced were because NSCodable serialization tells the deserializer what class should instantiated. Followed by remote code execution (woo!)
Similar problems have occurred with java serialization over the years, the python serialisation thing (that silly name I can’t recall).
I was recently learning swift and was getting frustrated by the verbosity/work for deserialisaing abstract classes when I realized the clunkiness was due to a design that made the deserialise attacker specified objects basically impossible. Obviously you could engineer a solution that would be exploitable but there’s only so much a platform can do to stop developer mistakes.
The best configuration file is simply source code that initializes whatever you want to run, and then runs it. That way, you can install hooks in the form of closures and make the program behave exactly like you want without the constraints that a simple "value-only" configuration file format has.
Relative to other languages, it is very easy to embed Lua and it is also easy to cut out Lua’s entire standard library to restrict what the scripts are capable of doing within your program.
Yaml despite its flaws works pretty well for Ansible playbooks and for storing localizations.
I think the syntax is actually what makes it more human readable, it's still 95% text/numbers just annotated with information that makes it clear what things actually are instead of hiding them behind confusing computer parsing rules nobody is going to think about while reading human-friendly text.
How does one write Kubernetes specs and Ansible without yaml?
Get to use parsers out of the box, validation tooling, support comments, IDE code completion and they are super easy to transform.
In a couple of years some trendy SV unicorn will make XML the best format of the world, as these cycles happen to be.
What's nice about YAML is the choice to use indentation, for the rest, I have a hard time following the language's choices.
Haven't used TOML yet, but it seems promising given that for most use cases you would only use a portion of the YAML language.
Instead just declare all config variables within code itself in a separate config class/module file, along with initialization to default values and provides dynamic getter/setter interface over a debug API (which can be enabled/disabled via a command line flag).
If you want, you can also provide a friendly cli tool to interact with the debug api. This tool could output help messages, show current config values - differentiate between default vs overwritten etc.
Of course, this can be written once as a utility library and cli and used consistently across all your programs.
1) it presents an obstacle to sharing config files in a multi-language environment.
2) writing user prefs in the host app’s native language is fraught with security issues
The first concern is admittedly niche, and both concerns are addressable with some care and thought, but they’re good to consider.
And we load the config by curling a JSON payload :) ?
Config files are fantastic. Trivial to read, write, copy, track in version control, diff, grep, generate with scripts, etc.
API-driven configuration has none of these properties.
Some Java application servers take this approach of API-driven configuration. It's an improvement over UI-driven configuration, which is what they had before. But it's still significantly worse than simple file-driven configuration.
If you want to provide a 'friendly' CLI tool, by all means do so - but provide a tool to interpret and generate config files, not something which replaces config files.
Code based configuration offers most of those benefits, just look at some of the suckless tools, if you've never used dwm or written a line of C I bet you can still guess how to configure some stuff here: https://git.suckless.org/dwm/file/config.def.h.html .
It doesn't work for all software of course, but for anything developer centric and/or anything with complex configuration done by experts it's a good choice, particularly for any software going down the path of creating a custom language (tmux, vim, etc). You get the power and flexibility of a full programming language, you get compile time config checks, you get excellent performance, you don't have to learn yet another config language and it's easy to write patches.
It's not great if you're targeting Joe Average, but neither is any configuration format, or any configuration at all.
If you need runtime reconfiguration in production, then it requires a runtime config management system tailored for operations folks with proper authn/authz, audit logs etc. This is a product by itself. The connection between your running program and the runtime configurator has to be intentional, secure etc.
IMO, config variables should have following bindings: 1. config variable in code. 2. program start env variable. 3. program start command argument flag. 4. runtime configuration.
All available configuration options are declared in the code. But not all configuration options should be accessible from #2 to #4. And the override preference order may not be same for all types of configuration variables.
Btw, this isn't related to Java or any particular programming language. I've seen this done in large C++ projects 20 years ago.
It just makes things run smooth in TypeScript because all the members and types can be statically analyzed giving you errors, auto completion and ability to jump to the definition when config files cannot do that but probably not good for large projects with multiple languages.
For small projects I'm just taking the benefit over small concerns.
This refrain just cheeses me right off every time. Nothing is human readable! Everything requires a program to read it, because no human being can read states of charge or states of magnetic polarization directly.
What makes something 'human readable' or not is a software tool. Underlying that tool is a data format that the tool can accept and display. What everyone means when they try to sound smart by saying 'human readable' is just 'plain text.' In other words, they know where to find the dumbest possible reader/editor for it. Text editors are the dumbest possible editors because they cannot constrain edits to conform with the grammar of the interface language; they allow bugs at a point in the development process where it is trivial to disallow bugs, especially considering that interface languages should probably be, at most, regular languages.
I'll step off my soapbox, now.
Personally, I think most editors are rather specialized. They deduce the character sets, often add syntax highlighting and provides paging for very large files. High level strongly typed languages, such as modern C#, Java, Scala are designed to be used in an IDE. You could view and edit it, but it provides a difficult situation (not unlike editing XML by hand).
"Human readable" is very subjective. It depends on the person, the task at hand, the intended recipients, etc. etc.
Every veteran in the field knows that data and data structure are the primary enablers for almost any solution. Nevertheless, we have regressed in the last decades w.r.t. that. Had we had better structured (AST-driven) editors and context dependent representation, XML (or sth similar) might have taken off.
Personally, I blame the exponential increase in inflow of new developers. Nothing but reinventing the wheel without knowing history.
- Meaningful version control
- dump it to SQL and version control that, if you must use Git
- Not everyone likes the extra step.
- not everyone likes the hoops you need to jump through with the alternatives either! That’s why we’re discussing this :-)
Two things: comments and multi-line strings.
I still prefer it over YAMLs awkward initial learning curve.
A simplified subset, .syml would be a good idea.
The implementation is in Python.
I've been doing a lot of research and development on language design for file-based content (e.g. for static site generators). I've found that YAML - although established as the go-to format for statically generated blogs, etc. - was never designed for these things as it by its nature does not support simple, essential features for this usecase like for instance unindented blocks of verbatim text (for which YAML frontmatter was invented as a very limited hack).
The result of all this R&D is a language called "eno notation" which is designed especially for file-based content usecases, and around which I've also built an entire ecosystem for many languages and editors - if you're working in that field, it might be worth taking a look!
iterations: 100000
evaluated: Fri Jul 06 2018 09:46:48 GMT+0200 (Central European Summer Time)
are tagged below as just a “Field”. Do client programs that read an Eno file need to run `int()`/`float()` or `.to_i`/`.to_f` on the field values they know should be numbers? That seems unergonomic.The idea is to have 2 levels: a simple, minimal syntax/notation (think binary) called Tree Notation, and then have higher level grammars on top of that, called tree languages.
It works for encoding data and also for programming languages, regardless of paradigm.
Here’s a BNf:
https://github.com/treenotation/jtree/issues/1
There’s a FAQ as well.
Docs needs work, in particular I’m hoping people will create their own explanations of the ideas in external places, as that might be a better way to understand it. Happy to provide help to anyone that is interested in that.
https://news.ycombinator.com/newsguidelines.html
Edit: we've had to ask you multiple times already not to be a jerk on HN. Would you please review the guidelines and take the spirit of this site more to heart?
Your impression is correct though, without highlighting it’s awful. Try the Tree Language designer app to see it with highlighting, type checking, autocomplete etc...there are interesting ways to accomplish everything without syntax, often not obvious. But tooling is essential to make it better than existing options. Help wanted!