Intercal, YAML, and Other Horrible Programming Languages
blog.earthly.dev
blog.earthly.dev
Pride in creating a "generic" system that can be configured to do all kinds of new things "without touching the code." Reality check: only one or two programmers understand how to modify the config file, and changes have to go through the same life cycle as a code change, so you haven't gained anything. You've only made it harder to onboard new programmers to the project.
Hope that if certain logic is encoded in config files, then it can never get complicated. Reality check: product requirements do not magically become simpler because of your implementation decisions. The config file will become as expressive as necessary to fulfill the requirements, and the code to translate the config file into runtime behavior will become much more complex than if you had coded the logic directly.
Hope that you can get non-programmers to code review your business logic. Reality check: the DSL you embedded in your config file isn't as "human readable" as you think it is. Also, they're not going to sign up for a Github account and learn how to review a PR so they can do your job for you.
Marketing your product as a "no code" solution. Reality check: none for you; this is great! Your customers, on the other hand, are going to find out that "no code" means "coding in something that was never meant to be a programming language."
If you write your config in a "full-blown" programming language, then your configs are full-blown programs in that programming language. This situation just plain sucks, or at least it comes very close to (or passes) the "suck" threshold every time I experience it. Code-as-configuration demands tremendous discipline from the team.
Whereas if you abuse YAML (or JSON or XML or whatever) to create a limited and hard/impossible-to-extend DSL, you still have much more control over what can and cannot be executed by the config engine, even if the DSL happens to accidentally become Turing complete. You can embed limited shell commands in the DSL as an escape hatch, but make it difficult enough that you really have to try to make a mess.
Another example of this DSL model done mostly-right is Make.
Once you accept that idea, whether to use JSON vs YAML vs TOML vs XML vs S-expressions is just bikeshedding over syntax.
As for "why YAML in 2021" specifically? Yes, YAML is a big spec and there are lot of ways to get strings wrong. But maybe you don't care or your team is unlikely to ever go near the darker corners of the spec. For simple config files, YAML is just really easy to read and write. And if you do need multi-line strings, it's a whole lot easier than doing it in JSON.
I'm personally a big fan of TOML, but maybe YAML is still better for highly-nested data.
Of course S-expressions are wonderful for many reasons, but they share the problem with JSON of being somewhat hard to diff and edit without support from tooling.
If you want support for trees xpath has been there for 25 years now.
The article actually mentions Dhall as a solution. This engineering problem has been resolved.
Yeah, I am really excited about Dhall. I think this is the future, it supports the types of abstractions that we need without the mess of full templating or full turing completeness.
The one downside to Dhall is you really want to have an implementation for it in each common language. You can use it to generate YAML, but I think it would be better if tools understood Dhall and that is a bigger ask because it is a more complicated implementation.
Let's build Dhall implementations for every major language, convince Gabe to format things in a way that makes it look more familiar to non-haskell people and consider this problem solved.
That sounds like fun, is there any effort underway for this? I'd be down to contribute.
I disagree with your point that it should supported by each language however, I think it's much better to use something simple like JSON as a "compilation" target since it's easy for machines to read and lets users pick the configuration backend.
Use a smart language like Dhall or bazel for managing configuration and use a mundane format like JSON for the machine, let the Dhall binary bridge the gap.
That is also my experience. TOML is really cool for simple key/value stores, but keeping everything linear in config makes nesting error-prone, eg with [[table.subtable.list]] to append an item something to table.subtable.list. It's really easy to miss a nesting level by accident.
Also related, newcomers in TOMLland find it really confusing that appending a single line to the configuration file will append it to the latest defined table, not as a top-level key.
This is probably not a common piece of YAML knowledge, but it's arguably better to use "folded" style for a password:
password: >-2\n
asdjoi'";j;oj;90\[2301@
Or use single quotes, which signals to the YAML parser not to treat any characters as special, but then you need to escape the literal SINGLE QUOTE character (') by doubling it: password: 'asdjoi''";j;oj;90\[2301@'
This is completely valid yaml that reduces to the JSON equivalent: {"password": "asdjoi'\";j;oj;90\\[2301@"}
And of course you can always write JSON syntax for when the text escaping gets hairy, because YAML is a superset of JSON.Here they are together:
password1: >-2
asdjoi'";j;oj;90\[2301@
password2: 'asdjoi''";j;oj;90\[2301@'
password3: "asdjoi'\";j;oj;90\\[2301@"
See here (https://yaml-multiline.info/) for a summary of the various multi-line text styles.Turing completeness is a red-herring when it comes to config languages IMHO. Purity is much more important consideration, e.g. to ensure it can't delete files, or vary its output based on random network calls.
So my preference is to have a simple language with explicit limit on number of operations and the amount of memory its interpreter can access before aborting rather than a complex config without any explicit limits on complexity leading to exploits with stack or memory overflow in pathological cases.
So things like DHall indeed.
This is not true. Turing complete languages are so because their halting problem is undecidable, it is irrelevant that the computer you run a python program has finite memory. Check out languages like Agda where you can do general purpose computing but are not Turing complete, since all programs can be proved to halt.
And conversely, a non-Turing-complete language can still allow the provably-halting program `for i in 2^256 { /*busy-wait*/ }`.
> [I]n many ways, XML and XSLT are better than an ad-hoc YAML based scripting language. XSLT is a documented and standardized thing, not just some ad-hoc format for specifying execution.
Standardization and reliable documentation really is an important risk mitigation compared to a “widespread” convention in YAML that might disappear (and even become confusing to new developers) if some new YAML-based API becomes more popular. In many cases this stability will not be worth the annoyances of XML, but it’s not a trivial concern.
Yes, which means we have well-documented functionality and tooling of that language to deal with various use cases. Which is not going to be the case with your ad-hoc format based on YAML or JSON.
>Code-as-configuration demands tremendous discipline from the team.
No more discipline than any other form of programming.
>Whereas if you abuse YAML (or JSON or XML or whatever) to create a limited and hard/impossible-to-extend DSL, you still have much more control
And here is the crux of the issue. Tools that are designed so that someone can keep "more control" rather than for tool users to solve real problems. The industry is sliding back towards bad old days of batch processing because of conceit and lack of lateral thinking.
Then if you need to, you can create complex or overly long configuration files in Python by inserting keys into a dictionary and dumping to ConfigParser (or however your favourite language does things). For example, its useful when writing a test for many permutations of something similar.
Meanwhile the parsing side is simple enough to be re-implemented in an hour when the time comes to rewrite your whole stack in C+Verilog for real ultimate performance.
The 2 main things are:
1) Using your own bespoke config format or some pet format that's not widely supported adds needless friction to writing little duct tape scripts, testing harnesses, and misc tools. It also adds unnecessary difficulty when porting parts of your program to new languages.
2) Using a Turing complete config format even if it's not bespoke makes all the drawbacks in (1) even more apparent.
Really? Unless it's written in Haskell or something else with a very strong type system, you won't do better than JSONSchema for validating the config file.
And here is the crux of the issue. Tools that are designed so that someone can keep "more control" rather than for tool users to solve real problems. The industry is sliding back towards bad old days of batch processing because of conceit and lack of lateral thinking.
Too much freedom is a bad thing. The industry is not "sliding" anywhere. We tried code-as-configuration, it required too much discipline, so the pendulum is swinging back. As pointed out elsewhere, hopefully Dhall will save us from all this by being the happy balance between expressive and chaos-limiting.
It has to be said there are a lot of things that are almost fully-scriptable, for example the "mutt" mail-client. It has a configuration language, but it isn't real in the sense that you can't define functions, use loops, etc. I eventually wrote my own mail-client so I could do complicated things with a real configuration language (lua in my case).
Seeing scripting languages grow up in an adhoc fashion often leaves you in the worst of all worlds. Once upon a time I decided I wanted to script the generation of GNU screen configuration files for example. I made a trivial patch:
* If the .screenrc file is non-executable - read/parse.
* Otherwise execute it, and parse the result.
Been a few years now, but I think the end result was that I wrote a configuration-generator in Perl that did the necessary things. (Of course this was before I submitted the "unbindall" primitive upstream, which was one small change that made custom use of screen more safer - using it as a login shell, for customers who shouldn't be able to run arbitrary things.)
I think you just internalized the pain of make. I used to be good at it, didn't program c for 20 years and came back to it for a few projects and wanted to tear my hair out.
The pls I'm currently working with have declarative build dsls in the same language (mix.exs for elixir and build.zig for zig) and this is fantastic.
So it should be for configs. Use a truly turing complete language if you need control flow. I think hashicorp got this right but by then everyone hated to have to learn ruby.
I think this is the real reason why yaml configs got popular. If you had a dsl in x language, programmers would get defensive that it was in blub and not their pl of choice. Yaml was a way of being a language agnostic neutral ground.
And now we have n+1 blubs.
Python ecosystem definitely suffers from this. I tried to do some machine learning experiments and basically all of the repos I wanted to use were on 2.x and after 30 minutes of faffing around I gave up and moved onto other packages. However, the biggest pain points for Python came in the 2-3 transition (and TensorFlow x->y in general). By then Python had too much momentum and popularity (and every undergrad learns python). TensorFlow, well at least there is a competitor (torch) and so we see that TF's popularity has basically been sucked dry, and I have no doubt that a large portion of it is just how awful Google+Nvidia have been in managing the TF releases.
1. arbitrary I/O -- can I read a file from disk, open a socket, make a DNS query, etc.
2. arbitrary computation -- can I do arithmetic, can I capitalize strings, can I write a (pure) Lisp interpreter, etc.
I claim that the first IS a problem but the second ISN'T.
Arbitrary I/O is a problem because it means the configuration isn't reproducible / deterministic, so it's not debuggable. Your deployed system could be in a state that depends on the developer's laptop, and then nobody else can debug it.
The second is NOT a problem. As long as the state of the deployed system is a FUNCTION of what you have versioned/configured, then it's no problem. Functions are useful. Pure functions can also be expressed in an imperative style (another design issue that's commonly confused).
Related thread: https://lobste.rs/s/gcfdnn/why_dhall_advertises_absence_turi...
and my other comment in this thread: https://news.ycombinator.com/item?id=26277812
See Starlark, which is a subset of Python used by the Bazel build system - https://github.com/bazelbuild/starlark
You know what's really interesting about YAML? !Tags (https://yaml.org/spec/1.2/spec.html#id2805019), when used as type hints.
In YAML, you can write: (forgive the contrived example)
!Person
name: Joe
login: joe
In Python, you can use "yaml.add_constructor" to automatically dispatch specific Python code when !Person tags are encountered in the YAML. You can transform text into an object tree (edit: graph, actually -- don't forget about anchors and aliases) automatically.In the right places, this is extremely handy. It's a middle ground between the "duck typing" you normally get with un-validated markup, and the quasi-"static typing" you get with the same inputs after a schema validation. There's an analogy with type annotations in Python.
As always, there are right and wrong places to apply this kind of thing.
Tagged YAML allowed us to replace an older system that had ten times as much code. Now, there are far fewer, much warmer code paths, and it's easier to use and far less scary to maintain as a result.
Other markups are not the only alternative to this kind of thing. I expect many developers would reach for a database (SQLite?) as a storage medium. However, keeping our data accessible in plain text has been a huge benefit.
example my-app.yaml
#!/usr/local/bin/kubectl -f
---
apiVersion: v1
kind: Namespace
metadata:
name: my-namespace
---
apiVersion: v1
kind: Pod
metadata:
name: my-pod
namespace: my-namespace
spec:
containers:
- name: nginx
image: docker.io/nginx:latest
ports:
- containerPort: 80
after making my-app.yaml executable you can now invoke the application's config to reconcile itself > chmod +x ./my-app.yaml
> ./my-app.yaml apply
namespace/my-namespace created
pod/my-pod created
> ./my-app.yaml get
NAME STATUS AGE
namespace/my-namespace Active 8s
NAME READY STATUS RESTARTS AGE
pod/my-pod 1/1 Running 0 7s
> ./my-app.yaml delete
namespace "my-namespace" deleted
pod "my-pod" deleted
obviously this is just a simple example but there is come really cool potential here.The problem is when you use YAML (or an ad hoc interpreter embedded in YAML, or a templating system built on top of YAML) for things that should really be a programming language. Things like imports and functions and composition are useful. Templating is a more ergonomic form of string concatenation.
However, if the config is a travisCI YAML file that is really just a glorified list of commands I don't think bash or a makefile is a bad solution. Take as much as possible out of the CI config and use a neutral standard format like a makefile to encapsulate much of your logic.
Then when TravisCI stops offering free usage for open source, its easy to move, because your format isn't specific to a vendor.
The visionary executive team at a certain company I worked at felt otherwise and poured millions and thousands of eng hours into recreating HTML/JavaScript MVC components, but with YAML
Why more languages don't adopt Tcl's concept of a safe/restricted interpreter for this exact use case is beyond me.
That is to say the configuration eventually becomes a program in itself, with a few key values... which then get pulled out into a simple config file.
See: autotools, sendmail, etc.
For example package.json in NPM packages -> that feels like a good fit for JSON (although it would be even better if JSON had comments). On the other hand, terraform, or build languages like Make or Meson -> they are complex enough that it probably makes sense to have a standalone DSL.
I was facing the same decision recently on my project while designing a declarative DSL for web app development (kind of like web framework). From simplest to most complex option:
- should I just let them define it all in JSON? There would be a lot of repetition at some point and it would become impractical, but it could be ok for the start.
- should I just implement JS library, that devs can use in JS to construct a config object that is then exported to JSON? That would be embedded DSL. Sounds flexible and easy to do, but it is also overly expressive and not "cool" (ok this is debatable).
- should I use something like Dhall? It is declarative and simple.
- should I come up with my own declarative, configuration-like DSL? It would probably end up similar to Dhall, but this means I can do whatever I want - I can make it as ergonomic and custom as I want to (which I guess is both good and bad :D!). It might also allow for nicer interop with Javascript and other languages.
In this case, we went for the last option, mostly because we felt the most important thing is ergonomics and interop, but well, I am still curious how would other directions play out. Plus at the end we didn't yet get to the point where language is more expressive than JSON (code example: https://github.com/wasp-lang/wasp/blob/master/examples/tutor...).
Maybe I am just missing a better design process, but it seems to me that with a language idea it is hard to say if it is good or not until you try using it.
What is funny is that a shell script does a pretty good job at giving you a nice, programmable way to invoke software. But, that is too complicated :-)
In fact I mention the spaces issue in the The Simplest Explanation of Oil:
https://www.oilshell.org/blog/2020/01/simplest-explanation.h...
More: http://www.oilshell.org/blog/2021/01/why-a-new-shell.html
You can use it right now as a dev tool. If you use ShellCheck to statically check it, then running your script under Oil is complementary (and it also has some static checks): https://www.oilshell.org/why.html
The cloud is basically built on top of Linux distros, so that part is easier to change and is rapidly evolving.
You could say that the = issue is deliberate, but I'd say the core problem is that shell didn't start out as a programming language, or at least it was a very impoverished one without variables.
The original paper from the 70's on the Thompson shell shows that. Shell had "goto" but no variables! So name=value had to be grafted on later without breaking too many things. Words were already split by spaces, so I guess they just made
name=value
a "pseudo-word" that becomes an assignment, whereas name = value
remained 3 words.I remember one project that started out with simple XML config, then I added conditionals. when I was starting o work on loops and reusable variables I realized that I was writing a programming language in XML. So I started to write config in C# and compiled that dynamically instead.
I implemented a simulation framework once, where the core simulation process took in all parameters as command line options. A runner program would read a YAML file describing the parameters in arbitrarily complex ways, and invoke the sim process potentially thousands of times (doing parameter sweeps). The intent was for the underlying sim process to always be directly runnable for debugging just by copy/pasting the generated list of options, regardless of how complex the config file got. It was a year or so before somebody started passing in an intermediate config file to the sim process by command line argument...
Ha, I agree, although I also think shell is a bit impoverished. It works but there are valid reasons people don't use it.
I hope to add the "missing declarative part" to shell in https://www.oilshell.org :
https://lobste.rs/s/6oxpe3/s_lot_yaml#c_mje209
Prediction: we'll see a lot more shell embedded in YAML in the coming years, with the same examples shown in the original article (Github Actions, didn't know about Helm Charts)
IMO we should get rid of the YAML!
https://lobste.rs/s/v4crap/crustaceans_2021_will_be_year_tec...
For example, there was an ANT library for JRuby that worked super well for using ANT constructs/libraries but in a sane language.
What's old is new again :)
Pulumi is an interesting tool in this direction. Rather than write in something like teraform, you just use your programming language of choice. Pulumi is just a library you use.
YAML does not have an if anymore than JSON have an if because you can write {"if": "foo", "then": "bar"} and some processor can process this with if-semantics.
{{ toYaml .Values.labels | indent 4 }}
{{- include "grafana.labels" . | nindent 4 }}
Building data structures by piercing together strings like this is a bit questionable.So yeah, if Helm is complicated, sure, don't blame yaml for that. Similarly, I'm not sure all of these declarative CI examples are all valid yaml either and not fed through front-end preprocessor first. They're essentially feature-identical with Jenkins declarative pipeline minus the ability to run arbitrary Groovy code in your build scripts, though of course you can get this exact behavior if you want by feeding a HEREDOC to a sh step that invokes the Groovy interpreter.
In any case, the author calling CI workflow definitions "config" is a little misleading. They're necessarily more complicated than a properties file and need to allow you to invoke external tools. Newer language ecosystems like Go and Rust are trying to solve this by putting dependency management, compiler, packager, and testing all into one tool provided with the language installation, but even there a lot of CI/CD needs to do a lot more than that, like deploy infrastructure, build container or VM images, etc.
People start with some YAML, and everything is fine. Then a simple condition is added, so ok, we are treating code as data, LISP style but in YAML. Then the logic and branching grows, and we introduce templates and so on.
You start with config, then the config ends up with its own config, and eventually, you are using Skaffold to configure Helm, which generates your YAML. That can't be the right solution, can it?
For most applications it would probably be overkill, but it seems useful for applications that need to change their own configuration (e.g. by the user through its GUI), or need its configuration changed by other tools.
I like for this a sorted name=value format, where non-scalars such as arrays and hashes are flattened into the names. E.g., if the config contains an array named "users" with 3 items, the name=value pairs would flatten to names users.0, users.1, and users.2. An array named "servers" whose entries are 3 hashes with the host name and port would flatten to names servers.0.host, servers.0.port, servers.1.host, servers.1.port, servers.2.host, servers.2.port.
That gives you diffs that tell you what has actually changed in the configuration itself rather than in the formatting of the configuration file.
cat my_configuration_text.sql | sqlite3 my_configuration_database.sqlite
Closely related, sqldiff ships with sqlite. It generates the difference between sqlite database files as a bunch of {insert,update,delete} statements.
I don't think much of that Dan Luu article applies to read-only files deployed with the application, which is often what configuration files are.
Definitely concurrency control is an issue with mutable files, that would all by itself make me hesitant to use them as a storage solution if that were in play. But it's not when the changes to the files are being done outside of the app being deployed, and they are deployed as read-only, any changes require app restart to pickup etc. That is how I have usually experienced configuration and I have personally never run into any problem with having it in the file system.
Most of the other stuff in the article also seems to me not to apply to the standard configuration use case.
(And if it did... it would apply to your source files too, right? Whether ruby or python or even JVM bytecode. Yet we obviously can and do put those in the filesystem. Ultimately read-only configuration are just another kind of source file they don't really have any special problems).
But "incremental remote update" (without requiring app restart especially!) is definitely not the standard configuration use case. I agree that a 'real database' seems reasonable for that use case, whether sqlite3 or something else. Whether you try to control the configuration (that will wind up in a db) in your version control system, or just use standard db backup/clone techniques instead.
It's as straightforward and readable as XSLT allows, leaving only a small refactoring on the table (putting "
" into an unconditional output instead of repeating it in the four cases).
for $i in 1 to 100
return if ($i mod 3 = 0 and $i mod 5 = 0) then "FizzBuzz"
else if ($i mod 3 = 0) then "Fizz"
else if ($i mod 5 = 0) then "Buzz"
else $i
Wrap it in string-join, if you need the output as one string with line breaks rather than a sequence list.Although one can also write horrible clever XPath:
for $i in 1 to 100
return (("Fizz"[$i mod 3 = 0] || "Buzz"[$i mod 5 = 0])[.], $i)[1] <xsl:stylesheet version="3.0" xmlns:xsl="http://www.w3.org/1999/XSL/Transform">
<xsl:template name="xsl:initial-template">
<xsl:sequence select='
for $i in 1 to 100
return (("Fizz"[$i mod 3 = 0] || "Buzz"[$i mod 5 = 0])[.], $i)[1]
' />
</xsl:template>
<xsl:stylesheet>
While I understand the ratio behind the article and agree, the author does not acknowledge, that XSL-T is not a "programming language" but a "templating language". E079 PROGRAMMER IS INSUFFICIENTLY POLITE
The balance between various statement identifiers is important. If less than approximately one fifth of the statement identifiers used are the polite versions containing PLEASE, that causes this error at compile time.
E099 PROGRAMMER IS OVERLY POLITE
Of course, the same problem can happen in the other direction; this error is caused at compile time if more than about one third of the statement identifiers are the polite form.
Some more fun details:- The compiler is called `ick`
- The compiler has a `-mystery` flag which is documented as
"This option is occasionally capable of doing something but is deliberately undocumented. Normally changing it will have no effect, but changing it is not recommended."
- Numbers have to be entered in English. 12345 would be written as `ONE TWO THREE FOUR FIVE` unless you put it in roman numeral mode where the characters ‘I’, ‘V’, ‘X’, ‘L’, ‘C’, ‘D’, and ‘M’ mean 1, 5, 10, 50, 100, 500 and 1000.
- The debugger is called `yuk` E134 PROGRAMMER'S POLITENESS IS TOO PREDICTABLE
Politeness, once made too predictable, is too easily overlooked. This error is
caused at compile time if the programmer's politeness level is too easily
predicted by the compiler as it encounters each statement in the code.1. Casing (some lisps are case sensitive others aren’t)
2. Lots of types of atom. Eg Common Lisp has like a dozen different types of number (short/long/single/double floats, integers, rationals, complex numbers thereof), weird syntax (eg put a decimal point at the end of an integer to make sure it’s read in base 10. The fact that you have to worry it won’t be is already a serious concern), strings and symbols.
3. Extra syntax/types, eg vectors, arrays, bitvectors, #., backtick and comma, quote, backslash rules (but not sure if there’s a standard way to escape characters), keywords, packages
4. Multiple similar things, eg vectors and lists, symbols and strings, alists and plists.
Many of these qualities may be useful in programming but if one treats a config representation as a thing which must be validated and parsed into the actual data, then all of this adds confusion. I think there should be only one kind of atom: the string which may be written without quotes. This way you can be flexible in parsing (some fields you might choose to parse 50% as 0.5 and other fields you might require that the number starts with a dollar sign so people don’t forget that it is referring so some amount of dollars) and make it easy for people to write config files (no errors about expecting a string and getting a symbol or vice versa).
(published
(type std.iso.8601
value 2020-02-26))
Or like this, where "date:" is something only your application knows: (published date:2020-02-26)
Instead, you could rely on an externally specified format (prefix @ is std.iso.8601)
And use it in your file to parse text so that your application can build values of the proper datatype: (published @2020-02-26)
The application language would register lexers for those formats or fail the parsing step.You could have fake parsers that just skip over the defined syntax if you don't need to process it in your code. Or, the way the syntax is defined could be such that it tells the lexer how to skip over a token even if the type is not useful for a tool's purpose (skip until a space, or parse exactly N characters, or "read until this delimiter with backslash being an escape character").
(probably reinventing the wheel)
(published 2021-02-26)
The program will get a list of two atoms, “published”, and “2021-02-26”. It can complain if you’ve eg written “published” when you should have written “submission_date” and it can complain if the next atom isn’t a valid date. You don’t need to tell the reader to parse something as a date because the reader isn’t best placed to know what should and shouldn’t be a date. Whereas the program should know exactly where it expects dates to be so it might as well handle parsing them so you don’t need to tag dates when you write your config.Any human can read that date and know what it means so why bother tagging it if the machine can also know it should be a date.
You can always write a dictionary into that config file, just like you would with normal yaml. It would actually be sort of nice if python had support for yaml style dictionaries exactly for this purpose.
{ok,Commands}=file:consult("functions.txt"),
[ apply(M,F,A) || {M,F,A} <- Commands ].
With contents of functions.txt: {io,format,["Hello World!~n"]}.Well, JSON without commas or braces would be restricted to single literal values. (Or, I guess, single-element lists, but what would the point of that be?)
Where JSON here is just a stand-in for "simple format with few types", not about it being parseable as JSON or anything.
But regardless, I feel like that exists right now; doesn't this describe protocol buffers?
E.g.: a list:
- foo
- bar
- buz and other things besides
list of dicts: persons:
- name: John Smith
age: 54
occupation: Obergruppenführer
- name: Frank Frink
age: 32
occupation: craftsman
(Those are already valid YAML iirc)No, he didn't, he said "JSON without commas or braces". That makes no sense.
You were the one who specified nothing but record separators:
>>> you could have multiple elemnts without commas or braces, just with whitespace and newlines as separators.
And sure, that approach has upsides and downsides compared to JSON, which is why it's already a widespread alternative to JSON.
Your proposal here is just JSON that looks slightly different. The reason we don't have that is that we do have that, with isomorphic syntax, and we call it "JSON". There's nothing at all interesting about the idea of "JSON without commas" if you satisfy it by saying "we've eliminated the comma by relabeling it as a 'hyphen' in some cases and a 'tab' in the rest".
Strictly interpreted, no, it doesn't.
But I picked his intention, and clarified it already as: "Where JSON here is just a stand-in for "simple format with few types", not about it being parseable as JSON or anything." to the comment you've respond to :-)
>Your proposal here is just JSON that looks slightly different. The reason we don't have that is that we do have that, with isomorphic syntax, and we call it "JSON". There's nothing at all interesting about the idea of "JSON without commas" if you satisfy it by saying "we've eliminated the comma by relabeling it as a 'hyphen' in some cases and a 'tab' in the rest".
Hey, one should at least try to infer what people mean from context. Not everybody is a native speaker or the best communicator.
What the parent asks for, and it is interesting, and we should have had it, is basically "minimal, sane, YAML subset".
TOML is somewhat it.
No, the biggest complain about YAML is the type coercion, unexpectedly changing your data, and incompatibility between parsers and versions...
That said, just because a language has these features, doesn't mean you need to use them. A lot of these issues are in part the fact they're there, but also that people don't take a stand / don't have the discipline to NOT use them.
It's why I'm now an opponent of Scala, because it's too free and feature rich.
There's a very specific, very common, use case that the people who write most markup parsers refuse to support...
A good config file format needs to: - be human readable - be machine modifiable while preserving comments and whitespace - be unable to run code - be unable to call a constructor on an object without being whitelisted
So for a new project I tried writing a spec in YAML instead and the indentation rules felt quite shaky for me as a beginner (ie i wasn't entirely sure where things ended up with all the modes) so it was a lot of trial-and-error for someone unfamiliar with YAML as me (and it explains why I've kinda stayed away from it because writing it was as hard as reading it sometimes).
CSON seems to have a sweet-spot between them with less clutter than JSON but more straightforward block rules than YAML.
https://github.com/lightbend/config/blob/master/HOCON.md
HOCON is the "human optimised config object notation" and is a superset of JSON. All valid JSON is valid HOCON but then it goes and adds lots of other features specifically designed for writing usable config files.
LOL, reminds me of the C++ as shell scripts hack where you stick something to the effect:
#!/usr/bin/tail -n +1 $0 | g++ -O -g -o ${0%.cpp} -x c++ - && ./${0%.cpp} $1 $2 && rm ./${0%.cpp} ; exit
at the top of your cpp file and chmod +x it.
Take a look at Dhall as an alternative.
Both of these are actually designed well for dealing with the problem of configuration and it really shows.
Cue needs to be adoption ready and is getting close. The big one will be the 'cue mod' command, when that lands we should be in good shape. I think v0.4 will be when people start looking hard at Cue, but maybe a later 0.3.x
Huh, didn't know urbit had a predecessor!
Anybody has direct impressions?
Just allow trailing commas at the end of lists. It's much less jarring and unfamiliar.
Trivial? Maybe. But my brain takes a while to adjust to new conventions and I do it way too frequently already.
However, `dhall format` is _very_ opinionated, and will remove it.
{ foo = 3, bar = 4, baz = 5, verylongproperty = 6, dunnoneedsmoremore = 7, howmanytoforcealinebreak = 8 }
And it forces the comma at the beginning of the line, Haskell-style: { foo = 3
, bar = 4
, baz = 5
, verylongproperty = 6
, dunnoneedsmoremore = 7
, howmanytoforcealinebreak = 8
}E.g. for Kubernetes: https://github.com/dhall-lang/dhall-kubernetes
A configuration file is just a bad programming language (or a good one...).
Some people have forced YAML to be a configuration file because they're lazy and don't want to write a parser/lexer. But most people do it because they don't know the difference.
The problem with these languages embedded into YAML is they are all one-off implementations. TravisCI conditionals have a TravisCI specific syntax, usage, and features. You can't use Travis's concat function or conditional regex in the YAML configuration for your ansible playbooks.
I remember some visionary people at a certain FANG org I worked at trying to reinvent HTML, but expressed as YAML.
Sometimes a screw is the correct or perhaps only fastener for a job and a hammer is the only screwdriver available or the only screwdriver capable of driving the screw. I've done it.
I'm currently in tradeoff mode as well. Do I spend a month copy / pasting some shit code in the existing codebase so that I can spend the rest of the year on the project rebuilding things, or should I pause the rebuild project and instead clean up the existing one (effectively a rebuild-in-place).
YAML though, ugh. Nothing worse than looking at a helm chart. JSON sucks too, but I’d rather a JSON config than a YAML one anyday.
This article feels more like a rant if anything and illustrates the author misunderstanding of the purpose of the language.
Otherwise, one could argue that some programming languages use newlines to separate commands, therefore the file type is NSV (newline separated values), and therefore NSV is a programming language, and any newline-separated file is a program. This is clearly nonsense.
Sometimes people confuse the medium with the language. YAML may be the medium in which programming instructions are conveyed, but it never makes YAML a programming language, irrespective of the file extension. If you put lines of bash into a YAML array, the YAML itself still only contains data. If you pass that data file to something that can take the bash lines out and make use of them, then great.
Essentially, it's a storage medium - just a slightly higher-level storage medium than we're used to thinking about. You could create a programming language syntax that is entwined with YAML, but then the language would be more correctly named something like Whatever-over-YAML (or Whatever for short)
People start with some YAML, and everything is fine. Then a simple condition is added, so ok, we are treating code as data, LISP style but in YAML. Then the logic and branching grows, and we introduce templates.
It not that anyone wants to get where we've ended up. It's that each step along the way seems to make sense until you end up trapped in complex templates, and scripts to configure your config and it's too late.
It is a vicious local optimum that everyone keeps falling into.
That's a function: anything -> text. It doesn't mean anything.
EDIT: Also, I typo'd Shall for Dhall. My bad.
JSON-LD or XML are perfectly good candidates for data with strict schema. But it took devops startup fanboys some time to realize schemas were useful in the first place.
Agreed, it's analogous to the way the programming language world has come around to realising the static type system folks had a point all along.
I don't have much experience with XML so I can't speak to whether the other criticisms of XML make sense, especially regarding complexity.
People can say what they want about JSON but anything beats the old world of rolling your own serializers and convincing your boss over the hours you've spent reinterpreting some "sporadic flavor" like YAML.
I think that is a good tradeoff, JSON being strict for data interop with the possibility of enabling comments for those cases where people use it for configuration (and specifying it as JWCC)
Hum... Why not? Serialized data can be written and read by humans too.
Besides, if comments are out of scope, there is no reason for them to be textual either.
You mean re-add them.
> I removed comments from JSON because I saw people were using them to hold parsing directives, a practice which would have destroyed interoperability.