StrictYAML
hitchdev.com
hitchdev.com
Even I tend to write "pseudo-YAML" and then add the JSON line noise programmatically if I'm entering a lot of data.
- Oh you used a tab instead of a space.
- What do you mean country code can be 'no'?
- Sure we are JSON compatible**!
- Yeah. 1e3 is an actual string. If it's not. We're not compatible with YAML 1.1
- Yes. I will import your remote exploit.
So like, Typescript but worse ?
I've had far more issues getting these YAML files correctly formatted than I should admit, and every test means pushing a commit and waiting to see what happens.
https://docs.github.com/en/actions/using-workflows/workflow-...
It refers you to this YAML guide, which itself never mentions what version it's helping you learn:
I use nekos act to run workflows locally (and it does not support reusable workflows), but it sure helped me to get the yaml right.
run: |
mkdir build && build
cmake ..
etc. This is great for CI.We seem to have very different ideas of what a "configuration language" should be capable of and what its parser should (be allowed to) do.
Look a little closer, dhall is probably the best option I've seen that preserves the properties I want out of a configuration language (correctness, determinism, turing incompleteness, easy projection into other formats, etc).
> However, when you protect an import with a semantic integrity check the import is permanently locally cached after the first request, so subsequent imports will no longer make outbound HTTP requests.
Also this PR to nixpkgs from 2020: https://github.com/NixOS/nixpkgs/pull/79900
> Many users have requested Dhall support for "offline" packages ... The goal of this change is to document what is the idiomatic way to implement "offline" Dhall builds ... The trick to implementing offline builds in Dhall is to take advantage of Dhall's support for semantic integrity checks. ... The offline nature of the builds are enforced by compiling the Haskell interpreter with the -f-with-http flag ...
https://docs.dhall-lang.org/discussions/Safety-guarantees.ht...
Any new tech is esoteric at first.
When I started Python 20 years ago, there were no job offer for it in my country.
Honest question - what is the example use case here? Which configuration files in YAML are currently being handled by people who don't know any coding, and would be impeded by curly braces (and anything like JSON or XML)? I think this is only an imagined problem.
Also, I have a met a non-programmer lady from HR, who was able to download a VB script into Outlook and adapt it to her needs (which was some kind of automation). Richard Stallman made a similar observation with Emacs configuration. I think you quite underestimate what non-programmers can do.
If I where to use it in Python, I would have to code a parser, then implement the entire type logic mysql.
And I bet this is why it's not more used: JSON or YAML are comparatively easy to implement because you just need the parser. You don't have a full featured language with a set based typing system on top.
Sometimes you want to let a user edit some fields in a map/dict/whatever, and that's it. I use TOML here, personally, but YAML has some advantages for more complex data, especially if string keys are long.
Not everyone is trying to set up a Kube cluster.
((name "Ford Prefect")
(age 42)
(possessions "Towel"))
(:name "Ford Prefect"
:age 42
:possessions ("Towel"))
(countries "GB" "IE" "FR" "DE" "NO")
(countries gb ie fr de no) ; if country codes are identifiers in a DSL
((first-name "Christopher")
(surname "Null"))
(:first-name "Christopher"
:surname "Null")Dhall and Cue are no doubt excellent and amazing for complicated use cases, but there are a lot of use cases where using them is akin to taking the VTOL jet out to pick up a few things at Trader Joe's.
Yeah, you always start out with K/V pairs but it never takes long until you need the same configuration but just with a slight tweak here or there. For example consider the same configuration but for different running environments.
People always come up with custom solutions per tool which some sort of metaprogramming, again in YAML. This is just bonkers IMO.
So why not just use a proper language from the beginning?
(server-addresses)
(server-addresses (ipv4 "10.255.5.5:801") (localhost) (ipv6 "1234::5678") (ipv4 "10.255.25.25:8001"))
(list
(server (address (ipv4 "10.255.5.5")) (port "801") (allowed-countries (list "se" "no" "dk" "fi")) )
(server (name "nil") (allowed-countries nil) (address (ipv6 "1234::5678") )))Don't use YAML 1.1
If you force 1.2 without the header I suppose it works, but breaks compatibility?
Don't use YAML 1.2, use the secure subset 1.1. Or even better StrictYAML
I mean, things like "NO" were fixed later.
StrictYAML got rid of all these, as all the safer YAML variants. E.g. perl5 still uses 1.0 for its cpan metadata abstraction, and still has to restrict these. I maintain the https://metacpan.org/pod/YAML::Safe module, which allows whitelisting of certain objects.
YAML spec is a huge with bunch of gotchas. E.g. C# has two listed parsers. One is 1.2 compatible and unmaintained.
There's no way to determine whether a `.yaml` file is YAML v1.1 or v1.2 without a version directive, and most YAML documents are v1.1 because most YAML parsers default to v1.1 semantics.
I used Ruby as an example since it's easily available on most platforms, but you could also use Python or C++ or Swift or whatever language[1] you prefer. The underlying issue -- YAML not being a subset of JSON -- is universal.
[1] Note that some libraries, such as go-yaml, do their own thing and don't conform to either v1.1 or v1.2 semantics.
- a git explainer centered around git internals, serving as an indictment of git's ux
- a parser for a restricted subset of yaml, serving as an indictment of yaml's excesses
on the front page yesterday:
- an rsync explainer centered around rsync internals, where part 1 details how rsync is wrapped in a dockerfile and a perl script in order to be made useful
- a sad thinkpiece on how baroque the web has become
on the front page the day before:
- a go utility weighing in at tens of source files that implements what should be a built-in feature of AWS
- an article on encapsulation in rust, the buried lede of which is "you must audit your transitive dependency graph in order to retain the benefits of rust"
why do so many technologies feel like self-harm?
Come to think of it, this probably applies to any product.
Please elucidate?
Did you mean something else?
> My purpose is to lay down criteria by which the manipulation of people for the sake of their tools can be immediately recognized, and thus to exclude those artifacts and institutions which inevitably extinguish a convivial life style. Paradoxically, a society of simple tools that allows men to achieve purposes with energy fully under their own control is now difficult to imagine.
> The hypothesis was that machines can replace slaves. The evidence shows that, used for this purpose, machines enslave men. Neither a dictatorial proletariat nor a leisure mass can escape the dominion of constantly expanding industrial tools
> The crisis can be solved only if we learn to invert the present deep structure of tools; if we give people tools that guarantee their right to work with high, independent efficiency, thus simultaneously eliminating the need for either slaves or masters and enhancing each person’s range of freedom. People need new tools to work with rather than tools that “work” for them. They need technology to make the most of the energy and imagination each has, rather than more well-programmed energy slaves.
0. https://archive.org/details/illich-conviviality
1. "After many doubts, and against the advice of friends whom I respect, I have chosen “convivial” as a technical term to designate a modern society of responsibly limited tools. In part this choice was conditioned by the desire to continue a discourse which had started with its Spanish cognate. The French cognate has been given technical meaning (for the kitchen) by Brillat-Savarin in his Physiology of Taste: Meditations on Transcendental Gastronomy. This specialized use of the term in French might explain why it has already proven effective in the unmistakably different and equally specialized context in which it will appear in this essay. I am aware that in English “convivial” now seeks the company of tipsy jollyness, which is distinct from that indicated by the OED and opposite to the austere meaning of modern “eutrapelia”, which I intend. By applying the term “convivial” to tools rather than to people, I hope to forestall confusion." (Illich)
1. Application configuration. That would be better off as environment variables and/or JSON(5) for structured data.
2. CI/CD configuration. Those should really be scripts, because basically every such configuration file I've seen invents its own way of making conditionals anyway.
Also scripts are very different from configs. Configs are declarative, (most) scripts are imperative. If your config is complex enough then maybe you should move some logic into imperative code, but you can still have a config on top of that.
Yes, if you want, you can write some terrible YAML. It's your project though, you don't need to make your YAML terrible. The same is true for everything else about your configuration, code, and infrastructure. I've seen some terrible JSON configuration that used the fact that some JSON parsers will overwrite subsequent reuses of keys as a method of documentation, for example.
I know that in YAML 1.2 you can just enter JSON into a YAML file and it'll work if that's what you want. I haven't found a real life example of why that would be better, though.
If you're going configuration format purist, use XML. It's actually good at requiring you to follow the schema and there are XML parsers for absolutely any platform. You can write conditions in an accompanying XSLT file for "active" configuration.
How did you even get that idea from GP’s comment?
Ci/cd is literally a different item from the json/config file one.
Hear hear. Every ci/cd is “what if we half-assed a terrible programming langage in yaml?” And apparently the world is ok with that.
For 3 months I have been working with a team leader that had the project of making all entrypoints in the program accessible from a JSON based DSL. I tried to suggest we first expose a nice regular ergonomic API, and if we can't satisfy our requirement, we then add the DSL on top.
Nope.
Got the PR this week. Now he leaves next month, so we will have to maintain that.
I don't think the cost of an additional standard is worth it in this case.
While YAML has issues, they aren't much of problem if you use a linter, such as yamllint [1]. This can be enforced as part of the CI process.
But the solution is to make it impossible to write 「null」 at all? I don't think I would have suggested that one...
It looks like you need to use an empty string and you can tell it to translate to None. That's better than nothing, but it's still basically an inability to use actual null here.
Here you can use `NullNone` to parse `null` into None.
from strictyaml import Map, NullNone, Int, load
schema = Map({"a": NullNone() | Int()})
load("a: null", schema) == {"a": None}
load("a: 7777", schema) == {"a": 7777}> Here you can use `NullNone` to parse `null` into None.
That's the worst possible outcome for poor Christopher Null.
The best you can do for him is use "a:" to mean null but at that point you're not really dealing with nulls, you've just gone with strings and used an empty string.
* Allow a hierarchical configuration scheme with support for key-value mappings and lists.
* Support cross-references between one part of the configuration and another.
* Provide a string interpolation facility to easily build up configuration values from other configuration values.
* Provide the ability to compose configurations (using include and merge facilities).
* Provide the ability to access real application objects safely, where supported by the platform.
* Be completely declarative.
It's similar to newer formats such as JSON5, HJSON, HOCON and similar but offers a number of features [0] which they don't, as indicated by the above list. It's not intended to occupy the niche where you find things like Cue, Jsonnet, Dhall and similar.
It was just never especially publicised when first implemented for use in Python projects, but it now also has implementations for the JVM, .NET, Go, Rust, D, JavaScript [1], Ruby and Elixir (all BSD-3-Clause licensed) and it would be great to get feedback on the project from the HN community.
[0]: https://docs.red-dove.com/cfg/intro.html - description of features and comparison with other similar systems
[1]: https://docs.red-dove.com/cfg/playground.html - uses the JS implementation to create an interactive playground
Sometimes I need to substitute variables.
Sometimes I need to generate similar blocks from templates.
Sometimes I need to read part of configuration from environment variable.
Simply importing Python module was always easier.
Terraform introduced a brand new language, Pulumi uses JavaScript.
strictyaml website has examples in Python, that's why Python.
If you use anything else from Python lang they drop like flies. def? loop? Out of the question. I've even seen so-called IT people not able to handle scripting.
.ini files are about the only thing you can count on, (almost) guaranteed that a non-developer can handle.
json, yaml et al are ways to declare literal data. this is good. they are fine.
the issues always come from what the data is used for. nailing your schema, making your data structures as simple as they can be and no simpler, this is where the engineering happens. this is the hard part. literally all that matters.
not validating arbitrary data inputs is obviously a bad idea. whether you validate them via a high level library or tediously by hand[3] isn’t very important.
what is important is that the data structures are sane, simple, and stable. if they are easy to describe, they might be a good idea. if the approach the complexity of general purpose pl, they probably aren’t.
most literal data schemas are too broadly scoped. too general. github actions, other ci, k8s, etc. they have too many knobs, too many permutations. this is not a feature, it is a failure of design.
good schema validation won’t fix broken design. it’s unrelated.
1. https://github.com/nathants/py-schema
2. https://github.com/nathants/clj-schema
3. https://github.com/nathants/libaws/blob/ae48040911bf2c0554da...
How are the comments preserved in the Python representation? Is it some kind of ordered dict?
I would rather avoid YAML all together.
Small human editable files: TOML.
Big machine editable files: JSON.
And I validate all that with pydantic.
I wish CUElang was ready for prime time, that would solve all this at once.
"Everyone hates YAML. Everyone writes a lot of YAML."