The Norway Problem
hitchdev.com
hitchdev.com
https://www.theverge.com/2020/8/6/21355674/human-genes-renam...
Edit: Apparently Excel has its own Norway Problem ... https://answers.microsoft.com/en-us/msoffice/forum/msoffice_...
The more general problem basically being sentinel values (which these sorts of inferences can be treated as) in stringly-typed contexts: if everything is a string and you match some of those for special consideration, you will eventually match them in a context where that's wholly incorrect, and break something.
x == “00.10”
You’ll get a type error that x is a decimal and the string literal is a string. So then you know you have to reimport it in the right way. So the type system told you that an assumption was violated.
This won’t always happen, though. E.g. sort by this field will happily do a decimal sort instead of the string 00.10.
The best approach is to ask the user at import time “here is my guess, feel free to correct me”. Excel/Inflex have this opportunity, but YAML doesn’t.
That is, aside from explicit schemas. Mostly, we don’t have a schema.
So that system is not consistent with type checking? How is this not considered a bug?
Types only help if you pick the right ones.
> sentinel values
Using in-band signaling always involves the risk of misinterpreting types.
> This is part of more general problem
DWIM ("Do What I Mean") was a terrible way to handle typos and spelling errors when Warren Teitelman tried it at Xerox PARC[1] over 50 years ago. From[2]:
>> In one notorious incident, Warren added a DWIM feature to the command interpreter used at Xerox PARC. One day another hacker there typed
delete *$
>> to free up some disk space. (The editor there named backup files by appending $ to the original file name, so he was trying to delete any backup files left over from old editing sessions.) It happened that there weren't any editor backup files, so DWIM helpfully reported *$ not found, assuming you meant 'delete *'
>> [...] The disgruntled victim later said he had been sorely tempted to go to Warren's office, tie Warren down in his chair in front of his workstation, and then type 'delete *$' twice.Trying to "automagically" interpret or fix input is always a terrible idea because you cannot discover the actual intent of an author from the text they wrote. In literary criticism they call this problem "Death of the Author"[3].
[1] https://en.wikipedia.org/wiki/DWIM
[2] http://www.catb.org/jargon/html/D/DWIM.html
[3] https://tvtropes.org/pmwiki/pmwiki.php/Main/DeathOfTheAuthor
Which can be a fun game, but is ultimately pointless.
Ironically, this did not render the way you intended because HN interpreted the asterisk as an emphasis marker in this line.
It works here:
... type 'delete *$' twice.
because the line is indented and so renders as code, but not here:> ... type 'delete $' twice.
because the subsequent line has emphasized text*. So the scoping of the asterisks is all screwed up.
Consider this headline in English: "Man attacks boy with knife". This can be read two ways, either the man is using a knife to attack the boy, or the boy had the knife and thus was being attacked.
The same sentence in Polish would make use of either genitive or instrumental case to disambiguate (although barely). However, a naive translation would only differ in the placement of a `z` (with) and so errors could still slip through. At least in this case the error would not introduce ambiguity, simply incorrectness.
Similar to language design we can also consider: does the inclusion/requirement of parity features reduce the expressivity of the language?
This was a real eye-opener for me when learning Latin in school: stylistic expressions such as meter, juxtaposition, symmetry are so much easier to include when the meaning of a sentence doesn't depend on word order.
Eh.... some things are easy and some things are hard in any language. The specifics differ, and so do the details of what kinds of things you're looking for in poetry. Traditional Germanic verse focuses on alliteration. Modern English verse focuses on rhyme. Latin verse focuses on neither. [1]
English divides poetically strong syllables from poetically weak syllables according to stress. It also has mechanisms for promoting weak syllables to strong ones if they're surrounded by other weak syllables.
In contrast, Latin divides strong syllables from weak syllables by length. Stress is irrelevant. But while stress can be changed easily, you're much more restricted when it comes to syllable length -- and so Publius Ovidius Naso is invariably referred to by cognomen in verse, because it isn't possible to fit his nomen, Ovidius, into a Latin metrical scheme. That's not a problem English has.
[1] I am aware of one exceptional Latin verse:
> O Tite, tute, Tati, tibi tanta, tyranne, tulisti.
OOH, this is a a typically human problem. We have a system. It's partly designed, partly evolved^. It's true enough to serve well in the contexts we use it in on most days. There are bugs in places (like norway, lol) that we didn't think of initially, and haven't encountered often enough to evolve around.
In code, we call it bugs. In bureaucracy, we just call it bureaucracy. Agency A needs institution B's document X, in a way that has bugs.
Obviously, it's also a typical machine problem. @hitchdev wants to tell pyyaml that Norway exists, and pyyaml doesn't understand. A user wants to enter "MARCH1" as text (or the name of a gene), and excel doesn't understand.
Even the most rigid bureaucracy is made of people and has fairly advanced comprehension ability though. If Agency A, institution B or document X are so rigid that "NO" or "MARCH1" break them... it probably means that there's a machine bug behind the human one.
Meanwhile... a human reading this blog (even if they don't program) can understand just fine from context and assumptions of intent.
IDK... maybe I'm losing my edge, but natural language programming is starting to seem like a possibility to me.
^I feel like we need a new word for these: versioned, maybe?
I can just about understand that "No" might cause a problem, but “Membrane Associated Ring-CH-Type Finger 1" being converted to MAR-1 defeats me.
No, that's not what's happening. To clarify...
If you type a 41 characters long string of "Membrane Associated Ring-CH-Type Finger 1" into a cell -- Excel will not convert that to a date of MAR-1.
On the other hand, it's if you type an 6-char abbreviation of "MARCH1" that looks like a realistic date -- Excel converts it to MAR-1.
No one in their right mind uses a spreadsheet for data analysis. Good for working out your ideas but not in a production environment. I figure excel was chosen as this the utility the scientists were most familiar with.
The proper tool for the job would be a database. I recall reading about a utility, a highly customized database with an interface that looks just like a spreadsheet.
A lot of tools operate on CSV files. People use Excel to peek at the results or prepare input for other tools, and that’s how the date coercion slips in.
Sometimes, people do use it to collate the results of small manual experiments, where a database might be overkill. Even so, the data is usually analyzed elsewhere (R, graphPad, etc).
The mistake was to believe that Excel can operate on CSV files. It doesn't support them in any meaningful way. It supports them in a "I can sort of pretend that I support CSV files" way.
The proper solution, in my opinion, is a lookup table stored in the database. It can be updated, it can be cached, it can be extended.
And for transfer of data, use formats to which you can attach a schema. This way type data is not lost on export. XML did this but everyone hates XML. And everyone hates XSD (the schema format) even more. However, if you use the proper tools with it, it is just wonderful.
Now, you and I know this problem is solved by prepending ‘ to the number and it will be treated as a string, but your average Excel user has no understanding of types or why they might matter. Many engineers will also look past this when generating Excel reports.
https://social.msdn.microsoft.com/Forums/vstudio/en-US/92e0a...
It's "user error" except that there is no way to set the default import to import as "Text" (as far as I know), so one has to remember to do the three step "Text" import every time instead of the default one step "General" import.
[0] The only thing you can safely do with CSV files is to interpret every value as text cell. CSV files always require out of band negotiation on everything, including delimiters, quotation, escape characters, the data type of each column.
Users BELIEVE Excel supports CSV file. That's the reality on the ground. Fighting against that is a losing battle.
Yes, yes, I see... This could be problematic, indeed. If only there were a logical solution.
TOML is fine for configuration, but not an adequate solution for representing arbitrary data.
JSON is a fine data exchange format, but is not particularly human-friendly, and is especially poor for editable content: Lacks comments, multi-line strings, is far too strict about unimportant syntax, etc.
Jsonnet (a derivative of Google's internal configuration language) is very good, but has failed to reach widespread adoption.
Cue is a newer Jsonnet-inspired language that ticks a lot of boxes for me (strict, schema support, human-readable, compact), but has not seen wide adoption.
Protobuf has a JSON-like text format that's friendlier, but I don't think it's widely adopted, and as I recall, it inherits a lot of Protobufisms.
Dhall is interesting, but a bit too complex to replace YAML.
Starlark is a neat language, but has the same problem as Dhall. It's essentially a stripped-down Python.
Amazon Ion [1] is neat, but I've not seen any adoption outside of AWS.
NestedText [2] looks promising, but it's just a Python library.
StrictYAML [3] is a nice attempt at cleaning up YAML. But we need a new language with wide adoption across many popular languages, and this is Python only.
Any others?
My own little JSON Next entry / format is called JSON 1.1 or JSONX, that is, JSON with eXtensions, see https://json-next.github.io
Also, there's no explanation what <..-..> and <..+..> do.
foo: |
"This is a string that goes across
multiple lines," he wrote.
In JSON5, you'd have to write: {
foo: \"This is a string that goes across \
multiple lines,\" he wrote."
}
This sort of ergonomic approach is why YAML is so well-liked, I think. (Granted, YAML's use of obscure Perl-like sigils to indicate whitespace mode is annoying, but it does cover a lot of situations.)YAML is also great at arrays, mimicking how you'd write a list in plaintext:
foo:
- "hello"
- 42
- trueMy feeling is that :keywords reduce the need and temptation to conflate strings and boolean/enumerations that occurs when there's no clear way to convey or distinguish between a string of data and a unique named 'symbol'. I miss them when I'm in Pythonland.
[1] https: https://www.compoundtheory.com/clojure-edn-walkthrough/
You get neat ways of nesting data, but that is not enough for a robust and mistake-resilient configuration language.
The problem isn't parsing in itself. The problem is having clear sematics, without devolving into full SGML DTDs (or worse still, XML schemas).
Hm, not sure that's true, S-expressions would only define the "shape" of how you're defining something, not the semantics of how you're defining something. EDN https://github.com/edn-format/edn for all purposes is S-expressions and have support for custom literals and more, to avoid "the trouble with data types from JSON"
The tag system is quite brilliant though.
a:
b:
- c: 1
- d:
- e: 2
- f:
g: 3
Turns into this, which is unreadable: [[a.b]]
c = 1
[[a.b]]
[[a.b.d]]
e = 2
[[a.b.d]]
[a.b.d.f]
g = 3
TOML also has a few restrictions, such as not supporting mixed-type arrays like [1, "hello", true], or arrays at the root of the data. JSON can represent any TOML value (as far as I know), but TOML cannot represent any JSON value.At my company we use YAML a lot for table-driven tests (e.g. [1]), and this not only means lots of nested arrays, but also having to represent pure data (i.e. the expected output of a test), which requires a format that supports encoding arbitrary "pure" data structures of arrays, numbers, strings, booleans, and objects.
[[a.b]]
c = 1
d = [
{ e = 2 },
{ f = { g = 3 } }
]A bit like JSON5, but I believe even more advanced.
Some of the neat features: Custom literals / tagged elements that can have their support added for them on runtime/compile time (dates can be represented, parsed and turned into proper dates in your language). Also being able to namespace data inside of it makes things a bit easier to manage without having to result to nesting or other hacks. Very human friendly, plus machine friendly.
Biggest drawback so far seems to be performance of parsing, although I'm not sure if that's actually about the format itself, or about the small adoption of the format and therefore not many parsers focusing on speed has been written.
The problem with most of these is they're useless to describe the data. Honestly, it is completely not useful to have the following to describe data:
email => string
name => string
dob => string
IMHO, it is akin to having a dictionary (like Oxford English) read like:
email - noun
name - noun
birthday - noun
It says next to nothing except, yes, they are nouns. All too often I waste time fighting nils and bullshit in fields or duplicating validation logic all over the place.
"Oh wow, this field... is a string..? That's great... smiles gently except... THERE SHOULD NOT BE EMOJI IN MY FUCKING UUID, SCHEMA-CHUD. GET THE FUCK OFF MY LAWN!"
{
"isTrue":false:Boolean,
"id":"123e4567-e89b-12d3-a456-426614174000":UUID
}thanks for the lolz
Not only are the constraints very hard to express (remember that one 2000 char regexp that really validates email addresses?), they are also contextual: the correct validation in an Android client is not the same as on the server side. Eg you might want to check uniqueness or foreign key constraints that you cannot check on the client. Sometimes you want to store and transmit invalid messages (eg partially completed user input). And then you have evolving validation requirements: what do you do with the messages from three years ago that don't have field X yet?
Unfortunately I don't think you can express what you need in a declarative format. Even minimal features such as regexp validation or enums have pitfalls.
I think it's better to bite the bullet and implement the contextually required validation on each system boundary, for any message crossing boundaries.
https://i.stack.imgur.com/YI6KR.png
Lua patterns have also shown up in other places, such as BSD's httpd, and an implementation for Rust:
https://www.gsp.com/cgi-bin/man.cgi?section=7&topic=PATTERNS
[1] https://amzn.github.io/ion-docs/ [2] https://amzn.github.io/ion-schema/
For situations like TFA you really want a configuration language that behaves exactly like you think it will, and since you don't have to interop with other organizations you don't really need a global standard.
Moreover, broadly used config languages can be somewhat counterproductive to that goal. Take JSON as an example; idiomatic JSON serdes in multiple programming languages has discrepancies in minint, maxfloat, datetime, timezone, round-tripping, max depth, and all kinds of other nuanced issues. Existing tooling is nice when it does what you expect, but for a no-frills, no-surprises configuration language I would almost always just prefer to use the programming language itself or otherwise write a parser if that doesn't suffice (e.g., in multilingual projects).
Mildly off-topic: The problem here, more or less, was that the configuration change didn't have the desired effect on an in-memory representation of that configuration. We can mitigate that at the language level, but as a sanity check it's also a good idea to just diff the in-memory objects and make sure the change looks kind of like what you'd expect.
For example, the fact that NestedText is a Python library means a Python team could use it, but it's a poor fit for an organization whose other teams use Go and JavaScript/TypeScript.
We use YAML for much more than configuration, by the way. I feel like YAML hits a nice sweet spot where it's usable for almost everything.
Until you have to, and all hell breaks loose ?
Now, the example of codepages maybe isn't really appropriate to companies, but is still a good enough metaphor ?
It's far more productive to push for incremental changes to the YAML spec (or even a fork of it) to make it more sane and better defined. Things like a StrictYAML subset mode for parsers in other popular languages.
The problems this article raises and strictyaml purports to address were addressed in YAML 1.2, already supported in python via ruamel.yaml; YAML 1.2 addresses much of this in the Core schema which is the closest successor to the default behavior of earlier spec versions, and does so more completely in the support for schemas more generally, which define both the supported “built-in" tags (roughly, types) and how they are matched from the low-level representation which consists only of strings, sequences, and maps (which, incidentally, are the only three tags of the “Failsafe” schema; there’s also a “JSON” Schema between Failsafe and Core, which has tags corresponding to the types supported by JSON.
I had a chance of using SOAP at one point. It was a F5 device and I used a python library. What I really liked is that when it connected to it it downloaded its schema, and then used that to generate an object. At that point you just communicated with device like you did with any object in Python.
We abandoned it for inferior technologies like REST and JSON, because they were harder to use from JS, as parent mentioned.
First of all, I was there 20 years ago. I had to deal with XML, XSLT, one kind of Java XML parsers that didn't fully do what I needed, another kind of Java XML parsers that didn't fully do what I needed. And oh boy was it a pain. I just wanted to get a few properties of a bunch of entities in a bigger XML document, that's all. Big fail.
Second, JSON always had a parser in JS, so I don't know where that eval nonsense is coming from.
Third, JS actually had the best dev UX for XML of all languages 20 years ago. Maybe you know JavaScript from Node.js, but 20 years ago it used to run excusively in web browsers, which even then were pretty good at parsing XML documents. The browser of course had a JS DOM traversal API known to every single JS developer, and very soon (Although TBH I can't remember if before or after JSON) it also had xpath querying functions, all built in.
XML was so bad, that its replacement came from the language where it was actually easiest to use. think about that for a second.
So the answer to the question "Why was XML replaced?" is not "Because webdevs lol".
I suspect it was because it has both content and attributes, which all but guarantees it's impossible to create a bunch of simple, common data structures from it (like JSON does).
Firstly, it sounds like XML ran over your dog or something. Sorry to hear about that. It wasn’t particularly hard to use at all, and if you’re dealing with the possibility of emojis in your JSON UUIDs in 2021, one might even say it’s easier to use.
If you’re referring to JSON.parse() in “had a parser” above, then you have a temporal problem. Regarding eval(), it’s suggested right in the original RFC for JSON. Check it out. Web developers at the time were following that advice.
website with grammar spec: https://tree-annotation.org/
prototype of a JSON/YAML alternative for JS: https://github.com/tree-annotation/tao-data-js
same thing, even less finished for C#: https://github.com/tree-annotation/tao-data-csharp
working on it constantly, more to come soon
The world desperately needs support for YAML 1.2, which solves the problems the article addresses fairly completely (largely in the “default” Core schema[0], but more completely with the support for schemas in general), plus a bunch of others, and has for more than a decade. But YAML 1.2 libraries aren’t available for most languages.
[0] not actually an official default, but reflects a cleanup of the YAML 1.1 behavior without optional types, so its defaultish. Back when it looked like YAML 1.3 might happen in some reasonably-near future, it was actually indicated by team members that the JSON Schema for YAML (not to be confused with the JSON Schema spec) would be the explicit default YAML Schema in 1.3, which has a lot to recommend it.
What advanced parts of YAML are you talking about that remain problems in YAML 1.2?
> The most tragic aspect of this bug, howevere, is that it is intended behavior according to the YAML 2.0 specification.
For the ease of entering time units YAML 1.1 parsed any set of two digits, separated by colons, as a number in sexagesimal (base 60). So 1:11:00 would parse to the integer 4260, as in 1 hour and 11 minutes equals 4260 seconds.
Now try plugging MAC addresses into that parser.
The most annoying part is that the MAC addresses would only be mis-parsed if there were no hex digits in the string. Like the bug in this post, it could only be reproduced with specific values.
Generally, if you're doing implicit typing, you need to keep the number of cases as low as possible, and preferably error out in case of ambiguity.
They treat the ":" like a sum of two sexagesimal numbers, rather than a sexagesimal digit separator.
If I enter 1-3-0-start, I get 90 seconds of cooking. If I enter 9-9-start, I get 99 seconds of cooking, so in that sense, 99 > 130.
If I want about 90 seconds, I’ll use 88 as it’s faster to enter (fewer finger movements).
Load soap into the dishwasher after emptying rather than after loading. If the soap dispenser is closed, the dishes are dirty.
insert code flame war here
If the dishwasher has dishes in it and it's not running, they're clean.
[0] https://www.reddit.com/r/self/comments/ayr9c/when_im_rich_im...
There was a list of AWS Account IDs that parsed just fine until someone added one that started with a 0 and had no numbers greater than 7 in it, after which our parser started spitting out decidedly different values than we were expecting. Fixing it was easy, but figuring out what in the heck was going on took some digging.
It had it literally at the same time as it had the problem in the article (the article refers to YAML 2.O, a nonexistent spec, and to PyYAML, a real parser which supports only YAML 1.1.)
Both the unquoted-YES/NO-as-boolean and sexagesimal literals were removed in YAML 1.2. (As was the 0-prefixed-number-as-octal mentioned in a sibling comment.)
This is a mind-boggling level of idiocy. Even leaving aside the MAC address problem, this conversion treats "11:15" (= 675) different from "11:15:00" (= 40500), even though those denote the same time, while treating "00:15:00" (15 minutes past midnight) and "15:00" (3 in the afternoon) the same.
And here's something else to keep you up at night: Just think of how many unintentional land mines lurk in your serialized data, waiting to blow up spectacularly (or even worse, silently) as soon as you attempt to change implementation technologies!
This is why I've been so anal about consistent decoder behavior in Concise Encoding https://github.com/kstenerud/concise-encoding/blob/master/ce...
...which is why it pains me to admit that in my own project for a Tcl-like scripting/config language[1] I missed the float v. string issue, so it'll currently "cleverly" return different types for 1.2 (float) v. 1.2.3 (atom). Coincidentally, I started work on a "stringy" alternative interpreter that hews closer to Tcl's philosophy (to fix a separate issue - namely, to avoid dynamically generating atoms, and therefore avoid crashing the Erlang VM when given potentially-adversarial input), so I'm gonna fix that case for at least the "stringy" mode (by emitting strings instead of numbers, too), knocking out two birds with one stone for the upcoming 0.3.0 release :)
----
[1]: https://otpcl.github.io, for those curious
Or give a schema to the parser, defining what type is expected in each field.
This is one of those great ideas that sadly one needs experience to realize are really bad ideas. Every new generation of programmers has to relearn it.
Other bad ideas that resurface constantly:
1. implicit declaration of variables
2. don't really need a ; as a statement terminator
3. assert should not abort because one can recover from assert failures
This is so true. I really like Julia and I know that explicitly declaring variables would be detrimental to adoption but I prefer it to the alternative, which is this: https://docs.julialang.org/en/v1/manual/variables-and-scopin...
1. If I scribble some one time code etc. the probability of having an error coming from implicit declarations is in the same order of magnitude as missing out edge cases or not getting the algorithm right for most people. The extra convenience may well be worth it.
2. I would relax this it should be clear to the programmer where a statement ends.
3. Go on with a warning is a sane strategy in some situations. I happily ruin my car engine to drive out of the dessert. The assert might have been to strict and i know something about the data so the program can ignore the assert failure.
.... and here is another entry for Walter's list of bad ideas:
4. "It's okay. I will use this code only once"
> 3. Go on with a warning is a sane strategy in some situations.
No, if its sometimes ok, to continue, than you should not assert it.
Assert means "I assert this will always be true, and if it's not our runtime is in unknown/bad state."
If you think you can recover, or partially recover, throw/return appropriate error, and go into emergency/recovery mode.
Downvote me if you want to open a bug ticket with the vendor and wait a week for the fix.
Upvote me if you’d give it a try to restart with a switch to ignore assertions.
You may abstain if you never shipped a bug.
Edit: not to forget that this website runs on lisp which violates all three. Was it really a bad choice for the website?
Several points:
1. Most of such critical components have several different and independent implementations, with analog backup (if possible).
2. You are arguing one specific safety critical case, that 99.999% or even more programmers will never face, should somehow inform decision about general purpose programming language.
3. Even if you are working in such safety critical situation, you should not really on assertion bypass, but have separate emergency procedure, which bypasses all the checks and try's to force the issue. (ever saw a --force flag ?)
Because what happens in reality, is developer encounters a bug (maybe while its still in development), notice you can bypass it by disabling assertions (or they are disabled by default), log it as a low priority bug, that never gets fixed.
Then a decade later me or someone like me is cursing you because you enterprise app just shit the bed, and is generating tons of assertion warnings, even when it running normally, so I have to figure out, which of them are "just normal" program flow, and which one just caused an outage.
I never experienced situation like you described, but I have experienced behavior like I wrote above, too many times.
Botom line is:
- don't assert if you don't mean it
- if you need bypass for various runtime checks, code one in explicitly.
Edit: Hacker News is written in ARC which is schema dialect. ARC doesn't have assertions as far as i can tell.
ARC doesn't have its own runtime and is run on racket language, that has optional assertion, that exit the runtime if they fail https://docs.racket-lang.org/ts-reference/Utilities.html
With most systems, the safest state is off. CNC machine making a weird noise? Smash that e-stop. Computer overheating? Unplug it. With this in mind, "assert" transitions the system from an undefined state to an inoperative state, which is safer.
That isn't to say that that you want bugs in your code, and that energizing some system is free of consequences. Your emergency stop of your mill just scrapped a $10,000 part. Unplugging your server made your website go down and you lost a million dollars in revenue. But, it didn't kill someone or burn the building down, so that's nice.
1. You're actually right if the entire program is less than about 20 lines. But bad programs always grow, and implicit declaration will inevitably lead you to have a bug which is really hard to find.
2. The trouble comes from programmer typos that turn out to be real syntax, so the compiler doesn't complain, and people tend to be blind to such mistakes so don't see it. My favorite actual real life C example:
for (i = 0; i < 10; ++i);
{
do_something();
}
My friend who coded this is an excellent, experienced programmer. He lost a day trying to debug this, and came to me sure it was a compiler bug. I pointed to the spurious ; and he just laughed.(I incorporated this lesson into D's design, spurious ; produce a compiler error.)
3. I used to work for Boeing on flight critical systems, so I speak about how these things are really designed. Critical systems always have a backup. An assert fail means the system is in an unknown, unanticipated state, and cannot be relied on. It is shut down and the backup is engaged. The proof of this working is how incredibly safe air travel is.
I ask you to reconsider your assumptions. How did this play out in the 737 MAX crashes? Was there a backup AoA sensor? Did MCAS properly shut down and backup engaged? Was manual overriding the system not vital knowledge to the crew?
You don’t have to answer. I probably wouldn’t get it anyway.
But rest assured that I won’t try to program flight control and I strongly appreciate your strive for better software.
They didn't follow the rule in the MCAS design that a single point of failure cannot lead to a crash.
> Was manual overriding the system not vital knowledge to the crew?
It was, and if the crew followed the procedure they wouldn't have crashed.
This is also true of Haskell btw.
But then it's inconsistent and has unnecessary complexity because now there's one (or more) exceptions to the rules to remember: when the ';' is needed. And of course if you get it wrong you'll only discover it at runtime.
"Consistent applications of a general rule" is preferable to "An easier general rule but with exceptions to the rule".
For the ';', perhaps not. For the token that is used to terminate (or separate) statements? Yes, the ';' is an exception to the general rule of how to terminate statements.
The semicolon also works on some sort of statements and not others, throwing errors only at runtime.
It's easier to remember one rule than many.
It's not a language in which you ever need be saving bytes on the source code. Just use a new line and indent. It's more readable and easier.
And I'd also add that it's something that you almost never do. One practical use is writing single line scripts that you pass to the interpreter on the command line. E.g. `python -c 'print("first command"); print("second command")'`
If you don't know about the `;` at all in python then you are 100% fine.
I find it much, much easier to look at code and parse blocks via indentation, than the many ways and exceptions of writing ; and {, }, while an extra or missing ';' or {} easily remains unspotted and leads to silly CVEs.
It even has semicolon insertion, but because the language is carefully designed, this doesn't cause problems, and most users can go a lifetime without knowing about it.
Our coding style requires semicolons for uninitialized variables, so you'll see
local x;
if flag then
x = 12
else
x = 24
end
As a way of marking that the lack of initialization is deliberate. `local x = nil` is used only if x might remain nil. -- Two assignment statements
x = 10 y = 20It's a bad idea because ASCII already includes dedicated characters for field separator, record separator and so on. These could easily be made displayable in a text editor if you wanted just as you can display newlines as ↲. Anyone who invents a format that involves using normal printable characters as delimiters and escaping them when you need them, is, I feel very confident in saying, grotesquely and malevolently incompetent and should be barred from writing software for life. CSV, JSON, XML, YAML, all guilty.
However, your tty driver, terminal or program are all likely to eat them or munge them. Also, virtually nothing actually uses these characters for these purposes.
Right. Which is why we have all these hilarious escaping and interpolation problems. Any why programmers will never be taken seriously by real engineers. It's like we have cement mixed and ready to go but we decide to go and forage for mud instead and think that makes us cleverer than the cement guys.
Maybe that has something to do with this?
ASCII is over 60 years old and separators haven't caught on yet; what's different now?
> These could easily be made displayable in a text editor if you wanted just as you can display newlines as ↲.
Can you name a common text editor with support for ASCII separators? It's a lot easier to use delimiters and escaping then change every text editor in the world.
> Anyone who invents a format that involves using normal printable characters as delimiters and escaping them when you need them, is, I feel very confident in saying, grotesquely and malevolently incompetent and should be barred from writing software for life. CSV, JSON, XML, YAML, all guilty.
All of the formats you rant about are widely used, well supported, and easy to edit with a text editor - none of these are true of ASCII separators. People chose formats they can edit today instead of formats they might be able to edit in the future. All of these formats have some issues but none of the designers were incompetent.
I think SGML (roll your own delimiters and nesting) was pretty close to the Right Thing,™ but ISO has the specs locked down so everyone had a second-hand understanding of it.
I feel like I prefer explicit
self.member = value
this.member = value
vs implicit member = value
But clearly C++/Java/C# people are happy with implicit ... though many of them try to make it explicit by using a naming convention.The added mental load of tracking variables' sources builds up.
But at least, to my knowledge, in Java these things can't turn out to be global vars. Having this ‘feature’ in JS or Python would be quite a pain in the butt.
It's actually a bit surprising that this is one thing that javascript does better than Java. In most other areas, it's Java that's (sometimes overly) explicit.
What I'm referring to is the notion that:
a = b c = d;
can be successfully parsed with no ; between b and c. This is true, it can be. But then it makes errors difficult to detect, such as: a = b
*p;
Is that one statement or two?I'm reminded of the discussion we had a few days ago about environment variables; one problem there is that env variables are always strings, and sometimes you do want different types in your config. But clearly having the system automatically interpret whether it's a string or something else is a major source of bugs. Maybe having an explicit definition of which field should be which type would help, but then you end up with the heavy-handed XML with its XSD schema.
Or you just use JSON, which is light-weight, easy to read, but unambiguous about its types. I guess there's a good reason it's so popular.
Maybe other systems like yaml and environment variables should only ever be used for strings, and not for anything else, and I suppose replacing regular yaml with 'strictyaml' could play a role there. Or cause unending confusion, because it does violate the spec.
The article didn't fully explain it but strictyaml requires a typed schema or defaults to string (or list or dict) if one is not provided. So it strictly follows the provided schema.
With the one exception that with floatig point values the precision is not specified in the JSON spec and thus is implementation defined[1] which may lead to its own issues and corner cases. It for sure is better than YAML's 'NO' problem, but depending on your needs JSON may have issues as well
[1]: https://stackoverflow.com/questions/35709595/why-would-you-u...
I couldn't get access to the original dataset but the column gave it away. Namibia's 2-letter ISO country code is NA—which happens to be in pandas' default list of NaN equivalent strings.
It was a headache and a half...
na_valuesscalar, str, list-like, or dict, default None
Additional strings to recognize as NA/NaN. If dict passed, specific per-column NA values. By default the following values are interpreted as NaN: ‘’, ‘#N/A’, ‘#N/A N/A’, ‘#NA’, ‘-1.#IND’, ‘-1.#QNAN’, ‘-NaN’, ‘-nan’, ‘1.#IND’, ‘1.#QNAN’, ‘<NA>’, ‘N/A’, ‘NA’, ‘NULL’, ‘NaN’, ‘n/a’, ‘nan’, ‘null’.
You fix it by using `keep_default_na=False`, by the way.tl;dr: there are a bunch of fields of various types that arrive as strings, and they get coerced but without paying attention to which field should have which type
Whenever an input accepts YAML you can actually pass in JSON there and it’ll be valid
It really surprised me when I found out and I use JSON Whenever possible since then since it’s much stricter
...unless your parser strictly implements YAML 1.1, in which case you should be careful to add whitespace around commas (and a few other minor things). This is a valid JSON that some YAML parsers will have problems with:
{"foo":"bar","\/":10e1}
The very first result Google gives me for "yaml parser" is https://yaml-online-parser.appspot.com, which breaks on the backslash-forward slash sequence.Strictly speaking, this is only true of YAML 1.2, not YAML 1.0-1.1 (the article here addresses YAML 1.1 behavior, the headline example od which was removed ib YAML 1.2 twelve years ago), though it calla YAML 1.1 “YAML 2.0”, which doesn’t actually exists.
Of course, there are lots of features, like custom types, that JSON doesn’t support, but you can still use YAML’s JSON-style syntax instead of actual JSON, for them.
I was with the article up until that point. I don't agree that diminishing returns with regards to type strictness applies linearly. Term-level Haskell is not massively harder than writing most equivalent code in JavaScript — in fact I'd say it's easier and you reap greater benefit. Perhaps it's a different story when you go all-in on type-level programming, but I'm not sure that's what the author was getting at. This smells of the Middle Ground logical fallacy to me. Or of course the comment was tongue-in-cheek and I'm overreacting.
First, he mentions "YAML 2.0" but there's no such reference about "2.0" from yaml.org or Google/Bing searches. Yaml.org and wikipedia says yaml is at 1.2. Apparently the other commenters in this thread clarified that the older "YAML 1.1" is what the author is referring to.
Ok, if we look at the official YAML 1.1 spec[1], it has this excerpt for implicit bool conversions:
y|Y|yes|Yes|YES|n|N|no|No|NO
|true|True|TRUE|false|False|FALSE
|on|On|ON|off|Off|OFF
But the pyyaml code excerpts[2][3] from resolver.py has this: u'tag:yaml.org,2002:bool',
re.compile(ur'''^(?:yes|Yes|YES|n|N|no|No|NO
|true|True|TRUE|false|False|FALSE
|on|On|ON|off|Off|OFF)$''', re.X),
The programmer omitted the single character options of 'y' and 'Y' but it still has 'n' and 'N' ?!? The lack of symmetry makes the parser inconsistent.And btw for trivia... PyYAML also converts strings with leading zeros to numbers like MS Excel: https://stackoverflow.com/questions/54820256/how-to-read-loa...
[1] https://yaml.org/type/bool.html
[2] 2020 latest: https://github.com/yaml/pyyaml/blob/ee37f4653c08fc07aecff69c...
[3] 2006 original : https://github.com/yaml/pyyaml/blob/4c570faa8bc4608609f0e531...
% cat countries.yml
---
countries:
- US
- GB
- NO
- FR
% yamllint countries.yml
countries.yml
5:4 warning truthy value should be one of [false, true] (truthy)My personal favorite is TOML, but I would even prefer plain JSON over YAML
The last thing I want at 2 AM when trying to look figure out if an outage is due to a configuration change is having to think if each line of my configuration is doing the thing I want.
YAML prizes making data look nicely formatted over simplicity or precision. That for me, is not a tradeoff, I am willing to make.
JSON:
- no comments, unless you fake them with fake properties, unless your configuration has a schema that doesn't allow extra fake properties
- no trailing commas; makes editing more annoying
- no raw strings
YAML:
- the automatic type coercion
- the many ways to encode strings ( https://yaml-multiline.info/ )
- the roulette wheel of whether this particular parser is anal about two-space indentation or accepts anything as long as it's used consistently
- the roulette wheel of whether this particular parser supports uncommon features like anchors
TOML:
- runtime footguns in automated serialization ( https://news.ycombinator.com/item?id=24853386 )
- hard to represent deeply-nested structures, unless you switch to inline tables which are like JSON but just different enough to be annoying
"You need this special tool to work" immediately and instantly rules out "easy to edit". Or makes the debate irrelevant: every format is easy to edit if you have "a convenient UI" to do it for you.
And plain text editor is a "widely deployed special tool to work". Actual data is
countries:\n- GB\n- IE\n- FR\n- DE\n- NO
Or 636f 756e 7472 6965 733a 0a2d 2047 420a
2d20 4945 0a2d 2046 520a 2d20 4445 0a2d- A proper editor was never around.
- Closing tags were verbose.
- Attributes vs tags was confusing.
- It didn't map "naturally" to common data types, like lists, maps, integers, float, etc.
XML is serialization. I hardly believe you was concerned about serialization while posting comment or thought about attributes-tags distinction.
This page utilizes request to server for multi-user editing. But it is easy to build truly serverless (like a file) document with same interface:
data:text/html,<html><ul>Host: <span class=host contenteditable>example.com
Change it, save it, done. Web handles input of lists, maps, integers, float and much more. 636f 756e 7472 6965 733a 0a2d 2047 420a
2d20 4945 0a2d 2046 520a 2d20 4445 0a2d
It is not practical to edit Excel documents in plain text: <?xml version="1.0"?>
<Workbook xmlns="urn:schemas-microsoft-com:office:spreadsheet"
xmlns:o="urn:schemas-microsoft-com:office:office"
xmlns:x="urn:schemas-microsoft-com:office:excel"
xmlns:ss="urn:schemas-microsoft-com:office:spreadsheet"
xmlns:html="http://www.w3.org/TR/REC-html40">
<Worksheet ss:Name="Sheet1">
<Table>
<Row>
<Cell><Data ss:Type="String">ID</Data></Cell>
Tim Berners-Lee browser was browser-editor. Can't you see parallels?* Text AND binary so that humans can edit easily, and machines can transmit energy and bandwidth efficiently.
* Carefully designed spec to avoid ambiguities (and their security implications).
* Strong type support so you're not using all kinds of incompatible hacks to serialize your data.
* Versioned, because there's no such thing as the perfect format.
* Also, the website is 32k bytes ;-)
So everything else needs some kind of initiator and/or container syntax to logically separate it from the other objects when interpreted by a human or machine.
Any reason for not using RFC2119 keywords in the spec? Using them should make the spec easier to read.
Unquoted strings are much nicer for humans to work with. All special keywords and object encodings are prefixed with sigils (@, &, $, #, etc), so any bare text starting with a letter is either a string or an invalid document, and any bare text starting with a numeral is either a number or an invalid document.
> Any reason for not using RFC2119 keywords in the spec? Using them should make the spec easier to read.
I use a superset of those keywords to give more precision in meaning: https://github.com/kstenerud/concise-encoding/blob/master/ce...
An important feature of RFC2119 keywords is that they're always capitalized (ie. the keyword is "MUST", not "Must", or "must"). This makes requirements and recommendations stand out amid explanatory text, improving legibility. For example, RFC2119 itself uses MUST and must with different meanings.
Because strings can contain whitespace and other structural characters that would confuse a parser.
> Having two representations for the same data means you can't normalize a document unambiguously.
The document will always be normalized unambiguously in binary format. The text format is a bit more lenient because humans are involved.
The idea is that the binary format is the source of truth, and is what is used in 90% of situations. The text format is only needed as a conduit for human input, or as a human readable representation of the binary data when you need to see what's going on.
> An important feature of RFC2119 keywords is that they're always capitalized (ie. the keyword is "MUST", not "Must", or "must").
Hmm good point. I'll add that.
+ Avoids ambiguities.
- The format seems to feel the need to support everything, including things I am not sure are actual usecases (what's the point of Markup element for example? What does Metadata save us compared to just including it in document, given that parsers must parse it anyway?). This must make implementation most complex and costly, and makes reading the text format more difficult.
- Not a fan of octal notation. At 3am not sure I can't confuse 0 and o given certain fonts. Does anyone even use it these days?
- Unquoted string were discussed in the thread, I'd like to point out that it's very easy to make an unquoted string not "text-safe" (according to the spec) without noticing it, at which point document is invalid.
Just add white-space (maybe a user pasted a string from somewhere without noticing whitespace at the end or forgot the rules), a dot, an exclamation or a question mark. Having surprises like that is IMHO worse than a consistent quoting method.
Basically all the things I don't like are about the format supporting a bit too much. YAML 1.1 should teach us more is sometimes less.
I put in octal because it was trivial to implement after the others. The canonical format when it's stored or being sent is binary, and a decoder shouldn't be presenting integers in octal (that would just be weird). But a human might want octal when inputting data that will be converted to the binary format.
Markup is for presentation data, UI layouts, etc, but with full type support rather than all the hacky XML+whatever solutions that many UI toolkits are adopting. Also, having presentation data in binary form is nice to have.
I saw there's a 'Media' type in the spec. It's seems the type is actually for serializing files. But there's no "name" (or we can call it "description") field. Of course we could accomplish this with a separate field - but than again the entire type's functionality could be accomplished with a u8x array and a string field. So if you're specifying this type at all, might as well add a name field to make it useful.
The media object is a way to embed media data directly into a document such that the receiving end will have some idea of how to deal with it (from its media type). It won't have or need a "file name" because it's not intended to be stored in a filesystem, but rather to be used directly by an application. Yes, it could be built up from the primitives, but then you lose the canonical "media" type, and everyone invents their own incompatible compound types (much like what happened with dates in JSON and XML).
- I'm removing the metadata type. You're right that it's not really gaining us anything.
- I'm changing strings so they always must be quoted. This actually simplifies a lot of things.
Thanks for the critique!
You wouldn't serialize data structures to jsonnet though, you'd just generate JSON.
Only when you "unmarshal" to an untyped data structure and then make assumptions about the type. I've used yaml with a go application, and it can't interpret NO as a bool when the field is a string.
When it gets hairy is that most programming languages have low entrance barrier. To write Haskell effectively you’ve got to unlearn a lot of rooted bad habits and you get to dive into the “mathematical” aspect of the language. Not only you got monads, but there’s plethora of other types you need to get comfortably onboard with and the whole branch of mathematics talking about types (you don’t need to even know that such a field as category theory exists to use it).
However, since most people just want to write X, or just want hire a dev team at price they can afford, Haskell rarely is the first choice language.
1. Really annoying to do any kind of i/o
2. Extremely poor interoperability with non-Haskell code
3. (opinion) Unpleasant, inconsistent, hairy syntax
https://news.ycombinator.com/item?id=26679728
> the article refers to YAML 2.O, a nonexistent spec, and to PyYAML, a real parser which supports only YAML 1.1.
> Both the unquoted-YES/NO-as-boolean and sexagesimal literals were removed in YAML 1.2.
And these sort of things happen time and time again.
And although officially JSON requires quoted strings, almost none of the parsers actually enforce that, and so you will find a huge amount of JSON out there that is not actually compliant with the official spec.
Just like browsers have huge hacks in them to handle misformed HTML.
What programming language? I'm not familiar with those parsers, the ones I know of very much do enforce quoted strings.
> you will find a huge amount of JSON out there that is not actually compliant with the official spec
The parsers I use all follow the current JSON RFC specification, and I've never encountered any JSON from APIs which they reject.
> Just like browsers have huge hacks in them to handle misformed HTML.
Web browsers do deal with a variety of things, not so much JSON parsers in my experience.
I've never seen one
If you want something to be a specific type, you better have an explicit way of indicating that. If you say quotes will always indicate a string, great. Of course we know it's not that simple, since there are character sets to consider.
The safest answer is to do something like XML with DTDs. But that imposes a LOT of overhead. Naturally we hate that, so we make some "convention over configuration" choices. But eventually, we hit a point where the invisible magic bites us.
This is one case where tests would catch the problem, if those tests are thorough enough - explicitly testing every possibility or better yet, generative testing.
Does Yaml have any sort of strict mode?
I imagine I could find a linter that disallows implicit strings.
I don’t think that applies to Python - it’s quite strongly (although not statically) typed. I agree that it does apply to JavaScript and PHP.
You're not looking really hard then, but really
> When things get hard on bash, you will start to see python scripts
That's kinda the thing innit? Unless the system specifically only allows shell scripts (something I don't think I've ever encountered though I'm sure it exists) it's quite easy to just use something else when bash sucks, so while people will absolutely complain about it they also have an escape: don't use bash.
When a piece of software uses YAML for its configuration though, you don't really have such an option.
Furthermore, bash being a relatively old technology people know to avoid it, or what the most common pitfalls are. Though they'll still fall into these pitfalls regularly.
Powershell has been working on linux for quite a while now and doesnt seem get any attention even when it has a nice IDE support and copy the good things about bash.
The reason people are comfortable with the POSIX shell is because you use the same syntax for typing commands manually as you do for scripts. But, you're going to have a hard time finding people who prefers writing:
Remove-Item some/directory -recursive
Rather than rm -fr some/directory
People who write shellscripts are often not seeing themselves writing a "program". They are just automating things they would do manually. Going to an IDE in this case is not something you'd consider.I happen to be very aware of all the pitfalls in POSIX shell, and it's rare that I see a shellscript where I cannot immediately point out multiple potential problems, and I definitely agree that most scripts should probably be written in a language that doesn't contain so many guns aimed at the user's feet. I'm just pointing out a likely reason why people are not adopting powershell in the huge numbers that Microsoft may have hoped for.
rm -r -f some/directory1. mainstream
2. programming language
(of course technically it is a programming language, but it is also more precisely a scripting language)
For example, Ruby's standard library only supports YAML 1.1. It relies on libyaml, which is not yet compliant with 1.2. Meanwhile, Python's popular PyYAML library only supports 1.1, and asks users to migrate to a newer fork called ruamel.yaml for 1.2 support.
This is an article justifying use of (and justifying design decisions of) a particular Python quasi-YAML parsing library. If you are in a position to select a non-YAML-1.1-compliant parsing library for Python, or to take the articles advice on design of a YAML(-ish) parsing library, you are, necessarily, not stuck with YAML 1.1.
> for whatever reason
Articles like this spreading misinformation about the current state of standard YAML are part of the reason. LibYAML lagging support is another since so much of the ecosystem depends on libYAML (though, while the documentation situation is terrible, it looks like maybe libYAML has some level of 1.2 support since 0.23.)
> For example, Ruby's standard library only supports YAML 1.1. It relies on libyaml, [...] Python's popular PyYAML library only supports 1.1
Which, also, is dependent on libYAML.
> and asks users to migrate to a newer fork called ruamel.yaml for 1.2 support.
Which makes a lot more sense than migrating to a library thar supports neither 1.1 nor 1.2, but a nonstandard variant that addresses some of the same issues resolved years ago in 1.2, especially when a library supporting 1.2 is available for the same language.
It is interesting how the standard of any language seems to diverge due to just the implementation from different parsers.
- “modified” numbers, e.g. $50, 35%, 1.2345568896347853246863477
- Dates. If your language tries to convert a date to Unix time or Julian Day, you can have problems with time zones or distant or historical dates.
- strings vs symbols. The person writing config shouldn’t have to care about this distinction.
- Automatic deduplication for fields of objects can be a problem.
The Stripe library has constants for which type of VAT number is supplied. One of those constants is 'NO_VAT'...
Needless to say, this caused me some grey hairs
https://www.theguardian.com/politics/2013/apr/18/uncovered-e...
https://www.bbc.co.uk/news/magazine-22223190
https://theconversation.com/the-reinhart-rogoff-error-or-how...
I cry and rage and rend my clothes when I stumble upon some new thing that makes me have to use it.
And that's why you have a staging environment. Or you debug in production, whatever you prefer.
https://mobile.twitter.com/stahnma/status/634849376343429120
I say this because two days ago I wrote a test that used all country codes as input. It took 15 minutes to write that test. During the whole testing session I found at least 5 mistakes of which 3 would have been quite dramatic.
And how many minutes to test all city/state/region/street/person names ?
It can also happen that you test s will become outdated, like when url standard changed and more characters codes were allowed.
Testing doesn't take too long on my machine (maybe 10 seconds), but even if it would, it would be totally acceptable as I run it pre commit only.
What matters is how you deal with incidents as an organisation, not that you should never release a bug.
The real problem is that YAML parsers in wide use have not been updated to the spec that was released TWELVE years ago.
So who's going to help the common YAML parser developers update their implementations to support version 1.2? I think that would be a big help. Maybe the Norwegian government can chip in some money & time to get them updated, that would probably quietly eliminate a number of problems.
> The primary objective of this revision is to bring YAML into compliance with JSON as an official subset. YAML 1.2 is compatible with 1.1 for most practical applications - this is a minor revision. An expected source of incompatibility with prior versions of YAML, especially the syck implementation, is the change in implicit typing rules. We have removed unique implicit typing rules and have updated these rules to align them with JSON's productions. In this version of YAML, boolean values may be serialized as “true” or “false”; the empty scalar as “null”. Unquoted numeric values are a superset of JSON's numeric production. Other changes in the specification were the removal of the Unicode line breaks and production bug fixes. We also define 3 built-in implicit typing rule sets: untyped, strict JSON, and a more flexible YAML rule set that extends JSON typing."
Since "no" is not the same as false, the Norway problem disappears. It's safer to always quote single-word strings like 'this', just like you always have to quote all strings in JSON.
As far as I can see, this has nothing to do with typing and everything to do with syntax (of literals). If strings were required to be quoted this problem wouldn’t appear.
This is the reason no programming language has this issue — regardless of type system (JS/Python/Java/Haskell). If you want a string here you need quotes.
Haskell could even be regarded as what the author calls “implicitly typed” — since types are derived from literals — and I’ve never heard a Haskeller complain about this issue.
I must say that I feel a little bit of relief to see that they have problems that nobody else has, besides insanely expensive alcohol that is only sold in "wine monopoly" stores that are more heavily guarded than banks.
Their product used an MVCC database (I think ObjectStore). One of their customers in Norway had a problem where updates to the database seemed to not show up. IIRC the problem was a bug in this company's software that caused MVCC to show an older version of the database content than expected.
Or… just quote your strings.
I used it for configuration of a Go program recently and found it pleasant to work with. I hope the language is declared stable soon, because it's a good model.
I understand that people don't like directly use JSON because it's not very friendly: no comments, no multi-line string, etc.
A great alternative IMHO is cson[0]. It's like JSON to JavaScript but for CoffeeScript (though nobody talks about it nowadays). It has indentation-based syntax, comments, and multiline string which usually don't need to escape. The advantage is it's close enough to JSON which is the canonical format that everybody can agree on nowadays. For YAML and TOML there are too many visual part-aways from JSON.
Or just create a JSON variant that enables comments and the backtick multiline string from JavaScript.
Aren't you setting yourself up for surprises if you write file formats such as TOML and YAML without reading the documentation, learning and experimenting first? How about unit testing? Or verifying the type in your config parser? Have you tried opening your site in the norway config on your development or testing environment? Or even in production? It all seems very basic and not at all blog post or even HN worthy.
I'm going to assume the authors still haven't learned their lesson and are going to experience many more surprises in the future working with plain text file formats.
YAML though is always a bad fit. If you want machine readable config, use JSON; human readable, use TOML. When does YAML ever fit?
This one made chuckle, and TIL that Null is a real life surname.
I finally solved it by exporting to csv, and using third-party software that handled its own import and did it correctly.
The problem is that if a YAML parser sees a string like this:
"0123e04"
It interprets it as a number: 123 * 10^4
Our hacky solution was to prefix the revision hashes like sha-0123e04, but still this was quite annoying.
After that experience, I have stopped using YAML for any of my own configuration. Have started preferring putting my configurations in code. And when I don't want that, have found JSON good enough for my purposes.
The point is that this behavior is sporadic. It doesn't apply consistently across all git hashes, which is the real problem. It is easy to be caught unawares by this behavior.
Literally the only thing we were doing was passing it between shell commands, helm charts, Kubernetes deployments and then back (if we needed to debug).
It sounds like you have a more attractive alternative in this case than to treat hashes as strings. Would love to hear it.
https://news.ycombinator.com/item?id=26679590
Concise encoding seems to have an hex-int type ?
The number one rule when creating a serialisation format should be that serialisation and deserialisation is predictable. It's quite remarkable that two of the most popular formats doesn't do this.
I'm actually surprised we haven't seen any major security issues caused by this.
Hopefully not a real story. If you’re trying out new configurations in production and have no mechanism to rollback problematic changes, you’ve got bigger problems than YAML.
To me, though, YAML, including “StrictYAML” doesn’t solve any problems JSON, perhaps w/comments, already solves.
In the first model entering
GB 9.3
gets you a string and a number.
But the second gets you two strings?
Both are wrong in my opinion.
"GB" 9.3
is the correct approach
Explicit beats implicit every time.
2020-03-25 -> datetime.date(2020, 3, 25), not "2020-03-25"
The article author hitchdev does not say it outright, but it is heavily implied that the YAML file was edited by hand. This is the immediate cause of the problem. The indirect root of the problem is that the spec authors chose a plain text serialisation format and thus created an affordance http://enwp.org/Affordance#As_perceived_action_possibilities to be edited by hand.
This turns out the be unsafe/source of bugs because YAML end-users are not capable of correctly applying the serialisation rules considering the edge cases detailed in the article because humans are creatures of habit, applying analogy and common sense, making assumptions and then sometimes go wrong, whereas a piece of software will not make the Norway, Null etc. mistakes. hitchdev even writes that quoting the string is "a fix for sure, but kind of a hack", but that's a grave misunderstanding. Quoting the string here is actually applying the serialisation rules correctly.
The tangential at the end of the article about typing is also orthogonal/irrelevant. YAML is strictly/strongly/unambiguously typed, and so is the mentioned variant Strict YAML. The difference is that Strict YAML has serialisation rules that are more amenable to or aligning with the human factors of habit etc. and thus work better in practice.
My personal recommendation is to never edit YAML by hand and always use a serialiser. This is less convenient, but safe.
In closing, I would like the reader of this comment to make an effort to distinguish between "what is" and "what ought to be" in their head, otherwise the ideas here will be very muddled.
I don’t follow this. If yaml is your config format, and you are not editing it by hand, what are you editing?
This is not some interesting trade-off, this problem is fixable on all axes by using non-ambiguous, non-overloaded typing rules for your config format.
Even JSON and XML got this right.
Yes, yes, I pointed that out. grep "immediate cause" and "indirect root"
> the serialization rules are quite terrible
Did that need to be said explicitly? I agree FWIW. I have already made a value judgement mildly against YAML, in case that's not clear. It's only mild because the problem can be worked around. I think this approach is more practical than moving the whole world over to a completely different thing.
> problem is fixable […] non-ambiguous […] rules
Is the implication here that you say YAML is ambiguous? It's not. I don't want sloppy analysis. To be precise, the ambiguity is imagined, it does not exist on the spec or software level, only in the head of people.
The article author also misidentifies the version of the YAML spec (calling it 2.0, which doesn’t exist; the behavior is from YAML 1.1, and this class of problems motivated a bunch of changes in YAML 1.2, which has been out since 2009.)
But the article author isn’t trying to analyze the problem, he’s trying to rationalize why what is notionally a YAML-processing library just ignores the spec.