…only because NestedText does not support numeric types at all. That seems like throwing out the baby with the bathwater.
…only because NestedText does not support numeric types at all. That seems like throwing out the baby with the bathwater.
One should remember that any sane application will be parsing the config file into internal data structures and validating it anyway so it gets little benefit from the numbers being already “parsed”.
There are also issues when something looks numeric but doesn’t parse (eg 1.2.3, 3/2, 12in, 4h30m2s, 2:30, 2020-02-29, etc). One way to deal with these is a tokenisation rule like in Common Lisp: if it is a valid number syntax then treat it as a number, otherwise it’s a symbol, but this can lead to issues (eg you would need to know that when your number needs more than float precision or otherwise doesn’t follow the rules, it should be in quotes. It seems crazy to pass that detail on to the poor sod who has to write the config file).
The benefit of standard numeric and boolean types is that different tools can exchange data in a well-understood way.
Getting rid of yaml's 30 ways to write "true" and "false" by making everything is a string just means that you now have 30 tool-specific ways to write "true" and "false".
The "everything is a string" approach already exists in shell scripts and TCL and it's not really that great.
Unless you have actual type annotations/tags (eg xml, jsonld, graphql), everything IS a string. There's no assumptions otherwise.
http://p3rl.org/guts#Magic-Variables
perl -MDevel::Peek=Dump -mTie::Scalar -e'
//g; Dump $_; tie $c => "Tie::StdScalar"; Dump $c; Dump \%ENV
' 2>&1 | grep MAGICPerl tied-variable magic just means there are (effectively) getter and setter properties attached to the variable. "Magic" is just the name that was chosen in the implementation, and it stuck.
It's used to implement variables with special, automatic meanings, like $$ for "current pid" and $! for "last error".
It's also used to implement variables with user-defined behaviours on access, which is quite handy for a lot of abstractions.
A lot of modern languages support both of these things, because they are useful, but it's not called magic in those languages, it's called something like "watchers", "proxies", "getters and setters" or "hooks".
No, the criticism of YAML-style "magic" is that it leads to entirely surprising behaviour from innocuous input. Perl magic is not that kind. If you're using a special variable, you already know why.
In fact in Perl, you can opt for longer, readable, lexicon over the terse single character variables; and that’s literally how modern Perl should be written.
Whereas the problems described with YAML is where it can automatically alter your data based on what the parser “thinks” the data should represent. Which is generally what people mean when they talk about “magic” in IT: systems that don’t honour your input and instead automatically convert it into something else. Perl doesn’t do this even in spite of it looking like executable line noise to many.
If you take a look at how string handling works in most programming languages there is a lot of "magic" going on there. Which isn't a bad thing necessarily, because most programmers don't want to deal with the intricates of strings unless they really have to. The key is that this magic doesn't get in your way and doesn't do too magical things nobody ever asked of it.
https://yaml.org/spec/1.2/spec.html#id2805071
Of course, even though YAML 1.2 is a decade old, there are still many parsers that accept YAML 1.1.
YAML has actual type annotations (tags).
I've never seen typed yaml, this is wild.
negative: !!int -12
zero: !!int 0
positive: !!int 34
Can't say I love the notation, but indeed that is type annotations. I guess neither "yaml type hints" nor "yaml type annotations" are the right query. Had to search explicitly for "yaml tags".https://docs.aws.amazon.com/AWSCloudFormation/latest/UserGui...
In Python:
>>> import json
>>> json.loads('{"x": 9007199254740993}')
{'x': 9007199254740993}
In my browser's JavaScript console: > JSON.parse('{"x": 9007199254740993}')
{x: 9007199254740992}
(Consider what happens if you try to send a tweet's ID, a perfectly normal number like 205052027259195393, through JSON. Or if you try to serialize a stack trace on a 64-bit system, where addresses are also perfectly normal numbers.)Javascript number is a floating point
Number.MAX_SAFE_INTEGER
// 9007199254740991
Number.MAX_SAFE_INTEGER + 1
// 9007199254740992
Number.MAX_SAFE_INTEGER + 2
// 9007199254740992
no automatic promotion to BigInt 9007199254740991n + 2n
// 9007199254740993nThe JSON standard doesn't place restrictions on size or precision of numbers, instead just noting that implementations can vary their treatment of and limits on numbers. While JS uses doubles for all numbers, many other languages emit an integer type for a JSON integer. So, once you go beyond the range where a double can accurately represent all integers, you run the risk of a mismatch in how the number is interpreted by different languages parsing the same JSON.
Of course the spec also allows you to create way too big or too precise numbers that would be problematic in most languages as well; it's just that this is a somewhat common bugbear.
I wouldn't necessarily call it a flaw in JSON though, more an issue with JSON.parse or really just a fact of life when dealing with numbers in JS. Alternatives to the built in JSON.parse exist to read large integers as strings or bigints.
https://developer.mozilla.org/en-US/docs/Web/JavaScript/Refe...
Described issue is not problem of a JSON but engine which parsed it and language which stands behind the parser. Any config format will eventually have same result and same issue.
So forcing programmer to parse every single piece of data for sake of "it's his responsibility" is not a case here.
I also disagree this is in any way programmer responsibility to create standardized way of creating parser for everything. This format gives you nothing but indentation so you are forced to create documentation for everything field, what type it's and what kind of values it takes. Lots of extra work for nothing when you have any other format.
I mean, there's a pretty clear argument here: Twitter themselves used to return these numbers as numbers in the API until they realized they were about to hit this problem. https://developer.twitter.com/en/docs/twitter-ids
That's an incredibly weird complaint when the real problem is that javascript's JSON.parse doesn't use BigInt for large numbers.
I don't think that's what this tool is for. This tool is for humans to read and edit data. That's a different use case from automated programs exchanging data, for which I agree you should be using standardized numeric and boolean data types and not making everything a string. But how that standardized data gets determined from data that humans enter should be up to the individual application.
> A key that requires quoting must not contain both single and double quote characters.
You can't really serialize user data with that restriction.
Maybe we don't need it, but it often helps.
Also, this data format isn't necessarily for humans to exchange data with other humans, but for humans to give data to applications in a format that's much easier for humans to use.
> if this is insufficient to communicate it to the machine
Not at all. Each specific application can easily parse this data format according to its own needs. What this data format doesn't specify is a single translation into application data that is the same for all applications. But applications don't need or want that, because they have different use cases.
If you make everything a string then the interpretation of "no" as boolean true or false is left to each tool, and there are even tools which have different interpretations of "yes/no" for each field.
Most likely if their input is meant to be machine interpreted they would need to be trained to provide specific inputs anyway. I like that NestedText doesn't hide that problem. It lets the user organization decide how it wants to manage that problem, and what symbols or words are understood by the people authoring the files.
ISO8601 joins the chat.
The nice thing about that is it solves the problem rather than hoping it doesn't matter or assuming each program's validator will think to note all of the possible data types not compatible with the program natively. If a u32 is defined in the file and you've only got doubles to work with it's a given you'll have to deal with it in your tool specific validation. For everyone else it's well defined.
The downside is it's a bit more verbose and if you have all of that info already it's pretty easy to jump to just using a binary format which will be more efficient anyways.
The spec requires that a full 64 bits of signed integer be parsed and understood, that hexadecimal, octal, and binary values be integers, and that floating point values be parsed as doubles.
It doesn't support hexadecimal float, however, which is a pity: having a guaranteed bit-identical format is a nice affordance.
Too many types could get overwhelming, but I like where amazon's Ion is [1]. It actually supports multiple number types, with decimal being the default for values with a dot.
> you would need to know that when your number needs more than float precision or otherwise doesn’t follow the rules, it should be in quotes
Not really. The configuration value should either be a number or not, which is determined by the application reading the config. As a config writer you only care to make the type match (so, in json, if the application uses number you make sure you use number, and if it expects a string you use that)
(disclaimer: I work for amazon, but have nothing to do with Ion other than having used it. Opinion is my own, not my employer's, yadda yadda)
Only if this format is intended for use-cases that never need to deal with numbers.
> One should remember that any sane application will be parsing the config file into internal data structures and validating it anyway so it gets little benefit from the numbers being already “parsed”.
That statement couldn't possibly be more wrong.
Number parsing (and encoding!) is a decidedly non-trivial problem. You need to concern yourself with -- at a minimum -- all of the following:
- Unsigned 64-bit numbers.
- A series of digits that would be bigger than a 64 bit whole number. Convert to float? Truncate in some way? Error?
- NaN
- Infinity
- Negative zero
- Denormal numbers.
- Differentiating between decimal/currency types and floating point numbers. Not all decimal values can be exactly represented as floats!
- Efficiently encoding floating point to use the minimum digits without losing precision.
- Parsing those minimal numbers with perfect "round-tripping".
- Doing the above efficiently.
- Securely too! Efficient parsers cut corners on sanity checks. I hoped you fuzzed your parser...
The above can easily amount to many kilobytes of extremely complex code. Look up "ryu" as an example of what Google came up with to make JSON number parsing reasonably efficient.
Meanwhile, reading a fixed-length number from a binary format can be done in a single machine instruction. One. It might not even take an entire CPU clock cycle! Okay, two, if you need to bounds-check your buffer, but there's ways to avoid that.
Afterwards, the bounds check is again literally just two machine instructions in complexity. That's not the difficult bit!
The difficult bit is the parsing.
1. Declare numbers as numbers in the configuration language. E.g. "decimal(1e1000)".
2. Parse declared numbers with a lossless format like Python's decimal.Decimal.
3. Let users decide at their own risk if they want to convert to a lossy format like float.
That doesn't matter at all. The author's aims will be ignored if this format is used for anything even vaguely important. Eventually it'll need tooling to both read and write it.
DevOps pipelines, applications with GUIs, or something will need to both parse and generate this format in a consistent way.
There is no such thing as a human-write-only format in widespread use.
Even programming languages are regularly generated by tools such as RPC API codegen tools, LINQ-to-SQL and the like.
One example you provide is decimals for currency values but I claim you would want such values to look like $1234 in config files so that when they are reviewed or written, the person reading the file knows they are looking at a dollar value and can be concerned if it is too large.
I’m not suggesting that applications write their own number parsing. Just do uint64::parse or parseInt or Double.of_string, or whatever else you need to access your language’s number parsing routines.
> Just do uint64::parse or parseInt or Double.of_string, or whatever else you need to access your language’s number parsing routines.
Okay, so the computer is doing the parsing.
Those functions are notoriously inconsistent in their behaviour, particularly across different programming languages. If you're not careful, you'll end up accidentally using the internationalised versions of those functions. Even if you're careful, other people won't be.
Remember, data formats are for interchange. They have to be language agnostic. They have to be well-defined, and it should be possible to write a parser for them without having to guess at the precise details.
The harmful consequences of the Robustness Principle are now well-recognised in computer science: https://tools.ietf.org/id/draft-thomson-postel-was-wrong-03....
Some things need to be done properly, nor not at all.
I am far more worried about localisation issues than language issues. If you are storing something central to multiple applications I'd argue a text file is the wrong tool
And - it is certainly OK in many instances to have fixed-width, fixed byte-order binary encoding as the format's basis. It comes with the twin downsides of wholly different categories of errors cropping up, and with the lack of a universally agreed upon tool for human entry.
Perhaps text was a fashion, though. I definitely have had thoughts in that vein lately. And in that case we shouldn't always be rushing to use it as the source of truth when we have many good, machine-level agreements about numeric formats.
It seems you're mixing up the language definition with implementations that try to follow the language definition.
If different implementations have different results then either they are buggy or the language has some important holes in the specification.
Either way,the solution to this problem is not less validation.
> One should remember that any sane application will be parsing the config file into internal data structures and validating it anyway so it gets little benefit from the numbers being already “parsed”.
That's the whole point, isn't it?
I mean, if you already acknowledge the fact that this parsing and validation is a basic requirement, why handle it as an afterthought and force developers to add their own hand-rollef absurd and unnecessary type checks and type coversions?
Wouldn't it simply easier to let the language and the parser do that already?
I mean, no one ever complained that JSON had string types. In fact, one of JSON's main complaints is that it doesn't support enough types, such as timestamps.
name(str): Dave
age(int4): 22
dob(date): 2020-02-01
photo(base64):TWFuIGl....It's less of a problem if you're using it for configuration files, where a program knows what key's values need to be cast to an integer or float.
But it seems disastrous if you wanted to use it for storing or transmitting data, above all between applications. You're immediately throwing out the possibility of being able to serialize and then deserialize data in basically any programming language.
I shudder at the idea of an API that accepted NestedText, where I'd need to worry about whether my floating-point output was compatible with its floating-point string parser. Yikes. I want the serialization format to handle that. Isn't a major criticism of JSON that it doens't have a built-in datetime representation?
As for stringification, JSON's data types are mismatched with pretty much everything that isn't JavaScript to some degree. If you need to serialize a 64-bit integer to JSON, you serialize it as a string because the parser on the other end is probably going to try to parse it as a double-precision floating point number. Once you've started serializing numbers as strings anyway, it's not too far to "serialize every scalar as a string".
Indeed, not being able to have lists of dictionaries, or lists of lists, is very restrictive. Seems to be for very simple configurations only. E.g. a set of preferences, but not a set of monitor calibrations. (Inventing arbitrary dictionary keys seems pretty hacky.)
I don’t think this is true. None of the examples have lists containing non-string objects, but the documentation doesn’t seem to draw a distinction between lists and dictionaries wrt what can be placed in them.
(Both lists and dictionaries are initially described as only containing strings, and later this description is expanded to include nesting; this counterintuitive arrangement may explain the confusion.)
Frankly, in most languages, this is better because you don't have the types of objects randomly change based on user input. (In a few languages, with a few libraries, you can specify the type of the document to the parser and have it fail to parse the entire document if it can't deserialize to the right type, in which case this is a little weaker. But you can still do that with NestedText, just one step after the parser - have your own function that takes a ComplexStructure<..., String, ...> and returns a ComplexStructure<..., int, ...> or throws an error.)
Only in untyped or dynamically-typed languages. But even in JavaScript one may write +obj.version instead of obj.version to make it numeric. Evading this is a straight way to hell.
In a statically typed one, conversion is typically generated from description and type checking applies just at reading.
The problem with vaguely specified format is in more simplex cases. Shall we accept 45x as number (and what it value will be, 45 or 0)? 045? 045x? What date is 1/2/3, 1-2-3? And so on.
I think the point is that the answers to those questions may well be application-specific. In which case it is better to not bake them into the file format.
- There's no way to say that you want an integer; you get a floating-point value.
- See https://news.ycombinator.com/item?id=24676484 , you can't reliably accept integers over 2^53 without taking them as strings.
- Someone can always specify something of the actual wrong type. (Imagine changing YAML "version: 1.9.1" to "version: 1.10". You can't just stringify 1.10, you'll get "1.1"!)
So, in a practical data format, the schema for your document needs to say something like "This is a number, which must be an integer between 0 and 2^16" or "This is a string, make sure to quote it" or whatever, and a generic statically-typed JSON- or YAML-parsing library isn't going to handle that for you. And telling your users "the input format is JSON" doesn't answer that question: you must make it explicit to users.
Fortunately, you can handle it just fine in a statically-typed language in one of two ways. One is to accept an object from your parser that consists of variant types and pass it through your own function that validates it against a schema, and then either returns a more-restrictively-typed object or throws an error. Such a function could easily do string conversion too if given NestedText input, as I mentioned. The other is to pass some information into your parser saying, don't act like a generic JSON/YAML parser, instead interpret these particular fields in this particular way and accept only things with this structure. If you're doing that, you can easily tell the parser to use this particular string-to-integer function on the strings in NestedText and then return an appropriately-typed object containing an integer to you.