Show HN: Weird JSON superset with BigInt, Infinity, TypedArray etc
github.com
github.com
For systems where you have modern stacks and control all the moving parts, something like Cap'n Proto will likely make everyone happier in the end. But you don't always have that luxury, and this project is very cool in its own right...
Although I am not sure I would go as far as to drop booleans. As for numeric types, that's a complicated issue, I think they should be supported for pragmatic reasons even though it's impossible to come up with a numeric type that satisfies everyone.
The real problems come in when you handle unexpected cases - and there's no real answer as to whether you return true, false or an error. Personally I think the responsibility should be thrown back at the coder. E.g. I will sometimes do:
if(!(jobj.get("key", "true").contains("f")))fun();
Where "true" is considered safe and default. Generally you need to make some effort to switch it off by making sure an "f" exists somewhere, but either case is fine.Imagine for example that the value "true" turns on something potentially dangerous, you might want to guard against accidentally switching it on with:
if(jobj.get("key", "false").equals("true"))fun();
So you for sure have to make sure the string is "true" to switch it on, nothing else.I would say so yes.
Pick null for example, depending on your language of choice, you might choose to utilize: "null", "NULL", "", "undefined", "\0", etc.
The same with booleans, are we speaking "true", "True", "TRUE", "1", "0xFFFFFFFF", or anything that isn't a false is true, etc.
Really how you want to handle these situations is entirely up to you. Personally I prefer some functions like:
string s = jobj.get(key, default)
The default value being used if a valid string could not be found. Your default could of course be null or equivalent (handle the error), or it could be some known safe value (the show must go on).That's an argument against your position not for. By having a type "null", there's no issue of interoperability between languages who write null in different ways.
For example: Did you set `{"address":null}` or did you set `{"address":0x00000000}`? It's important to know the difference, do you need to throw an error or warning? If your NULL value maps to something that is also valid input, how are you supposed to know how to handle the case? If you removed the "address" key on the other hand, you can simply do some exists() functions check.
I think putting types into the JSON spec only serves to create confusion. I think the JSON libraries should offer common helper functions for doing these conversions (e.g. string -> int), but ultimately it's the responsibility of the coder.
{ "name":"null", "age":"44" }
vs { "age":"44" }
In the first example, your get function would return "null" for "name" and in the second example it would return the value of default, which would be implementation specific.So "null" can no longer be someones username, which will cause weird edge cases...
> "null" is now a keyword that needs to be enforced.
So in the JSON itself, you literally remove the field. In another comment I make is the point that NULL itself could be considered a proper value, you're just pushing the problem down the line. The library making such a decision for you is what may ultimately trip you up.
Instead, you could have something like:
string name = jobj.get("name", NULL)
So that it does return NULL in this case if a name hasn't been given.But consider an example where you're loading some code into a VM's RAM:
int location = toint(jobj.get("address"))
NULL in this case is a perfectly valid place to load a program in a VM, or was it that there was an error? Or is it that no address was given in the JSON configuration?Instead, we can specify our own "NULL" value:
int location = toint(jobj.get("address", "-1"))
So now you're asking to load something before the first position of the VM memory, which indicates an error.> So in the JSON itself, you literally remove the field.
That's a good point, but in some cases having a `null` is an easier option, as you only need to determine if the value is valid, not if it exists AND is valid (but I suppose you could say it's always valid if it exists... that's if you trust the endpoint to provide valid values though).
> But consider an example where you're loading some code into a VM's RAM
Allocating memory locations from dynamically created JSON values is bound to lead to many different errors...
I guess my main point is that JSON is inherently untyped and suited for messy data or user input. This messiness makes data validation on the consumer of the JSON necessary, which means the consumer has to check if the value is valid before using it. In most cases `null` is considered invalid but if the field isn't there, then it could be an error in the JSON provider.
I can't for the life of me understand why protocol specs omit this.
That spec is quite large, and everyone wants different parts of it. And different languages have different and somewhat incompatible support for datetimes.
And, worse, people want different behavior for invalid datetimes.
So your parser either balloons in size and requirements and its test-suite or restricts it to a handful of cases, at which point for most users, your datetime type is useless.
This is why if you have a protocol like JSON that allows you to construct arbitrary types or simply parse strings, you let users figure out how to handle it.
Actually JavaScript already did pick a representation - milliseconds since 1970. It would be easy and unambiguous to add { "foo": T1606410158672 } to the grammar.
- You get automatic conversion to/from a known datatype, Timestamp, not to an arbitrary number
- You get automatic checks on the validity of this datatype
- You can specify exact precision of this datatype (for example, you may want to represent nanoseconds, not milliseconds, since epoch)
The support for datetimes is irrelevant if the protocol spec specifies how to serialize a date. Converters for languages will appear or already exist.
Datetimes is one of the most common data types and yes, everyone ends up "constructing arbitrary types or simply parsing strings and letting users guess how to handle stuff". Remember .NET's handling of datetime? [1]
ISO 8601 was published in 1988. You'd think that protocols designed in the 21st century would at least know about it. You don't have to require all of it. A sensible expectation of date and datetime would suffice the vast majority of use cases.
[1] https://www.hanselman.com/blog/on-the-nightmare-that-is-json...
Consider JSON and how numbers are handled right now.
Per the spec, it's an arbitrarily long decimal, and some libraries do deserialize it to a BigDecimal class.
But many libraries in a chain of processing will treat it as a double, often libraries outside your control.
So you can't rely on them preserving any accuracy beyond what a double can handle.
Likewise, if you have a complex date-time representation, the spec can require all kinds of detail, but in practice, some of that data will be thrown out.
And worse, they won't throw an error, they'll truncate that data quietly.
You: no-no-no, it won't work, let us make it significantly worse. Let us not specify this, and let everyone specify their own ad-hoc solutions with manual parsing and checks which will fail in every which way imaginable and let people serialise datetimes as `"\/Date(1320825600000-0800)\/"` for all I care because ... ?
Why use a static alphabet and not a dynamic one, e.g. with a magic header hidden somewhere that automatically changes the alphabet once parsed for the following content, including other magic alphabet headers? /s
edit: Can my comment make it into testimonials, too, plz?