Amazon Ion – A richly-typed, self-describing, hierarchical serialization format
amzn.github.io
amzn.github.io
In that world, lantencies were so low that the response to your order submission would land in your front-end before you've had time to lift your finger off the enter key. I now work with web based systems. On days like this, I miss the old ways.
Having been in industry only 2 decades it amuses me how many times this gets rediscovered.
Imagine a system every message is a UID or DID (https://www.w3.org/TR/did-core/) followed by raw binary data. The UID completely describes the shape of the rest of the message. You can also transmit messages to define new UIDs: these messages' UID is a shared global UID that everyone knows about.
Once a client learns a UID, messages are about as compact as possible. And the data defining UIDs can be much more descriptive than e.g. property names in JSON. You can send documentation and other excess data when defining the UID, because you don't have to worry about size, because you're only sending the UID once. And UIDs can reference other UIDs to reduce duplication.
After studying it a bit, I'm certain this is how it's used inside Google (might also be mentioned elsewhere).
All you'd really need to do is to compile all protos into a repository (you can spit out the binary descriptors from protoc), then fetch those and decode in the client.
It'd actually be quite straightforward to set up,
See the “@0xdbb9ad1f14bf0b36” at the top of this capnproto file for example: https://capnproto.org/language.html
It’s a 64bit random number so it’ll never have unintentional collisions.
Also note that a capnp schema is natively represented as a capnp message. Pretty convenient for the “You can also transmit messages to define new UIDs” part of your scheme :)
I wonder if giving it a name based on the hash of the definition has been explored; like Unison [0] where all code is content addressable, but for just capnproto definitions. Is there a reason not to?
Of course you can get intentional collisions. The security model here assumes that anyone that wants to know your message’s ID can just ask.
Did you know that the Internet Protocol uses a 4-bit header to specify the format (v4 or v6) of the rest of the message? They should have used 128 bits. What a bunch of fools.
It'll have unintentional collisions if you ever generate more than 4 billion of these random numbers. That's not inconceivable.
I’ve probably generated 100 IDs over my lifetime.
On top of that you would not only need to generate the same ID, you would need to USE it in the same system where that is could have some semantics to not cause an error.
Cap'n Proto type IDs are not really intended to be used in any sort of global database where you look up types by ID. Luckily no one really wants to do that anyway. In practice you always have a more restricted set of schemas you're interested in for your particular project.
(Plus if you actually created a global database, then you'd find out if there were any collisions...)
If it's 64 bit, doesn't that mean you'd need to generate ~10000000000000000000000000000000000000000000000000000000000000000 (2^64) of those numbers to have a collision, not 2^32?
The birthday paradox is named after the non-intuitive fact that with just 32 people in a room you have > 50% of 2 people having a birthday on the same day of the year.
Slight correction: only 23 people, actually. So in every second football ("soccer") game, you have two people on the field with the same birthday.
For example, if you generate 2 numbers and they are the same, but are different to the capnproto number, that’s a collision but doesn’t actually matter.
EDIT: It does apply, I misunderstood what the number was being used for.
Capn'Proto is more like a formalization of C structs in that new fields are only added to the end. If memory serves, on the wire there is no tag, type or length info (for fixed size field types), and everything is rooted at fixed offsets
Protobuf uses tag-type-values, i.e. each field is encoded with a tag specifying the field number and some basic type info before the value. The type info is only just enough information to be able to skip the field if you don't recognize it, e.g. it specifies "integer" vs. "byte blob". Some types (such as byte blob) also have a length, some (integer) do not. Nested messages are usually encoded as byte blobs with a length, but there's an alternate encoding where they have a start tag and an end tag instead ("start group" and "end group" are two of the basic types). On one hand, having a length for nested messages seems better because it means you can skip the message during deserialization if you aren't interested in it. On the other hand, it means that during serialization, you have to compute the length of the sub-message before actually serializing it, meaning the whole tree has to be traversed twice, which kind of sucks, especially when the message tree is larger than the L1/L2 cache. Ironically, most Protobuf decoders don't actually support skipping parsing of nested messages so the length that was so expensive to compute ends up being largely unused. Yet, most decoders only support length-delimited nested messages and therefore that's what everyone has to produce. Whoops.
Now on to Cap'n Proto. In a given Cap'n Proto "struct", there is a data section and a pointer section. Primitive types (integers, booleans, etc.) go into the data section. This is the part that looks like a C struct -- fields are identified solely by their offset from the start of the data section. Since new fields can be added over time, if you're reading old data, you may find the data section is too small. So, any fields that are out-of-bounds must be assumed to have default values. Fields that have complex variable-width types, like strings or nested structs, go into the pointer section. Each pointer is 64 bits, but does not work like a native pointer. Half of the pointer specifies an _offset_ of the pointed-to object, relative to the location of the pointer. The other half contains... type information! The pointer encodes enough information for you to know the basic size and shape of the destination object -- just enough information to make a copy of it even if you don't know the schema. This turns out to be super-important in practice for proxy servers and such that need to pass messages through without necessarily knowing the details of the application schema.
In short, both formats actually contain type information on the wire! But, not a full schema -- only the minimal information needed to deal with version skew and make copying possible without data loss.
Can you even distinguish floats and ints without a schema in protobufs? I don't remember.
I really enjoy capnproto, flatbuffers and Avro and bounce between them depending on the task at hand.
Well... I would call it type information, just not complete.
> 32 bit values, 64 bit values and length prefixed values
In protobuf, most integer types are actually encoded as varints, i.e. variable-width integers, not fixed 32-bit or 64-bit. varint encoding encodes 7 bits per byte, and uses the extra bit to indicate whether there are more bytes. (It's not a very good encoding, as it is very branch-heavy. Don't use this in new protocols.)
> Can you even distinguish floats and ints without a schema in protobufs?
You can't distinguish between float vs. fixed32. But int32 would be varint-encoded, while floats are always fixed-width, so you could distinguish between those. (You can't distinguish between int32, uint32, and sint32, though -- and sint32 in particular won't coerce "in the obvious way" to the others.)
The really unfortunate thing is you can't distinguish between strings vs. nested messages (of the length-delimited variety). So if you don't specifically know that something is a nested message then it's not safe to try parsing it...
If the order submission process depends on the manual press on the enter key (+/- 50ms) is there any point to that though?
I think that post must be a few years out of date - and moreover by its own admission doesn’t even test hardly any “gaming” keyboards. There is a tremendous amount of competition in keyboards that has been building for the past 10 years.
Input latency is now a marketing thing like horsepower, and there are reasonably reputable [1] places and countless small time YouTube reviewers that test these things. It’s not like it is difficult to improve latency, and now that it is something that is competitively marketed it is delivered on.
[1] https://www.rtings.com/keyboard/tests/latency
Personally I think it’s a bit ridiculous. This fetishization with minimizing latency to now sub-ms levels doesn’t necessarily lead to better performance as many top level gamers do not use the lowest latency level keyboards. But that doesn’t change the fact that modern mainstream gaming keyboards can hit a latency far below 50ms.
- apple magic keyboard (? vs 2017) 15ms vs 27ms
- das keyboard (3 vs S professional/4 Professional) 25 vs 11/10ms
- razer ornata (chroma vs chroma/chroma 2) 35 vs 11.4/10.1ms
Interestingly it is not some simple uniform difference: the Apple keyboard does much worse in the rtings test, perhaps getting not much of a bonus from key travel compensation. But the das keyboard vs the razer that are 10ms apart on my link perform equally on rtings (but maybe I found the wrong model). I don’t have a good explanation for that discrepancy.
I still use it in my UDP game servers, with added packet id if message exceeds max datagram length and has to be split
I also miss working on those systems.
Serialization is platform-dependent (to make it a simple memcpy most of the time), and the schema is sent up front (but can be updated later, with in-bound messages at will). See the User Guide (http://binlog.org/UserGuide.html) and the Internals (http://binlog.org/Internals.html) for more.
Even an unoptimized version of this managed to get throughput in the 300K/s range.
Somehow it's the endpoint of my journey into serialization. Basically, avoid it if you need to be super fast. For most things though, it's useful to have something that you can read by eye, so if you're not in that HFT bracket it might be nicer to just use JSON or whatever.
It is UB if the underlying dynamic type is not compatible from the access type
As the bytes are coming from the kernel, you get to decide the dynamic type.
I think it should be pretty obvious how these are helpful and why they are needed no?
I can see why sum types would be useful in a schema or for the elements of a collection that is required to be homogeneous (ie. List<Foo|Bar>).
For what use case would one use custom sum types in a schema-less data format?
as for the wire format, a variant struct where you've only instantiated a single field will encode down to just about the minimum amount of information required.
One can always choose not to use (native) sumtypes if they are interested in extreme performance or compatibility.
But logically speaking, it is _good_ that it's a restriction that a sumtype can't just turn into a multiple-fields type. Because while my software (as the consumer) might still be able to deserialize it, the assumption that only one field is set would be broken and my logic would now potentially broken. Much better if that happens at deserialization time then later one when I find out that my data is incorrect/corrupt.
Non union fields can even be upgraded to unions later
Personally I find the protobufs "everything is optional!" Behaviour fucking insane and awful to deal with, but it is true to the semantics of its underlying wire format.
{a:1} {a:{b:2}} {a:4} {a:{b:4}}
There's no static type layer over top of this, so it's inherently up to interpretation and whatever type system you want to use to describe this data, to be able to express that the values of `a` can be of type `number | {b: number}`
Yeah, that's the problem. I mean, hey, why json? We could just use unstructured plaintext for everything and now we are free to do everything. But obviously that has its own drawbacks.
Having built-in support for sumtypes means better and more ergonomic support from libraries, it means there is one standard and not different ways to encode things and it also means better performance and tooling.
[{}, "hi", 1, 2, 3, "yo", {a: "bc"}]
In the same sense "1e-12" is not a number, it's a string. Yes, it's a string that encodes a number in a certain notion, but for alle the tooling, the IDE, the libraries, etc. it will stay a string.
`data X = A | B
[A, B, ...]`
Is a list containing a sum type: list[X]
Now, I wouldn't actually call it a list of any, I would say you proved my point for me. Your example is functionally the same as mine. I would give this example:
`[A, B, ...]`
and say that that is a list of sum types. You may say "no no no! Only now is it a list of sum types!":
`data X = A | B
[A, B, ...]`
But my point is that there is no JSON/Ion equivalent of your `data X = A | B`. Everyone in this comment tree is confusing the data itself with out-of-band schema over that data. "Sumtype" is nothing more than a fiction, or a schema. Saying that JSON/Ions don't support sumtypes is like saying JSON doesn't support "NonNegativeInteger" type. Sure it does! Here are some: 1, 2, 3, 10. What tooling or type system you use outside of the data itself to enforce constraints on the data types is orthogonal to the data format itself.
No one disagrees - it's just that we complain about this. We _want_ to have such an equivalent.
> Saying that JSON/Ions don't support sumtypes is like saying JSON doesn't support "NonNegativeInteger" type.
Correct. But your conclusion is wrong. You seem to assume that no one has a problem with the fact that JSON doesn't support a "NonNegativeInteger" type. But I at least would happily use a format that explicitly supports that.
I mean, look at ION. Json doesn't support the concept of (restricted) integers, but ION extends JSON and offers this type. That's great, because it means if a library reads an integer field, it can map it to an integer and knows that there are constraints.
This is a _very_ relevant issue. Many json libraries in the past have had bugs or could be ddos-ed by feeding them json with large numbers, since the json spec does not constrain the size of numbers.
In that sense, ION could have _also_ added support for "NonNegativeInteger" or sumtypes, or other specific types, but they haven't. And since sumtypes are very fundamental, we complain about it more than we would complain about the lack of "NonNegativeInteger".
[5,"hello",3] has the type list (int ∪ string), not list (int + string). You can emulate the latter by manually adding a tag, but native support is much preferable.
[
{tag: "a", foo: 1},
{tag: "b", bar: "hi", baz: 2},
{tag: "a", foo: 3},
{tag: "a", foo: 4},
{tag: "a", foo: 5},
{tag: "b", bar: "yo", baz: 6}
]It seems that sum types are the norm, actually.
Avro had unions in version 1.0 [0], which is from 2012.
Capnproto had unions back in 2013 [1]. That's from the v0.1 days, or maybe even earlier.
Protobuf has had oneof support for about 7 years. They were added in version 2.6.0, from 2014-08-15 [2]. That's still 6 years after the initial public release in 2008, though, so this is maybe what you were thinking of? I don't know too many people who were using protobuf in those days outside of Google, though.
---
[0] https://avro.apache.org/docs/1.0.0/spec.html#Unions
[1] https://github.com/capnproto/capnproto/commit/eb8404a157e074...
[2] https://github.com/protocolbuffers/protobuf/blob/master/CHAN...
And yes, I definitely am primarily thinking of protobuf, as I struggled with this back with version 2.5. I had the (apparently mistakenly) impression that Avro and Cap'n Proto (which I think actually first came out in this timeframe) were about on par.
https://amzn.github.io/ion-schema/docs/spec.html#union
The ion data model doesn't describe a schema or type system. It's a data structure where values are of a known type. In the binary format values are preceded by a type id, in the text format the syntax declares the type - "" for string, {} for struct. The data model doesn't declare what types a value could have, only the type it does have.
Amazon Ion - https://news.ycombinator.com/item?id=23921610 - July 2020 (110 comments)
Amazon open-sources Ion – a binary and text interchangable, typed JSON-superset - https://news.ycombinator.com/item?id=11546098 - April 2016 (163 comments)
I also now remember back to using something that was akin to FaaS but wasn’t called that. I could give them a JAR of some code that would execute on some Ion data for the order data when it changed. Basically FaaS for an ETL pipeline…
Crazy how ahead of the times some companies were.
First thoughts are:
ION pros: - easy to skip around while reading a file - no need to write a schema - backed by amazon so major langs will have impls - good date support - better concatenation, probably better suited to logging than bare
ION cons - what's the text format even for?
BARE pros: - schemas keep things tightly versioned - smaller binaries (not self describing like ion) - simpler to implement so tons of devs have impl'ed for their favorite lang - better suited to small messages (think REST json api)
BARE cons: - no skip read - no date support
I might do an ion ruby implementation too, to really feel out the difference.
Configuration files?
Not sure if that's an intended use case, but being more flexible than JSON and stricter than YAML seems ideal for configuration.
Looks like no one’s even so much as commented on it in the last year, so it might have been abandoned.
This proposal hasn’t been abandoned. We hope to post an update soon!
But Douglas Crockford just don't want to innovate anything, just like Gruber didn't want to make a proper specification of the Markdown format.
Sometimes people are keeping innovation back. Fortunately this did not happend with html.
The main thing missing from the text format is a magic and version number. At least the binary format has it.
I don't know that's the case... I've used JSON in lots of non-JS languages because it just works, and errors rarely are caused by mismatches in how JSON behaves in language X and language Y. A lot of that is that it is simple, and rigid.
I do think JSON is the defacto standard, and it really does get the job done, but for some more advances uses something like this could really shine.
(Should have gone with 'rational' instead of 'decimal', though. Decimal will be too painful to implement accross languages and implementations. Java bias?)
This is the correct representation, and how Google or the blockchain do it.
Hex for when efficiency isn't paramount.
Base 85 or BasE91 for when efficiency is more of a concern. http://base91.sourceforge.net/
- Hex is human readable, case insensitive, not that "inefficient", and always aligns to bytes.
- Base 85 and basE91 are efficient.
- Bitcoin uses Base58 because they thought base 64 was too human unreadable. Ethereum uses Hex.
- Base 256 (bytes) is efficient and the native language of computers.
Base 64 is not efficient, not human readable, and not easy to encode.
The biggest problem with base 64 is that base 64 is not base 64. Are you doing base 64 with padding? Are you doing base 64 with URL safe characters or URL unsafe characters? Are you following the standard RFC 4648 bucket encoding, or are you using iterative divide by radix? I think a great place where the cracks show is JOSE, where for things like thumbprints there's a ton of conversion steps (UTF-8 key -> base 64 -> ASCII bytes -> digest (bytes) -> base 64 thumbprint).
My personal advise for 90% of projects considering looking at base 64 should just use Hex or bytes. If needing human readability, use Hex. Otherwise use binary.
animal: Tiger:: {
gender: 'F',
weight: 450
}
This solves the inheritance problem, i.e., if you have multiple subclasses how do you know which type to deserialize as?Unlike avro, ion doesn't require a schema.
Some fun things about Mir Ion:
- We can fully deserialize Ion at compile-time (via D’s CTFE functionality)
- We’re one of the fastest JSON parsing libraries (and one of the most memory efficient too — we actually store all JSON data in memory as Ion data, which is vastly more efficient)
- We’re nearly 100% compliant to all of the upstream test cases (our main issue is that we’re often too lax on spec, and allow files that are invalid through)
- The entire library is (nearly) all `@nogc`, thanks to the Mir standard library
If anyone has any questions on Mir Ion, feel free to shoot me a line at harrison (at) 0xcc.pw
That seems to be...a problem? How do you deal with archeological dates, of which there are many, in Ion?
On the other hand, representability of a given date becomes progressively less useful the further back in time you go, and stuff becomes really gnarly once you go past the Julian calendar in 45BC.
Also, simplifying to “no dates before Jan 1 0001” has very little impact on applications dealing with the modern-ish world (with “modern” generously defined as “anything after the collapse of the Roman Empire”), and I can only assume applications dealing with earlier times could do with a more specialised representation for dates anyway.
1 BC for some is not "-1" for everyone.
Ion text is like JSON, in fact all JSON is valid ion text. Ion text has comments, trailing commas, dates, and unquoted keys. It's a really good alternative to JSON, YAML, or TOML.
Ion binary is compact and fast to parse. Values are length prefixed so the parser can skip over unneeded fields or structs, saving time parsing and memory allocated. Common string values, like struct keys and enum values, are given numeric ids and stored once in a header table.
I think many text parsers are missing libraries that edit documents in place, preserving formatting and comments.
You can't say the same of, for example, YAML & JSON, since the former (if not the latter?) has constructs unrepresentable in the other.
It's slightly confused because an application might 'serialise to' JSON or YAML or Ion equivalently - but really that's saying the application's data being serialised fits a model that's a subset of the intersection between those formats.
You could call Ion two, but it's more than that in that it's also a promise that they're 1:1 (err, and onto if you like) - their intersection is their union.
For RPC the binary encoding compares poorly to external schema formats like protobuf. In this context binary ion is a poorly compressed text format.
I don't think the partial document read capability of the binary format is all that important, but I've never worked on an application that would benefit from it either.
Thank god.. JSON for config files without comments is so awful.
1. Considering that a goal of Ion is to be a strict superset of JSON, separate syntax ensures that any JSON value can be parsed without misinterpreting some field as an annotation--there are no reserved/"magic" field names.
2. Annotations can be applied to any type of value, not just objects, which are the only type that have fields.
Sure 99% of decoders convert them to and from binary doubles, but that's purely an implementation choice.
All JSON numbers are implemented as integers or floating point, and as a result, have to be cast as a decimal (a decimal type is generally something that meets this specification: http://speleotrove.com/decimal/) when you import them.
Decimal types differ from floating point types in three ways: they are accurate, and they take into account rounding rules and precision. Decimal math is slower, can have greater precision and is better suited to domains where finite precision is needed. Floating point is faster, but is not as precise, so it's good for some scientific uses... or where perfect precision isn't important but speed is... say 3d graphics.
I've billed lots of hours over the years fixing code where a developer used floats where they should have used decimals. For example, if you are dealing with money, you probably want decimal. It's one of those problems like trying to parse email addresses with a regex or rolling your own crypto... it will kind a work until someone finds out it really doesn't (think accounting going, our numbers are off by random amounts, WTF?).
And you're confusing JSON the format with typical implementations. Open a JSON file and you see decimal digits. There is no limit to the number of the digits in the grammar. Parsing these digits and converting them to binary doubles, for example, is actually slower than parsing them as decimals, because you have to do the latter anyway to accomplish the former. Almost all JSON libraries convert to binary (e.g. doubles) because of their ubiquitous hardware and software support...but some libraries like RapidJSON expose raw numeric strings out of the parser if you want to plug in a decimal library
JSON spec for numbers: integer or float (implemented as a double precision float). JSON libraries read numbers as double precision float because that is the correct type for JSON numbers, not for any other reason.
Also, symbols converted to integers means the receiving end has to already know exactly what they are.
Edit: I was going to paste in a relevant quote from Zarf (i.e. Andrew Plotkin) on naming. Some of his most important programs have total nonsense names like "glulx", and the reasoning was that at least it would be easy to search for when the name is unique. But ironically, "Zarf" is so common a term that I can't find the quote.
The binary parser is much faster. All fields are length-prefixed so a parser doesn't have to scan forward for the next syntax element.
The ion parsers (lexer? not sure the right vocab) I've worked with have a `JSON.parse` equivalent that returns a fully realized object, a Map, Array, Int, ect but they also have a streaming parser that yields value by value. You can skip over values you don't need, step over structs or into structs without creating a Map or Array. That can be much faster.
We have done some work on performance comparisons with the ion-java-benchmark-cli tool (https://github.com/amzn/ion-java-benchmark-cli). Right now you can compare JSON serialized with Jackson and there is a pull request (https://github.com/amzn/ion-java-benchmark-cli/pull/27) for comparing against CBOR that should be merged soon.
We are always happy to hear suggestions for what is useful in this area.
If you want to create an issue for it (the best repo is probably the ion-docs one: https://github.com/amzn/ion-docs/issues) that will help to show us there is demand for it. Providing information on your use case helps us prioritize.