Amazon Ion Specification
amazon-ion.github.io
amazon-ion.github.io
https://news.ycombinator.com/item?id=29284428 (2 years ago, 229 comments)
https://news.ycombinator.com/item?id=23921610 (3 years ago, 110 comments)
https://news.ycombinator.com/item?id=11546098 (7 years ago, 163 comments)
On the other hand, they're not using this for the boto schemas, which seems like a natural place to show that it's able to capture real-world schemas so that makes it hard for me to think this has any traction
It's the sort of thing where I'd advise exploring other options first and only using it if the whys[2] really resonate with you because it definitely comes with some overhead.
[1] https://smithy.io/2.0/index.html [2] https://amazon-ion.github.io/ion-docs/guides/why.html
AWS is starting support PartiQL (https://partiql.org/) queries in some places and PartiQL uses Ion's type system internally.
Ion has the option of using symbol tables to replace strings (e.g. in struct/map keys or in values). So, if you benchmark had a large number of records with similar structures, I would expect Ion to pull ahead. On the other hand, if each record had nothing in common, I'd expect them to perform similarly.
One feature of the Ion libraries that I've liked is the parser will take any of the formats and figure out what to do with it (text, binary, compressed binary). It's one less thing to worry about. You can switch encodings later without breaking consumers, you can write plain text Ion when you're testing, etc.
userBirthDay: null <-- ok, but what type is it? String? Int? Timestamp?
userBirthDay: null.timestamp <-- ok, it's a timestamp typed variable, but we don't know the value. Yay, happy programmer.
Ion originated 10+ years ago from the Amazon catalog team - the team that kept data about the hundreds of millions of items available on Amazon. Nearly every team in the company called the catalog to get information about items all the time - scanning the entire catalog, parts of the catalog, millions of individual item lookups every second, etc.
They did the math and some very large percentage of network traffic in Amazon Retails data centers was catalog data. If that data, currently in XML or JSON format, was sent in a more compact format it would save some ridiculous millions of dollars every year. So Ion was born and eventually open sourced.
https://en.wikipedia.org/wiki/ASN.1
which is forgotten but not gone.
I hadn’t thought about it in like a decade but yeah it’s still silently in the background…
- Avro and Ion are the only two that are labeled Textual/Binary
- They are in the same Big Data grouping
- They both are schema-embedded, and support some rich nested datastructures, though they deviate on many of the specifics
So I think it's reasonable to pick out Avro as an especially similar point of comparison.
My 2 cents: don't use it.
Honestly, if you’re in a case where you absolutely know none of these work for you and you can absolutely prove you need another, you’re probably just going to write your own. And that’s a fleetingly rare case.
You can write Ion by hand (like JSON) and share it without a schema (unlike protobufs). There’s fewer ways to express values than YAML, but more data types.
Having S-exps is convenient for writing DSLs in a data language that’s easily readable from other languages.
Wider range of data types - Ion supports decimals, symbols, blobs, and clobs which don't exist in CBOR. Optional schemas and annotations - Ion allows attaching type/schema information to data for validation purposes. CBOR has no schema support. Text format - Ion provides a human-readable text format for data interchange, CBOR is binary only. Maturity - Ion has been used in production at Amazon since 2009, CBOR is a newer standard (RFC 7049 in 2014). Language support - More mature library ecosystem around Ion vs CBOR which is still gaining adoption.
Pros of CBOR vs Ion:
Standardized - CBOR is an IETF standard, Ion is an Amazon-proprietary format. Simplicity - CBOR has a smaller set of basic data types making it simpler to implement. Used in other standards - CBOR is used in data formats like COSE for crypto operations and CWT for web tokens. Efficiency - The CBOR binary format can have a smaller encoding size than Ion's. JSON interoperability - CBOR is designed to be a JSON-compatible binary format. Ion is JSON-like but not fully compatible.
In summary, Ion has richer data typing and schema capabilities and a long production history. But CBOR is simpler, standardized, and gaining momentum - especially in crypto and web standards using it as a binary encoding basis.
So Ion may be better for applications dealing with complex, annotated data. But CBOR has advantages for an efficient binary interchange format, particularly when standards compatibility is important.