Amazon Ion
amzn.github.io
amzn.github.io
Ion never had nice code wrappers around serialized structures, and most of the time, especially with rich structures it was frustrating experience.
These were used in services and reactors which never touched raw Ion (at least not in any way different from Coral or BSF).
Full disclosure, I spent a lot of my free time working on Ion, both the supported implementations as well as my own. The additional data types are worth it alone, imho. Having to use JSON for most things now I’m frustrated at what is “missing”.
I suppose it's mostly an under-investment of time, not a shortcoming of the format itself.
Ion is readable and (seemingly) not very strict about schema. Seems like that would not readily incentivise additional tooling.
The format has much less extra syntactical noise than JSON.
For example,
name: "vii" # comments allowed!
id: 23923373
Pretty nifty as it allows readable configuration files with structured data.I think the Protobuf spec focuses on the binary serialization - the text format and JSON representations are not related to that at all, of course.
PartiQL is AWS's specification for a parser/query language that is compatible with standard SQL, but can query semi-structured or unstructured data (think JSON, Parquet, CSV/TSV etc)
https://aws.amazon.com/blogs/opensource/announcing-partiql-o...
PartiQL uses Ion as it's backbone and data format:
https://partiql.org/faqs.html#why-do-you-choose-ion-to-exten...
https://github.com/partiql/partiql-lang-kotlin/blob/master/e...
I looked pretty deeply into this, but failed a bit short of understanding what they meant when "if your query engine supports PartiQL." Does that mean writing a new DB that delegates incoming queries to PartiQL? Not sure.
Anyways, they use it in Quantum Ledger DB, and a few other internal projects:
https://docs.aws.amazon.com/qldb/latest/developerguide/ql-re...
So maybe that can give some more context around "what the hell is this, why does it exist, how would you use it?"
The problem is that it spreads, like an infection, to surrounding services. Inside Amazon there are literally hundreds of libraries that duplicate standard json libraries in various languages but support ion instead of json. All of this is just to deal with interoperability.
Ion is slower than protobufs and less universally understood than json. Honestly it's just an annoyance.
Point is AWS might not make up its mind yet on whether this does more harm to their DB business or not
When I was researching this area, it seems like Apache Calcite is the way to go, and already does this though?
It lets you use standard SQL and has adapters which translate from the abstract query AST to the particular implementation.
https://calcite.apache.org/docs/adapter.html
You can query anything with this, exactly the same way. It's kind of wild I've never heard it talked about tbh.
I went looking for solutions to multi-datastore querying when I fiddled with a business intelligence side project. Pretty useless to only be able to query one type of data, and too time consuming to implement individual mappings.
Apache Calcite and Metabase Query Language (quite the exact usecase there for Metabase, haha) were the only things I could find.
With an appropriate FDW, sure, and I'm pretty sure I've seen an FDW for parquet specifically, as well as other columnar formats.
Oh you mean advantage to you? haha...well..
But again, yeah it's a JVM thing so your options are that.
It’s roughly the same vintage as protobuf and thrift, from google and Facebook respectively, so perhaps it’s just Amazon’s equivalent, which they just never released as quick as the others did?
Obvious pros and cons, or yet another serialization format with no obvious benefits over anything else?
vs. protobuf: ion is self describing, vs needing a schema
vs. thrift: similar, thrift needs a schema to interpret a binary file
both thrift and protobuf are really binary formats, though they have a canonical textual representation, it's not actually used to serialize. Sounds like ion supports serializing as text as a first class concept.
vs. msgpack: ion has a corresponding text format, whereas msgpack is only binary. Additionally, ion has a symbol type, msgpack doesn't.
I think the biggest benefit here is that it's a new chance for a format that fixes some of json's rough edges to gain critical mass. There's probably nothing ultra special about it that hasn't been solved in other formats, but maybe the timing will be right and everyone will just adopt it as a json replacement (sort of how people just gave up on xml and switch to json seemingly overnight). It's impossible to predict stuff like that.
Edit: upon noticing that it was released in 2016, it seems less likely everyone will jump on the ion bandwagon ...
1) timestamp : I have had issues with a round-tripping timestamp representation quite a bit 2) decimal : currency is denoted in decimal rather than float and shows the Amazon retail heritage. This is very useful. 3) symbols : I've had cases where symbol table/dictionary would have made big difference in serialized size
https://github.com/tlocke/zish
It has timestamps and decimals. Full disclosure, I am the author.
In securities transactions, the quantity and quote are critical. You aren’t buying securities from Plaid, right?
If you try to liquidate or resize based on the Plaid quote, your brokerage or counterparty is going to provide a totally different quote, and one from a system engineered to provide quotes aligned exactly to the market standards.
I don’t see the risk/terror.
Ion is directly comparable to JSON/MessagePack/BSON/CBOR.
I would expect Ion will have slightly different time/space tradeoffs than the other binary schemaless formats.
Can anyone currently at Amazon shed some light on how prevalent Ion is internally?
[0] https://news.ycombinator.com/item?id=23922278
[1] https://github.com/amzn/ion-docs/commits/ https://github.com/amzn?q=ion&type=&language=
The support for S-Expressions is both a blessing and a curse. The ability to write logic with native data structures in it is fundamentally interesting, but it leads to lots of reinvention of somewhat crappy Lisp implementations.
The tooling ecosystem has been slowly improving outside of JVM, particularly the latest JS implementation.
In a vacuum, the support for type annotations, timestamps, decimals and binary serialization make it superior to JSON for use cases where self describing data is appropriate.
... why am I getting downvoted for offering direct experience as an AMZN engineer? Amazon InfoSec forbids PHP. See also: https://news.ycombinator.com/item?id=23030330
Google also bans PHP but has official PHP client libraries for all its APIs.
Both companies care about having and maintaining PHP SDKs so long as their paying customers want to consume their products/APIs using PHP.
You can't use PHP internally at Amazon. Downvotes and ignoring facts do not suddenly make my factual comment "untrue".
The AWS SDK in PHP helps generate web requests to Amazon's services, most written in Java.
Many AWS services use it as an interface language. Many AWS customers use PHP.
That's literally fact. That's literally how it works.
Additional resource: See also: https://news.ycombinator.com/item?id=23030330
But if you take the initiative to open source a client library in PHP and it gets the attention of AWS it absolutely could result in an interview.
If you are interviewing and you whiteboard your solution in PHP they won't hold it against you. The language is less important than the concepts. Granted, if the only language you know is PHP that could be a risk in your career. I think that holds true for any developer, though.
Source: Used to work and interview at AWS
> The following timestamp encoded as a JSON string requires 26 bytes
> ...
> This timestamp requires just 11 bytes when encoded in Ion binary
So, we just use JSON, and our solution to this problem has been to pass 64 bit unix timestamps around. It doesn't provide arbitrary precision, but for most use cases it is more than enough practical range & precision to get the job done. And of course we store & transmit everything as UTC, so there is no weirdness around needing to store additional timezone information. To give you an idea, our database columns are named things like CreatedUnixTimestamp.
It is also trivial to compare 64-bit timestamps without conversion, so any SQL storage of these as integers should yield massive speedups to queries against these types - Assuming you are coming from some more complex datatype like a string or byte array.
Passing an integer does not have the same semantics as passing a timestamp. Relying on out-of-band info to parse a document is a problem in the making.
> but for most use cases it is more than enough practical range & precision to get the job done.
Parsing s-expressions would also get the job done, even if it's a primitive s-expression that only supports cons cells and a string data type. However, people find value in enabling the parser to validate booleans, arrays, and objects.
ION is just a logical next step. Timestamps are quite naturally a fundamental data type in comm between web services, particularly in binary form.
For reference, MAX_SAFE_INTEGER can represent something around the year 285428751.
Everything on the server is just done in terms of UTC. I actually cannot think of a reason I would want to process a timestamp in terms of local time on the server.
(Though looks like Ion is not solely targeting JS, but I make an assumption it is nice to consume Ion data in frontend)
Edit: in Public API
// Field names
This experience has reminded me why JSON is such a great format.And having a whinge while I'm writing, "superset of JSON" is basically false advertising even though it is true; JSONs refusal to admit that line breaks are a thing is a major feature. I don't care it if it is technically correct and useful to some customers, if line breaks matter it is inappropriate to talk about a format's relation to JSON because people will get the wrong idea. The JSON brand is so strong because it is nigh-impossible to get wrong. This format gets screwed up - eg, for people who don't like JS.
fun x:
query = sql ::
select * from table
I'd be pretty happy. class SQL(str):
...
query = SQL("""
select * from table
""")(1) awesome but (2) 'key values' is a confusing way to say this
* int: arbitrary size integers
* decimal: arbitrary precision, base-10 encoded real numbers
* timestamp: arbitrary precision date / timestamps, with ISO 8601 format "2019-05-01T18:12:53.472-0800".
So exact same drawbacks as JSON basically:
* Large integers will be casted to 32 or 64 bits in most languages no matter what.
* Arbitrary decimal will be casted to float or double as well.
* ISO timestamps are not well specified when it comes to millisecond, microsecond and timezone.
The standard acommodates timezones as offsets from UTC, because it's a representation of a timestamp, not a local time at a particular geographical location. So things like daylight savings time periods are not relevant.
Also, this is the way any progress is made. Between 15 competing standards, some win over.
Were it not so, we'd still use whatever Cobol used for data serialization.
also
We had JSON5, now we have Ion. Google and Microsoft will probably run their own, too, soon.
Why the IT community always forks their standards and never merges baffles me since >20 years.