Please stop writing new serialization protocols
scottlocklin.wordpress.com
scottlocklin.wordpress.com
> Imagine if there were 40 competing and completely
> mutually unintelligible versions of html or text encodings
There are. > There really should be a one size fits all minimal
> serialization protocol
There can't be. > just the same way there is a one size fits all network
> protocol which moves data around the entire internet
There isn't.[0] https://www.joelonsoftware.com/2003/10/08/the-absolute-minim...
edit: scrolling down shows me this has already been posted 3 times in this thread (and will probably be posted a few more times before it's off the front page)
But data isn't like water, it's like 'chemicals'. You can't have a standard component for processing data the same way you can't have a single components that knows how to process sulphuric acid, crude oil, hydrazine, mercury and molten salt.
Data can be binary, delimited, fixed field, lossy, ASCII, Many variations of Unicode, executable, contain attack such as SQL injection, encrypted, have time sensitive delivery requirements, include checksums, require checksums to be applied in the protocol, be meaningless without metadata or other data, and have who knows what other constraints, limitations and requirements.
Building a data processing system is not at all like building a water works. It's more like building a chemicals processing plant.
It is my opinion that this will be the function AI serves in reality.
[0] https://lh6.googleusercontent.com/-QkYNnuYZs8Y/Uf0y_mV-1uI/A...
No it won't. It just puts a '>' in front of the line. You can also put asterisks around the text to italicize it (as I've done here). But the '>' has no special support.
> This is just normal body text.
https://news.ycombinator.com/formatdoc
> The indentation marks it as unwrapped preformatted text, which renders this text illegible on mobile.
Verbatim/code formatting does totally break legibility, especially on mobile.
Nothing is competing with HTML on being what HTML is. Nothing is even close.
All text encodings in common use build on top of ANSI. The de-facto standard (now) is UTF-8 and there's no reason to change that.
IP really does move all the data around the entire internet (that's kind of in the name right there!).
Yes, there are variations, extensions and incompatibilities, but these things are not being re-invented all the time, exactly because these things are ubiquitous standards.
One example of a tradeoff that is hard to eliminate is that you can reduce size and increase performance substantially if you pre-specify a schema like Cap'n Proto (and others) do. The downside is then if you just get a message without knowing what it is about it's difficult to figure out. The only way out of this tradeoff that I can see is having a global schema registry and every message having 8 bytes dedicated to schema ID, and that has downsides of its own, especially for small messages.
I do agree with the author though that we could do with more binary serialization protocols with tools to easily translate back and forth to a human-readable text format for debugging.
You're saying that as if there was a clerk reading the message at the other end. If a computer program gets a message it doesn't "understand", it's not going to "figure it out" either way.
Worse, if you're a programmer that doesn't know the actual schema and you're trying to figure out what the schema might be just by looking at the data, you'll probably run into trouble.
A fake example scenario of what I'm talking about: Someone notices a server is producing way too many errors, they check the logs and see a bunch of "invalid request: ...." messages. The messages are JSON like {"api_version":2,"cpu_usage":0.57}. They figure out that the server monitoring system is forwarding stats to the wrong IP address. Even though they don't know the full schema of the messages, they get the gist and it helped debug. With Cap'n Proto that would have been a few nonsensical bytes.
I don't think this is a very big downside, so in general I like protocols like Cap'n Proto.
Perhaps a much bigger downside is that in most languages integrating code generation into the build pipeline is a pain in the ass, so it is way easier to use dynamic serialization protocols that don't require compiling a schema.
Self-describing formats do make backwards compatibility a bit easier. With Cap'n Proto, adding a new field in the middle will break consumers.
Meanwhile, XML-RPC (which is not a serialization format!), JSON-RPC, SOAP, Swagger, are stacks that intentionally leave open the possibility that someone will come along with and consume the form-on-the-wire directly, outside of the tooling of the environment. Most in-the-wild JSON-responding APIs have the same expectation.
IDLs themselves are a very old idea, probably because we like declarative ways of specifying contracts that are then applicable across a heterogeneous environment, or in different languages and runtimes, and so on.
As for why there's dozens of offshoots of standalone serialization formats which are all predominantly occupied with the efficient packing of numbers while keeping the general data model of JSON, I can't answer [1].
Cap'n Proto https://capnproto.org/
Simple Binary Encoding (by Martin Thompson) https://github.com/real-logic/simple-binary-encoding/wiki/De...
and if neither of those will do, raw C-structs on the wire (basically what the other 2 are anyways).
Along the same sentiment, I'm not a fan of APIs using JSON and/or XML or some other overly-flexible textual encoding. Simple binary encodings, TLV-ish if necessary, are the best.
I was never really convinced by the "human readable" argument for textual encodings either --- you just need to get used to it, then you can read and write the bytes in a hexdump as easily as you can English. In fact I'd prefer working with hexdumps to XML. But unfortunately there's now a whole generation of developers who can't even count to 2 in binary and don't know what a hex editor is...
This sums up my issue with the current fashion in software development.
Man, when I read this, all I can think is "real programmers don't eat quiche." [1]
We will continue to see more abstraction in our tools. By and large, the changes we've seen so far have made us more productive as programmers. You should try quiche sometime. You might like it.
>"Java monkeys eventually noticed how slow XML was between garbage collects and wrote the slightly less shitty but still completely missing the point Avro."
I would like to know why the author feels that Avro misses the point. Can anyone hazard a guess?
and similar for:
>"Oh yeah, I do like Messagepack; it’s pretty cool."
It would be interesting to hear why they(or anyone else for that matter) consider Messagepack a worthwhile contribution to the serialization tool shed but Avro is not.
External Data Representation (XDR)
https://en.wikipedia.org/wiki/External_Data_Representation
https://tools.ietf.org/html/rfc4506
Abstract Syntax Notation One (ASN.1)
I hate ASN.1.
"human readable JSON like format"
"JSON style human semi-readable form"
> accidentally mentions JSON twice in the list
JSON is the most widely used serialization protocol in web application development today. The article's repeated mention of it is no accident at all: it's considered best practice across the industry. Frankly, I don't know why we are still talking about this.Similarly, JSON still has ambiguity around encoding (type of UTF). Dates are the same. Even integers, which are limited to JavaScript's 53-bit implementation (all numbers in JavaScript are doubles, so you're stuck with the mantissa limit). This is a huge pain if you're dealing with 64-bit values.
UTF-8 and JSON is a good place to start for Web Development, but there is no way it's the final or only answer.
According to what? JSON just says that a number is a series of digits.
If you ever need an argument for why "worse is better" actually rules the world of software, point to JSON and drop your mic.
1. Do you want an efficient format (i.e. binary) or a format that is easily readable/writable by humans (i.e. text)?
2. Do you want self-describing format (data schema is part of the standard) or schema-less format, to save space on transmission? And how much?
3. How to deal with mixed text data? Do you want primarily to store strings and escape any structured data within them (a la XML), or do you want to primarily store structured data and force all text to be escaped in them (a la JSON).
Personally, I like CBOR, which interestingly wasn't mentioned.
The speed of parsing case is definitely rare, but i've done some work where packets come in off a 10-gig ethernet port, and every nanosecond counts.
XDR is nice, though, apart from being big-endian and not having widely-supported 64-bit integers. It's a pity it's unfashionable.
These days, I've mostly thrown the towel and plunk in JSON everywhere. At least it's hip, and if I ever find a piece of code where the serialization is the bottleneck, early enough to start pushing XDR (not ASN.1 - I once tried to write a parser for _that_ and ended up with some extra grey hairs ;-)).
When was the last time you talk to people working on another software stack? Cooperations between different tribes need to be enforced by strong leadership or a Big Need like imminent extinction of the tribe. As long as that doesn't exist and the whole ecosystem is continuing to grow you can just sit there and watch people building the next silo and the next instead of getting to a higher step in evolution.
And it's actually the reasonable thing to do. I mean, would you rather have a miniscule share of a cake others baked, or do you want to have your own cake? When both is about the same effort, I'd rather have my own cake, even if I have to define a new serialization protocol to store it.
JSON may be everywhere, and it's tempting to look at its flaws and think, "we can do better" but it also has the great benefit of having decent serialization libraries already written in the vast majority of programming languages. That's one heck of a feature.
I'm tired of these state of affairs. All of the above can be done in real programming languages, with real syntax. There is no need for yet another external DSL when an embedded DSL in the form of a library will suffice.
The trick for adoption I imagine is designing a DSL in any LISP of one's choice that defines systems as s-expressions... and just turn it into its equivalent XML instead.
Once the project inevitably becomes ubiquitous in enterprise IT you make a blogpost thanking everyone for being part of the social experiment but the DSL is back to being LISP again, with an optional pricing model per-processor a la SQL Server for those companies wishing to continue using XML instead of passing their specs through an automated converter.
On an unrelated note, I'm looking for a cofounder with a FP background...
Maybe I like, I don't know, cobol. What if I feel happy with my non-DSL DSL being cobol? Would people like it? Or maybe, really, the dedicated DSL is a better approach? To give people, at least, some handlebars, a lowest common denominator, if you wish.
The point is, no matter what language is selected, someone is always going to be unhappy about the choice. Look at Chef, great technology, IMO. Gives all the power of Ruby. One can do totally whatever they want. Yet people were not happy with it. Alternatives came out, people started reinventing stuff all over again. But even with those, one can use whatever language they want to automate these tools even further. All these DSLs are just helpers, a starting point, make the initial adoption easy.
I think people are looking at this aspect from a wrong perspective. "Is my time worth learning this new DSL / am I compensated enough to make it worth learning". It does not matter, as long as it all calculates.
1) You define your configuration in (common lisp) s-expr's.
2) The configuration management system is a lisp program which reads the configuration defined in (1), generates bash code, uses ssh to contact the nodes to be configured and runs said bash code.
That would retain one of the big advantages of ansible, namely you don't need to have some agent running on each node, and you can leverage ssh for security/cert handling etc. And you don't need a lisp environment installed on the nodes either.
Programming languages and their libraries come with their idiomatic conventions of API design, calling conventions, data structures, etc, which tend to translate awkwardly to a "foreign" programming language.
The solution is usually one of two choices: pick a lowest-common-denominator programming language and implement everything in it such that all higher-level wrappers are only cosmetic shims; or, make a DSL that has clear and defined mapping to language-specific constructs and conventions in each programming language.
The former is the coder's solution, the latter is the designer's solution.
For example do you see Chef taking off and becoming popular if it just implemented Puppet's language. Or would Salt have any appeal if it just had the Ruby syntax like Chef does?
It just won't happen easily. One reason is because the API surface and the DSL is so varied that chances are there were significant design mistakes made in a previous version. And one of the reason people want the new shiny is that it fixes some of those fundamental design issues. Copying and emulating an existing thing, means sticking with broken fundamental design decisions.
The irony is of course the new shiny also has it own fundamental design flaws and so now there will be a new group of people implementing the newer and shinier thing. (With obligatory articles on HN on "Why we switched from X to Y").
People use those "extenal DSLs" because they are tired of bash and ssh for the things the want to do.
> superset of this with a JSON style human semi-readable form,
and an optional self-description field, and you’ve solved all
possible serialization problems
I'm really curious about Ion: https://amznlabs.github.io/ion-docs/Is it in use anywhere else besides Amazon?
"Facebook couldn’t possibly use something written at Google, so they built “Thrift,” which hardly lives up to its name, but at least has a less shitty version of RPC in it."
Thrift predates the public release of protobuf.
https://tools.ietf.org/html/rfc1832.html#section-6 has a good encoding example of XDR.
Contrasting that with Protocol Buffers is enlightening as it clearly demonstrates some differences in design goals and where tradeoffs are made. Feel free to correct my interpretations as I may have missed something!
1. XDR appears to have no equivalent of a Protocol Buffer field number => the format is not self describing. That is, one must have a schema to properly interpret the data
2. XDR appears to encode lengths as fixed width based on the block size => faster to read/write but larger on the wire than using a varint encoding
3. XDR's string data type is defined as ASCII => it does not support modern unicode outside of the variable-length opaque type
1 would seem to present difficulty for modern distributed systems as one cannot control the release process of every distinct binary in the ecosystem to ensure that they are all schema equivalent at the same time. This can be remediated by propagating the schema as a header to the underlying data for consumption on the other side, but that bloats the wire format. Header compression could help with this but may be problematic for systems with severely constrained networks that constantly reestablish new connections (ex. mobile).
1 also has implications for data persistence. One cannot ever remove or reorder an XDR struct member else they will incorrectly parse data that was written in the older format. This is in contrast with Protocol Buffers, where one can remove or reorder message members whenever they'd like, as long as they take care not to reuse a tag number (and the newer `reserved` feature can help with that).
2 is just a performance tradeoff: binary size or (de|en)code performance?
3 has implications for memory constrained systems. Ex. on Android we eagerly parse string fields to avoid doubling the allocation overhead (first as raw bytes, then as a String object). If we required all string datatypes in Protocol Buffers to be defined as bytes fields (the variable-length opaque data type equivalent), we wouldn't be able to provide this optimization.
Overall, XDR looks like a good fit for inter-process communication in a homogeneous environment. Protocol Buffers looks like it's a good fit for cross language communication across heterogenous and unversioned environments. Directly comparing the two, XDR is much more verbose on the wire (particularly if we mitigate the versioning issues by serializing the schema in a header) whereas it's likely significantly faster to (en|de)code. i.e. there's a tradeoff for networking/storage costs vs. CPU performance.
Scott makes a bunch of provocative declarations in his post but I think many of them betray a lack of background to appropriately understand the tradeoffs involved. As illustrated above, Protocol Buffers makes a bunch of design affordances for compactness on the wire which XDR does not accommodate. He also believes Google's RPC system to be "unarguably shitty" even though it has never been open sourced due to dependency issues (what is open sourced as part of Protocol Buffers is a shim, gRPC is the future here). His impression of why Facebook built Thrift is similarly misinformed as Protocol Buffers was not open source when Thrift was written.
RFC1832 has a rationale for decision 2 towards the end of the document. Here is an excerpt.
"(4) Why is the XDR unit four bytes wide?
There is a tradeoff in choosing the XDR unit size. Choosing a small
size such as two makes the encoded data small, but causes alignment
problems for machines that aren't aligned on these boundaries. A
large size such as eight means the data will be aligned on virtually
every machine, but causes the encoded data to grow too big. We chose
four as a compromise. Four is big enough to support most
architectures efficiently, except for rare machines such as the
eight-byte aligned Cray*. Four is also small enough to keep the
encoded data restricted to a reasonable size."
Most RISC architectures of the time did not support unaligned memory accesses.Decision 3 is a matter of timing. RFC1014 ( https://tools.ietf.org/html/rfc1014) was released in 1987. Unicode was still a work in progress.
Edited to fix formatting of the excerpt and a typo.
Would it be better if we still send HTML over ASN.1? I think the state of things would be much worse.
Not blaming anybody for skimming the post, which was a pretty typical blogrant, if it was from certain other people it would be clearly clickbait, but this seemed like at least genuine ranting.
I'll probably use messagepack instead of json in an upcoming project for performance reasons, and the fact it works better than the json thing in J.