Cap'n Proto 0.6 released – 2.5 years of improvements
capnproto.org
capnproto.org
Edit: Why does this matter? Well, first, it matters because so little software is capability-safe. Capn's RPC subsystem is based directly upon E's CapTP protocol. (E is the classic capability-safe language.) As a result, the security guarantees afforded by cap-safe construction are extended across the wire to cover the entire distributed system. This security guarantee holds even if not every component of individual nodes is cap-safe, and that's how Sandstorm works.
Continuing on, there's also historical stuff going on. HN loves JSON; JSON is based on E's DataL mini-language for data serialization. HN loves ECMAScript; ES's technical committee is steered by ex-E language designers who have been porting features from E into ES.
this is very similar to what the new bus1 ipc author wants to do: https://www.youtube.com/watch?v=6zN0b6BfgLY
I never minded these frameworks, but then I wanted to write a few parsers for some file formats and they used Thrift/Flatbuffers to encode a few ints, which seemed like a major overkill. There was no RPC, no packet loss, no nothing.
But I'd hazard a guess that you're spot on regards the normal needs of normal programs. The vast majority of programs utilizing message-passing are not serialization-bound. For these programs I'd recommend JSON because it is easy for humans to inspect and debug, and there are plenty of libraries to choose from if you are not performance-sensitive.
(YAML is hopeless regards security and XML is just horrid to look at when you're debugging).
But I can certainly understand where you are coming from. I've designed several text based protocols in my time for that exact reason. And I still do for stuff that I positively know will never require any sort of performance (mostly on embedded systems, somewhat ironically since those are sometimes constrained down to just a few hundred bytes of memory :-))
I'll be surprised if you use a language that has a good protobuf implementation that doesn't also have a good json/yaml/xml implementation.
Incidentally, as far as I can tell, protobuf also neglects performance. For high performance you want as much of your message to be predictably positioned in the bytestream as possible, but protobuf uses a key-value style system which means you always have to look at each key-value pair to determine what it is and how big it is. This makes good sense for compatibility, but it's a definite trade off that limits how fast deserialization can be.
I can't speak to XML, but can say that even with "type safe" json serialization, I've seen several bugs due to some json libraries treating all numbers as doubles internally, meaning bad things happen when you're actually serializing integers larger than 52 bits (say, a nano timestamp). Sure, maybe it's a bug in the json library because the json spec doesn't establish any max precision for numbers, but it's not hard to run into it when different components/languages with independently implemented json libraries try to talk to each other.
Shared schema approaches like proto also make it really hard to accidentally typo the name of the field, or accidentally try to read a field as a string rather than a list of strings.
Just like I wouldn't consider code to be type safe if it passed data around as raw Objects and each usage cast it to the expected type, I wouldn't consider services that communicate with each other using independently (manually) implemented parsers/extractors as being type safe, even if the "meat" of the code is dealing with strongly typed data on either end.
I don't see any reference to range/precision in the spec, other than the suggestion that numbers that can be represented as an IEEE double are likely to be interoperable, and integers within 53 bits of range will generally be interoperable due to exact agreement on their value.
https://golang.org/pkg/encoding/json/#Unmarshal
They have some reasonable examples there.
Mapping arbitrary JSON to Go types may or may not be simple, depending on the data format.
Personally I think it's easier to maintain a format defined in an IDL like Protobuf or Thrift than a JSON format. Everyone is always on about how JSON is a "human-readable" format, which is nice and probably 90% true. What we want is a human-writable format, which neither JSON nor YAML really are. It's easier to define a structure in Thrift than in JSON.
I still gotta review this release of Cap'n Proto, but after struggling with some buggy language implementations from the upstream Thrift project and evaluating Protobuf 3.0 + grpc, I think that Protobuf 3.0 is going to be the way forward for now.
If performance ever becomes an issue (unlikely in my case) I could probably switch to some binary serialization protocol eventually.
True, in many use cases, performance is not critical. However, when you're doing big data processing or serving at scale, serialization performance does in fact start to matter a lot. If Google Search passed all backend messages as JSON rather than Protobuf, it would probably require 10x the hardware (we may be talking about millions of machines here) and wouldn't be able to respond as quickly. Similarly, Cloudflare couldn't handle its logging load without Cap'n Proto.
Many times, companies who started small and thought performance didn't matter later found themselves needing to do massive, expensive rewrites in order to keep up with scale (e.g., famously, Twitter).
Hand-written binary serializers may be easy to write initially, but they are very hard to maintain. Over time, you'll need to add and remove fields, while maintaining compatibility with old binaries already running in the wild, or compatibility across languages. It's very easy to screw this up with a hand-written encoder, but very easy to get right with Protobuf or Cap'n Proto.
Finally, even if performance isn't an issue, you may find that type safety is useful. If you are writing code in a type-safe language, correctly dealing with JSON actually tends to be a huge pain involving lots of branching. By defining a schema upfront and generating code, you can get a much nicer interface. And if you're doing that anyway, then you might as well take advantage of the more-efficient encoding while you're at it -- it's easy to dump human-readable text for debugging when needed. (As of this release, Cap'n Proto ships with a library for converting to/from JSON, and there are lots of third-party libraries to do the same with Protobuf. Both libraries also have one-liner "dump a debug string" functions.)
I mean that when I write code to consume a message, I would like to detect (at compile time) if I typo a field name, and I also don't want to worry about what happens if a (possibly-malicious) sender sending me a different type than I expected.
I think binding is precisely what is going on when a compiler generates code for marshaling and unmarshaling messages based on a schema.
I'm merely pointing out that the term is vague and can cause confusion. That you make a strong association with one use of the term doesn't really govern what others associate with it.
Recent releases have it built in already, just FYI.
(Also, hi Kenton! I spent lots of time looking at your code when hacking on protobuf :))
Here's one example that has repeated itself a few times in a few companies. I kinda like having logging that works without strangling what you are monitoring and/or the server doing the monitoring. I also like having extensible log servers that you can write plugins for (although I'm not that fond of the idea anymore as people invariably will write _slow_ plugins, and then you are screwed. Do put in the APIs, but don't tell anyone it is easy to add processing chains). I like it to the point where people made fun of me for this "obsession" even 15 years ago.
So it goes like this: I use some fast serialization that can be parsed using next to no CPU per message -- or in later years, using Protobuf or similar. Because it is more convenient and I can throw in binary payloads in exceptional cases. I benchmark and I usually get high ridiculously high throughput without even making much of an effort. Great, this saves me the trouble of having lots of log server horsepower and I can get on with life. (Stuff will need to be sharded eventually, but since we keep everything dead simple and we actually have a plan for how to do this when needed I don't worry).
Then at some point someone who has never had to write fast code comes along. It used to be "let's use XML". Now it is invariably "let's use JSON". Because "it is fast enough". This is where it starts. And they'll use Ruby, or something else that is godawfully useless for anything that is supposed to be high throughput, because hey, it is what they understand.
And for a while, it is fast enough. Sure, you now have only a fraction of the original throughput, but the systems are lightly loaded and you are not bumping into any CPU limits. "See!? I told you it would work!".
Then people start to do serious logging and the stuff can't keep up anymore. So they start coming up with lots of schemes for dealing with the load. For every scheme implemented, stuff gets a bit more brittle and complex because now the code has all of these assumptions and mechanisms to preserve. And before you know it, you have a slow, complex logging system, that everyone depends on and nobody wants to touch. Well, sometimes someone will reimplement it, but they'll use whatever is hip at the moment. Like JavaScript and then ooh'ing and aah'ing when they double the throughput -- even though they kind of originally lost 2-3 orders of magnitude of performance.
And it isn't like making a performant version is a big task. With Cap'n Proto, Protobuffers and whatnot, someone has done the heavy lifting for you. Someone a lot smarter than you. And it is probably going to be faster to implement, more efficient, and easier to maintain in the long run.
Performance matters where it matters. And it _is_ stupid to forego it when it costs you little or no effort. And yes, being an old fart means I don't give a shit about people's egos anymore and I'll call them stupid to their faces when they are being stupid.
If you log using JSON (or XML, yech!) as your envelope format, you _are_ stupid.
Others reasons are:
- Starting from an IDL you get a nice service contract. Which means both sides of the service (client & server implementors) get a strong description on how the service should behave and what they need to implement. The more roles you have (architects, client-implementors, service-implementors, testers, mock-service implementors), the more you profit from a good specification.
- If you generate all the serialization/deserialization code you avoid the tedious handwriting of this code and the possibility of introducing errors in this. E.g. some misspelling of field names or functions or copy&paste errors can always occur and some of these errors are only detected late in the development process.
These advantages might not be visible if you have a small service and development team - you might even experience an overhead compared to just "changing the client and server code", which are both under your control. But once you are working on a system with thousands of APIs you are quite happy if you don't have to care about handwriting serialization and RPC code and instead can focus on business logic. I personally was responsible for deploying a similar solution for an in-vehicle infotainment system with hundreds of services and around 4000 functions, which worked out really well.
Out of curiosity, what format did your infotainment system for serialization?
For my use case in electronic trading, a format like this is great b/c you don't waste many CPU cycles in decoding/encoding while at the same time you can take complex message structures and easily generate code stubs to read/write them in various languages. You change something? Just regenerate the code stub. You can even guarantee backwards compatibility if needed.
I'd also disagree that performance on message passing doesn't matter; depends how many messages you're passing :) I've seen CPU-bound systems which spend most of their CPU time decoding JSON.
To me, formats like JSON or XML have one clear advantage: debugging is easier. I can fire up a decoder for a binary format, but that slows down the investigation.
That said, most binary formats have a schema file due to necessity of not serializing field names. That schema file makes it easy to have tools that spit out language bindings. And that reduces the tedium of actually writing code that uses (and validates) the data, so I generally prefer those formats for that reason.
But it's been slow.
On the other hand, now that 0.6 includes a JSON library, it's relatively easy to do browser<->server in JSON and then use Cap'n Proto on the back-end. But, obviously, it would be nicer to use Cap'n Proto through the whole stack.
Contributions are welcome!
Part of the problem is that there is a bit of an ideological difference, where I prefer dynamic code to adding build steps but all the existing tools and code assume you want to generate source. I also found it quite irritating to bootstrap too, because the tooling itself uses capnproto. It sounds neat, but then it means that you have to be able to read capnproto in order to be able to read it.
The documentation was also incredibly patchy at the time which made it quite hard to develop for, although I'm told that the kentonv & gang are very helpful.
In the end, I was just starting to struggle with capnproto generics when my requirement went away as the server I was connecting to added a msgpack option.
However, this is done through a C extension that calls into the C++ implementation. For browser-side Javascript, that's obviously problematic. (Emscripten isn't really the answer -- would be way too large a download.)
Could the emscripten result be slimmed-down?
(I fear the answer is probably “in theory, yes, in practice, it's >100KB”)
Going to protobuf from JSON saved us about 50% on bandwidth for the high volume real time data service that we develop. Love it.
What exactly would you be using? Flash?
In all seriousness, it's nice to have a format that works for all your clients (browser, iOS, Android, desktop, server).
https://en.wikipedia.org/wiki/Comparison_of_data_serializati...
If you read the Cap'n Proto RPC docs, everywhere where it mentions "traditional RPC", I specifically had Google's internal RPC in mind (having previously been the maintainer of Protobufs at Google). So, you can more-or-less substitute gRPC in there for a direct comparison. https://capnproto.org/rpc.html
There are two key differences:
1. Cap'n Proto treats references to RPC endpoints as a first-class type. So, you can introduce a new endpoint dynamically, and you can send someone a message containing a reference to that endpoint. Only the recipient of the message will be able to access the new endpoint, and when that recipient drops their reference or disconnects, you'll get notified so that you can clean it up. This is incredibly useful for modeling stateful interactions, where a client opens an object, performs a series of operations on it, then finally commits it. Put another way, this allows object-oriented programming over RPC. Also note that you can easily pass off object references from machine to machine -- currently this will set up transparent proxying, but in the future we plan to optimize it so that machines automatically form direct connections as needed, which will be really powerful for distributed computing scenarios.
2. Relatedly, Cap'n Proto supports "promise pipelining", which allows you to use the result of one RPC as an input to the next without waiting for a round-trip to the client. This makes it possible to use object-oriented interaction patterns with deep call sequences without introducing excessive round-trip latency. This is described in detail at the RPC link above.
So... CORBA?
Server->Client streaming is e.g. a powerful way to get realtime updates about some state on the server to the client without needing to poll (not performant) or needing to define callback services (ugly to maintain, service lifecycle questions, and if you need an extra connection for from the service to the callback service (client) there are also challenges around routing).
Native support for streaming also removes the need to model explicit flow control (backpressure) behavior in the user defined APIs (like functions for requesting some items in addition to the callback functions for delivering items).
Another difference is that grpc utilizes HTTP/2 as underlying transport protocol. However I'm thinking that's mainly an implementation detail. If Google had designed grpc slightly different (not relying on barely implemented features like HTTP trailers) it could have been a bigger differentiator, e.g. by allowing browsers to directly make grpc calls without proxies.
The drawbacks of callbacks that you mention don't apply in this scenario: The same network connection is utilized for the callback, solving the routing question. The callback object receives a notification when the caller is done with it (or disconnects), so it can clean up, solving the lifecycle question. Flow control / backpressure can be achieved by pausing calls to the callback if too many previous calls have not returned (I actually plan to bake this pattern into the library in the future to make it dead simple).
So, it seems to me streaming in gRPC is actually a narrow subset of what Cap'n Proto can express.
For web stack compatibility purposes, it's straightforward to implement Cap'n-Proto-over-WebSocket. I'm not sure that HTTP/2 buys much on top of that.
Regarding backpressure and HTTP/2: You get some different kind of backpressure behavior with your approach (application level flow control) and the grpc approach (transport level flow control). Let's say you have 2 functions which are called in parallel. For one the arguments are very big (let's say 100kB), for the other one they are small (some bytes). With HTTP/2 and transport level flow control the small function could get the bandwidth as the big one, which means more requests/s for the small function. With pure application level flow control the small function needs to wait until the big one is fully sent before it can be put on the wire.
Which means if I have 2 tasks that execute concurrently and look like
A: while (true) { await callBigFunction(); }
B: while (true) { await callSmallFunction(); }
then with grpc I get more calls for B and without Cap'N'Proto I get the same for both (correct me if I'm wrong).I don't think it's a huge disadvantage in practice, because if you have large data chunks you should probably design your API different - and if you want realtime behavior then probably both protocols are not optimal. But it's still something to keep in mind.
Besides building more support in the library it would maybe also be very worthwhile to have these patterns described together with examples on the website. E.g. how to achieve callback interfaces which don't user another connection, server->client streaming, etc.
For me as a potential user information about these kind of features would make Cap'n'Proto immediately look more interesting. When I look at the website I only see Promise Pipelining described as a main feature - which is novel for sure, but not something that attracts my interest if I look for the other features we discussed about.
What confuses me is, then what are the costs of migrating to this system? Am I essentially dumping my programming language's object model for my capnproto implementation's? When can this be annoying? Or does it vary from implementation to implementation?
In a similar tangent - how similar is this to apache arrow, not because of the columnar analytics part, but could I expect to just dump a bunch of data in shared memory and read it from another process to eliminate IPC serialization/copy costs?
For message-passing scenarios, Cap'n Proto is an incremental improvement over Protobufs -- faster, but still O(n), since you have to build the messages. For loading large data files from disk, though, Cap'n Proto is a paradigm shift, allowing O(1) random access.
Edit: Your question is addressed here: https://news.ycombinator.com/item?id=14249367
I discussed in more detail in reply to your first post, but just to be really clear on this:
No. In fact, for deeply-nested object trees, constructing a Cap'n Proto object can often be cheaper than a typical native object since it does less memory allocation. However, there are some limitations -- see my other reply.
(Constructing Protobuf objects, meanwhile, will usually be pretty much identical to POCS, since that's essentially what Protobuf objects are.)
There is a common myth that Cap'n Proto "just moves the serialization work to object-build time", but ultimately does the same amount of work. This is not true: Although you could describe Cap'n Proto as "doing the serialization at object build time", the work involved is not significantly different from building a regular in-memory object.
With that said, using protobufs for internal state is not an uncommon practice and if you don't care about cleanliness and just want to pound out some code quickly, sometimes it can work well.
Cap'n Proto has an additional disadvantage here in that its zero-copy nature requires arena allocation, in order to make sure all the objects are allocated contiguously so that they can be written out all at once. This actually make allocation memory for Cap'n Proto object much faster than for native objects -- but you can't delete anything except by deleting the entire message. So if you have a data structure that is gradually gaining and losing sub-objects over time, in Cap'n Proto you'll see a memory leak, as the old objects aren't freed up. You can work around this by occasionally copying the entire data structure into a new message and deleting the old one -- essentially "garbage collecting". But it's rather inconvenient.
This is actually one reason I want to extend the Cap'n Proto C++ API to generate POCS (Plain Old C Structs) for each type, in addition to the current zero-copy readers/builders. You could use the POCS for in-memory state that you mutate over time, then you could dump it into a message when needed (requiring one copy, but it should still be faster than protobuf encoding).
https://capnproto.org/roadmap.html#c-capn-proto-api-features
1. If you have a remote object reference, you can (explicitly) cast it to any interface type, and then attempt to call it. If it doesn't implement that interface (or that method), an "unimplemented" exception will be thrown back. It's relatively common to do feature detection this way.
2. It's easy for the application to define a Cap'n Proto interface to support fancier negotiation. For example, you could have all your RPC interfaces extend a common base interface which has a method getSchema() which returns the full interface schema, or a list of interface IDs, or whatever it is you want.
3. We actually plan to bake in schema queries in a future version, such that all objects will support some sort of getSchema() call implicitly. This would especially be useful for clients in dynamic languages that could potentially connect to a server without having a schema at all, and load everything dynamically.
[1] https://github.com/sandstorm-io/capnproto/blob/247e7f568b166...
C++'s standard library for some reason decided -- relatively recently -- to introduce the term "promise" to mean something different (what most people call a "resolver" or a "fulfiller" for a promise), which is unfortunate. It's C++ that is being inconsistent here.
There are also subtle historical differences between the meaning of "promise" and "future". Historically, futures have usually existed in multi-threaded designs rather than event-loop/callback designs; you wait() on a future, blocking the calling thread, whereas with promises you call promise.then() to register a callback to call when a promise completes. The C++ committee again seems to have gotten this confused with their definition of "future", which looks more like a traditional promise.
For a while, capnp-rpc-rust used `gj::Promise`, which is based directly on the C++ Cap'n Proto implementation of promises (i.e. `kj::Promise`). Back in January, capnp-rpc-rust was updated to use `futures::Future` instead, and it was a fairly straightforward transition, as described in this blog post: https://dwrensha.github.io/capnproto-rust/2017/01/04/rpc-fut...
The trickiest part of the transition was dealing with scheduling. The implementation of `kj::Promise` has a built-in scheduling queue that guarantees a certain form of deterministic FIFO semantics, and those semantics are heavily depended upon in the Cap'n Proto RPC implementation. Rust's `future::Future` is less batteries-included, requiring capnp-rpc-rust to explicitly create queues where deterministic scheduling is needed.
Confusing the terminology perhaps even more, in capnproto-rust there is a type `capnp::capability::Promise` that implements `futures::Future`.
That said, it probably would have been better to drop Windows temporarily than to go two years without a release. But it always seemed like I'd find time in the next month.
Also note that Windows wasn't the only thing. I have a huge test matrix that I run for every release and, again, to avoid confusion, I want the whole thing to pass for any release. Things like building with -fno-exceptions or 32-bit builds or Android or ancient GCC versions tend to break frequently as the code evolves, but we really ought to have all these working for a release.
But maybe I'm just too OCD about this... :)
We now have AppVeyor (and Travis-CI) set up to build a chunk of the test matrix on every commit, which should help a lot going forward.
It's not OCD, it's good practice and I wish more open-source projects would follow that! Good job!
> Cap’n Proto is an insanely fast data interchange format and capability-based RPC system. Think JSON, except binary. Or think Protocol Buffers, except faster.
EDIT: done
http://www.art.net/~hopkins/Don/lang/forth.html
"The first Forth system I used was Cap'n Software Forth, on the Apple ][, by John Draper. The first time I met John Draper was when Mike Grant brought him over to my house, because Mike's mother was fed up with Draper, and didn't want him staying over any longer. So Mike brought him over to stay at my house, instead. He had been attending some science fiction convention, was about to go to the Galopagos Islands, always insisted on doing back exercises with everyone, got very rude in an elevator when someone lit up a cigarette, and bragged he could smoke Mike's brother Greg under the table. In case you're ever at a party, and you have some pot that he wants to smoke and you just can't get rid of him, try filling up a bowl with some tobacco and offering it to him. It's a good idea to keep some "emergency tobacco" on your person at all times whenever attending raves in the bay area. My mom got fed up too, and ended up driving him all the way to the airport to get rid of him. On the way, he offered to sell us his extra can of peanuts, but my mom suggested that he might get hungry later, and that he had better hold onto them. What tact!"