If you read the Cap'n Proto RPC docs, everywhere where it mentions "traditional RPC", I specifically had Google's internal RPC in mind (having previously been the maintainer of Protobufs at Google). So, you can more-or-less substitute gRPC in there for a direct comparison. https://capnproto.org/rpc.html
There are two key differences:
1. Cap'n Proto treats references to RPC endpoints as a first-class type. So, you can introduce a new endpoint dynamically, and you can send someone a message containing a reference to that endpoint. Only the recipient of the message will be able to access the new endpoint, and when that recipient drops their reference or disconnects, you'll get notified so that you can clean it up. This is incredibly useful for modeling stateful interactions, where a client opens an object, performs a series of operations on it, then finally commits it. Put another way, this allows object-oriented programming over RPC. Also note that you can easily pass off object references from machine to machine -- currently this will set up transparent proxying, but in the future we plan to optimize it so that machines automatically form direct connections as needed, which will be really powerful for distributed computing scenarios.
2. Relatedly, Cap'n Proto supports "promise pipelining", which allows you to use the result of one RPC as an input to the next without waiting for a round-trip to the client. This makes it possible to use object-oriented interaction patterns with deep call sequences without introducing excessive round-trip latency. This is described in detail at the RPC link above.
Server->Client streaming is e.g. a powerful way to get realtime updates about some state on the server to the client without needing to poll (not performant) or needing to define callback services (ugly to maintain, service lifecycle questions, and if you need an extra connection for from the service to the callback service (client) there are also challenges around routing).
Native support for streaming also removes the need to model explicit flow control (backpressure) behavior in the user defined APIs (like functions for requesting some items in addition to the callback functions for delivering items).
Another difference is that grpc utilizes HTTP/2 as underlying transport protocol. However I'm thinking that's mainly an implementation detail. If Google had designed grpc slightly different (not relying on barely implemented features like HTTP trailers) it could have been a bigger differentiator, e.g. by allowing browsers to directly make grpc calls without proxies.
The drawbacks of callbacks that you mention don't apply in this scenario: The same network connection is utilized for the callback, solving the routing question. The callback object receives a notification when the caller is done with it (or disconnects), so it can clean up, solving the lifecycle question. Flow control / backpressure can be achieved by pausing calls to the callback if too many previous calls have not returned (I actually plan to bake this pattern into the library in the future to make it dead simple).
So, it seems to me streaming in gRPC is actually a narrow subset of what Cap'n Proto can express.
For web stack compatibility purposes, it's straightforward to implement Cap'n-Proto-over-WebSocket. I'm not sure that HTTP/2 buys much on top of that.
Regarding backpressure and HTTP/2: You get some different kind of backpressure behavior with your approach (application level flow control) and the grpc approach (transport level flow control). Let's say you have 2 functions which are called in parallel. For one the arguments are very big (let's say 100kB), for the other one they are small (some bytes). With HTTP/2 and transport level flow control the small function could get the bandwidth as the big one, which means more requests/s for the small function. With pure application level flow control the small function needs to wait until the big one is fully sent before it can be put on the wire.
Which means if I have 2 tasks that execute concurrently and look like
A: while (true) { await callBigFunction(); }
B: while (true) { await callSmallFunction(); }
then with grpc I get more calls for B and without Cap'N'Proto I get the same for both (correct me if I'm wrong).I don't think it's a huge disadvantage in practice, because if you have large data chunks you should probably design your API different - and if you want realtime behavior then probably both protocols are not optimal. But it's still something to keep in mind.
Besides building more support in the library it would maybe also be very worthwhile to have these patterns described together with examples on the website. E.g. how to achieve callback interfaces which don't user another connection, server->client streaming, etc.
For me as a potential user information about these kind of features would make Cap'n'Proto immediately look more interesting. When I look at the website I only see Promise Pipelining described as a main feature - which is novel for sure, but not something that attracts my interest if I look for the other features we discussed about.
So... CORBA?