gRPC request context which caries values across microservice boundaries
github.com
github.com
It's called KonigKontext, check it out here: https://github.com/konigsoftware/konig-kontext!
Example use cases include:
Propagating security principals, or user credentials and identifiers throughout an entire request lifetime across all of your microservices. Propogating distributed tracing information. Set a request trace id upon receiving a request and later access that id in any downstream microservice. However, Konig Kontext is built to support any type of context value, so it can be extended to fit any specific use cases as well.
Let me know what you think!
It allows for the opaque propagation of side channel information, where intermediates don't need to know what's there.
side-channel-with-well-defined-propagation-rules is pretty useful.
Maybe my reaction is a question of purity. I don't really see why one would think a request body should be in the schema but we should leave other data out of the schema. Wouldn't every single argument point to consistency?
I like gRPC and what it gives you. I personally would like that same explicit schema and type safety to apply to my tracing as well. Its interesting to me that others would draw a line.
This applies especially to grpc, where everything is optional.
To be clear, I prefer explicit over implicit. But that doesn't always scale well to large orgs.
The combination of out-of-band/out-of-schema data with APIs to set and get from app code seems like a bad match.
I'd prefer if it was completely invisible to app code and existed only in interceptors.
So, I’d argue the OP is using gRPC correctly, and I’d be using it wrong if I hadn’t given up on it long ago. I appreciate things like static type checking, and services that err on the side of rejecting unparsable requests in order to avoid data corruption. Protobufs/gRPC require heroic effort on the part of the RPC handler implementation if you want those things. In particular, it is easier to implement your own serialization format than typecheck the results returned by the APIs protoc emits.
It’s one of those things you eventually wish you had.
[1] https://github.com/konigsoftware/konig-kontext/blob/84faa627...
The library is convenient for sure, but I feel that if you had a need to propagate context within gRPC, you'd probably already discovered the API and implemented propagation with your own header keys.
[1] https://grpc.github.io/grpc-java/javadoc/io/grpc/Context.htm...
[2] https://grpc.github.io/grpc-java/javadoc/io/grpc/ServerInter...
[3] https://github.com/konigsoftware/konig-kontext/blob/84faa627...
e.g., I handle a request, get the incoming context, have to stash it because I might execute a coroutine that is suspended/resumed across different threads, and subsequently then execute another GRPC call in a Java library, that happens to start, get rescheduled and resume on receiving the response on a different thread, in a possibly different thread pool.
The OpenTelemetry handling for this is quite complex: it must be used as a javaagent so it can actually instrument underlying libraries with the necessary code for handling thread scheduling/context switches in both thread pools (e.g., ForkJoinPool), threads themselves, with cooperative scheduling in application code (e.g., Thread) and Kotlin's coroutine handling with is mostly codegen (e.g., async, suspend fun.)
Finally, in my own Ph.D. work, we did a similar thing to propagate trace identifiers for a dynamic analysis for fault injection, and we quickly ran into a problem that --- not only is the propagation difficult in itself --- but, you also run the risk of running out of header space if you store any (longish?) information when GRPC is run over HTTP2 because of the maximum allowed header size.
Do you actually mean that unless explicitly propagated to a subsequent downstream RPC the data is dropped? If so, that's by design.
However, most large-scale organizations that are doing distributed tracing (e.g., Twitter, Uber) have either invented, reproduced, or leveraged OpenTelemetry's design for this precise thing.
Naive context propagation isn't (really) the difficult part with most of these designs -- it's what you've done, using an interceptor, reading the data and assigning it automatically on subsequent requests -- the challenge is dealing with this under many different, real world conditions: a.) concurrency and thread scheduling; b.) not all services use the same version of downstream RPC libraries; c.) not all calls are GRPC, and some use HTTP (and, different HTTP libraries, at that.); and d.) you cross message passing boundaries: i.e., I receive request, write to Kafka queue/reliable workflow backend (e.g., Cadence, Temporal) and re-read the request and then execute a subsequent RPC as a result of that message.
If you're using Kotlin, I suspect you will run into these challenges. Tune your thread pools up/down, restrict your JVM's resources, and you'll suddenly see that if the thing that handles the request uses different threads/coroutines/etc. then the code block that issues the downstream RPC, you'll start dropping the context without explicit handling of that case.
In fact, a very simple test case in Java where you use several, concurrently executed CompleteableFuture's that each issue RPCs, in a very small thread pool should be enough to see the issue.
Disclaimer: I wrote the OTel one, and am sad to see yet more context implementations being made, including the OTel one, these really need to all be on the way out.
This was a solution that has worked well for my company that averages <1 req/s, so yes I have not tested it under more extreme conditions. This is version 1.0.0 so it is quite new and naive by design. I was posting here to get some feedback on the initial version and see how I can improve it, which you have given me!
Feel free to contribute to the project! It seems like your expertise applies nicely!
- User credentials e.g. JWT
- Distributed tracing identifier
Soon someone will find a way to keep it in the same process.
What if you have a nested stack of calls where microservice A calls microservice B which calls ... etc. Then you're looking at the schema of microservice F, and even the source code where it's called in microservice E, and can't figure out how to you can call it to get it to do the same thing. Little do you know, you need to set some things up that currently only A knows how to do.
public interface MyWellDefinedBoundary {
...
}
;)Edit: I should say this is what OSGi was good for. Now it's been replaced, its need is greater than ever.
Type systems and APIs do that, but only within a single language.
Think that's called Erlang.
> Imo microservices should be a logical boundary only.
Having option to split it freely is a costly abstraction to deal with, especially if it is cross-language.
I prefer to just leave "cutting lines" in monolith. Well defined modules and relations between them so if some feature needs to be spun off it's not too hard.
No, it's not.
As someone who programmed Erlang both professionally and published academically at Erlang venues for a long time, no.
These optimizations "for runtime" are not well supported by Erlang (i.e., cluster performance changes dramatically when behavioral characteristics of message passing switch from local to remote to remote cluster very quickly) and were long discussed in Waldo's paper back in the 90s, dynamic relocation is not supported well (i.e., unless you use global, which falls apart quickly under network anomalies, of which I, and several others, wrote paper(s) about), and the runtime hardly provides any information on introspection on cluster performance.
Sadly, distributed Erlang had the edge on programming distributed systems almost 20 years before they became pervasive, but has since been left to atrophy and hasn't seen any real innovation in quite a long time.