gRPC benchmark results
github.com
github.com
My guess is that little time was spent optimizing the C++ and Rust codebase, and Rust performs better because the code doesn't do copies.
https://www.reddit.com/r/grpc/comments/muy8dj/grpc_bench_ope...
Small clarification (to my understanding, I'm not a Java Guru) on why Java got on top - those Java implementations use something called Direct Executor. It's super performant when there's no chance of a blocking operation. But if you are to do anything more than echo service, you might be in trouble. Other implementations probably don't suffer from the same constraint
PR and discussions here- https://github.com/LesnyRumcajs/grpc_bench/pull/91
Java subs are celebrating it like they won a war. Java is a great language and will remain one of the top languages in the near future. But these silly benchmarks don't serve any purpose.
I am expecting to see this benchmark thrown around a lot from now on any language discussion. Whoever is doing these benchmarks should be more responsible and sensible.
If anything, to me it’s a clear example of why benchmarks like these are silly as they benchmark very uninteresting aspects that typically are never real world bottlenecks.
Instead here it seems to mean something more like "is capable of deadlocking if a thread pool is exhausted" or therefore roughly equivalent to "not lock-free". Is that correct? (But then the comments about `volatile` confuse me more - can you deadlock a thread with purely `volatile` access if your platform doesn't natively support it? That seems like a large failing of the JVM.)
==> Running benchmark for java_grpc_pgc_bench... Requests/sec: 25563.46
==> Running benchmark for cpp_grpc_mt_bench... Requests/sec: 31389.24
==> Running benchmark for dotnet_grpc_bench... Requests/sec: 25376.18
==> Running benchmark for go_grpc_bench... Requests/sec: 29158.60
==> Running benchmark for rust_tonic_st_bench... Requests/sec: 28120.25
Different images for the same language (java, rust, cpp) performed with similar if not worse results
Edit: what were your latencies and memory stats like?
I executed it giving both the client and the server 4 CPUs, now I'm running it again with 180s duration and gonna update the original comment if I get to see the report or there is any significant change
With C++ and Rust you will be able to implement this in a more optimized way as you can look for approaches that avoid copies which can be the main factor of a slow implementation, specially in multi-threaded scenarios.
It does, but the numbers are not too far. It's 3.5ms for Java versus 2.3ms for C++.
Even when running a lot of requests when 99% latency becomes the average one, that's like 1 millisecond difference, way under typical user's network latency.
These numbers are meaningless unless accompanies by specific usecases - don't just choose java or C++ just because the benchmark says its fast or the best.
High frequency trading needs that extra milli second, but a batch job or backend computation won't.
Latency histogram for fixed throughout (e.g. 1k RPS, 10k RPS, 50k RPS) are far more useful. You want to know if going from 1k to 10k increases latency.
These issues can be mitigated by carefully tuning the parameters of the Java virtual machine, but in practice, for most projects, it is not an issue.
No measurements are valuable all by themselves. None.
A long time ago I had to implement a real time service with Java. The best solution was to use whopping 16 megabytes of heap so that it would do a full GC multiple times per second but each of the GCs lasted less than a millisecond.
The plan is to parametrize all std collections on the allocator eventually, but that's not stable yet.
https://crates.io/crates/jemalloc-sys
"background_threads (disabled by default): enables background threads by default at run-time. When set to true, background threads are created on demand (the number of background threads will be no more than the number of CPUs or active arenas). Threads run periodically, and handle purging asynchronously. [...]"
> keep objects after the request ends
Yes. Can that be made in a way compatible with per request allocators?
Why would cloud provider be a business in the first place, do you think?
That point makes as much sense as complaining that if you were after performance you'd use a Formula1 car and not a Tesla.
People who live in the real world and have to do real work need to use real world tools, and one of which is AWS Lambda.
Let's put things in perspective: would it make any sense at all to advise a company to not only rewrite a whole application stack from scratch in your pet performant language but also jump head on to some boutique service provider? I mean, who in their right mind would get accountants involved in a goal to shave a few milliseconds over a few gRPC calls? Is a suggestion to improve performance expected to be considered even sane if it requires rewriting everything and change shop?
Another real tool is dedicated servers, something people used before "cloud" and something that people who care about performance still uses. AWS even offers dedicated servers themselves, so not sure why you would need to involve any "boutique service provider". Otherwise you have OVH, Hetzner and a range of others who compete well with AWS on dedicated instances as well, neither I'd say are "boutique".
> rewrite a whole application stack from scratch in your pet performant language
Not sure where this comes from, which one of these languages are "pet performant (SIC) languages"?
> who in their right mind would get accountants involved in a goal to shave a few milliseconds over a few gRPC calls
Hmm, unless the accountants are involved somehow in the API design (not sure what you're building), I don't know what the accountants have to do with anything here.
In the end, AWS and Lambda absolutely does not fit every use case. Depending on your use case, and if it's important a few ms here and there, you chose different solutions. Since this benchmark is about throughoutput, I thought we were discussing the use case of needing the best throughoutput, otherwise this is all off-topic. And with that, I'm just sharing that if that is your focus, you would probably not be using Lambda in the first place, as you'll get very shitty throughoutput and you have to pay a lot, compared to other mature solutions for this that we already had for many many years.
First time I heard of that and also doesn't match my own experience using AWS. AWS tends to have a lot of noise from neighbors but might be because of the region I was using. Could you link the official statement you got this from?
> The difference with metal is that you get the whole box for yourself.
Yes, + you normally avoid virtualization as that can have impact on your performance too. I'm glad we agree there is a difference that is worth mentioning when it comes to performance :)
At 15 minutes. Afaik it's all but T type instances that get reserved resources. Might have been different in early AWS, pre nitro etc.
Something that i just remembered is that in AWS dedicated means that the hardware is dedicated to you, so no other customer vm on it, but you can have multiple vm's on it.
What do you count as "large number of users"? Worked on projects with millions of users, thousands of requests per second and when we refactored based on increasing throughoutput/decreasing latency we had exactly 0 accountants involved, even if the company had accountants in-house and full-time.
I think it depends more on the company size than the number of users you have.
> Engineering history is littered with projects to shave milliseconds or cents off a very small component which is used extremely frequently.
Yup, agree and been there myself, hence my comments about staying away from Lambda and VPS for this kind of focus and go for dedicated instances where this matter.
Maybe interacting with the finance department was above your pay grade. The bigger the operation, the more costs matter. That's a thumb-rule you can find anywhere.
It was not, we were a nimble and lightweight team with everyone doing everything they could possibly do. I had insight into accounting and helped them with implementing some stuff and some of the designers helped with frontend as some of them knew HTML and CSS and so on.
> The bigger the operation, the more costs matter. That's a thumb-rule you can find anywhere.
Yeah... I think I agree? "The bigger the operation" is referring to the employees working for the company, not the number of users right? If so, that was exactly my point. It's about the number of employees that dictate if accountants gets involved or not, not the number of users.
You should really rephrase that as "I was totally unaware that there were accountants involved" because it is simply inconceivable that a business activity involving allocating resources and changing operational needs would not be tracked. Either you are grossly misrepresenting your personal anecdote or you are filling in quite a few blindspots.
Whatever happened to "Assume good faith"?
The whole point is that you don't. You literally pay for the memory your application requires, proportionally to the amount of memory. The difference in the resources required to run is around two orders of magnitude.
Let's put things in perspective: in some platforms such as AWS Lambda, you are charged per memory used per second.
Paying for the amount of memory you use down to the MB, is hardly something everyone does, which makes it weird that you are now the third reply that assume this runs on AWS/Lambda.
It's not an edge case. It is a clear and irrefutable example that you pay for the memory you use.
Let's be very clear here: when you provision a VM anywhere in the world, you have to pick how much memory you require. You are charged for that memory, proportionally to the memory you require. If your app requires over 100x memory to run, that comes out of your wallet.
Never said it was an edge case....
> when you provision a VM anywhere in the world, you have to pick how much memory you require
Yes, thank you! That's exactly my point! You create a VM (or a dedicated instance) somewhere, they ask you for the memory usage up front. Usually they start at 128MB and go their way up from there.
Even if you take the instance with the smallest amount of memory, you'll fit any of the benchmarked programs, effective making the "115.41 MiB Java vs 4.15 MiB Rust" statement not as important anymore.
> You are charged for that memory, proportionally to the memory you require. If your app requires over 100x memory to run, that comes out of your wallet
Hm, maybe on Lambda you pay more if the memory your app uses grows. But this is certainly not standard for "normal" hosting where you rent either a VM or proper instance. Then the memory grows until it cannot take any more memory, and the process either crashes, gets killed by OOM or does whatever operation you've designed it to take when running out of memory.
The whole point is that memory costs money, and the more memory you require, the more you pay.
AWS Lambdas is a clear example that demonstrates this, but is not an isolated case. This is the case for all cloud providers selling VM time. All of them.
Even bare metal providers charge you more for more memory.
Isn't it clear that if your app requires more and more memory, that comes out of your pocket? Is this something that really warrants a debate?
For example, anecdotally a small, semi-complex JavaFX 2D game I wrote uses by default something like 250 MiB of RAM, but manually limiting it to max 80 was possible (but in the latter case, the GC had to run quite often)
Not just that, but they were very sensitive to fluctuations.
You have exactly 0 information available and yet you make a comment like this? Is this why we have cargo-culting?
Not every solution to a problem needs to use as little memory as possible. And also not every solution can ignore memory usage. It depends, of course.
Since this thread is about throughoutput and optimizing for that, I could tell you that most times I've been put in charge for optimizing throughoutput, having low memory usage have been pretty far down the list.
Those problems are called "one-offs" relative to things I design for scale.
Yes, "relative to things I design"
You think that applies to everyone?
You think that comes close to applying to people focusing on improving throughoutput specifically?
But that's not the point. People aren't suggesting micro managing memory. Java has a well deserved reputation of being a memory hog and its using 20x what is needed in this benchmark.
The explanations could be: a) the bytecode generated by the Kotlin compiler is much worse and b) the benchmark Kotlin code is not as good or optimized as the Java one.
* Exception is supporting things like default methods on JVM 1.6 bytecode
Additionally it is stuck on Java 8 view of the world, otherwise those .class files won't be usable on Android toolchain thanks Google.
I know for a fact it uses different bytecode features if the level is >= 1.8, not sure how smart it is above that.
Going forward while for Java code there is no worry about using SIMD, JNI replacement, value types, Kotlin code will need to make use of KMM for code that is supposed to target both JVM and Android.
It can do stuff android doesn't support if you set the bytecode target level to > 1.8
It's like using modern Javascript features but providing a polyfill for older browsers.
Also stuff like value classes will have different semantics in memory consumption and performance across targets.
Maybe some learning required?
You are either being deliberately obtuse or you genuinely have difficulty to understand that Java version, bytecode version, Kotlin compiler output and what bytecode Android supports are totally independent concepts.
It is going to be fun to port back stuff to Java.
https://github.com/LesnyRumcajs/grpc_bench/blob/master/kotli...
https://github.com/LesnyRumcajs/grpc_bench/blob/master/java_...
Looking at the build files, different versions of the libraries appears to be used. The Java implementation is configurable but defaults to something called "direct executor".
Maybe that explains the difference.
In my case, I have a programming language interpreter implemented in Kotlin, and as an experiment I made the entire interpreter using suspending calls so that I could call asynchronous functions, and my performance tests dropped by about 20%.
Irreducible loops make many optimizations much more complex. Bytecode generated by javac never contains irreducible loops, and since bytecode generated by javac is the number 1 use case targeted by JVM JIT compilers, they probably just don't bother trying to be that smart about irreducibility.
Since openjdk has had to cope with lack of (user defined) value types and the garbage heavy ecosystem it has a very well optimised GC for handling heap allocations.
I wonder if the results will be different if other language benchmarks (C++ or rust) use different allocators (eg: bump allocator, per request collectors) for cheaper allocation.
pgc: ParallelGC sgc: SerialGC g1gc: G1GC she: ShenandoahGC zgc: ZGC
EDIT: Someone else noted (https://news.ycombinator.com/item?id=27085507) a discussion on Reddit where a <5ms latency was achieved in 99.9% cases, so perhaps this is indeed a subpar result.
For the single core test, we can infer an upper bound of 20-33us of server cpu time per request for most of the relevant languages for gRPC. That seems pretty good.
https://github.com/fujita/greeter-bpf
2-3x faster than GRPC-Go.
Not really anything here:
https://github.com/LesnyRumcajs/grpc_bench/tree/master/java_...
COPY java_grpc_sgc_bench /app
Directory here: https://github.com/LesnyRumcajs/grpc_bench/tree/master/java_...https://github.com/LesnyRumcajs/grpc_bench/tree/master/java_...
90k req/sec with p99 under 1ms
ghz --proto=/proto/helloworld/helloworld.proto --call=helloworld.Greeter.SayHello --insecure --concurrency="50" --connections="5" --duration "60s" --data-file 1kb.json 127.0.0.1:50051 --cpus=4
Summary:
Count: 5411979
Total: 60.00 s
Slowest: 15.24 ms
Fastest: 0.03 ms
Average: 0.34 ms
Requests/sec: 90194.36
Response time histogram:
0.026 [1] |
1.548 [997969] |∎∎∎∎∎∎∎∎∎∎∎∎∎∎∎∎∎∎∎∎∎∎∎∎∎∎∎∎∎∎∎∎∎∎∎∎∎∎∎∎
3.070 [1426] |
4.592 [251] |
6.114 [108] |
7.636 [87] |
9.157 [80] |
10.679 [17] |
12.201 [40] |
13.723 [11] |
15.245 [10] |
Latency distribution:
10 % in 0.12 ms
25 % in 0.19 ms
50 % in 0.29 ms
75 % in 0.42 ms
90 % in 0.58 ms
95 % in 0.72 ms
99 % in 1.04 ms
Status code distribution:
[OK] 5411959 responses
[Canceled] 3 responses
[Unavailable] 17 responses
Error distribution:
[3] rpc error: code = Canceled desc = grpc: the
client connection is closing
[17] rpc error: code = Unavailable desc = transport
is closingAlso the tool used in that benchmark ghz is also written in Go, so somehow the client is able to have high troughput and verify requests but the server can't?
If I remember correctly the bytecode is interpreted and executed by the runtime. This means the bytecode is fully in memory as bytecode when the jvm loads .class files. The native code for the JVM interpreter and runtime execution is also present in memory.
At this point execution happens without "duplicate code" in memory.
Note: the JVM could compile the bytecode fully into native code before execution but that's an implementation detail for a specific JVM implementation.
Then we add the JIT. The JIT will notice hot paths during interpretation/execution and compile a method to native code and inject it so it calls the jit_native version instead of the bytecode_interpreted version. The JIT does not overwrite original bytecode but stores the compilation result separately.
So to come back to your statement "both bytecode and native code generated from bytecode in memory at the same time": yes, but only for the JIT-compiled hot paths.
This to me is the "power" of the JVM + JIT. Run interpreted bytecode and compile to native for the hot paths.
Seriously, every now and then there's a story about some fintech company who use Java and use large enough heap that there's no garbage collection during stock market opening hours.
Why isn't teh SDK/Language major versions not mentioned in wiki page ?
Am I making mistake in reading this .... As the number of CPUs increases req/s and latency does not change, except for few java and rust readings. Network connections and IO isn't seem to scale as number of CPUs increases ?
Also, I believe the heavily editorialized title should be replaced with something that is close to the title of the original post, such as "2021-04-13 gRPC benchmark results".
> Otherwise please use the original title, unless it is misleading or linkbait; don't editorialize.
Edit: it’s not even php, it’s roadrunner (a golang server) proxying for php. By default, it’s probably not configured for production.
Unfortunately I've never found a best-in-class modern RPC system for C++ that embraces integration with other libraries (for I/O, concurrency, etc)
Thrift (Apache not Facebook) and CapnProto are the only other contenders that I know of.
I know the Cloud APIs use gRPC, but only on the client side. But I'm under the impression that gRPC and Stubby (the internal RPC framework?) are completely different codebases (even if they are allegedly similar).
Results for 2 and 3 CPUs are strange too. So little scaling, what's the point...