The jvm and v8 is just another layer of abstraction that gets in the way when performances problems like this arise at scale.
The tools used in this article require an understanding of what the hardware is actually doing.
The jvm and v8 is just another layer of abstraction that gets in the way when performances problems like this arise at scale.
The tools used in this article require an understanding of what the hardware is actually doing.
However, the issue in the Netflix blogpost is in the JVM C++ code. I think it's entirely possible to encounter the problem in any language if you're writing performance-critical code.
What I find a little odd is why those variables were only on different cache lines 1/8th (12.5%) of the time. What linker behavior or feature would result in randomly shifting those objects while preserving their adjacency? ASLR is the first thing that comes to mind, randomizing the base address of their shared region. But heap allocators on 64-bit architectures usually use 16-byte alignment rather than the random 8-byte alignment that would account for this behavior. Similarly, mmap normally would be page-aligned, even when randomized; or certainly at least 16-byte aligned?
edit: Pointers will be 8-aligned. Random 16-byte allocation if one pointer is at 8-offset and the next is at 0-offset will sometimes give you a cache-line crossing. Admittedly it should be 25%, not 12%... Maybe Java's allocator is only 8-aligned?
Which doesn't apply to 99% of business applications out there.
I think those are both probably legitimate approaches, but if it's the latter you really want to understand how the hardware works.
I think that’s more what OP meant.
I’ll also note that while I don’t know what this service is doing, 300 RPS seems mighty low. That means each core is doing 30 RPS if you’re 12 wide. Maybe 60 since I don’t know what that 50% autoscaling piece means. Think about that. Each request is burning 16ms of CPU time. While I don’t know what this service does, this does seem like a lot when you consider just how fast a modern CPU is. That being said, it’s not impossible that even tweaking the Java code itself might unlock more CPU and that would be a more cost effective solution than using native code which is typically a bad fit for most web services (even rust).
Possibly. Equally, maybe that runtime type information is what enables monomorphism that wouldn't be possible in C (Yes, C doesn't explicitly have virtual functions, but you end up needing to do the same thing, so you pass function pointers around and it's the same thing at runtime), and the result is better performance.
> That’s a pretty significant RAM and performance saving and gets you within shouting distance of hand tuned assembly, especially since hot paths ARE frequently hand-tuned assembly.
In my experience higher-level languages significantly outperform lower-level languages for realistic-sized business problems, because programmer effort is almost always the limiting factor. Most of the people using C/zig/rust because "muh fast" don't even know how to use a profiler, yet alone the level of multiple-tool analysis that we see in the article. And again, this is the kind of thing you'd need to be doing in C to address performance issues in the same kind of code. Sure, maybe you don't have reflective type information so it shows up as a mispredicted branch rather than a cache issue - but guess what, you also need this kind of low-level tool to get that information, C won't tell you anything about branch prediction either.
You set up a .NET/JVM project and it takes off, and you find out you're facing massive memory bloat of the managed heap. What do you do? The answer is basically: try using this slightly different allocation pattern by flipping a switch in the runtime, OR try to do a bunch of manual memory management anyway. You quickly discover the allocation patterns don't help much, so you turn to manual management.
Often enough, its simple to understand where the memory allocation comes from, and you could probably fix it with an arena allocator or something simple. So you try that, but then you find out that many library functions in .NET/JVM that you need, don't allow you to pass in preallocated buffers, leaving your ability to solve the memory problem crippled. You now where the memory is needed, when it expires, and when it can be reused, but you don't have the tools to apply this information anywhere.
At that point, you can either leave it be and buy more RAM, or rewrite from scratch in another language. Would be cool to have languages that are more of a hybrid, kind of like .NET unsafe, but then in a non-optional way.
A) The business needs evolve quickly and/or cost cutting dev time is more valuable than having the application perform quickly.
B) the performance of the app isn’t that critical and developers will have a more pleasant experience in a better ecosystem
C) the managed language is mainly responsible for the control plane of a data path.
Anything beyond that, especially high performance code, will struggle. That’s why you see dev cycles spent on encryption algorithms, common low level routines, to the point of hand vectorizing assembly, etc. It’s a very well defined problem that’s fixed and has an outsized payoff because those things are used so often and used in places where the performance will matter.
There’s probably a 10x to 100x reduction in worldwide compute usage possible if you could wave a magic wand and have everything run as optimally as if the best engineers built everything and optimized everything to the levels we know how to do things now (computers have never been faster and never felt slower because of this).
However, it’s just not economical in terms of dev cycles per output and complaining otherwise is tilting at the windmills of market forces that give the edge in many scenarios to those languages (even when you factor in inefficiencies and/or extra time optimizing ). That’s what the person you’re replying to is stating and is something I’m 100% agreed with. This is someone who is a systems engineer who codes primarily in lower level languages and is generally a fan, especially Rust. I have probably written non trivial code in every popular language out there at this point and they’re just faster to get shit done in. Picking the right language is a mixture of figuring out what kind of talent you can attract, how much you can pay them, and what will satisfy your business needs within those constraints. A lot of people do choose incorrectly or suboptimally due to ignorance of how to choose, ignorance of the alternatives, picking because of personal familiarity the original author has, etc. Those choices are often more fatal when you choose a native language if your competitors choose better whereas the converse is less often true.
However, what is in the realm of achievable engineering is improving the performance of the code you write yourself, but a large part of that performance is opaquely hidden inside the runtime, with no way for you to change anything. If we were to create a language that has a smooth path from "managed-runtime" to "low-level-freedom" we would be able to adapt our codebases as they become more popular and performance starts to matter. What that would look like, no idea.
Consider however: * when you’re small your business bottlenecks are other things
* when you get larger you have more resources and can make a choice to switch or hire experts to remove the major bottlenecks
* improvements to the runtime improve your scaling for free
I’m not speaking theoretically. I worked at a startup that indoor positioning. Our stack end to end was written in pure Java. We struggled to attain great performance but we did what we needed to and focused on algorithmic improvements. We managed to get far enough along the way to get acquired by Apple. Then I spent about 4 months porting the entire codebase verbatim to C++. It ran maybe about as fast (maybe slightly faster but not much). Switching to the native math libraries for the ML hot path gave the biggest speed up but that’s more a problem with the Android ecosystem lacking those libraries (at least at the time or us failing to use them if they did exist). Over the course of the next two years we eeked out maybe an overall 5-10x CPU efficiency gain vs the equivalent Java version (if I’m remembering correctly - I regret not taking notes now) but towards the end it was definitely diminishing returns territory (eg changing a vector of shared_ptr objects on the critical path of the particle simulator netted something like a 5% speed up).
This was important work here because we got to a point where battery life actually started to meaningfully matter for the success of the project whereas as a startup we were trying to survive. But we were always conscious that Java was the better tradeoff for velocity and writing the localizer in c++ carried logistical challenges in growing it. In fact, starting in Java meant that we had an easier time because a lot of the initial figuring out of the structure of the code (via many refactorings and optimizations) had already happened in Java where it’s easier to move faster. Even within the startup we had discussions about migrating to C++ and it never felt like the right time.
My point is, good engineers know when to pick the right tool for the job and know what their risks are and what their contingencies are. If there’s necessity you’ll either change tools or fix your existing ones. Of course, not everyone does that, but my hunch is only those that are going to succeed anyway end up being fine. Kind of how the invisible hand of the market ends up working.
I think the idea of a smooth transition is a fantasy. Sometimes the architectures are so different that you’d have to fundamentally restructure your application to get that jump. It’s a map with many mountain ranges and valleys. There’s plenty of local optima and you can easily get stuck which requires you to fundamentally rearchitect things. For example, io_uring is very different. If you want to eke out optimal performance out of it, you need to build your application around that concept. It’s rare you get something like Project Loom in native landed but that’s a point in favor of Java managing things for you - free architectural speed up without you changing anything.
That's just a hair over 7 requests per hyper thread per second, or 135ms per request. That second number can be seen in the last graph, where the last graph shows an average request latency of about the same value, I'm guessing 120 milliseconds...
Without knowing what the code is doing, it's unfair to judge, but you've got to wonder what Netflix does to burn through $1B-per-annum in cloud infrastructure expenditure. Consider that at that rate they must have a truly epic deep discount, so that's probably equivalent to $3 to $5 billion at on-demand or retail pricing. Bonkers!
> We tend to think of modern JVMs as highly optimized runtime environments, in many cases rivaling more “performance-oriented” languages like C++. While it holds true for the majority of workloads
What you say more aligns with my thoughts. I've never heard someone express the opinion in this article. Then again, I don't know anyone quite as knowledgeable about the jvm as this article is.
I remember there being a writeup about paint.net and what they had to do to keep performance. They were worrying a lot about memory and how to manipulate .net to manage it.
If you want performance, you don't get to ignore memory no matter the language, but some languages/platforms make it easier and some get in the way.