Shenandoah in OpenJDK 17: Sub-millisecond GC pauses
developers.redhat.com
developers.redhat.com
Caching is often not such a bad idea regardless of GC.
Cache effects might play a role here as well. Deallocating memory immediately after use allows you to allocate it again when it is still hot in the cache. So eg if you're iterating a structure and doing lot of interleaved allocations/deallocations they are likely to not cause cache misses (assuming the allocator is not dumb and uses some kind of LRU strategy).
With GC and deferred deallocation, you're moving a lot more data between the main memory and the caches, because memory for sure gets pushed out of cache by the time it is reclaimed. And additionally, GC has to touch quite a big number of objects when tracing (and move unneeded stuff to cache, and push needed stuff out of cache). Memory bandwidth is a scarce resource these days.
If you want to see these effects in extreme, try running a GC based program in an environment that is low on memory but has swap enabled. Hitting the first full GC is basically performance game-over, regardless concurrent or not.
On the other hand, manually managed apps can often deal with big chunks of their heap swapped out without terrible consequences.
Java: pauses are bad, let’s make the garbage collector better.
Much credit to the Java community for ignoring the noise and building something great that lets most applications have the best of both worlds.
[0] https://awesomeopensource.com/project/ixy-languages/ixy-lang...
But I’m by no means anywhere close to an expert in the topic.
Contrary to the WebAssembly speech, CLR was one of the very first VMs to have support for C like languages, but those capabilities were only fully surfaced to the original Managed C++ (.NET 1.0) and C++/CLI (2.0 onwards).
So now those capabilities are being exposed into C# as well.
For example stack allocation of arrays is now also valid in safe contexts.
You can apply using to any type that has Dispose() without inheriting from IDispose.
Memory slices are now also a thing and native function pointers for example.
Last I saw, this was still being debated by the LDT and weighed against similar proposals such as introducing a `defer` keyword like Swift/Go.
I’m working on a website that has been growing since 1996… The chances of us ever migrating to .NET 5 and onwards is basically zero.
As for C and C++, yes, of course. I meant "everyone else" as in the trendy new languages.
But no matter what, it is literally a GC algorithm, described in every book about GCs.
I worked on one which was required to respond to a packet from the network within 5 microseconds of the packet arriving.
Please, don't assume all applications are mobile uis, webapps and corporate backends.
Neither is software running in all your electronics.
All these require software that has some different requirements that typically are incompatible with unpredictability of a garbage collected dynamic language.
Garbage collection gets a bad rep from garbage collectors that are tuned for throughput rather than latency, but there is no lower limit to latency for garbage collection and there are GCs out there that prove it (like Azul). If you can afford dynamic memory allocation, you can afford garbage collection. And if you can't afford it, you can just preallocate and turn off the garbage collector.
This is 100% false. It's easy to configure Linux itself to only occupy 1 or 2 cores (one real + the hyper-thread for that real core), and then pin your own application threads to the other cores where no other code will run. This setup has identical application performance (other than not having those two Linux cores) to having no OS at all. 100% predictable, ZERO jitter application code (including user-space networking).
Except…you can still SSH in to your "no OS" Linux box, you have a file system, can run cron jobs, git, gdb, WireGuard, etc. So it's much, much better in practice and what actual high-performance network developers actually do.
Essentially no one has run a literal "no OS" box for at least a decade; everyone who is running on normal hardware does what I just described with Linux because the cost is dirt cheap and it's way too convenient.
https://www.ptc.com/en/blogs/plm/ptc-perc-virtual-machine-te...
Better deflect that incoming missile on time before the next GC pause.
Aegis is built for defending against missiles going at most mach 4-5. Mach 5 is ~1700 m/s, so during a millisecond GC pause it will travel slightly under 2 meters. Ship-based SAMs will have proximity-fused fragmentation warheads and even back in the 70s those had a kill radius of >100m (see for example https://en.wikipedia.org/wiki/S-75_Dvina). So even a GC pause at the worst possible moment would not significantly impact the probability to hit the target. Even really fast things in the real world are really slow compared to computers.
(Of course, in practice the Aegis system will only provide midcourse corrections to the missile. The fuse and end phase guidance control software would not interface much with the rest of the combat management system anyway.)
The rockets close head on (so the speeds add together) and then the rocket trying to kill the other one explodes at right distance to the side of the other rocket, exploding in a cone of debris.
That cone of debris must intersect the other rocket at a right point. It can't just hit anywhere, to reliably kill it must hit a specific part of the other rocket.
So, realistically, you have to time your explosion to something like tens of microseconds.
Also getting there is a control loop problem and it requires very precise control. The more jitter in execution of the control loop the worse steering it will be and less chance it will get to the right place at the right time.
First of all, a SM-2 will not use the flight path you describe as they tend to "dive down" on a sea skimming missile. But even the shorter range Sea Sparrow missiles that do behave as you describe use multiple guidance phases. In the early and mid-phase guidance phases, the radar reflections from the target missile will not be strong enough for the relatively small receiver in the missile, so the system must depend on midcourse guidance updates (little bit to the right, little bit up, etc) from the combat management system on board the launching vessel to navigate to a "close enough" location from where the sensors on board the missile can acquire the target. The end-phase guidance and fuse timings are quite critical, as you say, but they are also not really under command of the CMS but rather done by embedded systems on board the missile. Generation of the mid-course guidance updates is not nearly as time critical and won't suffer much from a 1 ms pause. The error in location estimation of the target due to sensor imperfections (limited angular resolution, clock jitter, athmospheric effects, etc etc) will be way larger than the 2 meter from the GC pause.
I can't imagine a scenario in which an enemy would be able to saturate the defense system with a continuous onslaught that would make the system run out of memory. You'll sooner run out of SAMs and ammo for that fancy gun that's used to shoot down missiles at close range. And if the enemy can mount such an onslaught, then the ultimate GC will soon run anyway - all memory will be released when the ship sinks.
(I think back to that famous story of a missile with a memory leak, whose designers figured the missile will hit its target or run out of fuel faster than it'll run out of memory.)
But to answer your question, there are cases where the JVM is used without GC and it runs for a day, and gets restarted at night. Afaik some bank does it for some form of HFT.
https://devblogs.microsoft.com/oldnewthing/20180228-00/?p=98...
"[...] Since the missile will explode when it hits its target or at the end of its flight, the ultimate in garbage collection is performed without programmer intervention."
Dev A "You should use X"
Dev B "We tried X, it was way to slow"
Dev A "You must be doing it wrong, I use X all the time and it's very fast"
This went back and forth like this for a good 15 minutes before they realized they where both getting the same performance, but just had a very different definition of "fast" and "slow".
Sounds like HFT? But then your competitors are probably using an FPGA to respond in 300 nanoseconds, so good luck with that 5 microseconds tick-to-trade.
I don't even know how you came to the conclusion that I think that "everything has to be either manually submitted order or HFT" but that's on you.
I hope this means we can get more games in Java since I'd like to code one or two in my favorite language :)
Ah, good times. I currently work in Hadoop and it's honestly stupid how much tuning a Spark job is like tuning a Minecraft server (from the JVM level anyway). Life's path is strange sometimes.
But then again, that's a problem you'd get without GC too, just maybe a bit less harsh/unpredictable.
Perhaps a nitpick: 'moving' garbage collectors [0] like Shenandoah don't 'eat through' memory occupied by now-unreachable objects. They work by copying/moving the still-reachable objects. Their performance is a function of the number of still-reachable objects, and their size, rather than the unreachable objects.
[0] https://en.wikipedia.org/wiki/Tracing_garbage_collection#Mov...
Gaming in Java has always been possible.
And, I don't "do my own memory management". I just don't rely on a GC to do it.
Reference-counting GC is available when it is not too costly, which is common, and more convenient, which is rare. Ordinary automatically-generated destructors handle almost everything, almost all the time. Once in a great while, performance demands a concession such as an arena allocator, which also is not "manual memory management", and also not GC, and is radically cheaper than either one.
I'd use Java if it let me annotate when I want things to be deleted.
Taking a dogmatic view on an engineering trade-off isn't good engineering, and closes you off from exploring other, potentially better, ways of solving the problems you actually care about. (Which probably have nothing to do with the details of memory management)
Something tells me you wouldn't say "I prefer exactly no borrow checker [in Rust], and I don't think I'm giving anything up by skipping them".
(Comment was edited while I was writing a response, original comment I responded to above)
I'm glad to hear that you've taken the time to understand how GC work, and the trade-offs that languages make when use them.
I would be curious to learn more about what you're research as has shown with regards to using a GC in your problem space. It's always interesting to learn about areas of engineering that unique requirements.
My understanding is that there are some lock-free algorithms that we do not know how to write without garbage collection [Keir, 2004], so you strictly are giving things up, because you can't use those algorithms.
(But I'm not completely up to date on latest lock-free work.)
A garbage collector can do that work concurrently.
CppCon 2016: Herb Sutter “Leak-Freedom in C++... By Default.”
https://www.youtube.com/watch?v=JfmTagWcqoE
Better making use of the best practices advised by Herb, otherwise those destructors are going to surprise you.
> It has been more than thirty years since a destructor surprised me. Herb was unlikely to have been programming at that time.
So apparently you have skills that top one of the major C++ community figures.
Being so, the C++ community at large would appreciate those valuable insights.
Same applies to the "program terminates before it runs out of memory" approach.
C++ has destructors, and owning pointers (called "unique_ptr"), and also arena allocators for when those are useful. Rust has its Drop trait and analogs of the other things. (Use of arenas with Rust standard library containers is approaching maturity.) They make memory management automatic and wholly painless. Both languages offer reference-counted GC for places where that is helpful, but such places are rare, and where used typically burn an unmeasurably small fraction of runtime, with no "pauses".
> burn an unmeasurably small fraction of runtime
That’s also known as a very short pause.
Your "GC pause" simply happens on scope-exit.
At least bring some alternatives to the table or explain what you are doing and why you mean it's so much better..
For every RAII implementation I've seen, the recursive part is not within the free() implementation; the recursion happens outside the free() (or equivalent) calls, and free() (or equivalent) is called separately for each step, to release the memory for just that step. The time taken by each free() call is independent from the size of the object graph being released.
You're not using caches, then?
Not using a GC does not, in fact, make more work or more bugs, given a language that provides resource management facilities, such as C++ and Rust.
Programs using GC do typically leak, and most GCs make leaks much harder to find and fix. Most GCs even make it hard to determine if you have a leak.
But that experience does not generalize. Am I spectacularly lucky in my generally competent colleagues? I doubt it.
https://docs.google.com/document/d/e/2PACX-1vRZr-HJcYmf2Y76D...
https://msrc-blog.microsoft.com/2019/07/18/we-need-a-safer-s...
https://support.apple.com/en-us/HT212805
https://support.apple.com/en-us/HT212622
https://support.apple.com/en-us/HT212531
But maybe Apple, Google and Microsoft just aren't able to hire the right kind of highly skilled C++ devs that would write such perfect code, despite having a seat at ISO, and being clang/LLVM contributors.
Perfection is not needed. Ordinary good code suffices. Good code using modern C++, or current Rust, is easier than in older languages (among which count older C++). When bad code is extra work, it becomes an unattractive alternative even to the lazy, leaving its production mainly to the masochistic and the aggressively incompetent, who are often the same.
Here is the thing, when the ISO C++ leaders, and major C++ contributors, push for a change, it is time to start doing some self reflection.
In any case, the tradeoff you're making is that you have deal with memory. Even in a language like Rust, you need to care about references, ownership, struct vs heap and possibly even reference counting. You might even be bitten by "running out of memory" due to memory defragmentation.
In most cases, you don't need to care about these things in a GC'ed language, at the expense of larger memory usage and possibly noticable pauses.
GC has other drawbacks as well. The whole tracing workload (which still exists even in low-latency, concurrent GC's) messes up your locality of references and interacts badly with CPU- and OS-level caching and memory management. Plus the mutator part of your program is heavily constrained in how it can layout objects in memory, since the tracing GC must be enabled to select references to other objects unambiguously. It's a non-trivial drain on performance on memory-bandwidth limited workloads, which tends to be most of them these days.
Reference counting does away with most of these issues, and gives you deterministic finalization of all resources not just memory. Alternately, you can selectively use arenas to defer the freeing of some objects, while still being deterministic elsewhere.
Reference counting is a form of GC to use when performance doesn't matter. Performance never matters except where it does. But automated memory management without reference counting is usual practice in a modern non-GC language.
Freeing an array of objects only involves a series of calls to a deallocator when the array elements have pointers in them. That is typically unavoidable in Java, but not in non-GC languages.
Arena allocation, where memory in a subsystem is allocated using a specific allocator object, deallocation there is an in-line no-op, and all of subsystem memory is reclaimed en bloc at a chosen event boundary, is a common alternative where more control is needed. It is still not "manual memory management"; memory for objects is managed invisibly, and still without reference counting.
The usual advertising for GC is that it means you don't need to think about managing memory. The actual experience is that, where it matters at all, you have to think about it a great deal more.
Only if elements of the array are separately allocated (boxed) or have destructors.
Thus, it is incorrect to talk about multiple calls to a deallocator. You get just one for the whole array.
If I were to port the app to C#, I have no way of telling the C# GC to never pause the audio thread, and to only pause the GUI thread. I'm not sure if it's possible for a tracing GC to coexist with a never-paused thread which owns manually memory-managed types, and for the two threads to exchange manually freed or GC'd objects without FFI boilerplate, serialization, or copying. If it exists (and there's a GUI framework written in the GC portion of the language), let me know!
Admittedly manual memory management does mean a lot more things to worry about (ranging from manual memory management to lock-free wait-free programming) Personally I find deterministic lifetimes and single ownership which can be moved/swapped to be elegant. However, Qt's QObject system combines the practical disadvantages of manual freeing (having to track complex semantics in your head, and check whether each method call transfers ownership or not, with leaks or use-after-free if you get it wrong) with the inelegance of GC (pervasive aliasing and mutability, unclear ownership, a magical runtime-like system).
My point was that there are tradeoffs, which you admit. The comment i responded to indicated that there wasn’t one.
But the fundamental thing is that you still have to think about memory, memory layout will have to be considered during refactors, etc.
And "tradeoffs" is a misleading term to apply to a process that does not involve giving up anything in exchange for the benefit of pause-free, fully predictable operation.
And you do absolutely give up something with not using a GC: speed of development, much bigger teams can work on the project at the same time while not stepping as much on each other’s foot, etc. It’s not an accident that perhaps the majority of all software development is in managed languages, where it is not absolutely crucial to control the exact memory layout.
That is the advertising claim, not substantiated.
A more plausible explanation for "the majority ... in managed languages" is: it enables employing lower-skilled, thus more easily obtained, labor. It doesn't so much matter how fast development is. Such practices commonly split a simple job among five or more people (a "team") who, collectively, probably cost more than one more-skilled employee who could do it all in much less time, but who is hard to attract to do it at all, and anyway better used elsewhere.
And frankly, you are not getting ahead with your arbitrary gatekeeping on “who is a real programmer”.
The statement about enabling lower-skilled employers to be hired is pretty much false. Embedded development payed significantly less than most places that uses Java today (at least in Norway), and there's really no difference in the skill of the people I work with.
There's a _significant_ difference in the problem being solved though. The project I'm working at now is, code-wise, bigger by an order of magnitude. As is the problem domain (ticket-selling services, vs firmware of payment terminal in old job).
When doing embedded, I never had to care about racing conditions in threads or across several microservices. The embedded device was single-core anyways, and only talked to one server.
Dealing with memory has been replaced by handling concurrency. Of those I'd argue that the latter is significantly harder than the first. But both add to development time, as it becomes a core point in most design discussions.
In our case, we truly do not need to care about memory. If we run out of it, we either upgrade the instance or add another instance. The monthly hardware cost is lower than the hourly cost of a developer, anyways. Having one less thing to worry about decreases developer time, as there's one less thing to include in our designs, and one less thing to worry about going wrong.
Not sure if you can disable the GC entirely, but you can set the initial heap high enough that it won’t trigger for the lifetime of the process.
It is a fact that the majority of programming in modern non-GC languages, such as C++ and Rust, involves no manual memory management, no GC, and no reference counting (which is also GC). The memory-management automation provided by the core language and the standard library are equal to almost all challenges, all by themselves.
Thus, users of modern non-GC languages do not experience the pauses seen in GC languages, or the unavoidable pointer chasing, or the cache poisoning, or the "reachable leaks" that are GC languages' dirty secret. They are not, in fact, obliged to "trade off" anything at all for being free of those failings.
That is not to say that all design goals are easy to achieve, but overcoming a GC's failings is is not among the activities needed achieve them.
It is extremely rare, nowadays, to "create a destructor" to manage memory. The destructors generated automatically from templates in the Standard Library suffice.
Creating classes is just programming.
So, no, that is not manual memory management.
It's OK that these issues exist - as you say it's just programming - but don't delude yourself into thinking that's automatic memory management. It's implicit, but it's not automatic and that means it's (at least partly) manual. When you say you don't use manual memory management but then say you let standard-library templates handle it, that's a contradiction and tantamount to a lie.
/doubt
Once you're facing, for example, the deallocation of an entire tree that you own and that's going out of scope, your pause is unbounded. Compared to that, a GC'd language could delay the deallocation, or even in case of stuff like C4 have a dedicated thread deal with it.
Sure, you can always write code in a way that avoids this, but non-consing is also an option for many of the GC'd languages (at least the more advanced ones, like Common Lisp).
So, doubt all you like, unavoidable random millisecond pauses are not a problem in non-GC languages, no matter how much you wish otherwise.
Java(JikesRVM)/Oberon/Go/D/Nim: Let rewrite the whole toolchain in X, including the memory manager itself.
Go’s GC has it easier due to the language having native support for stack-based allocation which can sometimes be used in otherwise garbage-heavy places.
A claim without evidence isn't worth very much.
> Go’s GC has it easier due to the language having native support for stack-based allocation which can sometimes be used in otherwise garbage-heavy places.
Java (i.e., "modern JVMs") also has "native" stack-based allocation (indeed, it has an escape analyzer whose sole purpose is putting things on the stack); however, it lacks value types.
One can easily add a single line and “break” the escape analysis, while it is compile time checked with value types.
Of course, if one does GC tuning like a pro, uses a sufficiently smart GC and maybe a custom VM (and unicorns) one may reach the mythical native speeds (typically at the cost of using a lot more memory).
C diehards: Pauses are bad, we need manual memory management.
Everyone else: Does not have manual memory management.
That is a misunderstanding.
If you would like to avoid pauses in c, you must be careful about how you allocate, what you put where, &c; you cannot simply use malloc and free. Similarly, if you would like to avoid pauses in a gc language, you must be careful about how you allocate, what you put where, &c; you cannot simply allocate.
These techniques are actionable in any language. Java 'has' automatic memory management, but it also 'has' manual memory management (if you implement it).
In almost any language with automatic memory management, I can portably make my own freelists, regions, temporary (reused) buffers, ...
I was looking at gRPC benchmarks the other day (https://github.com/LesnyRumcajs/grpc_bench/wiki/2021-08-30-b...) and 6 of the top 7 performing gRPC implementations on 3 core CPUS were on the JVM.
Say what you will about Java (the language), the JVM is a seriously good piece of kit.
It may lack a lot of the fancier features of some newer languages but its still a very effective language.
Coming from JavaScript, TypeScript, Python and Go back to Java I can tell you that’s actually a good thing.
Cleverness in code starts working against you as a code base and organizations scale.
What do you find magical in terms of Python in particular?
> What do you find magical in terms of Python in particular?
Lots of libraries are magical. SQLAlchemy, most of the data science libraries, even the standard library returns different types based on the values of inputs (e.g., open() returns a binary file or a text file depending on the string you pass into it). Additionally, people often go crazy with various dynamic typing features (rather than refactoring into a common interface, people will do a bunch of `if isinstance(...)`es all over the place. Similarly people tend to go crazy with metaclasses and so on. I could go on and on.
We have code like (stupid example)
something.map( {a, b, ...rest} => ({ ...rest, ...a, ...[b, c, d]}))
I know what it does and how it works, but I argue that's hard to understand if you haven't written it and know the context.My Python experience is limited to only a handful of projects, but I looked into meta classes at the beginning of learning Python and that was some crazy stuff.
What are some things thing you can do in Java with less code than in Javascript?
We also do not use XML to configure DI. Just add a annotation to a class and you can inject it automatically in any constructor without any additional config/annotations.
If this is too much "magic" for someone's taste, I will not argue even if I have a different view. But please do not make absolute statements based on a bad experience in some companies.
Modern Java and frameworks have evolved a lot in recent years.
I understand the need for DI, sometimes the order of instantiation is necessary but for everything else I get lost in the DI constructor injection stuff.
In a functional language you could use a reader or environment monad to abstract the dependencies, but you don’t have that ability in a language like Java (and it’s not worth torturing the language to do it because it’s not idiomatic). So DI in Java ends up constructor-injecting service dependencies and using the method parameters to declare data dependencies.
Edit: another benefit of DI is that it allows for multiple lifetime scopes beyond application and method call. Spring has request and transaction scopes, for instance. Managing all that in Java without DI is a nightmare.
Now, I can go ahead and make the concrete implementation dependent on configuration, e.g. by providing multiple implementations annotating them with a condition that is evaluated using the application configuration. I can also switch to a factory method that creates and sets up the concrete implementation for the interface, I have not to change any place in the application that is using that piece, though.
The same should be possible in Go with reflect and an IoC container, but I am not sure if there's a solid implementation for that out there.
"null checks are ugly and error handling is messy" are your opinions, not objective facts. I actually prefer Rust-like enums, but short of that, I think Go's error and nil handling seems quite a lot nicer in practice than Java's `Optional` facilities. Note that these are my opinions, and I'm not posing them as objective facts.
The rest of them I've been enthusiastic about for a few weeks until reality catches up with me and I realize what I am missing out on.
foo{ bar=true, baz=true }
Another thing I'd like to see is opt-in null safety. The Optional type is a very good mechanism, but when every reference is potentially null, its ability to enhance language safety is limited.
Other than that, it has everything I want. And of course, the ecosystem, tooling, etc. is simply the best in class.
My personal preference is for ML-style languages, so Java will never be my true love. But as far as languages go, you can do much worse.
Even then we won't have fully reified generics, but some of the stuff they're doing with type projections is insane. I'm pretty sure it's going to be a game changer.
If we're considering other JVM languages, Scala and Clojure are also on the table.
Additionally, Spring Boot exists and is lightweight and super easy to understand, jooQ is an ORM/query builder that makes Hibernate look like a toy (and it super simple).
Just because Java gets tainted by overengineering doesn't mean the language itself isn't simple.
Even using the built-in JDBC support is an option with modern Java, PreparedStatement and ResultSet implement AutoClosable, which means you can write code like this:
try (
final var rs = conn.prepareStatement("""
<SQL here>
""").executeQuery();
) {
while (rs.next()) {
// do your thing
}
} catch (SQLException e) {
// your code here
}
Which is just the standard library and no more verbose than most other frameworks.As for web frameworks, there's a million of them. Pick a non-magical one if you don't like magic.
Spring had good documentation that you can read on an ereader. It's also very mature, so there's lots of help on Stackoverflow.
Hibernate was a bit more difficult, but Vlad Mihalcea has improved the documentation side a lot.
Hibernate also allows you to drop down to native queries, which coupled with Spring Data's "projections" are useful.
Also, I think it is overblown how magical those annotations/code generation are. It is simply class path scanning and code generation under the hood, so that a proxy object is used instead of the interface. I would take this path any number of times instead of what abomination n developer hacked up over the years with no proper abstraction and documentation.
There are good things and bad ones across both of them.
Both also provide enough rant material. :)
You know that thing, which WebAssembly is trying to clone.
I've recently written a pretty compute-heavy app in .NET 5 that generated some garbage (mostly nursery stuff), and I tried to run it on a 16-core CPU.
The results were disappointing, even with the Server GC, mainly due to the fact that the GC still does a ton of stop-the-world pauses, and the amount of garbage basically scaled linearly with the number of threads.
My benchmarking results indicated that adding more than 8 threads were basically pointless, since the increased garbage meant that more time was spent in STW GC mode.
I 'fixed' the perf issue, by eliminating all GC allocations, which alone resulted in a 2x speedup per thread, and allowed basically perfect scaling across cores.
In the day and age where you can rent 64 core CPU VMs for $2/hour, I wish they would spend more engineering effort on fixing this.
You can see this in practice when following up which tech gets out and how it evolves across its lifetime.
I bet many teams before the whole C# 7 effort, would just write a couple of COM stuff in ATL, call them, from .NET and be done with it.
This is also why WinRT was designed the way it was, given that Windows Dev was driving it.
The gRPC libraries are primarily developed by Google. Google pours most of its development resources on its core languages: Java, Javascript, Go, C++, and Python. Google then allocates a developer to port the implementation to other languages and platforms. C++, unlike Java, is not typically used for cloud development, so I don't think a lot of resources were allocated for the language.
You can see how the difference in implementation effort can impact performance by comparing dotnet_grpc (Microsoft's fork) against csharp_grpc (Google's original implementation). There's more than a sevenfold improvement in req/s in Microsoft's implementation in the 1 CPU server case (35070 vs 5337), outperforming nearly all the Java benchmarks.
Also, many of those top performing JVM implementations are the same code running under a different garbage collector. .NET has two GCs each with a parallel option, but we only see one benchmark using likely the slower GC (Workstation GC instead of ServerGC).
Scala still runs on the JVM.
I though most of Google's infrastructure (such as Borg or things like that) was in C++.
I think actually its the opposite, and the C++ implementation had seen most activity. Some other implementations (e.g. the node.js one, the initial version of .NET support, and also the rust_grpcio flavor) had been built on top of Googles C++ gRPC core library. I am not sure however if some of those now have moved to more native implementations recently, since I haven't followed the development closely.
Disclaimer: contributor to the gRPC Java benchs.
We've got some graphs [0] comparing throughput and long-tail latency for the various Java GCs available in OpenJDK 12. Shenandoah's worst case pause time was ~45µs which is the same as disabling the GC (Epsilon GC) which is pretty impressive. Overall performance did suffer under Shenandoah a little bit back then, though. However, I've heard that this improved recently.
[0] https://github.com/ixy-languages/ixy-languages/blob/master/J...
- Removal of weak references (finalizer, phantom, soft) can help reduce the GC pause time further as these warrant stop the world phase
- Shenandoah likes ReentrantLocks better than synchronised classes the latter bloats up its monitors and increases the size of the rootset
- The reflection problem may initial Shenandoah to run GC cycles even when there is no memory pressure and thus may reduce performance over time. It is better to limit it using reflection inflation flag
- On Amazon ECS, provide the CPU share count explicitly; Shenandoah likes more threads
Other GCs will do the same for you (G1, ZGC) which you can then put into a tool like gceasy to pretty graphs.
I'm in way disputing this, but why would a weak reference be different than any other reference, beyond the added flexibility it gives the GC (you may or may not GC this reference if you please)?
I admit I don't give GC much thought beyond every coupe of years when I switch jobs :)
The easiest way to handle this conflict is to just pause the world when doing the processing of weakrefs after the concurrent GC finishes.
Can you explain this further? I'm missing some connection here between ECS configuration and how the JVM sees the container.
[0]https://docs.aws.amazon.com/AmazonECS/latest/developerguide/...
BTW, you can specify the number of GC threads that you like to use -XX:ConcGCThreads=X -XX:ParallelGCThreads=Y (X for concurrent GC, Y for STW pauses - which should no longer be very relevant).
Cheers!
This is the power of JVM we do not get from other languages/VMs. Kudos to JVM (previous and current) architects, developers, and everyone involved with it.
sub-millisecond gc pause has been available to Golang users for ages.
I think I know what you mean. If a couple of guys quit, the company is in major trouble? :-)
Forgot this: /s
Given the similar feature sets, does anyone with experience running both ZGC and shenandoah have any differentiators?
https://www.oracle.com/java/technologies/javase/17all-relnot...
It's really just 'usage logging', which is for Oracle's management tools. Their stated aim is to eliminate the functional differences, so presumably they are awaiting a non-proprietary replacement for that feature in JDK.
Are there any differences in the targeted workloads or capabilities between ZGC and Shenandoah GC?
Both seem to be marketed for large heaps and low pause times.
It is my understanding that D's garbage collector isn't that good, could they pick up an OpenJDK collector?
Also, I think D allows for pointers which is a no-go with compacting GCs (as those move objects around). So all in all, the algorithm used itself can be ported, but it has many runtime specific logic, which is just as important for great performance.
Not sure what that means in practice, but if someone is putting money on the line, it must be worth _something_.