OTOH, if you think other languages let you do away with a GC without pretty significant extra work, especially in concurrent systems, well, then you haven't had experience with those languages.
OTOH, if you think other languages let you do away with a GC without pretty significant extra work, especially in concurrent systems, well, then you haven't had experience with those languages.
I'd say the opposite is true. Relying on GC requires significant extra work. Because you always need to think about memory (exception: small script-like applications). The only thing a GC does is that it enable you to not think about it, but the moment you don't you will write bad code and realize it was a disservice all along. And by then it is too late.
So in a GC language you need to constantly be aware of when you take something for granted. Which is more work than just doing it manually yourself.
generally you design your application the way you normally would, do some testing, and then go back through the hot path and make some adjustments.
Second, while there have been similar claims made in the past, they always apply for a certain target throughput and latency. GCs are constantly making great strides in that regard. Next year, ZGC will have a <1ms worst-case latency for pretty serious allocation rates, and with an acceptable hit to throughput. As you can see in this post and previous ones, G1 offers great throughput with acceptable latencies.
Either it moves it into PagerDuty, like G1, or it moves it into your GCP bill like Shenandoa and ZGC.
The trade-off has always been latency-throughput-footprint, nothing I've seen yet has changed that. The innovation in Rust is realizing you can do all the tracing work at compile time.
> The innovation in Rust is realizing you can do all the tracing work at compile time.
This is spectacularly false. For all non-trivial allocation/deallocation patterns, Rust also uses a runtime reference-counting GC, which is significantly slower than the tracing GCs you find in OpenJDK. The benefit comes from not relying on it too much, but this comes at a considerable cost of lost abstraction, which means more costly maintenance over the years. This is the same for all low-level languages; the difference Rust brings is that it (conservatively!) checks for memory access errors.
Another difference is that most people who talk about Rust haven't actually written a significant application in it and had to maintain it for years. I'm not saying it's impossible -- people do this for C and C++, which make a similar tradeoff in this regard -- but it does come at a substantial cost.
I used to be a hard-core zealot for the JVM as the performance platform of the future - and did talks arguing just as you are here why HotSpot outperforms $LANGUAGE in real applications. I feel like I'm hearing myself in your argument..
I wrote a signficant portion of the Neo4j storage engine, which is in Java. Now I'm writing another database engine in Rust (sidenote: not for replacing the Neo4j engine, just because it's interesting). Arguably database engines qualify as "significant applications".
I find - subjectively:
- Maintaining performant code in Rust is easier. I do the same patterns as I did in Java, except it doesn't rely on easy-to-break assumptions of how HotSpot happens to work (ex: stack allocations)
- Like you said elsewhere, the issue is often about fragmentation, ultimately stemming from object churn. I find that Rust makes it, culturally perhaps, easier to maintain low allocation code than Java. (sidenote: I think these two points is also why Go code generally has "better GC behavior"; Go doesn't have a better GC, it has a language that encourages less heap allocation)
- The engine I'm writing in Rust is faster than anything I've written in Java and - critically - runs for days and months without notable stalls.
Come to the dark side!
The JVM works exceptionally well when you are dealing with long lived apps which deal with a lot of allocations and the machine running it has a substantial amount of memory. Things like webservers, for example, are near perfect fits for the JVM.
Rust does really well when you need high performance, low memory, and ultra fast startup times. It won't necessarily outperform the JVM when you talk about doing a lot of heap allocations (due to heap fragmentation) and it unfortunately suffers from the same heap fragmentation if you are dealing with a long lived server that does a lot of allocations. But then, maybe that performance loss is acceptable for the ability to very quickly scale up and down servers.
Now, the JVM is making great strides towards getting faster startup times and even fast performance at startup (AppCDS). However, those strides often involve trade offs with either build complexity or performance losses (such as Graal's AOT). The benefit for rust is that it is as fast as it ever will be without any special build steps or tweaks.
Oh, and let's not forget diagnostics. Flight recorder for the JVM is simply AMAZING. The ability to hook up to a poorly behaving production server, start flight recording, and getting detailed information about things like "where are allocations happening" or "what are the hot methods" is simply amazing. No other platform that I know of has the level of detail you can get right out of the box with flight recorder. Certainly not without restarting the application with additional configuration. For example, you'd need a special build of rust with profiling turned on to even start to get the same level of info, doing such also significantly negatively impacts performance.
What is happening now is that the free beer Java users are also getting those features on the package.
Are you talking about `Rc` and similar smart pointer types? If so, the twist is that in Rust almost all allocations in Rust are trivial in this sense.
Lest anyone read this and think it's true, it's not. Using Rc is a design choice, and not one that is a given. I have written tens of thousands of lines of Rust code doing very heavy data processing and used Rc only a handful of times. In fact, I find using Rc without a very good reason to usually be a bad idea that enables lazy thinking.
This is not really a good comparison because you can't write the same code in either language. In Java practically everything gets allocated on the heap barring some optimizations. Meanwhile Rust programs can selectively allocate memory on the stack when it makes sense to do so. Reference counting is is just one of many different allocation strategies available to Rust. It is not the first tool you grab when you want to allocate memory in Rust, therefore absolute throughput of reference counting might not be as relevant in Rust as the absolute performance of the GC in Java.
In our benchmarks we never saw a GC pause of more than 2 ms on either ZGC or Shenandoah, but the end-to-end latency, the one the user cares about, is impacted by much more than a single GC pause. Sometimes there would be several pauses in a rapid sequence, or just the background GC thread would do too much work at once.
Even after dedicating a core or two to the GC, you still face the issues of cache pollution and RAM throughput stealing that heap walking incurs.
You say that, but there seem to be no end to the stories of people spending enormous amount of time fighting the gc. It is not difficult to avoid heap allocations in other languages as well as freeing them deterministically.
Second, as this blog post series shows, the "fight" is not what it used to be. You don't really need to control allocation any more until your rates are really high. Java's GCs have just gotten so much better in JDK 14 and beyond.
I'm extremely skeptical about that. My experience is that with modern C++ you lose very little elegance and gain a huge amount of control by giving up a garbage collector. Memory management becomes a very minor problem. The vast majority of memory allocations are avoided and those that need to be there can be done ahead of time.
I have never heard anyone writing a latency sensitive program in C++ (games, trading etc.) say that their life would be easier if they were using a gc or that they wished they could do it in java. I have however seen decades of people talking about all the extreme lengths and rabbit holes they go down to deal with the java gc.
From a broader perspective, pretty much any language with a gc ends up having a perpetual conversation around how the next gc will solve the problems with the current gc. You can see it in java, go, julia, and D. The only one I never hear about is LuaJIT, but maybe I just haven't seen it or maybe the expectations are lower.
Do you have any references handy on the julia bit there? I actually have seen very little conversation in the julia community about replacing or upgrading the GC. Mostly just the ocassional post from an inexperienced user who thinks that a borrow checker would be a good fit for julia.
Discussion seems to almost always revolve around showing users who need it, how to manually manage memory when necessary by pre-allocating arrays, using in-place operations or writing stack allocated code so that they avoid the GC in performance critical code.
I've never seriously used a language without a GC, but my feeling in Julia has always been that I never really had gripes about the GC because it's so easy to avoid the GC and take memory management into my own hands.
One complaint I will say I've heard though is that while these people find they can get the allocation behavior they want (i.e. none), some of them would like semantic guarentees that they will not hit allocations, rather than needing to test and make sure their code doesn't start hitting the GC if they switch Julia versions. Someday we might be able to provide such guarantees, but for now GC behavior is just an implementation detail that can change across minor versions.
I can sympathize with those who find that uncomfortable for sure, but in practice, Julia versions have been consistently better at getting more automatic stack allocations, not less, so it hasn't really been a problem.
The cost of maintaining a large C++ application (>1MLOC) with a large team over years is very significantly higher than a similar Java application. In some cases the footprint and/or performance benefits are worth that extra cost, but in the vast majority of cases they're not.
> I have however seen decades of people talking about all the extreme lengths and rabbit holes they go down to deal with the java gc.
Again, 1. Java's GCs changed dramatically in the last two years -- the GCs described in the post, are brand new/recently revised and 2. that's because that's Java's particular rabbit hole. C++'s rabbit holes, from undefined behaviour, through partial evaluation with templates and constexprs, to compilation times and sheer language complexity are far deeper.
> I have never heard anyone writing a latency sensitive program in C++ (games, trading etc.) say that their life would be easier if they were using a gc or that they wished they could do it in java.
Their lives would be easier if they could do it in Java, but sometimes they can't. I think that the changes in the last couple of years and the upcoming changes in the next few years will make Java more appropriate even in domains where it hasn't been used before, but it's fine if not. Its market reach is so huge as it is. But games are not often maintained for many years, and telemetry isn't that important, so Java's benefits are not as big as for servers.
I am very skeptical of this, I don't know why it would be the case. My experience is that with modern C++ and avoiding inheritance programs end up much more direct and clear since a type isn't fragmented into multiple classes and base objects don't need to be used for generic programming and data structures.
> C++'s rabbit holes, from undefined behaviour, through partial evaluation with templates and constexprs, to compilation times and sheer language complexity are far deeper.
This seems like what would be said by someone who has just read a few comments on C++ here and there but not actually used it for non-trivial projects. These are rarely issues. I don't know what 'partial evaluation with templates' means and constexprs didn't even exist until recently. Compilation times do seem to be a big problem, mostly because many projects don't do anything with their structure to mitigate them.
Rust or other upcoming or future languages might change that.
† The threshold varies between VM and GC, but usually <10ms is easily achievable
Steve lays it out far better than I could, coining the term "static garbage collection": https://steveklabnik.com/writing/borrow-checking-escape-anal....
And to your "without pretty significant extra work" qualifier: I really don't find that to be true with Rust. The initial learning curve was a bit rough, but certainly far less so than other languages/platforms I've picked up (hello, ML). In the end, I find that it's just a nice, helpful, productive, and ludicrously performant language.
Since the automatic deallocation in rust, per default, happens at the point where the value goes out of scope, if a resource destruction could block, or perform expensive operations on destruction, it can not be allowed to go out of scope in a latency sensitive thread.
Usually not that much of an issue unless you are chasing really low latencies.
But things can still happen in rust that catch you off guard. Say a value goes out of scope, and because of reasons, it does so with a destructor doing a logging call, which has accidentally become blocking on a socket send call, which it did because nobody realized that the AWS/GCP/whatever logger adaptor actually didn't perform all IO in a thread with which was only communicated with locklessly, which nobody noticed before because it was only if a buffer was full, which only happened today because ....
Not a big deal, it's almost all the same things which mess upp latencies in C++ code. And that's the thing. It's not necessarily easier to get low latency in rust than in C++, but the work required for hitting a quality/performance/latency target in rust is probably still lower than for C++. Unless you are lucky to have a very mature C++ low-latency stack, together with all the utility functionality you need, which seems to be exceedingly rare. Is the work required lower than Java on a custom/tuned JVM?
We'll have to wait and see. It probably is, but it's a complex balance between access to utilities, language complexity, and several more parameters which ultimately decide which platform provides the best environment for low latency code, especially if the complexity is non trivial.
The thing is that most of what you said about Rust is true of C++ as well. It really isn't hard to write a good fast program in C++ once you know it. The problem -- as those of us, like me, who have been writing in low-level languages for a couple of decades now know -- is that low-level programs in any low level language are necessarily rigid. Changing something in one place often has a much bigger impact on the codebase than in a high level language, making the overall cost much higher.
Low level languages have their place, but they won't replace high-level ones for "ordinary" application development. People will only pay the price when that extra 5% is important or when running in a constrained environment. This is the same equation that's been around for twenty years and there's no sign it is changing.
This is still a lot of work for a video game (because you never want any latency, the only way to achieve this is with an arena allocator or going full @nogc).
But for apps where the latency requirement is bounded, D doesn't make it hard and the language is nice and ergonomic.
Isn't that how most GCs work?
Why would you do anything else? To release memory to the OS? That's not really a priority in most runtime systems, and I think wanting to do that is a pretty niche requirement.
You can definitely write useful Java applications where nothing gets heap allocated, through using existing objects, and through using scalar-replacement-of-aggregates.
I guess it depends whether GCs are always scheduled in an allocation or can be triggered another way. Either way that should be easy to disable.
I read somewhere that D doesn't have write barriers, so I would assume they have a hard time implementing more advanced GC features like generational collection or concurrent marking. It's not suprising that the GCs in the JVM achieve much better pause time.
The memory work remains the same, you can do it yourself or let the GC handle it. For 99% of applications, the GCs are good enough and getting better every year, but low-latency still needs predictability and would ideally choose manual memory management.
The realities of the job market and IT deployments are different though and that's why we still have JVMs involved with low-latency scenarios because of talent, tooling and productivity.
10µs ? You can't even afford that many function calls.
Seriously, 10µs is not a sensible target for a general purpose OS, maybe not even for a general purpose CPU. Achievable? Perhaps. But sensible? Not really.
I suppose that if you're working on small amounts of data every time your code executes, then this becomes vastly more reasonable.
But beyond that, the actual code you write is fairly natural for those languages. You can't allocate, and your code and data have to fit in the cache. But you can use normal language constructs, and most of the standard library - neither of which is true for Java.