On Java/JVM: Loom and Thread Fairness
morling.dev
morling.dev
Speaking as someone who has worked intricately on an async runtime in another language, I've more and more started to question the premises around the explicitly async paradigm. Lots of things we were "promised" with async turned out to be just as complicated in the threaded (i.e. blocking) programming models, and now we have two incompatible programming models. I'm now of the opinion that we should improve (instead of replace) our threaded models and do under-the-hood optimizations to sort out the performance issues.
If I had a dime for everytime someone made a blocking call from an async function.. And who can blame them? Most blocking functions aren't even documented as such, you just have to know.
I'd like to see how those platforms will be able to get past that legacy.
This one has been open for 3 years now (and it is probably not the first one), and the completion is nowhere in sight.
It also only targets runtime, but does not yet solve the library compatibility issue.
In order to do that, you'd have to have an abstraction over the syscalls. I'm not entirely sure this is feasible with all the ins and outs of Rust, but seeing the success of Loom, I'm at least optimistic.
Isn't Go proof that this works just fine, arguably even better, with the blocking model?
Regarding cancelation, I think Go got it near perfect with contexts, modulo the explicitness of it. What don't you like about it?
> If I had a dime for everytime someone made a blocking call from an async function.. And who can blame them? Most blocking functions aren't even documented as such, you just have to know.
Agreed. I'd like to see blocking represented in the type system the same way we do for async, so that you can't mix and match the two unless you deliberately choose to.
Function coloring as a feature?
[1]: Runtime Support for Multicore Haskell, Simon Marlow, Simon Peyton Jones, Satnam Singh (https://www.cs.tufts.edu/~nr/cs257/archive/simon-peyton-jone...)
Doing this on the JVM is not possible without actually modifying the running bytecode, so that's why they have that caveat for Loom.
Since the JVM comes with all the blocking call in its own standard libraries, it can rewrite them all internally. Every third-party library that uses the normal Java standard library will benefit from this under the hood replacement of blocking systems call with non-blocking system calls. Only when you use other native libraries you have to be careful to use mostly non-blocking functions.
Therefore Loom allows programmers to think and program with blocking calls, but they get the performance characteristics of non-blocking calls. That is awesome. And there is no need to use function colouring or type-level encoding of blocking/non-blocking code.
Microsoft tried to make everyone program the right way in WinRT and it was one of the reasons why many developers hate it (among the other reasons vs Win32), because you get async everything.
Isn't that exactly what loom aims to resolve?
The Loom scheduler does NOT guarantee resource starvation avoidance for high CPU usage. You need to Thread.yield.
I also thought of creating a mapping from synchronous code and rewriting it to a tree of LMAX disruptors. I ported a wait-free ringbuffer from C++ to Java today - it was written by Alexander Krizhanovsky[1]
In LMAX disruptor you split each IO request in half - you have a disruptor that enqueuing events events to request the IO and to pass the event on to a callback you enqueue the response on a different disruptor.
If a synchronous request handler has 27 lines it represents a tree of 27 disruptor threads independently scaled.
We can pipeline the request with 27 threads all in an event loop that pipeline each task × core count and without blocking the servicing thread or other work. So event loops without blocking! So CPU heavy tasks do not block other heavy CPU tasks and do not block the servicing thread and are not blocked by IO.
https://github.com/samsquire/ideas4#51-rewrite-synchronous-c...
[1]: https://www.linuxjournal.com/content/lock-free-multi-produce...
Currently you have to Thread.sleep(1) to force Loom to unmount the OS thread in a CPU-bound green thread.
https://cr.openjdk.java.net/~rpressler/loom/loom/sol1_part2....
Some GCs need to reach a safepoint before doing their work; and when a classic Thread is chugging along in a tight loop, that thread is blocking the whole system for a pending GC.
One hacky fix is to provide an opportunity for a safepoint (System.out.print every 10k iterations). Other tools like JVM command line options allow for observing safepoint triggering.
The same goes for virtual threads: inserting a Thread.sleep(1L) every nth iteration of a tight loop does it.
Also there was talk specifying a custom scheduler for the virtual threads. That would not help in inserting switching opportunities, but it would give a way to define what kind of fairness you want.
It's neither. The frequency of non-blocking CPU bound sections of most applications is low and so the need to introduce scheduling hints is also infrequent. Aside from an endless, non-blocking loop that never yields omitting possibly beneficial yield points isn't an error; scheduling will not be ideal but the process will ultimately produce the same result.
enum Hint {
ALWAYS,
FREQUENT,
SOMETIMES,
INFREQUENT,
NEVER
}
Thread::yield(Hint);
The method could be intrinsic to the compiler+runtime to provide runtime optimization (loop unrolling, tuning, etc.) of the 2nd - 4th cases. This would neatly avoid the need to: long n = 0;
while (true) {
// stuff
if (++n > SOME_LIMIT) {
n = 0;
// yield here
}
}
... or similar.Virtual threads are normally run on a fork join pool or other executor service, and can run on any OS thread in that pool. They are not tied to any single OS thread.
Yes, it is easy to understand in a debugger as support has been added to the debugger interface and to JVMTI, along with some Useful JFR events. We hope that in the future structured concurrency will make it even easier to understand huge numbers of virtual threads in production.
> We can forcefully preempt virtual threads at any safepoint poll point
What is a "safepoint poll point"? A sleep call?
A safepoint poll points are points in a threads execution where it's able to go into a safepoint and transfer control to the JVM so it polls to see if it should.
https://openjdk.java.net/groups/hotspot/docs/HotSpotGlossary...
When bytecode is being interpreted, every bytecode instruction has a safepoint after it. In compiled code polls are scattered thoughout the code, for example on method entries and loop back edges (so the comment above about needing to add code in tight loops to help the GC is wrong).
This is something where the answer is "it depends". Hotspot will omit the safe point poll in counted loops since it knows those loops will terminate.