Using Java's Project Loom to build more reliable distributed systems
jbaker.io
jbaker.io
I'm really looking forward to being able to use Loom to run much more efficient cluster simulations.
[1] https://cwiki.apache.org/confluence/display/CASSANDRA/CEP-10...
Have you found that the approach has made it easier for you to ship improvements to Cassandra?
This kind of approach is in my opinion a _requirement_ for work on distributed consensus protocols, which the Cassandra community is improving at the moment: the upcoming 4.1 release has some major improvements to its Paxos implementation that were validated by this simulation, and the following release is expected to contain a novel distributed consensus protocol we have named Accord, for supporting multi-key transactions.
So, I'm not sure it has affected the speed or ease of shipping improvements, rather than enabled a class of work that would have previously been impossible to do safely (IMO).
Stardog, enterprise knowledge graph platform, (stardog.com), is an early adopter of Antithesis and we've found it to be very helpful in HA Cluster, distributed consensus, etc.
Not a stakeholder, just a satisfied early adopter.
Can you describe what their tools are doing for you and how you're integrating with them?
Of course all projects will benefit from Loom. I am merely positing that Clojure in particular could leverage this both most quickly and to very deep positive effect, due to the nature of the language itself putting immutability first, which if I recall correctly they built up a pattern of concurrent execution around already.
Kotlin too will be be able to optimize their concurrency story.
I simply wanted to posit an observation I hadn't seen elsewhere on the web about Project Loom yet.
Edit to add: I hadn’t initially thought to speculate on the language itself taking advantage, but it occurred to me immediately after posting this comment. Another potential benefit is that many of its fundamental abstractions (most collection APIs are declarative) lend really well to making those operations concurrent, which is a common FP/lisp hypothetical but becomes more tangible if coroutines are less expensive. This isn’t dissimilar to how RDBMS query planners can provide wild performance improvements without any semantic changes, because SQL queries express “what” now “how”.
Maybe through some kind of non-preemptive user-space scheduling you mean?
Why? If you know your thread won't be interleaved with any other until you a well-defined point, how's that different to using a mutex or semaphore?
Close Encounters of the Java Memory Model Kind - https://news.ycombinator.com/item?id=11955392 - June 2016 (65 comments)
I really enjoy his anatomy quarks series:
Think this will obsolete go over the next few decades.
The executor services referred to in the blog are for the order of execution of tasks on the virtual thread pool. For a "virtualThreadExecutor" service, every task will get a virtual thread and scheduling will happen internally.
You can still use a fixed thread pool with a custom task scheduler if you like, but probably not exactly what you are after.
Maybe a little disappointing for low level nuts and other languages like kotlin, but the right move IMO. Virtual threads alone will be a huge benefit to the ecosystem. The other stuff will help, but won't have near the same impact.
I'd be surprised if Go adoption plummeted because of this, but who knows, I sure don't have a crystal ball.
I think it's all a moot point though, as it basically just demonstrates the next iteration of Paul Graham's Blub Paradox. With every iteration of new improvements for the JVM it reinforces the belief of many that the JVM is the best tool for every job (after all, it now just got cool feature y they just now learned about and can use and OMG Blub-er-Java is so cool, who needs anything else?!), and reinforces the belief of many others that the JVM is playing catchup with other languages (it only just -now- got feature y) and there are often better tools out there.
The path of least resistance in Scala leads to immutable collections.
In Go/Kotlin, one can atomically take from/send to exactly one channel with the `select` call.
Then I agree with "weirder" as well, in the case of Go channels. A send on a nil channel blocks forever. Why?
Channels exhibit the following properties:
* send to a nil channel blocks forever
* receive from a nil channel blocks forever
* send to a closed channel panics
* receive from a closed channel returns the zero value immediately
I have this bookmarked because I don't write Go enough to remember it by heart.If anything, if Loom is great, then it will keep Go on its toes and hopefully Go will also evolved due to external pressure.
You can kinda do this with futures but I suspect it'll be wildly inefficient. I really hope Java get's something to fill this niche. We already have a menagerie of Queue, BlockingQueue, and TransferQueue implementations. What's a few more?
Kotlin coroutines have the bare minimum in the language, and implement the rest (e.g. channel, select, `go`/`launch`) in libraries. Could you explain what the dividends for Go are?
https://kotlin.github.io/kotlinx.coroutines/kotlinx-coroutin...
https://openjdk.java.net/jeps/8277129
But frankly I'm afraid of how these changes affect garbage collection since more and more vthread stacks are going to be in the heap (I hope they are contemplating some form of deterministic stack destruction along with the above JEP).
People don't pick up Go over Java because of goroutines, Java is still and will forever be an "enterprise" language behind many layers of abstractions.
Writing a “hello world”-scoped microservice is a tiny niche.
Also have you proof that Go will die under large memory usage? It's FUD.
The Debian binary-tree test is designed to create a ton of allocations and stress the GC. Go comes in at 12.23 seconds with Java at 2.65 [0]
Discords famous article about moving a service off Go because of GC issues [1]
I think there is definitely enough evidence to suggest that Go’s GC does have performance issues and doesn’t give you the knobs to tune it.
[0] https://benchmarksgame-team.pages.debian.net/benchmarksgame/...
[1] https://discord.com/blog/why-discord-is-switching-from-go-to...
https://benchmarksgame-team.pages.debian.net/benchmarksgame/...
As for Discord it was a specific use case, it does not means Go has specific GC issues overall. How Java would have compared against Rust? 100% sure it would have performed worse, but you can't say for sure that Java would have performed better, especially after a re-write, rewriting the Go Discord Go in Go could have fixed the issue, no one knows.
The JVM is probably heavily optimized for what a binary tree is doing, does not mean the JVM overall is better for all use cases.
And by “what a binary tree is doing” you mean like.. garbage collecting no longer used objects? Like, why is it hard to believe that the runtime on which perhaps the majority of serious, huge web services run (twitter, apple’s web services, but google as well are huge java shops), the likes of which handle 325,000 transactions per second (Alibaba) underwent a tremendous amount of engineering and in the GC category is definitely the queen?
Where you can get away with barely any allocations is a much smaller niche even for microservices. And Java’s GC is in an entirely other generation of GCs compared to Go’s.
- Garbage collectors have required far less tuning; with G1, Shenandoah, ZGC it's likely that your application will need little tuning on normal sized heaps.
- modules, jlink etc allow one to build a much smaller Java application by including only those parts of the JVM that one needs.
- Graal native images are real. These boast a far lower startup overhead and much lower steady state memory usage for simpler applications.
Probably my counterexample of choice is this: https://github.com/dainiusjocas/lucene-grep - it uses Lucene, one of the best search libraries (core of Elasticsearch, Solr, most websites), which is notoriously not simple code, to implement grep-like functionality. In simple cases, they demonstrate a 30ms whole process runtime with no more than 32MB of RAM used (which looks suspiciously like a default).
The JVM is fast becoming a bit like Postgres... one of those 'second best at everything' pieces of tech.
/usr/bin/time -l ./hello-world
Hello World!
0.00 real 0.00 user 0.00 sys
3231744 maximum resident set size
0 average shared memory size
0 average unshared data size
0 average unshared stack size
841 page reclaims
1 page faults
0 swaps
0 block input operations
0 block output operations
0 messages sent
0 messages received
0 signals received
2 voluntary context switches
4 involuntary context switches
22395110 instructions retired
18507246 cycles elapsed
1294336 peak memory footprint
So "peak memory footprint" for hello world is 1.2 MB and it starts instantly.Now, not everyone can/will use AOT compilation. It's slow to compile and peak performance is lower unless you set up PGO, plus it may need a bit of work in cases where apps assume the ability to do things like generate code on the fly. But Go can't do runtime code generation easily at all, and if you are OK with those constraints, you get C-like results.
At which point would you say Java has improved enough to catch up with Kotlin (supposing Kotlin does not also keep improving)? As a long-term user of Kotlin, I would say I would not reach out for Kotlin anymore for new projects. The last remaining big thing Kotlin gives is non-nullability, but with simple tools, Java also has that already.
My point being, Kotlin vs Java isn't just about language features, it's about community, ecosystem, use cases etc.
(Fwiw, personally I prefer Kotlin because it's more expression oriented than Java.)
Imitation can only get you so far. Java is changing, and in many cases for the better, by absorbing features from other languages. However, I still think several other languages do a better job curating features to fit a niche.
That said, Loom appears to be a serious upgrade for JVM languages. Now, if startup could get an order of magnitude faster...
I like the OP's idea of using a virtual thread implementation to parallelize the application-layer, while having the test code implement a custom executor to control the interleaving of the units-of-work being scheduled. I can see a few ways this approach could be used more widely; for example in Django you could write a DB driver wrapper and have this as the "test executor". Or for remote API requests, run the test code in green threads and then have a test executor that intercepts and then then chooses how to interleave the requests/responses.
Java upgraded to native threads, and then you could have N java threads bound to N kernel threads. This was way better, but had downsides: You're limited on how many threads you can spawn, you need a threadpool to help manage, and any long-running tasks could effectively deplete your thread-pool.
With Loom, now you have M green threads mapped to N kernel threads. These green threads are way cheaper to spawn, so you could have thousands (millions even?) of green threads. Blocking calls won't tie up a kernel thread. So if you have many long-running IO tasks, they aren't going to waste a kernel thread and have it sit around idle waiting on IO. This is similar to async libraries, but without the mental overhead. You should be able to just code synchronously and the JVM will take care of the rest.
This is pretty much how JavaScript operates. You can consider calling an async function as spawning a user-level "thread"; chained-up callbacks are the same thing, but with manual CPS transform.
This has always perplexed me. Why N:1 threading is hated, but Node.JS is so loved?
javascript has exactly 1 thread (with the exception of web workers, but they are a pain to use)
the whole idea is to not have to relies on "async" stuff.
> This is pretty much how JavaScript operates
> javascript has exactly 1 thread
No disagreement here. I understand what you are saying. Would be great if you had tried to understand what I said.
Maybe my explanation is lacking. Let me quote pron.
> Again, threads — at least in this context — [...] refer only to the abstraction allowing programmers to write sequences of code that can run and pause.
https://cr.openjdk.java.net/~rpressler/loom/Loom-Proposal.ht...
An kernel thread running code line-by-line is a "thread"; callbacks that are sequenced (and have the dreaded pyramid indentation hell) forms a "thread"; when an async function is called there is also a "thread".
In JavaScript the latter two kinds of "threads" are run by one single kernel thread, i.e. N:1 threading.
If you agree that Project Loom virtual threads are threads, consider this. When pluggable executor is available, you start a few of virtual threads, confining them in the UI thread. Semantically, how are they different from calling async functions in JavaScript?
---
>> abstraction allowing programmers to write sequences of code that can run and pause.
> No. Callbacks are not a form of threads. Async is not a form of threads.
Simply a "No" adds nothing to the conversation. Do you not see that they are "sequences of code that can run and pause"? Or do you not agree with the wider meaning of the word "thread"?
You can see the implementation of VirtualThread here: https://github.com/openjdk/loom/blob/fibers/src/java.base/sh...
This uses the internal 'one-shot delimited continuation': https://github.com/openjdk/loom/blob/fibers/src/java.base/sh...
So, at least in principle, there is scope for other styles of concurrency to be implemented over this.
Explainer here: https://innovation.microsoft.com/en-us/exploring-project-coy...