Java Virtual Threads Preview
openjdk.java.net
openjdk.java.net
- continuation-passing style (CPS), hand-coded or compiler
- preemptive threading / processes
In between lie various solutions, like async/await (closer to CPS), and green threads (closer to preemptive threading).
The key difference between the two ends of this spectrum is memory footprint. With CPS you manually compress state into tuned structures. With threads you store state all over the stack in very inefficient ways because all those stack frames take extra space.
Any solution towards the thread side of the spectrum will yield significantly larger memory footprints than solutions towards the CPS side of the spectrum.
On the other hand, solutions towards the CPS side of the spectrum can require writing application-specific schedulers if fairness issues arise.
On the whole, solutions on the CPS side of the spectrum are better, IMO.
Java is kinda stuck with threads, so green threads make some sense. Of course, you can get pathological issues in M:N threading, so be careful about that.
The CPS approach is certainly more space efficient, but but I'm not sure how much of a difference it really makes in the end. Go seems to be doing well too, after some attempts with linked stack segments and then moving to copying stacks while growing them.
What is interesting is that that some of the CPS like implementations have other performance drawbacks. E.g. since it makes the virtual stack more distributed over memory, the cache efficiency of such an approach might be lower. In Rust one limitation of the CPS approach is that the coroutine state is first stack allocated before being moved onto the heap, and this operation has shown itself to be costly for some applications. So right now I'm not sure if there is any implementation which is superior in all possible benchmarks. But the Java one definitely seems to make a lot of sense for what they want to offer!
Now, if you need to serve 1e6 clients with threads, and those are 1MB stack threads, then you'll be using 1TB of your VM space, which... is almost certainly going to have some performance issues (MMU table size issues at the very least). If you splay your stacks on the heap as linked lists of stack chunks then you might get away with having a very large (and fragmented) heap with large page table entries, which might be a win.
I think approaches on the CPS side of the spectrum will be generally better than this. No, I don't study this and I don't have numbers. Yes, CPS in general means allocating closures on the heap so that some state does live splayed all over rather than compressed, but it doesn't have to be so. But often you'll have only a handful of such closures, and the language could understand that they are one-time use closures (hello Rust) so that no GC is needed.
I've written a small (proprietary) HTTP server that is hand-coded CPS -- specifically it supports hanging-GETs with Range: bytes=0- of regular files as a form of tail -f over HTTP, which is great for log files. That implementation has a single object per GET that has all the state needed, and the only other place state lives is in epoll event registrations (which essentially are the closures, and they are very small, and only one per-connection). Granted, this is a very simple application, and it would be a lot more complicated if, for example, it had to do async I/O directly on a block device to implement a filesystem in the same process -- that would require more care to keep the state compressed.
So in general I'm for CPS. But it's generally true that CPS solutions cost more dev time, and that can be prohibitive. The memory footprint cost difference will be a linear factor, which does not trivially justify the additional dev cost. Then again, if you'll be running lots and lots of instances with lots and lots of clients, the run-time savings can then easily be gargantuan compared to the dev costs -- but no one measures this, and by the time you wish you'd used CPS it will be too late and reimplementation costs prohibitive. Then again, async/await might fit the bill well enough most of the time.
async/await == the syntax and compiler help you manage the callback hell
The community-at-large decided that hand-tuning garbage collection was too finicky and not worth it, even though it obviously 'costs' memory.
I'm frankly at a loss as to why so, so, so many blogposts and tech experts are all-in on the CPS-side of this argument; it seems quite obvious to me that in the vast majority of cases, the considerably simpler* model of (green) threads means you're making the exact same trade-off: Simpler to write and debug code at the cost of needing more memory when running the app you write.
*) For sequential/imperative-style languages, that is. If you're writing in a language that is definitely and clearly intended to be written in a functional style, I can see how the gap between CPS-style and threading-style is far narrower. However, java, python, javascript - these are languages where the significant majority of lines of code written are sequential and imperative in nature.
Also note that in e.g. java you can actually configure stack sizes as you make threads. Thus, your choice of words of "'significantly' more memory footprint" is debatable.
A GC does not obviously cost memory. It might, or it might not. Both a GC and a traditional memory allocator have hidden costs. A GC with movable objects can sometimes do better, because it can manage fragmentation.
I prefer a message-passing style; on which side of the spectrum would this fall?
My experience with green threads is libraries that make I/O operations look like a regular function call. This is similar to RPC where a remote call and a local can look the same, even though the remote call is much slower. This can result in surprising performance characteristics. Even worse, a remote call can time out or take indefinitely long to complete; the same is not true of local calls. Message passing is more onerous but makes surprises more obvious.
I've often fancied writing for an architecture where main memory is treated as fast remote storage, accessible with message passing. I know such architectures exist but I've never had the opportunity to write for one. I wonder if the change in style would have a positive or negative effect on performance.
I don't find it's quite that simple.
My experience is that the complexity of CPS tends to scale linearly with use, whereas threads scale exponentially. For small uses threads are easier, but CPS quickly catches up.
CPS forces you to actually declare a dependency tree for your data. Things depend on other things, and that exists in your code. It's very easy for threads to end up a mess, where it's not clear how data is passing through the code, which causes bugs like deadlocks and race conditions.
It's deceptively easy to write code where thread A tries to lock mutexes X and Y, and thread C tries to lock mutexes Y and X, and it deadlocks because neither thread can get both locks.
It would be much harder and more arcane to do that in Javascript or in Python's async. I'm not saying it's impossible, but I don't think I've ever accidentally created a race condition or deadlock in their CPS engines.
TL;DR if your functions are only marked async so you can await something, threading probably is simpler. If you're actually passing promises around, things become much more favorable to CPS.
This isn't really what's happening here. Firstly, you can implement CPS on the JVM no problem. Kotlin Coroutines do exactly that. Loom's design is a very, very explicit design choice. Ron Pressler - the lead and designer of Loom - has talked about this extensively in many videos. He has argued persuasively for the way Loom works as not only a good way but the best possible way, one which is not a requirement of Java's previous design choices but rather, is only actually possible due to Java's prior design choices.
A recent talk on this topic is here:
https://www.youtube.com/watch?v=KmMU5Y_r0Uk
It's highly recommended. I'll try to summarize the basic argument.
The ideal, from a developer's perspective, is to have the programming model of threads with the efficiency of hand-coded CPS or state machines. Why: because threads naturally provide useful debugging and profiling information via their stacks, they provide backpressure, because there are tons of libraries that work with them and already use them, and most critically because it avoids the "colored function" problem which splits your ecosystem.
Why do most languages not provide that ideal? Mostly for implementation reasons. It's not due to theoretical disagreements or anything. Providing what Loom does is very difficult and is possible largely only because so much of Java and the Java ecosystem is written in Java itself. One reason native threads are relatively heavy is because the kernel can't assume anything about how the code in a process was compiled or what it will do. The JVM on the other hand is compiling code on the fly, and knows much more about the stack. In particular it knows about the (absence of) interior pointers, it knows it has a garbage collector, it controls the synchronization and mutex mechanisms, it controls debugging and profiling engines.
This allows it to very efficiently move data back and forth between a native thread stack and compressed encodings on a GCd heap. It's also why Loom has some weaknesses around calling into native code. Once you're outside the JVM controlled world it can no longer make these assumptions anymore and must revert to a much more conservative approach (this is "pinning" the "carrier thread"). Note, though, that this situation is not worse than async/await colored functions, which have exactly the same issue.
For Java it may still be possible not to allocate the whole stack as a single chunk and instead have smaller chunks like one per few frames. But I really doubt that it can reduce memory pressure compared with CSP in real applications especially given how good GC became in Java.
The big advantage of this over CSP is that you can take existing blocking code and run it on a virtual thread and get all the advantages, there is no function colouring limiting what you can call (give or take a couple of restrictions related to calling native code).
The main difference then between allocating stack chunks on the heap as needed, and stacks grown by the virtual memory subsystem, has to do with virtual memory management matters. If you can use huge pages for your heap, then allocating stack chunks on the heap will be cheaper than traditional stacks.
However, there's generally only a very small number of such closures -- typically only one -- and they are generally one-time use only. That means they can be freed as soon as they tail-call out. Hello Rust.
With threads, green or not, you have a much clearer failure model and is easier to debug.
The point is that e.g. futures are just some threads with a global synchronization mechanism for obtaining the result. Whatever makes the future stall will also make a low-level thread + your own synchronization stall. Or do you mean some more advanced failure-tolerant threading like in Erlang as compared to less advanced threading primitives like futures?
On the other hand, CPS usually is a bit noisier from the developer's perspective; either your continuations are callbacks (whence callback hell) or your continuations are, as you say, manually compressed, tuned structures, which requires a fair amount of manual labor.
I believe Rust uses a CPS transform (well, more of a continuation-returning style, no?), but it automatically generates the tuned structures ("futures"). The cognitive overhead isn't all gone, but it definitely helps.
For me it is the cleanest style of writing concurrent code. And more and more I find I can also replace state machines with it, which makes sense because the compiler generates state machines under the hood usually.
You know, the kind of code where you have to communicate with some outside device and it is easy to do blockingly but devolves to state machine madness if you need to do other things concurrently. For example it would be really nice if I could use async/await in C on a microcontroller to read from a serial port...
1. The "colored function" problem: http://journal.stuffwithstuff.com/2015/02/01/what-color-is-y...
2. Poor interaction with debuggers, profilers, and other tools that expect to be working with normal stacks.
Loom solves this because it lets you work with normal threads, but suddenly you can have millions of them in a process without blowing out your memory or other related problems.
Speaking personally, I've found Lua's coroutines to have the nicest experience for modeling flows like that. The big issue with async/await is the function color problem [0] -- writing async functions is perfectly fine, but mixing them with non-async functions can be extremely frustrating. Especially if you're doing anything with higher-order functions.
[0] https://journal.stuffwithstuff.com/2015/02/01/what-color-is-...
A different way of looking at it is that in asyncs functions you should only do things that have negligible runtime (compared to the response time of your GUI or network service). If your task needs more time, you mark the call site and the called function "async" and the task will suspend somewhere "down in the call stack". (Without looking into it too much, I think something similar actually happens with these virtual threads. They modified IO functions to do cooperative multitasking under the hood?)
As to async functions being contagious, I found it helps to split "imperative" procedures and "pure" functions, and the async color mostly applies to the previous.
As for loom, due to it running all in a runtime, a blocking ‘read’ call for example is not actually a blocking system call (everything uses non-blocking APIs at that level) so the runtime is free to suspend execution at such a blocking site and continue useful work elsewhere until that finishes. So for some “async” functionality you can just fire up a new virtual task with easy to understand blocking calls and that’s it, it will do the right thing automagically, and it will throw exception where it actually make sense, you will be able to debug it line by line, no callback hell, etc.
Loom will also provide something called structured concurrency where you can fire up semantically related threads and easily wait for their finish at one place.
As for pureness, I don’t think it maps that cleanly to async/blocking. What about doing the same function on each pixel of a picture in memory where you subdivide it into n smaller chunks and run it in parallel?
However in other languages, having functions be of a different 'color' is far more painful. In Python for example, a synchronous function has to setup an event loop manually before it can run an asynchronous function. The call works, but nothing is 'running' without the event loop. Additionally, the asynchronous function may have been written to work with a particular eventloop (e.g. trio vs curio), and thus you have to use that type.
If non-blocking code has a standardized control state like Javascript, I think it's better to be explicit about async vs sync.
The reason I say functional languages don't get bit by this as bad is because functional languages rely far less on the specific notion of a call stack, and it's usually much easier to work with continuations (either via primitives like shift/reset or via syntax like do-notation).
On the other hand, these stack frames can be thought of as a large arena allocator for what would otherwise be lots of smaller objects allocated on the heap.
Just because you think they can be more efficient?
After all, a big reason that NodeJS won a lot of popularity on the server is that, for many types of common webserver workloads (i.e. lots of IO, relatively minor CPU usage), NodeJS can actually scale much better than Java with its thread-per-request model.
With these virtual threads, though, you could get the best of all possible worlds - a webserver that scales like NodeJS, but without some of the "CPU starvation" issues you can hit in Node if one executing request doesn't yield, and also without having to worry about "function coloring" like you do in Node with async vs. non-async functions.
Really, really fantastic development, have been waiting to see when this would come out.
Linux can handle a ginormous amount of threads quite well, would be interesting to see a deeper investigation to this theory.
Loom solves this by moving stacks to and from the heap, where there's a compacting concurrent GC to clean up the unused space.
> The JDK implements virtual threads by storing their state, including the stack, on the Java heap. Virtual threads are scheduled by a scheduler in the Java class libraries, whose worker threads mount virtual threads on their backs when the virtual threads are executing, thus becoming their carriers. When a virtual thread parks -- say, when it blocks on some I/O operation or a java.util.concurrent synchronization construct -- it suspends, and the virtual thread's carrier is free to run any other task. When a virtual thread is unparked -- say, by an I/O operation completing -- it is submitted to the scheduler, which, when available, will mount and resume the virtual thread on some carrier thread, not necessarily the same one it ran on previously. In this way, when a virtual thread performs a blocking operation, instead of parking an OS thread, it is suspended by the JVM and another one scheduled in its place, all without blocking any OS threads (see the Limitations section).
Also, there's no new syntax, so you're stuck with all the same thread pool concurrency we've been using for decades.
EDIT: It looks like I'm wrong about this:
> My understanding is that you won't have to worry about blocking a virtual thread, because all IO APIs are being modified to park when executed in the context of a virtual thread.
That said, you'd still need to worry about unsafe code, like JNA/JNI or other such thing that could still block. And I'm not sure there will be a way to prevent long running CPU task from clogging up the virtual thread executor threads.
And, from what I read in the original JEP, the underlying system thread pool (which all virtual threads float between as needed) will be expanded when a virtual thread gets pinned, so you don't have to worry about exhausting your pool. (If you pin too many threads, obviously you'll be consuming more OS resources than you may have expected, but that's a different problem.)
> Some blocking APIs temporarily pin the carrier thread, e.g.most file I/O operations. The implementations of these APIs will compensate for the pinning by temporarily expanding parallelism by means of the ForkJoinPool "managed blocker" mechanism. Consequentially, the number of carrier threads may temporarily exceed the number of available processors.
> The implementation of the networking APIs defined in the java.net and java.nio.channels API packages have been updated to work with virtual threads. An operation that blocks, e.g. establishing a network connection or reading from a socket, will release the underlying carrier thread to do other work.
:D
"File I/O is problematic. Internally, the JDK uses buffered I/O for files, which always reports available bytes even when a read will block. On Linux, we plan to use io_uring for asynchronous file I/O, and in the meantime we’re using the ForkJoinPool.ManagedBlocker mechanism to smooth over blocking file I/O operations by adding more OS threads to the worker pool when a worker is blocked."
Always go with the platforms languages, and the IDEs from the platform owners, even if others are more shinny.
Long term it always pays off to be the turtle, as the platforms move into directions not forseen by the shinny objects, and 3rd party IDEs keep playing catching up with SDK features.
IBM does Java and the IDE (Eclipse).
Red-Hat and Microsoft do Java and the IDE (VSCode).
Sure, but my particular complaint isn't with the functionality; it's with the UI. Yes, VS Code absolutely improves the experience.
In the real Kotlin world of taking a random Kotlin library and call it from Java, most likely "it depends".
Kotlin can call a Java API to spawn a lightweight thread. There's no reason to use coroutines when you can do that.
Though loom doesn't have support for preempting green threads that are blocking the scheduler like go does, I think.
Node.js doesn't create a thread per request; it's single-threaded with evented I/O. You can use node-cluster to start more than a single thread to saturate multi-core CPUs and load-balance HTTP requests across these, but that doesn't make it thread-per-request.
> a big reason that NodeJS won a lot of popularity on the server is that, for many types of common webserver workloads [...], NodeJS can actually scale much better than Java with ~~it's~~ [Java's] thread-per-request model.
Why are you calling that out? The original "its" was correct without the apostrophe.
Adding in the 's is 100% my mistake. I've been guilty of using "it's" as the possessive form for most of my life, but that changes today! :)
Exciting! :)
No clue on the apostrophe.
One thing I say to people using "it's" is that by analogy, you also need to say: "He got he's skills. She missed she's ride. They have they's meeting."
For most words, the possessive form is "<word>'s"
For pronouns (including it) there are different rules. He becomes his, she goes to hers, it goes to its.
Also, words that already end in s don't get the " 's " treatment.
(Question - for words that end in "s", we put the apostrophe after the existing, ending 's', yes?)
Thanks again for posting this - viewing the possessive form of it as (yet another English language) exception to the normal rule of " 's " is really helpful.
This is a great distillation of the intuition I've always had, but never quite verbalized.
The challenge is normally that if any of the threads in the pool, as part of processing a request, needs to itself make an IO call, it will block. Ideally you'd want to park the request processing, return the thread to the pool, pick up the next request, until the IO is done where then on the next thread available from the pool you'd resume that request instead of picking another one. This is what the virtual threads will make really easy I think.
Maybe not _all_ possible worlds. You still have original Threads for things that need an actual OS thread. Its not a solution for UI threading.
There will be code that needs a native thread or non-preemptive threading and shouldn't be run on a virtual thread. In that sense there is method coloring but its yet to be seen how common a problem that will be.
Library writers and frameworks will need to sort out patterns for how to call Runnables in a safe way.
Still, its a nice tool to have.
W.R.T. code that needs a native thread: at the moment there's only two types of such code. One is code that uses Java's synchronized statement. That's supposedly just a, ehm, small matter of programming to fix. The other is calling into non-JVM controlled code. That's fundamental and no approach to scalable concurrency can fix it, not CPS/async/await or anything else because it's a foreign compiler.
But fortunately the JVM has some really interesting tricks up its sleeve there. For instance you can compile your native code using LLVM and then execute the bitcode on the JVM. Well, OK, currently GraalVM doesn't support Loom but hopefully Graal will be upgraded to do so as Loom gets integrated into HotSpot. And when it does, you will be able to call into code written in C/C++/Objective-C/Rust as long as that code can be recompiled with your own toolchain and as long as you can tolerate it being JITCd, also whilst benefiting from Loom's scalability.
Sorta kinda but not when you're working in a framework that will call your code or working in some library where the abstracted code is non-obvious or uneasy to configure.
Maybe its not function coloring, although I wouldn't know what else to call it and I think its quite similar. What would you call the problem?
Java is fine if you don't care about RAM and start time, though.
You should try Quarkus. It is a production framework built by Redhat. It uses Java-GraalVM under the cover to compile your entire webapp to an executable (like golang does).
It's just as fast.
Java is the highest performance and most tuned VM there is. I think you're really thinking of java from a long time ago, if ur thinking this
Not defending the opposite argument, but V8 is also pretty impressive. It's rooted in work done for Smalltalk long before JavaScript was a thing.
The same is true of HotSpot. https://en.wikipedia.org/wiki/HotSpot_(virtual_machine)#Hist...
Your "belief" is putting you at risk of ignoring a wide range of Java use cases unnecessarily.
[1] https://docs.spring.io/spring-native/docs/current/reference/...
Also, Java is working on reducing ram usage: https://openjdk.java.net/projects/lilliput/
May be Graal would save us all. Until then Java is beyond salvation.
And no, you can't configure Java to target <100 MB of RAM. I configured it with -Xmx64m and it still eats around 300 MB. Java just fat and you can't do nothing about it at this time.
1: https://www.youtube.com/watch?v=KmMU5Y_r0Uk
2: https://assets.ctfassets.net/oxjq45e8ilak/5QM86VAnN9XJ9HUIs2...
http://www.open-std.org/JTC1/SC22/WG21/docs/papers/2018/p136...
Perhaps it's easier to address the problems in a managed environment and I really do hope they pull it off. Also it's unclear whether virtual threads will support async file I/O out of the box, or ever. (C# does have Async methods on files.)
Java supports executors since Java 5, and async IO exists since ages with NIO.
Nowadays C++ gets lost discussing language minutiae that it isn't as much fun as it used to be.
Yeah, I agree. C++ is no longer fun at all.
https://en.wikipedia.org/wiki/Green_threads
Which makes me wonder how this is new:
> In Java 1.1, green threads were the only threading model used by the Java virtual machine (JVM),[8] at least on Solaris. As green threads have some limitations compared to native threads, subsequent Java versions dropped them in favor of native threads.[9][10]
So is the "new" part that green threads are coming back to Java?
On Windows NT as well. On my first job around 2000 I did a little bit of Java programming. As far as I recall, JVM scheduled all their threads on top of a single OS thread.
We've been hearing out openjdk's project loom for a while, but we haven't gotten to try this out in Java mainline. I am guessing this will take at least two previews before an initial release. And given the speed the ecosystem moves at, we may not see this reaching widespread use for quite a while.
Structured Concurrency - https://openjdk.java.net/jeps/8277129
Scope Locals - https://openjdk.java.net/jeps/8263012
Something like:
coroutine foo
while queue not full
put something in queue
when full
yield bar
coroutine bar
while queue not empty
take from queue
do something with what was taken
when empty
yield foo
Each time the coroutine yields, it removers it's state and execution resume another coroutine, and when execution is yield back it too resume from the yield point.As I understand, in Java, they are not adding coroutines, but something that is a virtual thread, which is more like a green or lightweight thread. It means that it can be pre-emptively paused and resumed, it doesn't have to voluntarily yield. There is some scheduler that could decide when to execute which virtual thread and so on.
The alternative for point 1 is cooperative scheduling (announcing explicitly when they yield), as you've described.
The alternative for point 2 is "stackless" continuations, where the task yields by returning a callback describing the next step of the task -- or, equivalently, returning some data describing the state of the task, which the task's primary entry point can use to decide where to continue from. (For instance, imagine a function with a big `switch` statement that, when invoked, decides which case to jump into based on its argument, which the caller got from the return value of the last time it got invoked.) Either way, every step of the task constructs and then returns out of its call stack, which can be much more memory efficient, but is also more painful to model tasks in without help from the compiler/language.
(*) Technically, the JVM could take on the role of scheduler as well; it could count the number of bytecode instructions executed, say, and pre-empt a task after some number. This is what Lua supports for some of its sandboxing capabilities. But I think that would be counter to Java's use of an OS thread pool to execute these virtual threads; its own scheduler would be interleaving somewhat unpredictably with the OS'. You'd want to do that if you have long-running jobs that do little I/O (so they hog an OS thread)... but then you'd probably rather put those jobs on actual background worker threads.
coroutine producer
forever
while no full packet
resume recv into buffer
on eof return null
yield (extract packet)
coroutine consumer
while packet = (resume producer)
if packet is null break
process packet
which is more like a green or lightweight thread. It means that it can be pre-emptively paused and resumed, it doesn't have to voluntarily yieldI’ve never heard of this meaning of coro/green/light distinction. Can you please point to some literature?
In your example, it seems to still be cooperative, you simply yield to the scheduler which is itself a coroutine and will then decide what other coroutine to yield back too. Here's a naive coroutine scheduler :
ArrayList coroutines;
coroutine scheduler
for i = 0;; i = i++ % coroutines.size()
yield coroutines[i]
It's still voluntary yielding though, preemptive would be that the scheduler can at any time interupt the task, but here it can't, it will still only be possible to schedule another task ounce a yield point voluntarily yields back to the scheduler.Actually, your example is simpler then that: (resume producer) is the same as: yield producer. And the yield with a return value is the same as: yield consumer. For the latter, the language probably allows yielding to the previous coroutine under the hood or like I said maybe it yields to a scheduler.
I was also showing that you can even do something like yield to a scheduler which will then pick the next coroutine to resume, which makes it even more "thread like", but still cooperative.
The coroutine's cooperative nature has an advantage, it naturally models coordination. With a preemptive scheme like Java virtual thread, you will still have to protect shared data and have ways to coordinate and synchronize like mutex, locks and all that.
As far as I know there is nothing preventing race conditions and dead/live locks in case of coroutines either, isn’t there? Like of course if you have 1 thread these issues won’t come up, but with true parallelism, this model in itself doesn’t protect anything.
If you have two coroutines writing to the same variable, but they yield to each other, you know they won't ever both run at the same time.
You also know if you spawn multiple coroutines that they won't yield except where they call yield, so everything before and after the yield you know will be atomic.
But that won’t be parallel just concurrent, and in case of cooperative “threads”, you could have probably written it in a more readable single threaded way, as that’s pretty much just calling two functions back and forth.
Your second point also only works when you have a single thread of execution, otherwise concurrency will entail parallelism and all the usual problems will become apparent.
That is, assuming you had a single core CPU, with threads you'd still need to synchronize things when implementing concurrency. Coroutines have a more explicit synchronization from their natural ping/pong as you yield which could be said to tend to be safer in the average case.
I think you're maybe conflating something. If two things write to the same global variable for example, that can never be parallel, but it can be concurrent. With threads, the writes to the variables need to be guarded with some synchronization mechanisms, if you forget you'll have bugs.
With coroutines, they will be naturally synchronized by the yield points.
> you could have probably written it in a more readable single threaded way, as that’s pretty much just calling two functions back and forth
It's not just calling two functions back and forth, the coroutines retain state and continue where they yielded. Each time they yield they do not consume additional stack frames.
(...aside from those that are perpetually stuck on Java 8 anyway.)
My team has been happily tracking the twice-yearly JDK bumps. We started development three years ago against Java 8 and made a series of jumps (9, 11, and then 14 onward) and never really had an issue.
I'm not sure I can live without `var`, `record`, and pattern-matching `instanceof` anymore. (With `sealed` interfaces and records, the visitor pattern is long gone... I can only wait with baited breath for exhaustive pattern-matching `switch` expressions.)
https://old.reddit.com/r/programming/comments/lsuojl/jdk_16_...
pron> Assuming you've already made the last ever major upgrade past 8 (which was a relatively tough one), the reason people pay for LTS isn't because upgrades are overall cheaper -- they're costlier, actually -- but because they're willing to pay to not get new features. We've designed the LTS model mostly for legacy applications that don't see much maintenance, and want their dependencies, the JDK included, to change as little as possible.
https://www.reddit.com/r/java/comments/o0m6g8/the_state_of_p...
pron> People who want a new feature to land in LTS still misunderstand what LTS is. People who upgrade from LTS to LTS every three years also misunderstand LTS, and probably get the worst of both worlds.
Personally, I only found Java 9 to be anything like a stumbling block, and that's solely because the module system (Jigsaw) threw all the tooling for a loop. You can easily avoid Jigsaw and never worry about it.
The Java folks try really hard not to break backwards compatibility in general, and modules (+ JDK internals encapsulation) are the only major bugbears to worry about. If you can upgrade, I've found it extremely worthwhile.
> Particularly for non-LTS versions there may be experimental features that are not going to be compatible with subsequent versions, increasing my risk that a migration to the next version is occasionally not quickly.
As for this, the experimental features may as well not exist if you don't enable them. You absolutely should kick the tires with them if you can, but their presence is feature-flagged off by default. I'm on a small team myself, and it's been painless for us ever since jumping to 11.
In algebraic type notation, this pattern replaces a function returning a sum type, X -> A + B + C, with a function accepting a callback that accepts a sum type, `X -> (A + B + C -> Y) -> Y`. But function accepting a sum type is the same as a product of functions, so you have `X -> (A -> Y, B -> Y, C -> Y) -> Y`. The product of functions is the visitor, and `X` is the thing you're visiting.
Traditional wisdom is correct when you have an open family of subclasses (i.e. you don't know, and shouldn't know, precisely how many subclasses there are). But for a closed family, it's just unnecessary; you're blinding yourself from information you already possessed.
The other case is general OSS software, Java reaches a wider audience in my field (distributed databases, streaming data, etc). Java is pretty much considered the lingua-franca of Big Data with some very small Scala footprint and much less fluency in Kotlin.
I generally write all my own stuff ontop of these on Kotlin but drop down into Java where I need to be able to share things.
Virtual threads, project lilliput and valhalla is likely to be a great benefit for Clojure, which has great thread primitives and also spawn a _lot_ of objects that (mostly) don't care about identity.
> There are situations when the VM cannot suspend a virtual thread, in which case it is said to be pinned. Currently, there are two:
> When a native method is currently executing in the virtual thread (even if it is calling back into Java)
Does that mean any kind of native code is currently paying some extra cost due to the possibility of being blocking? What if I e.g. want to call a library that is known to be non-blocking, or make a syscall that is non-blocking which is not pre-wrapped by the Java standard library? E.g. a library that allows to offer interacting with a BPF map comes to my mind. Is there maybe an escape hatch for virtual thread aware java libraries, where they can tell the runtime that they want to call native code without extra guardrails and overhead?
If you have "short" native methods, like in a typical async I/O library, this is not a problem. They cannot be suspended while in there, but they repeatedly go back to Java code where they can be suspended.
So your scenario is only really a concern with a long-running but non-blocking native method, say, a physics library that does lots of computations. The answer is most likely: Don't run that in a virtual thread.
Also all of the existing JVM IO is done with JNI. How do you think java.io is implemented itself? Of course Loom can change those implementations, but strictly speaking JNI is currently extremely common for Java IO. How big of a task supporting virtual threads for the Java libraries remains to be seen, especially for the unofficial extensions like the sun.nio.* package
In general most Java I/O is native because it's "sufficiently fast" for such things. Netty is the exception rather than the rule in this regard.
My understanding is that pinning only enters the equation if you need to yield while a native frame is on the stack. If that call is non-blocking, then by definition you'll call into it and return without needing to yield the current task.
A non-blocking call should give you some way to tell when the job you've requested has completed, of course, and then you need to either poll for it or arrange to be told when it's done. You don't want to spinlock in a virtual thread (you're just hogging an OS thread continuously, which is exactly what pinning is), so either way, you'll end up blocking -- but as long as you're blocking after returning from the native call, you should be fine.
> Does that mean any kind of native code is currently paying some extra cost due to the possibility of being blocking?
I have no special insight, but I imagine any costs are only incurred if a yield actually occurs with native code on the stack. Only then would the yield logic pin the current task to the current thread.
Besides, this hasn't been targeted to a release yet, so it might not come before Java 19, which is a year from now. Even then, it will still be a preview, which likely means another year before it's a stable feature.
Potentially, virtual threads enable them (assuming a serializable environment) (co routines are also enough for this, the serialization capability is what I find interesting).
This will enable a complete freeze of an execution to be stored and even send over the network to be completed somewhere else.
If you can just send the computation around your basically don't need the entire message/protocol boilerplate in your execution code.
For example, think of a game engine that allows to write code with loops and calling function that may wait for an event or a rule to be true. Imagine that your game can be paused stored in case of a connection lost, synced between different server or both at a server and at a client for fast response. Now without this language feature your basic game logic code gets, you can just have a wait statement inside a loop etc., How will you return to the same place on resume? You need more code, storing the entire execution state.. on every condition or a loop you either need to store something or have code that looks like a state machine etc., With this feature you don't need special design pattern, just write it, and the mess is in the language level.
It's solving the same problem that async does, but it does it with virtual threads instead. The idea is that functions aren't coloured, and that normal threaded code will "just work".
I see some benefits of this approach, but I feel that what all of the solutions (Java, C#, Rust, etc...) are missing is structured concurrency[1], without which madness and eldritch horrors of late-night concurrent code debugging are guaranteed.
[1]: https://vorpus.org/blog/notes-on-structured-concurrency-or-g...
If you read the JEP though, you'll see that Executors are auto-closable now, which means you can use try-with-resources to wait for all spawned threads to stop before continuing execution.
But the scape hatch works by.... coloring functions! Specifically if a function needs to spawn a longer lived background task, it needs to take a nursery parameter.
If a function wants to call a function that might spawn a long lived function, it needs to either own the lifetime of said nursery, or more commonly accept a nursery as a a parameter, and pass in the supplied one.
In practice with java, the nursery concept (StructuredExecutor) will only be used for those cases where it is actually helpful, (i.e. where you really want the function call to not return until all concurrent tasks are finished), and everywhere else, like background tasks, existing primitives will be used.
And all nurseries/StructuredExecutor is buying you is the ability structurally enforcing joining of the relevant virtual threads. It lets you avoid some of the common mistakes in structuring such code, but I'm not convinced that is where the eldritch horrors of concurrent code debugging live.
I think the real eldritch horrors come from buggy attempts to implement low lock code, failing to realize that locks or other synchronization is needed when accessing a certain variable, etc. Basically race condition type situations.
I personally almost never have had substantial concurrency issues related to failing to join my concurrent tasks.
"Provide syntactic stackless coroutines (async/await) in the Java language. These are easier to implement than user-mode threads and would provide a unifying construct representing the context of a sequence of operations, but that construct would be new, separate from threads while being similar to them in many respect yet different in some nuanced ways, would still split the world of APIs between those designed for threads and those designed for async/await, and would require the new thread-like construct to be introduced into all aspects of the platform and its tooling, resulting in something that would take longer for the ecosystem to adopt while not being as elegant and harmonious with the platform as user-mode threads. Most languages that have chosen to adopt async/await have done so due to an inability to implement user-mode threads (Koltin), legacy semantic guarantees (JavaScript), or language-specific technical constraints (C++). These do not apply to Java."
I wonder if Rust should also be included in the final sentences.
I can understand Java programmer want the goodies offered in languages that are more geared towards concurrency. But multiparadigmatic languages really suck because they are no longer a single language. You get islands of different practices and a partitioned set of practitioners. And they can't always use each other's code (C++ being the most extreme and painful example).
This makes me kind of glad I switched to Go 5-6 years ago. And it makes me wonder when the (good) intentions of the Go designers to not absorb every idea that comes along will go out the window and Go will start to grow knobbly bits all over.
Java was one step along the way, but let's say it had adequate representation of heavy-handed tools we already had in C/C++ that made some forms of concurrency somewhat easier. But it was still some ways from promoting concurrency in that threads were pretty costly and you still depended on locking to move state between threads. And it isn't like CSP hadn't been thought of.
After about 20 years of programming Java and 5-6 years programming Go I wouldn't really list concurrency as a main feature of Java. Because you kind of go at it the way you go at it in C/C++. I think someone who has programmed (for instance) Erlang would feel much the same way.
It has nothing to do with language. It has everything to do with how other libraries (especially standard libraries) structure their code.
Pay them no mind though. Java is already the premier server side language for serious work and this is just another tool in the toolbox, hopefully I will see less RxJava in my future. :)
Hence why JetBrains is so eager with creating duplicates from every Java library in Kotlin.
This reads heavily of FUD.
Last I checked though, Kotlin was threading coroutine suspend/resume points into methods as part of bytecode generation (it’s been a while, please do let me know if I’m wrong on this) which is not something most engineers are ready to read in the simple case, much less when trying to interpret the compiler’s name mangling scheme.
In either case, the implementation that ships with the JVM will be more capable because of it’s privileged position of integrating with the runtime, so good news! Eventually Kotlin might be able to use the native facility.
Virtual Threads are a somewhat lower level primitive than what co-routines provide. It will allow other frameworks on the JVM to integrate and benefit from it in a similar way. E.g. Vert.x, RX Java, Spring's Flux, etc. Probably using this directly is not a great idea as there are so many nice frameworks to choose from already that will protect you from doing silly things. But it is nice to have a good implementation of this built into the JVM.
Kotlin's co-routines is somewhat unique in how it can work with and seamlessly integrate code written for other concurrency and asynchronous frameworks. So technically, this is just yet another thing that they can work with. If it has something that resembles a callback, a promise, a future, etc. you simply wrap it with a suspendCoRoutine and the resulting function is a nice suspending kotlin co-routine friendly function. Kotlin's co-routine library ships with extension functions for Spring Flux, RX Java, and a few more things and it is easy to write your own ones. It's a great way of taking away the pain of using those frameworks directly. In the browser, you have similar extension functions on javascript promises, and so on.
I've been doing a lot of Browser UI programming using Kotlin JS lately. Co-routines are very nice for that. Works very similar to how I use them with Spring Boot to implement non blocking APIs. We actually share a lot of kotlin code between client and server.
This change would make the base primitives non blocking (really cheap Blocking which is better) by default, so now you can use any~ library and it should just work. Waaay better.
And even there "worked" is a very liberal definition. Callbacks, promises and whatever else you can find will imho never be as intuitive as threads.
and then almost immediately went in 5 different directions around how to program around it.. callback hell, async/await, generators, promises etc..
Using non-blocking read/write pretty quickly expands to writing your own scheduler with all the needed quirks/boosts/etc.
The other cost is context switching, which is much cheaper with virtual threads.
The context switching sucks so much that an alternate approach, swapping threads for throwing an exception on yield and re-running the method and ignoring side-effects until we get back to the resumption point, saves us a significant amount of time.
Sarcasm aside, it's probably to steer away from kotlin and make it easier for users when seeking help/docs online.
"Virtual" is a little less standard, but it draws on existing patterns: they've virtualized threads in the same way that the OS virtualizes the CPU (timeslicing) and virtualizes memory (literally, virtual memory).
Preemptive scheduling is more about whether the scheduler (in this case, the OS scheduling the carrier threads) can pause a thread no matter where it is in its processing, and since virtual threads are executed by OS threads (the virtual part is just the JVM tracking the context and where to resume at), they're preemptive.
[0] https://stackoverflow.com/questions/671049/how-do-you-kill-a...
If you're going to do a bunch of computation without blocking, I'm not sure virtual threads are the tool you want to apply in the first place. They're meant to provide cheap blocking and cheap context switching. If you've got a job that's more CPU-bound than I/O bound, it's probably worth using OS threads instead.
Why does Oracle keep throwing money down this pit?
Non-Goals:- It is not a goal to change the existing implementation of platform threads, that represent Operating System (OS) threads. It is not a goal to automatically convert existing thread construction to virtual threads. It is not a goal to change the Java Memory Model. It is not a goal to add new inter-thread communication mechanisms. It is not a goal to offer a new data-parallelism construct in addition to parallel streams.
- Native threads are great. They have a lot of uses, why touch them
- Automatic conversion will break a lot of stuff and eliminate some of the benefits of native threads. Being able to use an API/implementation that is *almost* the same is a huge beneift
- The memory model is a completely separate thing that Java worked on for decades. You don't want to touch that and you don't need to
- Inter-thread communication is a separate thing. There's no reason to go after that. Same is true for the data parallelism stuff.
The focus should be on hitting the 98% of what matters and getting it out without breaking everything we already have.
On the memory model: I'd argue that technically the memory model would be implicitly updated to treat these virtual threads the same as classic threads. Which is a very minor update. Beyond that, in order to maintain that memory model, the implementation will need to ensure proper barriers are used on pausing and resuming virtual threads, so that a virtual thread resumed on a different physical core is guaranteed to see any of its previous writes (as would be expected within a "single thread").
> It is not a goal to automatically convert existing thread construction to virtual threads.
Key word: automatically. If you want an OS thread, you should be able to get one. But if you simply spawn a task into a virtual thread, that task should behave basically the same as in a platform thread, but with less context switching overhead (and less strict claiming of OS resources).
The locus of choice is at whichever part of the code constructs the threads to begin with; if you change it to use virtual threads, you shouldn't have to change anything else.
> It is not a goal to change the Java Memory Model.
Was there something you wanted changed? I've heard that OCaml's memory model is pretty stellar, but I don't know much about it myself, and I think a memory model is a sufficiently fundamental thing that changing it might cause backcompat issues -- but it depends on the change you'd like to see.
> It is not a goal to add new inter-thread communication mechanisms.
Is there something new you'd like to see? Supposedly, the existing inter-thread communication mechanisms will work just as well for virtual threads.
> It is not a goal to offer a new data-parallelism construct in addition to parallel streams.
Virtual threads are explicitly task-parallelism constructs, so this makes sense. Maybe you would build data-parallel constructs on top of them, but it's not something the JVM would be directly responsible for here.
https://github.com/Kotlin/kotlinx.coroutines
In fact, Kotlin Coroutines are an brilliant on the android platform. We are talking severely memory and CPU constrained architectures here.
That said, Kotlin Coroutines are popularly used in production on server side - https://vertx.io/docs/vertx-lang-kotlin-coroutines/kotlin/
I doubt anyone would switch to Java Virtual Threads anytime soon, unless via Kotlin.
Maybe Kotlin will leverage this in its underlying infrastructure.
Switching to virtual threads will, at least in the microservices I work with, involve changing a few lines (Executors.newUncachedThreadPool() -> Executors.newVirtualThreadPool()).
Switching to Kotlin coroutines will involve re-writing a large part of our codebase, which is why it hasn't been done.
It would surprise me if people didn't switch to Java Virtual Threads by the thousands when it comes out.
I think you're mistaken. The Java community is just so tremendously larger relative to the Kotlin one, that this will have more users within months. I really liked what Kotlin was doing, but Java since got lambdas, they have closed the biggest gaps that drove migration.
it seems to me that kotlin is leading the way and java is following, but the delta between them is quite large and possibly growing.
What Kotliners used to say about Java back then:
- no data classes. Now there's records in Java. Check.
- no lightweight threads. Boom, project Loom came in.
- no support for FP. Now there are lambdas and :: operator in Java. Also, Stream API.
- no type inference. Then came 'var'.
And some minor things that are in Java now as well: pattern matching; sealed classes as a way to get ADTs in the foreseeable future; kind of immutable collections; etc.
Yeah, you can argue that these features are not-so-native and painful to use in Java comparing to other languages. And, as one who programs in Scala, I 100% agree with you here. But you can't argue that people used to switch to Kotlin and Scala without hesitation because it was worth it summing up all the switching pros and cons. Nowadays, this is not the case anymore. You don't have to risk and adopt a new language technology stack.
Delta is objectively decreasing. For sure.