Networking I/O with Virtual Threads – Under the hood
inside.java
inside.java
What is that breaking point where developers should switch to more scalable models? From the article:
> There is a threshold beyond which the use of synchronous APIs just doesn’t scale
What is a good rule of thumb these days?
This “mostly idle” assumption is gonna break the moment you hit a load spike and wake up most of those “idle” threads.
The success of async is often (mis)attributed to eliminating unnecessary resources, but in practice it's actually because async model implicitly enforces an invisible job queue, where usually only a single process/thread per core is heavy lifting the actual work, and async I/O will feed processing pressure back to the load source via network (e.g. TCP ACK).
Async NetworkStreams or Sockets in .NET is a well-designed API over traditional BSD sockets’ select() function which is how you scale to tens of thousands of concurrent connections using only as many threads as you have hardware DoP - it’s not that I’m wowed that .NET has modernised it, it’s that Java’s solution is to wallpaper over a gaping hole in their fundamental design instead.
That is why you will find out lots of code bases full of Task.Run() and direct calls to .Result, instead of using async/await operators.
Having worked with both platform since their yearly days, I am not sure which design is actually the best one, considering brown field development.
What I am describing is what a large majority of .NET consulting projects do, because async/await taints the whole call stack up to Main() or event handlers, and almost no one is doing .NET Core projects from scratch.
Just because it is discouraged at conference talks and blog posts, doesn't mean devs aren't doing it on the trenches when dealing with the crossfire of project deadlines and massive codebases.
Oh, I'm sure of that, don't worry.
I'm just fortunate that I've never had to experience that myself.
What's being introduced now is a way to use that API without needing to touch the whole concept of async code. From the developer's perspective blocking calls and threads is pretty nice. It's async that's awkward and hard to use. Making the existing threads facility scale much better is a very clean approach: virtual threads introduces "virtually" no new concepts at all, in fact, if you aren't working with native code via JNI or Panama then you can act as if virtual and physical threads are the same. You literally just upgrade Java and maybe flip a couple of switches and you can now assign one thread per connection yet handle millions of connections (assuming you don't run out of other resources of course).
Even with Loom's green-threads (not the original green-threads) async IO in Java will be too otherworldly until Java's language designers cease their intrasigence about adding first-class support for `await`.
I'm sorry but I can't take you seriously after you say something like that.
Besides that, it’s pretty hard to decide you want to go async later on. In languages where async is not omnipresent, you usually get some form of explicit continuations; async/await sugar if you’re lucky. But you can always accidentally call something not async and hang, or just spend a lot of CPU time with no ability to preempt.
I think Erlang, Go and probably a couple of others basically stand alone in this regard; everything is just async by default. It’s a shame that it requires a decent amount of runtime baggage to accomplish this.
Ruby and Python can only ever have one thread running at a time due to global interpreter lock. Though to keep in mind when I/O happens on a thread it can be parked and another thread can be picked up to continue processing. But in essence one thread is running at a time and these languages built billion dollar businesses.
Just scale horizontally computers are cheap is the message I get and remember to keep it simple.
I'd be very surprised if Robinhood's trade matching engine is done in Python. I'd bet they use an off-the-shelf engine in C++/Java or another language with good multi-threading support.
> computers are cheap is the message I get and remember to keep it simple.
Fully agree! Getting to product-market fit is probably easier with python and ruby.
But, it’s only efficient in convenient, shared-nothing type architectures. If you are dealing with longer lived, more stateful stuff (like a game server perhaps) it’s not going to be so easy.
Of course, it’s easy for fancy VC funded stuff to defer optimizations later. If you are working solo you are probably going to want to, counter intuitively maybe, spend more time thinking about how to keep things optimal and simple. There’s probably startups out there worth real money whose entire stack could run on a $20/mo DigitalOcean box if it were better optimized, but actually runs on a $1000/mo Amazon or GCP account. Of course there is a lot to be said about that (obviously, the latter set up should, if done well, be a lot more resilient and prepared to deal with scale) but it does go a long way to demonstrate why someone might care about efficient async; the scale of money you can burn when you are bootstrapping or doing solo hobby projects is a lot different from well-funded startup or established business...
Everything has tradeoffs. I do believe Go wins in context switching, but the GC latency and CPU usage has definitely caused some problems. The Go developers have done an impressive job with GC latency and performance, but it’s hard to be cheaper than free, so sometimes Rust is a better option, for example.
If you start off with a synchronous API and thousands of threads, migrating to something else later is very expensive.
I started writing some fancy asynch code, but on checking with a profiler, the thread overhead was pretty negligable, so I left it alone.
To not use async. is really bad, you are wasting so much CPU. Even in a async. scenario just copying memory from user space to kernel and back takes 30% of the CPU at full 100% utilization, imagine having all those context switches too!
Threads are a huge mistake 9/10 times. I don't judge when the OS uses them, but for most of my applications I prefer to get everything into a nice ringbuffer so I can tear through tens or hundreds of millions of transactions per second with that single blessed core.
Most of this is meaningless if you use a database engine to serialize your transactions for you, but its still worth considering from an abstract perspective IMO.
This means you get an actual meaningful stacktrace when debugging, and not something stemming from a mysterious event-loop thread.
In Java you could save your coroutine state to disk, and wake it up later in theory.
----
EDIT: This being said, I'm 99% sure Kotlin is going to pass Loom's goodness onto their developers when it's available, probably reusing the existing coroutine API.
So for every JVM/Java feature post Java 6 that gets introduced, they will have the dilemma of how to integrate them into a way that keeps language semantics across compilation targets, having multiple solutions to the same problem (Kotlin's one and what each platform later introduced), or just expose them via KMM and leave the #ifdef burden to the community.
That is why platform languages always carry the trophy, even if they are the turtle most of the time.
Virtual threads are superior to coroutines, which are still a pain in the ass in Kotlin, cause issues with mocking and you can't even evaluate them in the repl.
PMD, SonarQube and nullable annotations have long sorted that problem in our Java projects.
Not a fan of any of those tools, tbh. Still better than nothing I suppose, but I'd rather use Kotlin in the meantime.
Right now there are 20 years of having null as default and turning it on just generates endless warnings.
So it is left for complete new projects where almost no third party libraries are being used, which right now is very little.
I also hope that C# 10 doesn't get !! for checked parameters, but it might already be too late.
Value types are perhaps a better example. Kotlin has them already with nearly identical semantics to Valhalla, but without the ability for them to have more than one field due to the need for erasure. Once Valhalla arrives, Kotlin can simply remove that restriction when targeting the JVM, perhaps add another annotation or compiler flag to say "make this a real Java value type". No language changes needed beyond that.
Kotlin is semantically so close to Java already that they aren't really growing apart, they're growing together. It works well enough to justify its usage, for me.
Just like those, it will slowly adopt whatever is more appealing from the guests and then carry on its merry way, while the other slowly lose their relevance while newcomers try yet again to challenge the place of the host language on the platform.
This makes the same concept radically cheaper by not involving the kernel, which is great because Java hasn't had green threads for a long time, and doing everything async in small worker pools worked but has admittedly been pretty painful.