Async File IO
cmhteixeira.com
cmhteixeira.com
Post from mailing list from a couple years ago, for example: https://mail.openjdk.org/pipermail/loom-dev/2021-November/00...
I suppose it's possible to create a new set of API or augment the existing ones not to use thread pool, still rather unlikely. Also it'd require a general widely supported async from the OS, incl. Windows
You can't expect that with AsynchronousFileChannel::write at the moment, though, because the buffer is consumed on a background thread. As the docs say [1]:
> Buffers are not safe for use by multiple concurrent threads so care should be taken to not access the buffer until the operation has completed.
And as the article says, on Windows, these writes are "truly" async at the OS level, so portable code already needs to deal with that.
[1] https://docs.oracle.com/en/java/javase/17/docs/api/java.base...
* AsynchronousFileChannel was released in jdk7 on 2011-7-28, works by submitting io to a JVM thread pool
* Linux io_uring (AIO) released on 2019-5-5, works at the kernel level
The article's point is that AsynchronousFileChannel does not use io_uring, which makes sense given the API was implemented roughly 8 years previous. Yes, AsynchronousFileChannel is Asynchronous, however, it operates in user space and does not use io_uring.
My observation (and praise) is the JVM implementation is unlikely to break backward compatibility. If io_uring support is desired in Java, you have to jump through a few hoops with the current APIs. Not a big deal in my opinion... if you truely need those few extra percentage points of performance none of this is probably news to you.
IOCP merely allows IO completion to happen on worker threads in a threadpool rather than the particular thread that initiated the IO request. Many concepts in the NT kernel come from a long line of production kernels in mainstream DEC operating systems going back to RSX-11M and VMS. That entire lineage of operating systems (culminating in NT) all have true asynchronous IO using IO Request Packets (IRPs) within the kernel to initiate/queue IO requests and immediately return. There are individual cases within NTFS where blocking will occur but those are special cases rather than the rule as the kernel IO system in general is entirely async.
A thread pool was also more reliable than the kernel APIs at ensuring overlapped I/O for throughout. libaio (kernel API) calls sometimes turned into blocking calls depending on filesystem internals, in a way that thread pools don't.
POSIX AIO, on Linux, just uses a userspace thread pool. It's implemented in libc and is not particularly fast. You can do better with your own thread pool.
Perhaps Linux threads got faster, faster than the APIs got better.
Even with io_uring, in my tests on fast NVMe storage RAIDs where it should make the most difference, I found io_uring wasn't noticably faster than a well-optimised (for Linux) thread pool.
io_uring can potentally adapt better to shared workload and different cache situations, because it has access to kernel information that userspace thread pools are not granted. But I was surprised to find no significant random-access I/O performance increase between my thread pool and io_uring, when I was trying to optimise both methods to get the best throughput out of a system.
Tokio's filesystem handling in Rust is the same way by default - a pool of IO threads.
Edit: I just remembered that Avi Kivity (of KVM and ScyllaDB fame) wrote a tool to detect this: https://github.com/avikivity/fsqual
You need completion notifications instead (like IOCP or io_uring or POSIX aio). You could in principle start an async splice from an FD into a pipe and then poll for the pipe to be ready to read, but IIRC there is no async splice.
It includes some pretty big cases like "if the filesystem is encrypted" or "if the write extends the file": https://learn.microsoft.com/en-us/previous-versions/troubles...
The file, not the file system. NTFS encryption is applied at the file level (and should be fairly rare these days). Bitlocker should not be affected as it is an encrypted volume with a normal file system.
This part is close to 17 years old now (java 1.7, 2007). The NIO for files that came with 1.4 (20y+) was not async for files either. I wonder where the 'truly async' idea has come from as the docs are rather explicit about it.
> When not specified, the following system applies:
> - JDK: openjdk-17.0-amd64
> - OS: Ubuntu 20.04.6
> - Scala: 2.13.10
> - Physical disks: USB 2.0 16GB, FAT32
Generally you'd construct a test to be as similar to a real-world situation as possible to get results that resemble those you'd encounter in reality. Ubuntu on a FAT32 thumb stick is, well, unusual.
First, some framing:
Asynchronous scheduling is about sharing scarce resources efficiently. It allows you to keep your hardware busy without spending gobs of memory on OS threads. Specifically, it allows the system to make progress on many suspendable tasks in parallel. When a task needs to wait for something time-consuming to happen, it can be suspended, freeing up a CPU core to work on other tasks that are not waiting. The time-consuming work can be executed by some sort of centralized shared worker that has a broader view of the work that many tasks are waiting on. Often, that worker is a single thread or a small thread pool that performs blocking IO against a given type of resource (file, network socket, etc.).
About the JVM:
On the JVM, this type of sharing can currently only occur within a single process (at least on Linux). This is because the JVM doesn't currently take advantage of any OS-level asynchronous scheduling primitives. Therefore, the Java standard library itself manages some central IO worker threads within each JVM process, which suspendable tasks can offload work to.
Implementing io_uring, as is being done for the JVM now, will move to using the kernel's own asynchronous scheduling primitives. This will allow sharing to occur cross-process, since the shared workers will now be inside of the kernel. That'll be a nice efficiency gain for systems where many tiny processes run on the same kernel, but it likely won't help that much for big single-purpose machines.
Also of note: The virtual threads adopted in Project Loom allow programs written in a blocking style to behave more like lightweight suspendable tasks, so it has a lot of synergy with an io_uring implementation.
If anything the author just discovered how IO works on Linux.
Is this true? Wouldn't they still consume an OS thread in that case; possibly consuming the pool of OS threads that are servicing the virtual threads?
How does loom handle that?
Regardless of any future updates, this class will probably always use a thread pool/ExecutorService. It is part of the documented behaviour and API.
That said, it's on their wiki as something they plan to address. Virtual threads aren't the end of the dev work here.