POSIX AIO has usability problems also outlined in the previously linked thread.
Remember, all I/O in NT is async at the kernel level. It's not a "bolt-on".
IoRing is limited to file reads, unlike io_uring.
POSIX AIO has usability problems also outlined in the previously linked thread.
Remember, all I/O in NT is async at the kernel level. It's not a "bolt-on".
IoRing is limited to file reads, unlike io_uring.
Unfortunately that's only half-true. You can (and will) still get blockage sometimes with IOCP, it depends on a lot of factors, like how loaded the system is, I think. There is absolutely no guarantee that your I/O will actually occur asynchronously, only that you will be notified of its completion asynchronously.
Also, opening a file is also always synchronous, which is quite annoying if you're trying not to block e.g. a UI thread.
The implication of both of these is you still need dedicated I/O threads. I love IOCP as much as anyone, but it does have these flaws, and was very much designed to be used from multiple threads.
The only workaround I'm aware of was User-Mode Scheduling, which effectively notified you as soon as your thread got blocked, but it still required multiple logical threads, and Microsoft removed support for it in Windows 11.
All I/O in Linux is also async at the kernel level! The problem has always been expressing that asynchronicity to userspace in a sane way.
Arguably asyc/await could help with this; obviously it didn't exist in 1991 when Linux was created but it would be interesting to revisit this topic.
Wouldn't that just consist of I/O operations returning futures and then having an await() block the calling thread until the future is done (i.e. put it on a waitqueue)?
With Rust in the kernel this becomes somewhat possible to conceptualize.
You’d probably want to use either some sort of code generation to do the requisite CPS transform[1] or the Duff’s-device-like preprocessor trick[2], but it’s definitely doable with some compiler support. Not in an existing codebase, though.
(Brought to you by working on a C codebase that does express stuff like this as explicit callbacks and context structures. Ugh.)
[1] https://dx.doi.org/10.1007/s10990-012-9084-5, https://www.irif.fr/~jch/research/cpc-2012.pdf, https://github.com/kerneis/cpc
[2] https://www.chiark.greenend.org.uk/~sgtatham/coroutines.html
> As I was looking at the raw system calls related to I/O, something immediately popped out: CloseFile() operations were frequently taking 1-5 milliseconds whereas other operations like opening, reading, and writing files only took 1-5 microseconds. That's a 1000x difference!
This is why DevDrive was introduced[1]. You can either have Defender operate in async mode (default) or remove it entirely from the volume at your own risk.
The performance issue isn't related to sync or async I/O.
[0] https://gregoryszorc.com/blog/2015/10/22/append-i/o-performa...
In Linux, filesystem paths are super-optimized, with all the filtering (e.g. for SELinux) special-cased if needed.
But even still, Windows also had to cheat to avoid completely cratering the performance, there's a shortcut called "FastIO": https://learn.microsoft.com/en-us/windows-hardware/drivers/i...
I wrote a filesystem for Windows around 25 years ago, and I still remember how I implemented all the required prototypes and everything in Explorer worked. But notepad.exe was just showing me empty data. It took me several days to find a note tucked into MSDN that you need to implement FastIO for memory mapped files to work (which Notepad.exe used).
Robert Collins explains that performance is just as good as Linux and the performance loss on Windows is due to file system filters (Defender)[0].
This is what DevDrive intends (and does) fix.
It used to be _much_ slower, like orders of magnitude slower, especially for directories with a large number of files.
To get more of the abstractions out of the way, you want DevDrive. And don't use Explorer.exe as a test bed which has shims and hooks and god knows what else.
https://github.com/maharmstone/btrfs
It's complete enough you can boot and run Windows from Btrfs.
https://lilysthings.org/blog/windows-on-btrfs/
I wonder how many of these limitations affect this, notably performance.
If WSL1 could natively talk to Btrfs without the performance-sapping translation to NTFS, would that resolve the poor Git performance etc?
This impacts performance particularly when working with a ton of tiny files (like git does).
https://speakerdeck.com/trent/pyparallel-how-we-removed-the-...
And yeah, IOCP has implicit awareness of concurrency, and can schedule optimal threads to service a port automatically. There hasn’t been a way to do that on UNIX until io_uring.
In a nicely wrapped PDF :-)