Actually I'm on btrfs so reflink copy is like instant I think? I should test that better.
Actually I'm on btrfs so reflink copy is like instant I think? I should test that better.
That, or your algorithm could be optimal for concurrency already - and you would see an immediate performance improvement.
With blocking I/O and parallelism you have a thread ready to go when the operation is complete. You have N threads for N iops. With concurrency you have to dequeue completed work, and then delegate that work to (usually) fewer than N threads. Dequeuing completed iops takes time (it's an extra syscall), and there may not be a thread ready hand the completed iop. More latency.
Running 1000s of threads isn't realistic because your OS would typically grind to a halt, so concurrency is unavoidable. It does have a cost, though.
To be clear, the added latency here is better than the work never happening at all (which would be the result of running 1000s of threads on modern mainstream operating systems), but there is unavoidable latency if you are handling >N iops with N threads (which is intrinsic to the definition of concurrency).
I am referring to the broad, general case, much like big-O works. You can find numerous exception to big-O, such as preferring arrays over hashes when the set is very small. Let's invent big-L notation, N is the number of threads, M is number of iops. With pure parallelism you have L(N), with pure concurrency you have L(M), and with a hybrid you have L(M-N).
The whole point of io_uring is to drastically decrease the number of sys calls, creating channels where more requests can be filed with lower than traditional cost of a readFile syscall for example, and where completion can also be lower overhead delivery of events.
So historically I kind of would have agreed with the parent. Today, we don't really know! Hence my excitement.
In theory. I did some work on high performance filesystem I/O on Linux about a year ago, doing intensive random-access to fast SSDs, and found io_uring to be slightly slower than a well-tuned thread pool with an appropriate queue depth.
That was a little surprising as the thread pool has to do system calls for each I/O operation and io_uring does not. Perhaps it is faster with newer kernels or other access patterns.
io_uring is better able to adapt autonatically to different numbers of cores, device queue depth and amount of filesystem data cache residency. That comes from it having access to kernel state which is not made available to userspace on Linux, to guide thread offloading decisions, rather than from the ringbuffer communication.
not to fanboy out TOO much, but your posting on HyBi was & is greatly influential to me. see, https://github.com/rektide/pipe-layer#essence
(sorely neglected project to me but also still very near & dear, still a core principle & value in my pantheon of beliefs)]
no particular comment on io_uring. thankfully jens keeps making it better. the numbers he posts for his synthetics keep seeming impossibly good. but i fully am ready to believe the real situation is more complicated.
i do wish we'd see some uptake from the usual suspects. both Deno and Node have delayed/deferred work on these topics. but supposedly slowly happening in node. https://github.com/libuv/libuv/issues/1947 https://github.com/denoland/deno/issues/16232