I/O Access Methods for Linux
scylladb.com
scylladb.com
> This usually happens when the ratio of storage size to RAM size is significantly higher than 1:1. Every page that is brought into cache causes another page to be evicted.
While true one can optimize these things by unmapping larger ranges in bulk (to reduce cache pressure) or prefetching them (to reduce blocking) with madvise, allowing the kernel doing the loading asynchronously while you're accessing the previously prefetched ones.
If you know your read and write patterns very well you can effectively use this for nearly asynchronous IO without the pains of AIO and few to none extra threads.
This lets you play around with policy for blocking reads in user space. Examples: try to make progress on partially read data & queue up the rest on another thread; try performing reads from your network thread and if unavailable queue up the blocking read onto another disk thread & serve a different request.
There's no RWF_NOWAIT support for write in pwritev2. But it's technically possible to implement it and if your write will cause a writeout to disk return EWOULDBLOCK.
LWN description: https://lwn.net/Articles/612483/
LWN summary of my talk: https://lwn.net/Articles/636967/
Final patch set by Christoph: https://lwn.net/Articles/731700/
Disclaimer: I'm the author of the preadv2/pwritev2 syscalls. And the original author of the support for RWF_NOWAIT, which Christoph Hellwig took over. So feel free to consider this to be self congratulatory.
Correction: It looks like there's a patch set floating around for adding support for RWF_NOWAIT in pwritev2 as well. https://patchwork.kernel.org/patch/9787271/ (comment from Christoph)
If I understand this correctly one can use RWF_NOWAIT to do what one would naively expect from O_NONBLOCK. Am I correct?
How does this differ from recv(..., MSG_DONTWAIT)?
Besides some technical reasons, the big reason was an "impedance mismatch" of the recv API which was working with stream data. In that context you're depending on another party send data. Also, you can wait on this to happen using another API (select et. al) so the buffer will be filled not by your actions. On the other hand preadv2(..., RWF_NOWAIT) is not going to trigger any more read in and theres no wait to wait on the data. Although the previous statement is not a 100% true... preadv2 may or may not trigger readahead (if it's enabled).
Here's the whole thread about it... if you're interested in the history of how this came to be: https://lkml.org/lkml/2014/7/24/787
I would assume the intent is similar but applied to a different type/source of IO.
I can't speak for Avi though.
I suspect he is writing from the perspective of a database implementor. For a database the page cache is not predictable in ways you want it to be.
For writes the issue is IO scheduling, when flushing occurs, and how much is flushed at a time and at what rate.
For reads if a page is not in memory you have to block a thread to read it. If you are thread per core this means you either have a thread pool on the side to do IO or you block an event loop preventing other work from making progress. Even having the thread pool on the side is problematic because checking whether a page is in the cache is not cheap either and requires a system call which is kind of self defeating.
My understanding is that those normally don't incur faults that would switch to the kernel context. They're just slow due to several memory lookups that have to be done to satisfy the primary memory access.
SPARC and MIPS for example worked like this, and server POWER chips had a hashed page table that many OSes populated lazily, effectively treating it like a (more efficient) software TLB.
This may work for sequential reads, but not for random reads.
The same is true of plain read/write, barring extra data copies to/from the kernel.
The only reason to ever use AIO is if you're doing direct I/O, in which case mmap doesn't apply.
I have seen Windows File Explorer lock up over and over again for the same assumption. They assume the dir scan is fast operation. A slow, bad USB drive, NFS, Samba mount, or trying to scan directory with thousands of files will lockup the GUI for long period of time.
Dragging a file over slow/busy system icon with SMB mount will just spin / hang the whole GUI - not a very good UX.
Noticed this bit "block size which is typically 512 or 4096 bytes" and was wondering how would the application know how to align. Does it query the file descriptor for block size? Is there an ioctl call for that?
When it comes to IO it's also possible to differentiate between blocking/non-blocking and synchronous/asynchronous, and those categories are orthogonal in general.
So there is blocking and synchronous: read, write, readv, writev. The calling thread blocks until data is ready and while it gets copied to user memory.
non-blocking synchronous: using non-blocking file descriptors with select/poll/epoll/kqueue. Checking when data is ready is done asynchronously but then read, write still happens inline and the thread waits for it to be copied to user space. This works for socket IO in Linux but not for disk.
non-blocking asynchronous: using AIO on Linux. Here both waiting till data is ready to be transferred and transferring happens asynchronously. aio_read returns right way before read finished. Then have to use aio_error to check for its status. This works for disk but not socket IO on Linux.
blocking asynchronous: nothing here
It also provides cross-platform aligned buffers for DMA.
Note that block size can be logical or physical sector size. IIRC st_blksize will typically report logical sector size (usually 512 bytes), but physical sector size (usually 4096 bytes) is often larger for Advanced Format drives, and is what you want to be using for optimal alignment. Any multiple of 4096 bytes is usually a safe bet for block size. Note that some SSD drives have 8192 or 16384 bytes as a "physical sector" page size / block size.
In the benchmark what was the physical block size, 16KB? it seems with O_DIRECT after 16KB the speed tapers off and stays at about 130MB/s. And there is a significant jump from 4K (60Mb/s) and 8K (90Mb/s).
Thanks for the input, glad you found it useful!
Regarding the benchmark, the physical block size of the drive itself was 512 bytes. The memory page size on the Ubuntu system was 4K. It would be good to update the benchmark script to include these kinds of numbers.
I'm not sure why O_DIRECT speed increased with block size=4K (60MB/s) to 8K (90MB/s) and tapered off at 16KB (130MB/s). Perhaps system call overhead is more noticeable at 4K and 8K and less so at 16KB? Any other ideas?
130MB/s is the raw drive write speed and matches what I got from dd for the same flags for the same drive.
(vm)splice maybe?
Some other advantages:
The kernel has a global view of what is going on with all the different applications running on the system, whereas your application only knows about itself.
The cache can be shared amongst different applications.
You can restart applications and the cache will stay warm.
Another big disadvantage is that the page cache is only useful for caching data of actual files and nothing else.
In the context of this article, many databases use the page cache to cache their data-structures, either in combination with private pages (ie postgres) or entirely (ie LMDB).
Some applications do use the page cache for non-caching purposes though, for instance IPC, or ensuring their working-data will survive an application restart. The memory is still associated with files naturally, but this is often regarded as a feature, not a problem. Some applications use files on tmpfs so that their page cache data is backed by swap (or nothing), the same as anonymous pages.
A little off topic, but I have been waiting over a decade for Linux to merge all async waiting into one system call.
Wouldn't it be nice if there was a kqueue system call in posix? It would then force Linux to finally implement it.
I was always confused - still I am - with regards to different caching layers that involve in an IO operation.
http://haifux.org/lectures/253/alice_and_bob_in_io_land/alic...
But this raises the question: isn't this ratio decreasing with cheaper and cheaper RAM these days?
(I've seen a lot of systems already moving to completely in-memory databases, taking this to the extreme; so this is actually a reality)
Human-generated metadata (i.e. OLTP data) grows roughly with respect to the population of internet connected humans - humans only can produce so much content in their finite time. Advances in computers far exceed this growth rate, so more of these kinds of problems fit into RAM each year.
The volume of media data, which depends more on the increasing fidelity of the media (videos going from 320K to 4K, photos going from 1MP to 30MP) have been increasing rapidly but that trend won't continue forever because there are diminishing returns. Do you care about the difference between 4K and 8K video? TV companies hope you do. At some point (likely soon) you will stop caring.
OLAP data, which includes data produced by IoT devices, sensors, etc is increasing geometrically. But this kind of data is not best stored in RAM anyway, and is served best by scale-out column-oriented and time-series databases.
In that case, a fast data access library will not buy you much, because the disk will be slow anyway.
(By the way it is common that an in-memory database is combined with on-disk storage for large blobs; those blobs are just flat files, and the OS already offers the fastest way to access them without any library)
We've been using Cassandra for many years. As the business has grown, we've refreshed/updated/resized/tweaked/hacked the cluster. Unfortunately past a certain point, there were no more performance wins to be had - further - we started experiencing stability issues.
We came across ScyllaDB. I had heard about the underlying SeaStar framework and was excited to test something based on it that was addressing a real pain point.
We are not (yet) customers of ScyllaDB - however we're pretty advanced into our POC. To summarize: * Extremely low latency compared to cassandra * During delivery valeys, peaks, and even burst batch loads * It delivers the above running on a fraction of the hardware Cassandra does * Very good no-bs responsive engineering support
Scylla's having a conference in SF later this month where AdGear will be presenting our use case more in depth. If it sounds like the product may address some pain points you're working on, might be worth attending.
/not affiliated with Scylla beyond the potential-customer-in-POC case described above
Our use case is web-scale and real-time IoT.
Our experience with the product (as well as the team behind it) has been extremely positive.