Obviously the kernel fix is the right thing to do but until that's vetted may be something like the above can help.
Obviously the kernel fix is the right thing to do but until that's vetted may be something like the above can help.
In practice tough, we use PostgreSQL, and we don't have any control over how Postgres reads its pages. So for our customers, we'd still have this problem.
Postgres could do this though if they detect broken kernel version and the right workload and many users might auto benefit from that.
dd if=clickstream.csv.1 iflag=nocache bs=1M | wc -l
A more common technique is to bypass the page cache altogether and is often use to avoid the many unfortunate characteristics of the current Linux VM. This is done usually with directIO: dd if=clickstream.csv.1 iflag=direct bs=1M | wc -l
Now postgres might be able to use directIO as an option?Another related problem with too much caching when writing to slow device can be seen in this thread: http://thread.gmane.org/gmane.linux.kernel.mm/108708 That thread actually describes two problems. 1. That Linux waits too long before writing 2. When it does write large amounts to a slow device it locks out everything else
Also, there seems to be confusion as to whether MADV_DONTNEED can be 'destructive'. I think this is a difference between anonymous and file-backed mmap(). Do you know what the actual case is?
If the app tells the kernel it is done with the range - it is telling that it doesn't care about the data in that range anymore. So MADV_DONTNEED will not flush dirty pages to backing file store without msync() - if you access that range again it will reload the pages from the backing file or zero-filled ones for the anon case.
Do you see anything in the Linux kernel code that says otherwise?
I really doubt that if two processes are mapping the same file, and one calls madvise(MADV_DONTNEED), it'll drop the pages from memory entirely. That seems like a great way to let one process DoS another. If the other process has marked it MADV_WILLNEED, that would be especially bad.
If they've both mapped it MAP_PRIVATE, then the mappings should be entirely separate anyway (though copy-on-write semantics are presumably used), and a madvise() on one shouldn't affect the other.