The Four Month Bug: JVM statistics cause garbage collection pauses
evanjones.ca
evanjones.ca
What is happening is that the JVM is dirtying a previously clean page (this also happens in your test program, because the dirty pages in your mmap'ed file are being regularly written out - and therefore made clean - by background writeback).
If, at this point, the global dirty limits are exceeded (/proc/sys/vm/dirty_bytes and /proc/sys/vm/dirty_ratio) then the task will be paused to throttle the generation of dirty pages.
In some circumstances, the kernel has to ensure that pages aren't modified between initiating and completing the writeback for that specific page.
Btrfs is a copy-on-write filesystem, so it ends up needing to use this guarantee more often than the others. This is something the btrfs developers are actively working on improving.
Here's a good article describing this in a bit more detail: https://lwn.net/Articles/442355/
EDIT: Fix a typo
There isn't an infinite buffer in memory for disk writes either.
The only guarantee when you write to an mmaped page is that your wrote to the memory, whether or not it makes it do disk is up to many different things. So before you can write to the memory, it needs to have the right contents, that can mean a read has to finish, it can also mean pages have to get murdered to free up for you to have the memory to read in to. I can't think of how the write itself can actually block unless a read is required which hasn't finished (like the file is in read/write mode or something) in fact, other than a pagefault, there is no way it can be a blocking operation, the pages are stitched in to the processes page tables. At least I can't think of how it can block on a write right now, I've had a couple glasses of wine with dinner though. [edit] m_time update makes some sense, that blocks though?
In write only mode there are optimizations to not require the read.
You've almost got it. The mmaped pages may be in the process page table, but they may be in the page table as read-only: if the process tries to write to the page, the process traps into the kernel. If there are few dirty pages, the kernel will mark the page dirty, make it writable from the process and make the process runnable again. Apparently, if there are a lot of dirty pages, the kernel will not fill the request immediately, it will wait. While it's waiting the process is not runnable (other processes with the same memory space would continue to be runnable)
https://www.youtube.com/watch?feature=player_detailpage&v=JM...
Turns on the profiling data where there was a discrepancy between CPU times and Wall times was only OS X because of a problem with a trap() call on OS X specifically, not any other platform. His moral of the story: even profilers have bugs.
https://www.youtube.com/watch?feature=player_detailpage&v=JM...
I think I am seeing a pattern today.
Before that usually it spawns a bunch of pdflush processes to flush data out in the background, but if those can't keep up then it moves to blocking the process. On older systems and older spinning drives, blocks could take seconds even.
See /proc/meminfo for these two entries:
Dirty: 4 kB
Writeback: 0 kB
Dirty are the current dirty pages, and writeback is the current amount being written out.This would make a good case study in Heisenbugs. GC delays that only happen when you're collecting GC statistics.
That was the day i learned about _NO_DEBUG_HEAP !
What is reported?
I have been wondering if Zing handles threads blocked in memory mapped files better.
I noticed jankiness in my eclipse and adding that flag seemed to help.
http://lwn.net/Articles/396561/
But thanks for the snark anyway!
[1] BTW, if you're mentioning a term that isn't widely known, it's helpful to link to a definition.