ZFS on Linux still has annoying issues with ARC size
utcc.utoronto.ca
utcc.utoronto.ca
I have had to re-learn this lesson over and over again with my own software. "Tom, why did the system just do that?!" scream my users. "Er, let me check", i respond, already feeling that sinking feeling. "EDUNNOMATE" says the log. So, i add some logging around the decision (the data feeding into it, the choices made, the actions resulting), redeploy, and wait for my users to start screaming again, hoping that this time, i will be able to give them an answer.
https://www.usenix.org/legacy/events/fast03/tech/full_papers...
...that said the author has been writing on the ARC for more than 10 years judging from his blog links so perhaps that paper did not answer his questions.
Setting zfs_arc_min to something like 50% of arc_max stopped it from dumping the ARC every 10 minutes.
YMMV.
Not resilient on a system level, but refilling the cache is cheap.
ZFS was generally pleasant from an operability viewpoint once we ironed out the quirks, but the perf hit from no sendfile was too much.
It’d be really nice to see that fixed like the recent DIRECT_IO additions.
(I'll just point of that using sendfile means that traffic is unencrypted... which is probably fine on an internal network, but I've started adopting the stance that even internal network traffic should be encrypted unless there's a very good reason not to do that. An absolute requirement for performance might be a good reason.)
Now I’m more curious about the actual threshold where not having sendfile begins causing noticeable performance problems… at what point before you become Netflix?
Of course sharing resources between a couple services would be good, as NICs and switch ports are sill a way from free.
The default settings for L2ARC fill rate are also super low.
I haven’t had time to track down exactly why it’s so slow, yet.
I have an ubuntu mirror on the machine that's around 150gb, and doing a `tar -c $MIRRORPATH | pv > /dev/null` shows lots of reads from the HDDs, even on second, third, fourth runs. It confuses me.
Of course, if the L2ARC dies, you shrug, while if allocation_classes vdevs die, your pool is gone, so there's that tradeoff to be aware of too.
I personally can't decide whether I think it's a bug or not, since if the MRU is all old items there is an argument to be had that you don't want it in cache any more...but dumping 100% of it strikes me as a bug either way. :)
Page faults from NFS client side aren't served by the server when they should (readonly map, reading a page). I could imagine this is related.