Optimizing Linux Memory Management for Low-latency, High-throughput Databases
engineering.linkedin.com
engineering.linkedin.com
p.s. I can recently created a screencast about control groups (cgroups) for anyone interested @ http://sysadmincasts.com/episodes/14-introduction-to-linux-c...
[1] http://www.oracle.com/technetwork/articles/servers-storage-a...
[2] https://www.kernel.org/doc/Documentation/cgroups/memory.txt
[3] https://www.kernel.org/doc/Documentation/cgroups/cpusets.txt
NUMA can be a real pain. You can get a 40% hit on direct memory access, and far worse if you're modifying a cacheline in another processor. On one of our VoIP workloads, we noticed major (250%+) increase in performance and CPU stability after splitting a very thread-intensive process into multiple processes, each set with affinity to a particular core.
OSes try to help you, but it seems like they're primarily concerned with multiple processes, not huge processes like databases. Such processes should become NUMA aware and handle things themselves for best performance.
It might even make sense to ask if you can split the machine on NUMA boundaries and just act like they're separate systems. RAM's getting very cheap, and RAM/core is going up faster than CPU power is (it seems to me, anyways).
Also, is there a reason not to use large pages directly for the mmap'd sets if you know you're going to have them hot at all times? (I assume they read the entire file on start?)
> Also, is there a reason not to use large pages directly for the mmap'd sets if you know you're going to have them hot at all times? (I assume they read the entire file on start?)
We could use large pages directly. But, as I mentioned in the article, the performance gains would be negligible compared to the gains that come from having things in memory in the first place. These are not very large memory systems and the page table / TLB miss overhead doesn't seem to be biting us. We are just following the mantra 'pre-mature optimization is the root of all evil' :)
It's only when you start getting to the metal to see what your hardware is actually capable of that the TLB stands out as a glaring source of inefficiency.
Put another way: yeah, the TLB is making your app slow, but it's doing so always, so you don't notice. Instead, you mistakenly think your hardware is just slower than it really is.
There is some good shared knowledge in the post (unlike this comment, to be fair), but what does drop by 400% mean?
If a rate drops by 100% it becomes zero. I get that.
If it increases by 400%, the outcome is slightly ambiguous (do we add 400% for 500% total or do we multiply up to 400% of the original value).
But a rate decreasing by 400% - am I the only person who finds that (not uncommon) expression hard to conceptualize?
https://github.com/apache/cassandra/blob/trunk/src/java/org/...
I didn't have LI-size databases -- just a dozen Python processes allocating each perhaps 300MB and all restarting at the same time were enough to trigger it, taking 10 minutes rather than 2 seconds to start up.
http://lwn.net/Articles/568870/ (subscriber-only now, will be free in a week)
Should that be 80%?
Edited to add: Apparently it should be 75%, per comments elsewhere.
should be...
you know what it should be ;)