Linux with “memory folios”: a 7% performance boost when compiling the kernel
lore.kernel.org
lore.kernel.org
Linux future seems bright with regard to performance.
[0] https://lkml.org/lkml/2021/1/11/98
[1] https://lore.kernel.org/lkml/20210608115254.11930-1-sj38.par...
It then needs someone to do a big parameter tuning to select optimal settings.
Too many decent algorithms don't make it into the kernel because there are too many tunables, and the ones that do typically arent well tuned for anyone's use case.
Even big projects like Ubuntu typically don't change many tunables in the kernel.
But a more modern language would benefit from defining some optimizations like inlining as mandatory, the same way tail calls are (and of course GCC has this.) Then there isn't a chance of deploying a build with extra-slow behavior.
That would average out into a few percentage points of performance improvement in every application across the globe which runs on linux / android. Merging a patch like that in linux would be like quietly reaching into every device across the planet and giving them a small, free CPU improvement and drop in power usage. Not game changing for any individual user, but huge in aggregate.
I heard a story - years ago Google hired an engineer who happened to be an expert from a previous life in video codecs. He wasn’t working on YouTube at Google, but out of interest he pulled up the YouTube source code to see what it did. They were just using the defaults for some of the encoding parameters. He tweaked a few of the encoding parameters and in doing so saved Google millions of dollars per year in compute/storage/network traffic. A few hours of his work was probably more beneficial for Google than years spent in his primary role.
And if tuning opportunities like this abound at Google, you know they’re everywhere. It’d be much better if Linux distributions simply shipped kernels which are already compiled with PGO, with a reasonable profile based on normalish use.
The correct answer would've been for ffmpeg to not ship that way.
Anyway, I am an expert on video codecs just like that guy is, and I said what I said ;)
Running your DB with the default parameters.
Not running Linux kernel parameters depending on your workload. (+10Gbit NICs, lots of connections...)
There are reasons, why those knobs are there, as there are always trade-off, and having them automatically adjust is very hard.
But maybe that was what you're referring to.
I use it to encode my DVD and Blu Ray TV/movies to HEVC. Doing so reduces the filesize on DVD's by roughly 80% and Blu Rays by 60%
The downside is the encoding is an incredibly CPU intensive process. Hardware encoders like Nvidia's Nvenc or Intel's QuicSync look absolutely terrible and are a non-starter for archival storage
On a stock Fedora XFCE install, I would get roughly 0.5 FPS for a 1080p Blu Ray file (29.97 FPS at 1920x1080)
A Gentoo Linux installation with a global -O3 -march=native as well as LTO, PGO and Graphite enabled globally boosts it from 0.5 to roughly 1.3 FPS. Still slower than realtime, but an absolutely massive improvement
I do a probably overkill CRF of 20 just to be safe. But everything looks absolutely perfect, even blown up on my 65 inch (1080p) TV
No hard feelings if you don't want to though :)
When I was using Eigen on my code, the biggest performance boost came from -O3. -march and -mtune did minimal improvements on the systems which I've ran benchmarks on.
OTOH, I'd like to underline that heavily optimized scientific code and libraries are neither naive (in terms of algorithmic complexity/implementation) nor straightforward :D
E.g.: This is how Eigen configures its internal vectorization parameters: https://gitlab.com/libeigen/eigen/-/blob/master/Eigen/src/Co...
Of course on the scale of you or I, such optimisations may be little more than a curiosity most of the time. A process that normally takes 100 minutes (a video transcode, perhaps) now taking 95 followed by the machine being idle while it waits for us to look back and see the result, is not really benefiting from the benefit.
> without actually needing to learn how to change it
This is certainly a problem sometimes, but not an argument for tuning bring overrated in general. It usually comes down to tuning the wrong thing, like someone playing with mysql engine & kernel IO parameters to eek out a fraction of a of % bonus when they could improve index structures, or fix queries with no sargable predicates, or both in unison, to get benefits measured in orders of magnitude.
It can be overrated in the case of average gamers tweaking their hardware to get a few extra FPS on top of the many tens they already get, but again this not being worth the time for some, even many, doesn't mean it can make a huge difference to a pro-gamer or someone using old kit where "a few extra" is a relatively large gain.
If you're using a few optimized libraries and designed your code for high-speed parallel execution, the compiled thing becomes almost impossible to trace in a practical manner.
Some of the Matrix and solver code coming from Eigen is heavily optimized for SIMD operations and using it with -O0 is just painful.
Moreover, I had to verify its memory sanity and used Valgrind for that. Using a full size problem meant it had to churn for days to finish execution.
Having deadlines doesn't always help in this stuff.
I also think Perf is phenomenal for its scope. It's showing performance metrics of the code without any processor dependency and instrumentation.
I suspect your broader point stands - certainly "performance is multivarious and means different things in different contexts" has sunk in for me. But I suspect the actual difficulty of reasoning about performance deters most people from doing it beyond big O notation. The fact that I'm aware that performance is complicated doesn't drive me to understand the complexity, I mostly just say "good enough" and move on.
Nothing can hit 200fps at 4K, which you'd need to reach low enough latencies to be imperceptible.
Have you seen phoronix framework + benchmarks? They use common tools, but did some good work on making the tests repeatable and accessible to anyone else.
And that's just the phone. I could come up with many more questions regarding desktop and server usage as I'm more familiar with those.
I'm afraid, there is no one optimal set of configuration options. Not even two or three.
I don't quite follow why it's all being done as one huge set of patches -- rather than first merge the groundwork and then all the conversions.
It is. The patchset submitted contains only 33 patches, while the full patchset contans about 200 patches.
This is equivalent to saving ~16 hours when you have a week of rendertime.
Rendering is much more 'pure cpu' work, so you most likely won't see much difference there due to this work.
The most sensible hypothesis I've heard about it is that THP really helps with large rendering loads, and Windows doesn't do that yet.
As for bug reports, well https://twitter.com/bgolus/status/1080213166116597760
In my experience in interacting with them, most of these small indy dev teams use Unity or Unreal or Godot these days. Outside of graphic design, the graphics don't really cause big issues any more. Having a team of a handful enthusiasts but often slightly inexperienced programmers figure out the flaws in their own game logic is.
That is to say, we at most might have a bit of logic to tune whether we do *1.5 or *2 on realloc, but why not more?
There must be patterns we can exploit in common use cases to be sneaky and do less malloc-ing. Profile guided? Runtime? I might have some results by Christmas, I have some ideas.
Food for thought: Your container has a few bytes of state to make decisions with, your branch predictor has a few megabytes these days.
About the best you can do is if you know beforehand roughly how big it will be, is to reserve that capacity with std::vector::reserve().
Think bigger than changing the coefficient, there probably is no optimal factor.
But be careful not to call reserve() in a loop adding data to a vector - will allocated exactly what you request and not do exponential growth.
http://www.open-std.org/jtc1/sc22/wg21/docs/papers/2021/p040...
Does superpage management get easier when each superpage is composed of only 32x of the next size down? When I first stumbled on this idea, it seemed like it would have many more opportunities for forming intermediate-sized superpages.
2MB vs 4kB isn't quite the same ratio as 4MB -> 32GB, but it's still a lot less pages to cache in the TLB, and it's not too big to manage when you need to copy on write or swap out (or compress with zram) and whatever else needs to be done at the page level.
* Blow-ups in various kernel data structures. There was some virtio code which was allocating N pages per driver queue.
* Problems with GPUs, either the driver or the firmware assumed 4k pages. (Edit: This actually affected Power, not ARM, but the issue is caused by page size: https://lists.fedoraproject.org/archives/list/devel@lists.fe...)
* Filesystems make assumptions about page size versus block size.
* Processes generally take more RAM, with RAM wasted because of internal fragmentation.
Last I checked, at least CentOS 8 (which should be the same as RHEL 8) is still using 64k pages (search at https://git.centos.org/rpms/kernel/blob/c8/f/SOURCES/kernel-... for CONFIG_ARM64_64K_PAGES=y).
And AFAIK, to access the maximum amount of physical memory in AARCH64 (52-bit physical addresses, instead of 48 bit physical addresses), you must use 64k pages. Since RHEL is normally used on servers, it makes sense to want to be able to access huge amounts of physical memory; that's probably the true reason (or even the sole reason) RHEL uses 64k pages on AARCH64.
You have: 2^48 byte
You want: tebibyte
2^48 byte = 256 tebibyte
2^48 byte = (1 / 0.00390625) tebibyte
how many servers do you have with more than 256 TiB of RAM?If you add up an optimisation of just a nanosecond in like openSSH, how much would that do globally?
Makes me wonder whether this alternate major release cycle was a good idea. If you delay all feature development for a year, you'll get a barrage of features once the performance-only OS version is out the door, and there's not enough time to do all of them properly, so you get buggy and slow versions.
Maybe doing performance improvements and feature development at the same time would have been the better choice? How is it being done at Apple nowadays?
I believed optimizations like that at a global scale will not have any impact.
Lets say that this nanosecond will be saved trillions of time a day. Resulting in minutes to an hour a day saved globally.
* Not a single user will notice. * In 99.99% of cases the CPU will not be fully pegged and thus that one nanosecond of compute will not be used to do something else at all. * CPU throttling isn't that fast so, you won't even save that much power.
If we bump it up by 6 orders of magnitude to a millisecond that all remains true. Even though you are potentially saving 100s of years of computing time a day. Extremely small gains distributed across very large number of machines don't tend to be as impactful as you would hope on a global scale.
This is not to say that small gains are worthless. Many small gains added together can be substantial.
But aren't we talking about extremely mature kernel code here? My impression is that all kernel distros in high use are optimized but they are general use software. The degree to which you may optimize software is constrained but the diversity of use cases you must support.
Isn't this what __attribute__((pure)) [0] is for?
[0] https://gcc.gnu.org/onlinedocs/gcc/Common-Function-Attribute...
y=f(x);z=f(x) implies y=z
What they want is something different: y= f(x) implies y=f(y)
This means if you give something the head of a list of pages, it won't try to go to the head again and again, it knows its already there.The 'folio' idea as I understand it is roughly an alias for the existing 'page' structure, but code knows it is already at a head AND it should do the work on the whole list, not only on the head.