Linux Performance
brendangregg.com
brendangregg.com
https://access.redhat.com/documentation/en-us/red_hat_enterp...
sudo apt install tunedhttps://en.wikipedia.org/wiki/Brendan_Gregg
He is also the star of the Shouting in the Data Center viral video
https://www.youtube.com/watch?v=tDacjrSCeq4
True genius!
Another latency metric that you'll see, often w/respect to web apps and microservices is "P99" and similar. This is the amount of time in which 99% of requests get their response. For a higher percentile, you get a better idea of worst-case performance.
0. Another thing people may want to optimise for is performance per watt, but I won’t say much more about it.
1. There are cases where bandwidth optimisations are latency optimisations, eg if you can fit more of your processes onto one box, you are reducing the average distance between the processes and whatever they talk to and hence average latency
2. A very obvious thing to do when optimising for latency is increase bandwidth enough that the bandwidth doesn’t throttle you
3. I feel like mostly if you are aggressively optimising latency, there isn’t much Linux tuning to do. Maybe I’m wrong – I don’t really know much about this – but I think it’s mostly pinning to a core, running tickles, doing user space networking, and then hardwarey things like tuning page size, SMT, power-saving settings, and other things like choice of hardware.
[1] https://pdfs.semanticscholar.org/bce7/5f78d340cac32dccd8631f...
Sometimes you can improve both, but often it's a tradeoff.
As of Linux 6.5, the scheduler understands that when one SMT "core" is busy, that means it might not be the best idea to schedule something on the the other "core", since it's really just a single core with a very low cost context switch. This makes certain very-parallel things noticeably snappier for me, and I can see it on the CPU usage graphs.
YMMV due to cache coherency and NUMA issues. :D
Same goes for my dual core laptop vs quacore laptop, both same generation, but other one runs definitely hotter (4 core)
Linux Performance - https://news.ycombinator.com/item?id=13498485 - Jan 2017 (64 comments)
Linux Performance - https://news.ycombinator.com/item?id=8205057 - Aug 2014 (22 comments)
“Chesterton’s tuneable parameter”.
Side note, I’ve found btop a super useful replacement for glances, to have an all-in-one TUI view of system performance and loading. Wonder how much those dev(s) are leveraging this, and whether anything out there’s motivated to build better TUI monitoring tools.
Every server I go on, first thing is, start up tmux, dedicate one window to btop.
I haven't read all the slides yet but one thing I was wondering was if you ever found any significant performance increases from kernel build options. In my Gentoo days when I would play around with build flags I would change kernel Makefile to use -O3 and apply a patch for -march=native. In hindsight, looking at some Phoronix benchmarks it appears this is actually harmful to a number of workloads. Curious if you ever found any cases otherwise.
This is such a depth subject, with a long list of variety of observability tools. At minimum, make sure you know deeply uptime, dmesg, and iostat. These are your friends to give you a glimpse into various system aspects like load, memory, CPU, and more, enabling a diagnostic overview of system health. This is what I call, the “let me take a look at it” check list, 1st page of 100!
When emphasizing methodologies for performance analysis I recommend careful benchmarking to holistically evaluate system behavior and workload characteristics. with before and after scenarios. Make smaller changes first, then gradually compound what you think will provide benefits. Remember, labs and production never behave the same.
This is where it gets tricky, CPU profiling with tools like “perf” and visual aids like flame graphs enable targeted analysis of CPU activity, along with tracking hardware events to optimize computational efficiency. You need to know more than “it’s the app man, was fine until the latest release from development”
When you are the admin and speaking to a developer; Linux, tools like ftrace and BPF come into play, allowing for detailed tracking of kernel function execution and system calls, which can be vital in troubleshooting and performance optimization. You can also be the developer, varying the admin’s intuition… as the saying goes, trust but verify.
When it’s your code, then you better know BPF! It not only facilitates efficient in-kernel tracing but also propels the development of advanced custom profiling tools through bcc and bpftrace, offering deeper insights into system performance.
Last comment, it’s %$$% hard! Tuning means you need to navigate through adjusting a myriad of system components and kernel parameters, from CPUs and memory to network settings, aiming to optimize performance and reliability across various system workloads, else you can blame it on the network! :D
Really, you need to have a good behavioral attitude at change management, as chasing code or kernel parameters could be a daunting task that just overwhelms everyone in a moment where you might be time constrained and the preasure could lead to a higher degree of human errors.
Trying to squeeze a little more juice out of something is bound to come at the detriment of something else, or worse, break something else in unexpected ways.
Basically, if the tunables aren't obvious in whatever default config you're using, the issue isn't in that config, it's that you're asking too much of your hardware and just need better hardware.
That's .. yeah, that's completely false. I can think of dozens of things that are not right out of the box on any distro, on any hardware, in common practice. For example suppose you roll out Ubuntu on an EC2 instance, say a c6i.16xlarge, a 32C/64T, single-socket x86 server with a Nitro ENA. Where are the netrx interrupts delivered? Is RSS on/working? RPS, XPS? Interrupt coalescing? The distro can't make optimal choices for all use cases, but what they ship by default is a config that's not optimal for any use case. Literally nobody would consciously choose the defaults after thinking it over.
... this still hasn't made it to the distributions we're using, though. More specifically: the kernel releases. There's a lot in service pre 5.4.
Hanging on by a thread. Relief is near.
Of course there's no reason to tune if stock works fine. Plenty of people buy or a rent a reasonable computer and it has more than enough capacity for their work with default tuning. That's fine.
But when you run out of CPU or memory or X, it's often a good idea to see if there's reasonable things you can do to get more out of the hardware you already have. Depending on what you're doing, there's often a lot of room for improvement.
For some networking tasks, doing proper alignment of threads and work with Receive Side Scaling or similar can make a tremendous improvement in capacity versus naive threads. In some environments, the bandwidth costs when you're using enough capacity to see that mean that machine costs of doing it well versus naively don't matter, so you may as well do it naively and spend your engineering time elsewhere. In other environments, gettingthe same work done with 10% of the nodes is valuable.
Also, in many cases, better hardware needs more tuning, rather than less. You don't need to spend a lot of time avoiding cross core communication on an 8-core single socket machine. But if you get a dual-socket, 128-core per socket machine and you're not careful about cross socket communication, you'll spend a lot of CPU on memory arbitration (which you'll have to know or learn how to look for)
Now bookmarked this immediately.
Have not read deep, but from first view look good!
For both OS and services.
Would any one happen to know a similar set of resources for Windows tuning (preferably Windows 2019 AWS EC2s)?
The truth is usually there are tradeoffs and the defaults fit a broad general case.
If you want throughput there are tunables for that, if you want low latency then usually those are inversely correlated. Same for tuning for low data loss after failure and so on.
You have to spend time learning the tradeoffs, which sysadmins used to do- now nobody has time as they have been munged into one role at many places.
No, Brendan Gregg is certainly no junior. The issue is that people take his advice as law and do not read further.
I, myself, remember flipping random switches because a (book) resource said that it would unlock performance.
I’m suggesting that taking the time to understand performance properly is ideal and I am attempting to urge people to properly invest the time.
Then I am lamenting the fact that we do not actually have time over because there’s so much smushing of responsibilities.
That's because Brendan Gregg doesn't flip switches, he shouts at hard drives[1]. ;-)
In the past I've been a sr dev (but a jr sysadmin) and was tasked with improving the performance of an upgraded database server. The problem turned out to be with NUMA on the larger server which a combination of reading and semi-random config fiddling of both Linux and MySQL parameters (plus a bit of BIOS/CMOS tweaking) brought up to expected levels.
There's often no better way than learning on the job as there is so much to know that you can't simply learn them up front for when you will need it. What we can do is learn what there is to know and remember to look into those if it seems relevant. I mean everyone who's well experienced now probably started config twiddling somewhere to get there.
This isn't a issue if you use `relatime` which is the default.
BTW I used to continue to mount everything noatime anyway, since having the atime field be set upon file creation and not anymore afterwards was a way to get file creation time, and I found that more useful than the access time. This isn't necessary either anymore since the introduction of an actual file creation time.
The right way is to follow the scientific method. Collect data. Make a hypothesis with a plausible mechanism of action. Test the hypothesis. Arrive at a solution. Record how you arrived at your new found knowledge so that those who follow you understand why you made the change. The person who follows you is your future self as often as not.
Part of the problem is that reasonable defaults for performance is a somewhat new phenomenon. It used to be that the defaults for Linux kernel settings, Apache, MySQL, etc... were terrible for production use.
So there's a lot of history that "you have to change them" burned into people minds, documents, etc.
Some defaults are just historical curiosities, some defaults were configured 20 years ago and nobody took the crusade to update them, some might be bad, but changing them would break too much stuff in the wild.
Now I'm not suggesting that everyone should change everything. I almost never change defaults, myself. But I just don't agree than defaults are good. They're probably not bad, and that's about it.
Yeah, they fit the general case of a MIPS R4400 with a 1mbps network adapter, a situation that nobody faces today. I think the most glaring example is `rmem_max` which is never sufficient to support a coast-to-coast 1gbps flow, and every individual Linux user in history has needed to independently discover this stupid sysctl.
"Junior X starts flipping knobs thinking Y is somehow holding him back"
This is sadly how many of us have to learn. If you don't have a mentor or someone to check your work, it's a way to learn.
Experience will tell you not to make bad decision. You get experience by making the wrong decisions.
He also wrote a lot about Solaris, but I won't hold that against him /s.
Looks like the second edition of his book almost addresses that :P
"The second edition adds content on BPF, BCC, bpftrace, perf, and Ftrace, mostly removes Solaris" https://www.brendangregg.com/systems-performance-2nd-edition...