With 32 bit code this was absolutely not the case. I have not looked so closely at 64 bit code, but in 32 bit land -fomit-frame-pointer was often closer to a 20-30% performance boost. Of all of gcc's optimization options it was the only one that I ever found to make a significant difference in real code.
More often than not, they're invalidated by some advancement in either computer hardware or software. Not just "no longer relevant", but often counterproductive.
Worse still, these internalised tricks and rules are often related to Moore's law scaling, making them exponentially wrong over time. Not in the figurative sense, but a literal one.
You see this with people claiming that I'm exaggerating when I say all code should be parallel. Meanwhile AMD is about to sell a 128 core / 256 thread processor. You can have two per host for a staggering 512 hardware threads. Single threaded code is capable of utilising just 0.2% of that machine!
A similar "tuning best practice" that I still see applied in the field is database servers with dedicated data, log, temp, and templog drives... on cloud VMs where this is guaranteed to have worse performance than simply pooling the same amount of disk into a single logical drive.
People pick these things up and hold on to them when they're wrong by orders of magnitude.
I can't think of many (any?) other industries where more experienced senior staff can be this wrong about as many things...
-https://manpages.ubuntu.com/manpages/xenial/man1/pxz.1.html
-https://lzip.nongnu.org/plzip.html
And so on, however the Compression-ratio is slightly worse
tar -cf - input/ data/ | zstdmt > out.tar.zstWe’ve been promised these magical ultra parallel machines for a decade now and either they’re much harder to build than the alarmists say, or there's simply no market.
What we are seeing is a hierarchy of processing cores. It's already visible even in the desktop market. Intel has their efficiency cores. Apple has the power and efficiency cores. And the GPU on top. So we already have 3 different stages here, at the top the least amount of cores but the best single threaded performance going all the way down to, let me check how many threads, almost 10000 threads on a flagship consumer GPU running simultaneously right now.
Intel tried with their Larrabee on what happens if they just toss in ton of traditional low performance CPU cores. It failed to perform. It's really hard to beat the modern SIMT style GPUs when it comes to massively parallel computation.
Now whether or not all 128 threads need to operate on the same chunk of memory or not is somewhat up for debate. It is very good for some applications (Let's Encrypt issues all of the Internet's TLS certificates with one Postgres instance), and irrelevant for others (just run 128 copies of your node app and send requests to them at random).
And on AWS there are instance types with 448 VCPUs [1] and in fact 128 VCPUs is a very common instance type. In fact with the growth of Kubernetes these larger configurations are increasingly popular because it's more cost efficient to pack the containers into fewer/larger nodes.
Also would add that Scala is in my opinion by far the best language for writing safe, highly concurrent code when you pair it with frameworks like ZIO. [2]
[1] https://instances.vantage.sh
[2] https://zio.dev
That kind of is defeating your argument though: If people pack many containers on large boxes, the individual container probably is not making use of the possible parallelization across all cores.
If you have an app that runs 24/7 at 100% bursting across all cores then Kubernetes will simply schedule a single container on that instance.
Edit to clarify: SwapPING bad, having swap space good.
> Disabling swap doesn't prevent pathological behaviour at near-OOM, although it's true that having swap may prolong it.
Basically, the memory manager will just bulldoze through moving all the other applications to the swap to let the leaking program pursue its leak.
> you are left with a system in an unpredictable state. Having no swap doesn't avoid this.
Yes, it does: the program that leaked my 16GB is killed at the very beginning of the situation, and I don't have to either wait dozens of minutes for it to fill the swap on top of that to finally be killed, or just hard reboot the machine and lose the whole state.
How, and by whom? In a situation where you are running out of memory, it is programs that allocate the most often that have the highest chance of seeing an OOM, not those that allocate the most total amount.
Also, on Linux, just because you don't have a swap partition doesn't mean things won't get swapped out to disk. Any memory mapped files will be swapped back to disk before issuing OOM problems - and the most common case of memory mapped files is executables and libraries. Which is of course the very thing that leads to a kind of livelock of swapless Linux systems that live close to 100% memory utilization: every time a thread is woken up, the instruction pipeline stalls until the next CPU instruction can be read from disk, slowing everything to a crawl.
By the kernel OOM killer? Is that supposed to be a trick question?
> In a situation where you are running out of memory, it is programs that allocate the most often that have the highest chance of seeing an OOM, not those that allocate the most total amount.
I don't really care; it still beats the machine being unresponsive for hours, especially if it's a remote one I dont have access to.
> Also, on Linux, just because you don't have a swap partition doesn't mean things won't get swapped out to disk. Any memory mapped files will be swapped back to disk before issuing OOM problems
That's not what is commonly understood by swapping though; that's a cache purge.
> every time a thread is woken up, the instruction pipeline stalls until the next CPU instruction can be read from disk, slowing everything to a crawl.
When you have swap enabled, yes; that the whole problem. However, IME, OoM just leads to the problematic program being quickly killed, then the user being able to go back to work.
I think this is the divide. In my experience, OOM killer rarely kills the culprit. It doesn't actually know which programs are behaving irregularly (that's the 'trick' in his 'trick question') and the OOM score heuristic isn't reliable in desktop usecases. If you know ahead of time which process will likely cause you trouble, you can adjust oom_score_adj for that process to make it the first target, but this is not great for common desktop use.
> That's not what is commonly understood by swapping though; that's a cache purge.
No, it's exactly a form of swapping. I'm not talking about purging Linux's IO caches.
I am talking about taking a memory resident page of your process and unmapping it, such that the next it is accessed it causes a page fault and the VMM loads it back from disk. When you have swap space, this can happen to any page. When you don't, this can still happen to read-only pages that already have a corresponding file on disk - which is usually your executable and any shared libraries you loaded.
Now, whether you'll really hit this situation or not really depends on how much of your memory is being taken up by the code of running processes. In many common server workloads, this should be very rare. However, in desktop workloads, where you might have lots of programs running at the same time, the total amount of memory being taken up by all their code may end up high enough to enter this state before calling the OOMKiller.
If you've got something leaking a lot and fast, 512MB swap is going to fill up in seconds, not minutes. If you have a slower leak, swap usage provides a clear signal something is wrong and may give you a chance to intervene before the OOM killer. Even in fast death, swap usage may be useful forensics, if you record it at a high enough time resolution.
This is exactly why I don't have swap for over a decade on my machines (since my machines have 16+ GB RAM). The mere 1/32 improvement is not worth the headaches of swap.
I actually have been forced to go back to desktop Linux again recently after almost 20 years of avoiding it. It's pretty disappointing to see that Linux's OOM handling behaviour is still completely broken (at least on desktop). Windows and Mac behave much more sanely when you run out of memory, gradually slowing down rather than suddenly freezing for like a minute while it figures out what to kill.
I did eventually find a Reddit thread that suggested enabling zram (compressed ram swap disk) which does appear to have helped but why isn't that the default? Also I now get a kernel panic about once a week but that could be unrelated.
You can still server 512 different clients at once and still utilize 100% of the CPU.
The parallelism is only a requirement if you want to compute one thing faster, not a bunch of things at once. Very much the desirable property for databases (although even that can be worked around with sharding if you're stubborn enough) but many developers code could be painfully serial yet still be perfectly fine to scale.
Lawyer enters the thread.
It's more common than you think. When complex work is done within power structures, it's just not enough to be "right" about something. You also have to be able to sell your solution to people who are wrong and powerful, which is sometimes not possible.
A lawyer from 10 years ago that memorised some good "rules of thumb" about generic matters such as liability or court procedure won't have to throw out everything they know to be relevant today or another decade into the future.
Meanwhile in IT, I regularly see network engineers upgrading links from 500 megabits to 1 gibabit when they should be looking at upgrading to 100 or even 200 gigabits.
Their mental model of what's fast, relevant, or desirable is "off" by two orders of magnitude. Their knowledge was accurate just a decade ago, but is hilariously wrong now.
I can't think of any other mainstream profession where institutional knowledge is deprecated as quickly.
It's not guaranteed to be worse. For latency sensitive OLTP workloads dedicated, smaller, disks for the journal can result in substantially lower & predictable latency. Random reads for the data exhausting iops, causing stalls of journal writes isn't fun. There's also per-disk queuing on a bunch of levels, and avoiding intermingling data and journal writes can be quite beneficial.
That's not to say it's always better, at all. But making it out to be a stupid thing one should never do isn't wise either.
I mean, you're clearly exaggerating. My ad-hoc Python scripts don't have to be parallel.
Promise? That's fantastic!
Because few things deserve to use more than 0.2% percent of that machine.
Sure, you could parallel it, and now it takes 0.3% to do what could easily be done quicker and saner by 0.1%.
I too like to anthropomorphize scripts and punish them for their transgressions by constraining their resources.
> I love these "rules of thumb", "black magic sorcery", and "best practices" that grey-haired old wizards pick up.
Just quoting this to point out what I mean by addressing your tone. It's fine to disagree, but this feels like you're setting up that everyone who disagrees with you must be condescending and stuck in the past. Yes, performance advice is often tied to the specific situation, but that's also true for the things you state.
> Worse still, these internalised tricks and rules are often related to Moore's law scaling, making them exponentially wrong over time. Not in the figurative sense, but a literal one.
Implying that Moore's Law makes anything exponentially faster is an anachronism (talk about old knowledge!), it only makes things exponentially more parallel nowadays. See next point why that's relevant.
As others have already clarified, -fomit-frame-pointer was sometimes 20% on x86. For some applications, people would kill for a 20% boost in single-threaded performance, even today!
And importantly, that this specific optimization has become useless on x64 says nothing about the value of optimizations in general. Different ones are still useful, to give an example let's say SIMD (where it applies).
> You see this with people claiming that I'm exaggerating when I say all code should be parallel. Meanwhile AMD is about to sell a 128 core / 256 thread processor. You can have two per host for a staggering 512 hardware threads. Single threaded code is capable of utilising just 0.2% of that machine!
I'm not sure which strawman "grey-haired" developer you have in mind that buys a 512-thread server to run single-threaded code on it and then needs you to explain to them that they should parallelize it. Maybe you've met that person. But in general: Doesn't most code live in web servers nowadays anyway? If so, it's either already running in parallel, or can be made to do so with very little effort.
BTW, parallelizing code in a web server can be harmful (!) to total throughput due to communication overheads. This is something that I find people don't always realize when they naively parallelize something, and is also sometimes missed by benchmarks that don't put the server under heavy load.
> A similar "tuning best practice" that I still see applied in the field is database servers with dedicated data, log, temp, and templog drives... on cloud VMs where this is guaranteed to have worse performance than simply pooling the same amount of disk into a single logical drive.
A statement like this depends on many concrete details about the specific cloud infrastructure (you didn't say AWS, you said "cloud"), and cloud VMs are usually not the best performing way to run a database anyway. So again, this is very context-specific.
Also, if database performance is critical enough that you consider tuning these things, then renting bare metal should at least be a consideration as well. People who have never measured it are sometimes surprised by how large the virtualization overhead is (and how cheap it is to rent physical servers). On actual hardware, this advice might still be relevant.
> I can't think of many (any?) other industries where more experienced senior staff can be this wrong about as many things...
The actual point is that you should measure things in concrete cases, not make general claims. In that sense, I unfortunately don't find your "updated" way to state what supposedly is and isn't relevant not much of an improvement to the "outdated" claims you argue against so vehemently.
So what? Who cares? I find the whiny tone of your post to be offensive.
Whether or not this argument holds water depends entirely on the context of the situation; what is the relationship between the program being considered and the computer? If the program is the reason for that hardware to exist, then obviously you want to use that hardware to it's fullest potential. If you've bought a 256 core machine specifically to run this program, then obviously you want to use all those cores. Some sort of threaded or multiprocess architecture should be an easy sell in this sort of circumstance. It shouldn't even be a discussion.
But on the other hand, what if the program under consideration is ancillary? In that case, the program should use the bare minimum resources necessary, so as to not get in the way of whatever the primary program may be. If you have say.. a music playing daemon, then odds are your program should not be trying to use the host machine to it's fullest because that hardware actually exists to get some other job done. Ancillary programs should be designed to survive on leftover scraps, not with the assumption that they'll have first dibs on the use of the hardware.
Even on a laptop, I now have 16 threads and almost never see more than 1 utilised even when I'm waiting.
Any time a human waits for a computer that is just 6% utilised, a developer screwed up.
Just this minute, I downloaded a program installer using a 10 Gbps Internet link.
Yes. Ten gigabits, dedicated to me, personally! Your idea of "what is a fast Internet link" is wrong.
I downloaded a 600 MB software package in less than a second, and then had to wait for 30 seconds while the anti-malware tools scanned it... with a single thread.
Extracting the file was 10 seconds because 'zip' is by default single threaded.
Installing it took a solid minute during which my CPU usage never peaked above 2 cores because the installer is also single threaded. The only reason it uses more than 1 core is because the anti-malware scans it... again... on a second thread.
Similarly, any time I see anyone manipulating ~1 GB of data, invariably they do it with GUI or CLI tools that are single threaded. This was desirable in an era when everyone had a single mechanical HHD. Multiple read or write threads would have caused disk head seeks, slowing down the process. These days when everyone has NVMe SSDs, the only way to fully utilise the 500K IOPS available from a $150 device is by throwing dozens of parallel threads at it!
Even as someone who tries to actively question assumptions that may no longer hold I still find myself constantly flat-footed.
I can't pretend like I know all the various advancements in processor architecture of the past 5 years (since 2018) or how modern compilers exploit them and also the latest techniques in GPU computing.
Heck even old rules of thumb like the speed of platter drives are being challenged - mechanical drives with 500M/s throughput are hitting the shelves this year. I just upgraded a platter RAID where I'm getting ~1000M/s. The SSD delta has changed (that's of course assuming SSD hasn't moved which I'm sure is also false). I'm sure there's new advancement in material science and manufacturing that I'm also not privy too which has made this possible.
Storage alone is another vast world. I don't know all the features of ext4, xfs, zfs, hammer, and btrfs or how to competently compare and tune them for various workloads or how that analysis has changed in say the past 24 months. This is something I became intimately aware of when upgrading the array. I didn't know any of the modern tools or approaches.
All knowledge in this space eventually becomes vestigial and outdated and it's physically impossible to keep up. I haven't even kept up with the latest C features. Then there's Go, Rust, Typescript and countless other languages that each require at minimum, a hundred hours a year each to competently follow.
The best we can do is have humility and continually approach things as a student because we never know what table leg has been knocked out underneath us without our noticing.
X86. No other 32b architecture is that register-starved even if the larger number of registers is somewhat mitigated by being load-store.
Nowadays X64 has a decent amounts of registers + the HW internally has several times more the registers, so saving one is almost meaningless.
Not only do you get more complete and correct backtraces with DWARF, you also get preservation of stack-saved registers and the ability to reconstitute a subset of local state in a given frame.
libunwind is MIT licensed, not GPL.
I’ve also written a full DWARF stack unwinder myself, including a full DWARF expression interpreter — it’s nowhere near 10k lines of code.
There's a related problem of compilers generating imperfect debuginfo that fails to reconstruct local variables from registers/memory due to an incomplete DWARF representation, but I don't think that plays a roll in unwinding. (But maybe something similar does.)
Unless DWARF can handle it fine and compilers are thoroughly broken, in which case it's a tossup between "DWARF is fundamentally flawed" and "compiler designers are lazy for some unknown reason"?
It _could_, but the kernel folks have made it pretty clear they don’t want to put any code related to dwarf in the kernel.
DWARF is useful for so much more than flame graphs. Flame graphs are cool and useful, but they're over hyped. Anybody who has ever spent any significant portion of their career debugging compiled binaries knows that proper debug information is infinitely more useful than flame graphs could ever be. Flame graphs are also useful... very useful. But that's my point: if they advocated even half as much for ensuring everybody made debug info more readily accessible by default as they did crying for frame pointers, the world would be a much better place.
Is DWARF too complex? There are many layers to DWARF, and there are compromises everybody could make so we can get at an optimal balance for which DWARF info we should ensure is always there. And we could also spend more time fixing tooling so it works properly with that minimal debug data.
AFAIK there are two issues at play here.
1. There's a bug in the Linux kernel where it sometimes return an invalid RBP register through `perf_event_open`.
2. The compiler doesn't generate the necessary DWARF info to support unwinding asynchronously, so depending on where exactly you land in a function you might not be able to unwind the stack. (This is why unwinding from *within* the program always works correctly but unwinding from *outside* the program with perf doesn't.)
Source: I wrote my own sampling profiler.
And there's of course the issue that you'd need everything to be compiled with it (because the program you're profiling is going to call into external dynamically linked libraries), which on an average system you most likely won't have.
But leaving that aside, for profiling, dwarf is just very expensive (perf, where the data volume explodes, due to copying stacks) or not available (bpf stacks), because the unwinding has to happen in the kernel. Even if the kernel had a dwarf unwinder, it's vastly more expensive to do that, compared to unwinding via frame pointers.
Try running a system wide profile on larger and busy box. Dwarf based call graphs are basically not usable. The profile quickly is ginormous, and viewing profiles is extremely slow.
I want both really badly but the fact is that flamegraphs drive performance optimizations which save companies millions of dollars, while symbols sometimes fix bugs and sometimes get blamed for making their software easy to reverse engineer. You can understand why one gets more focus than the other.
> The thing is that DWARF unwinding (the only userspace option that doesn’t use frame pointers) is really unreliable. In fact it has such serious problems that it’s not that usable at all.
I'm not that familiar with the problem space - can you elaborate on why DWARF is better? or speculate on why the author is experiencing (apparent) reliability issues with it?
> Compiling the kernel turned out to be 2.4% slower, and a Blender test case regressed by 2%. The worst case appears to be Python programs, which can see as much as a 10% performance hit. To many, these costs were seen as unacceptable.
Sacrificing that much for the absolute minority of people that need to dig that much into performance doesn't seem to be worth it for distro. Would be interesting to see why some cases still get that much regression tho.
That's quite a stretch, and I'm saying this as someone who probably believes in things others would consider conspiracy theories. Freedom is irrelevant when you already have the source and the binaries and can do whatever you want with them... and the latter applies even if you're analysing closed-source software.
All for a pointless ~1% performance boost.
I can rebut with "All for a pointless waste of a register for something that will only be useful the <1% of the time the code is actually run."
If I wanted to do the same with DWARF I'd need to link some GPL library or write 10,000 lines of byzantine code to figure out which function called the one that's crashing.
Or just learn how to read a memory dump. If your code is somehow crashing so often enough that it becomes such a burden, then I'd say the solution is elsewhere. It's not hard to walk the stack manually.
When I worked at Google, profiling was done so often that it certainly exceeded 1%. There's something called Google-Wide Profiling[1] but teams often opt into much more frequent profiling of important micro services they deploy. It's not just for identifying crashes, but more for performance optimization. It's liberating to debug performance bottlenecks when you can just pull up a flame graph from half an hour ago and compare that against a flame graph from one week ago without worrying about re-deploying the old version to get profiling data.
Btw I think this is somewhat of a cultural bias depending on your previous experience. Some people simply value performance at all costs. Some people value more observability and debugability.
So a 10% use case for 0.001% of engineers out there. And I'm pretty sure you didn't run stock Fedora distro on servers either ?
If you're that high end that profiling code for few % gains is worth millions you can just compile the stuff with whatever flags and optimization you need instead of relying on distro to reduce performance for everyone to appease to few engineers at big corp
I've read more about this inconsistency, and it turned out that at some point libc++ changed something in the string class and it caused a 1% performance regression in "an important google-wide benchmark". That 1% of performance was so important, that since that up to at least few years ago every C++ programmer had to put up with that. I guess it makes some sense - 1% of CPU cost across the whole Google fleet is a lot of saved CPU hours (and that equals money).
This is a skill that is not only hard to learn, but also takes time to apply properly. Yes, I can read a memory dump. But I'd sacrifice more than 1% of performance to not do it manually. And the infra cost to account for that is also lower than the value of my time.
Telling basically all devs without such experience to "just" learn an advanced skill is not a good solution.
That may be true, but having proper stack traces will probably cumulate to optimizations that gain much more than the 1% loss of enabling frame pointers [1].
Ideally, it would be much easier to rebuild whole Linux systems. Then people who need it could enable frame pointers. Sadly, rebuilding a mainstream Linux distribution is much harder than rebuilding e.g. Gentoo or NixOS. Though, in practice even rebuilding a NixOS system is fraught with issues. I have done large rebuilds, but it would often fail because e.g. unit tests have issues that only manifest themselves with enough build parallelism.
[1] Note that on x86_32, the performance improvement was larger due to register pressure.
If a binary really wanted to hide its behavior (at least to the same extent that it could with -fomit-frame-pointer), it could just make fake stack frames beyond the current PC.
Do they? It seems like everyone is moving away from scripting languages for new projects these days, at least where web browser-related requirements doesn't push them into using Javascript.
Indeed, before this current pendulum swing, scripting languages had a good run. I'm not convinced it was because of stack traces, though. They won hearts and minds on promises of no longer having to do "XML sit-ups" and whatnot.
-fomit-frame-pointer is pretty high on the list of most destructive micro-optimizations in the history of open source. When you repurpose the backtrace pointer (RBP) for other things, it totally savages our ability to have native backtraces
Funny. I always learned that the *BP registers stood for Base Pointer and not Backtrace. And then I would use BP to reference function temp variables from that part of the call frame.The Intel “ASM86 Language Reference Manual” calls it that as well.
https://community.intel.com/cipcp26785/attachments/cipcp2678...
I mean, presumably you just answered your own question: while corporate software might not be able to link GPL code, Linux software — that mostly already is GPLed — certainly can. In fact, that's presumably why the library you're referencing (libdwarf?) is under the GPL in the first place — it was created by a Linux developer, for use in their own and others' FOSS Linux software, where it easily solves this problem for them. And because they have this easy solution to this problem on hand, they don't really see the need to turn on frame pointers in their projects' Makefiles.
Don't tell the golang folks!
Besides the inbuilt panic behavior which prints a very helpful backtrace, non-fatal error handling is spinkled all over with fmt.Sprintf "I am here at 134738djfjrsj so you find me in the logs"
Don't confuse problems in Go language with your own ineptitude to use it.
This results in a net saving of money over time.