On the Madness of Optimizing Compilers
prog21.dadgum.com
prog21.dadgum.com
I don't know if there's a name for it, but I've started reading "Don't do X" as "Don't do X unless you have a realllllly good reason to." Nothing in life is absolute, after all.
Similar advice is "Use a GC". There are (similar) use cases where you shouldn't use a GC. But if you don't have a strong case for paying the price of not using a GC (in terms of potential bugs and code complexity), then using a GC by default is pretty good advice.
But yeah, if you're writing VMWare, you're hopefully also skilled enough in the domain to know that you should try to get maximum performance. But if you're writing some shell scripts or webapp, a lot of the time your bit-twiddling optimisations will be noise relative to the variation from I/O alone.
Saying that C++ categorically cannot be used to create good machine code is ignoring much evidence to the contrary.
Not false. He seems sometimes contradictory.
> The golden rule of programming has always been that clarity and correctness matter much more than the utmost speed. Very few people will argue with that. And yet do we really believe it? If we did, then 99% of all programs would be written in something like Python. Or Erlang. Even traditional disclaimers such as "except for video games, which need to stay close to the machine level" usually don't hold water any more.
He claims that you can usually choose any language you want to develop your software and that it will not make any visible difference. Systems are so fast nowadays that the trying to optimize for speed will need to an effort that is probably not justified.
I don't really get the point the article is trying to get across anyway. I don't care about the complexity of the compiler at all, and 99.99999% of the time I don't care about what the assembly it spews out looks like. I'm just using the compiler as a tool, I don't need or want to know how difficult it was to get it to compile my code. The only thing I care about is that it produces valid code, and that the code gets a speedup when I compile with optimization enabled. That's about the extent of my interest in the compiler, and modern compilers have served me very well over the years. Why complain that they need to be complex to produce faster code?
Maybe because the resources could be invested in something (like Erlang).
I share similar views but for another subject (web browsers).
BTW I like a lot his blog.
This mentality is quite common in certain programming circles.
Or are you trying to dictate what compiler writers should be focusing on? They're free to choose what to work on, so don't use their compiler if you disagree with it.
If we are talking about C and C++, the binary doesn't have to work according to the UB rules.
> Or are you trying to dictate what compiler writers should be focusing on? They're free to choose what to work on, so don't use their compiler if you disagree with it
Hence why I never code in C unless there isn't any option, and avoid C style code as much as possible when using C++.
I bet very few C developers, less than 1% outside ISO tables and compiler vendor teams, know ANSI C from cover-to-cover and are able to write UB free code when one sticks to the letter of the standard.
For any other language except C(Haskell, Java, C++, C#, Go, etc.), I don't actually care that much about compiler performance as long as it does the obvious optimizations. Generally you can look at performance at a higher semantic level as soon as your code starts implying calls to malloc() - and conversely, you can start looking at the assembly as soon as you've removed all calls to malloc() in the offending section.
The other thing that transcends compilers or languages are the actual bits that get stored in memory or on disk. Even if I'm writing Perl, it makes a difference if I'm packing something into a contiguous array or just packing a bunch of pointers or references to all over my heap. Or if I can shave a couple bytes off every record in a hundred million records.
Javascript might be the one exception, since people using it don't have any other choice. I don't know - is there such a thing as an FFI for Javascript, using a browser extension or something?
This is a good approach, as are finer-grained optimisation pragmas & attributes, #ifdeffed as needed.
Ideally, I’d like the compiler to do only cursory optimisations by default, just enough that I can read the assembly and not say “wow, this is stupid”.
For anything else, I’d rather have a dialogue with the profiler & compiler heuristics, wherein they suggest optimisations, but I choose which ones I want using optimisation pragmas, and the compiler does the grunt work of program transformation.
For example, the compiler would say “this function is only ever called in cold paths; I suggest adding the ‘cold’ attribute” or “this loop can be vectorized; I suggest adding the ‘vectorize’ attribute”.
This is a strawman. The usual argument is that 1% better performance is, well, 1% better. That's a whole server if you have 100 servers. People are not desperate, but do want an additional server for free.
> Assuming that's true, then they should be crafting assembly code by hand.
This, and "a few cycles inside a loop", is another strawman. What usually happens is a good compiler saves a bit everywhere, which helps when the program is optimized to flat profile without hotspot. Since saving is everywhere, "crafting assembly" amounts to lots and lots of assembly, not just a few in strategic places.
Take game development for instance, if your code runs a bit faster that's more frames you're rendering per second. This can mean the difference between a laggy, unplayable mess and a good gaming experience on an average machine (a laptop, or a non-gaming rig).
Do we know that most new software runs in batteries? Do you have a reference for that? It would be useful.
Other question is - does a sleeping CPU always win, or are other optimisations such as memory usage, screen usage, and communications, more important? I know for sure that the CPU is not the main consumer of power on a bunch of Android phones that were used in an academic study.
Probably the CPU is not the main power sink on mobile devices (screen and radio are I think). But cranking up the optimization level is free for the developer.
It's not so simple. If you're trading off simplicity in the compiler for that speed,for example, it may be a net loss.
If speed was the only concern, we'd all go out and buy super computers.
You can optimise as much as you like but then John Q Madbugger comes along with their insane Javascript (not looking at Flickr here, no sir) and suddenly your performance is in the toilet.
Thinking over it some more, browsers themselves were probably not a good piece of software to call out on performance considering how much optimization is actually done to increase their performance. The horrid things people do on top of them deserve all the blame they get though, and then some.
I do however agree with underlying premise of this post: Code that works 100% of the time is better than higher performant code that explodes in your face every once in a while.
No they don't. Code size is rarely a problem (embedded probably cares..). Optimizations will in most cases lead to larger code, loop unrolling, inline expansion etc. In c the compiler cant't, in most cases do anything about memory usage. (Probably could pack structs.. , but that is not something the compiler would do I think, you would break ABI with non optimized code). In managed languages, memory usages is probably more depended on the runtime system GC, then compilers.
GCC has a flag for code size optimization -Os https://gcc.gnu.org/onlinedocs/gcc/Optimize-Options.html
"With -O, the compiler tries to reduce code size and execution time"
So actually yes, The compiler optimizes for both size AND speed. Sometimes you can optimize for maximum speed (-O3), and sometimes optimize for minimum code size (-Os). You can even optimize for best debugging or best compilation times, and forget about size/speed.
As for memory usage -
1. First of all - code uses up memory. smaller programs = less memory used for holding the exectuable code.
2. Consider the various ways you could compile a switch(-case) statement. You could build it as a set of conditional branches one after the other, you could compile it into a binary tree search (or a hash dictionary lookup), and you could jump to an address found in a lookup array table.
This former method is also called a "branch table". https://en.wikipedia.org/wiki/Branch_table
A branch table would be the fastest method for implementing a switch with a few sequential values (0,1,2,3,4,5) But would be a really bad idea for sparse values (0,200,5000,10000,200000).
Choosing not to use a branch table for sparse values is an optimization for Memory usage. And compilers make this decision all the time.
That depends. In game programming, if I had a choice between a render that will crash once every 10 hours and produce 60 fps (in some model envrionment) vs a render that never crashes but produces 45 fps, I would choose the first one.
Hell, I'd take a 1% chance of blowup for a 10% speedup any day.
Are you serious? Would you base your research career on software that you don't trust? I wouldn't.
Please, notice that 1% chance of full blowup means a not-null probability of nice-looking-but-actually-false results from time to time, and that's a risk no researcher is allowed to accept.
I guess you haven't seen the quality of academic code recently. The choices are basing your credibility on slow badly written code or fast badly written code.
Renting an extra server costs practically nothing comparing to wage costs for a tech company.
So I'd say they'd care more about the compiler being reliable and predictable.
I have a team that regularly generates savings of XX% of the entire datacenter CPU cycles of a very large fleet through better compiler optimization (that pretty much never breaks any of our programs. In fact, we break them more through almost anything other than better optimization. Shocking, i know, but you can test this stuff really well it turns out, and write sanitizers to find behavior you wouldn't discover at compile time and ...)
I know at least another 100 large companies where the cost savings of doing this kind of thing also totals in the tens to hundreds of millions, per year.
and that's just the normal X% a year, not the "we put a team of experts targeting your use case and they got you 100x" type of savings.
So yeah, there are people for whom this matters ... a lot, and dismissing them casually is not a great idea :)
Truthfully, IME, compiler-optimization induced bugs are so infrequent compared to normal you-wrote-really-wrong code bugs that IMHO it's missing the forest for the trees. (now, maybe people suck more at tracking them down, but ...)
Heck, I had once screwed up some code in GCC that caused it to essentially remove random stores and loads from the program being compiled. It only caused one visible failure across a bunch of real world apps that had serious validation suites.
From reading hacker news, you'd get the idea compiler optimizations follows your code down alleyways to mug it. In truth, people are so busy overflowing buffers and whatnot that we are lucky if we can break their code due to optimizing.
Performance gains of 1% just don't matter to most programmers.
This guy, what a fucking joke. Yeah to avoid bugs instead of using these fancy compilers (which rarely if ever are the cause of bugs) we should be assembling our programs by hand.
At some point, for extremely hot loops, there may be no choice but to write a few pages of assembly by hand. If this saves tens of thousands of lines of compiler code that doesn't generalize, and is used on thousands or millions of machines, it's worth it, IMO.
Citation needed. Pulling numbers out of your ass doesn't an argument make.
And with that philosophy, I believe it was within 2-3x of C++ in most benchmarks, despite being in a higher level language (though no doubt it's much slower on some benchmarks).
That said, they continue to optimize the compilers (e.g. https://news.ycombinator.com/item?id=9099744). People obviously do want speed; 2x slower isn't acceptable in many cases.
"The results of our experiment suggest that Proebsting’s Law is probably true. The reality is somewhat grimmer than Proebsting initially supposed."
"Proebsting has suggested that the compiler and programming language research communities should focus less on optimization and more on programmer productivity"
Perhaps languages could be designed so as to make this task easier?
In language design we end up torn between making things easier for the programmer and making things easier for the compiler-designed-for-performance. Doing both things well is perhaps something that language design will be able to improve upon in the future.
https://news.ycombinator.com/item?id=11092067
It's great reading. thanks.
> The programmer using such a system will write his beautifully-structured, but possibly inefficient, program P; then he will interactively specify transformations that make it efficient.
[2] https://people.csail.mit.edu/jrk/jrkthesis.pdf "Decoupling Algorithms from the Organization of Computation for High Performance Image Processing"
"inline" and "register" are two very known examples. Other examples are function attributes in gcc: https://gcc.gnu.org/onlinedocs/gcc/Function-Attributes.html
And most current wisdom is that "inline" no longer serves as a useful performance suggestion to the compiler, but instead should only be used as a linkage directive to allow multiple declarations of the same function: http://stackoverflow.com/questions/29796264/is-there-still-a...
But those are still 2 well known features that were added intentionally to help improve performance.
Compiler bugs are pretty rare, a bigger problem are insane language standards that have undefined behaviour lingering behind every corner.
For example, GPGPU compilers are awful:
http://www.doc.ic.ac.uk/~afd/homepages/papers/pdfs/2015/PLDI...
This is one of the reasons gcc/llvm won: they don't care about that stuff. I mean, they care about performance, but not at all costs like these guys care.
At some point, people wanted compilers that did less stupid things. So they moved to the vendors who did that.
The reason it's so awful in the GPGPU space is because none of these vendors produce compilers yet. Instead, the vendors you have are those whose main goal in life is to generate better performance numbers than the other guy.
Ask anyone about the old days of fortran compilers and early optimizing C compilers and what have you.
What exists now is essentially "what happened after the market pushed back". Sadly, it didn't go as far as some people want (in either direction).
Really? http://www.complang.tuwien.ac.at/kps2015/proceedings/KPS_201... https://news.ycombinator.com/item?id=11219874
While GCC and LLVM do care about implementing the standard correctly, they show little mercy to the poor developer who failed to avoid one of the many sharp corners of undefined behaviour.
There was a long thread about this just a few days ago (https://news.ycombinator.com/item?id=11219874)
I already commented there, and since i didn't always have pleasant things to say, i'm not going to repeat them here.
[Indeed,] There's more to it than bugs.
Undefined behaviour for instance is generally understood by the programmer as being runtime driven craziness. Divide by zero for instance will raise a clean error (generally resulting in a clean crash) on most platform.
Popular C compilers such as GCC and LLVM however, have a different understanding, based on the wording of the standard. They are allowed to assume that undefined behaviour simply doesn't happen, and use that assumptions in optimizations such as dead code elimination or alias analysis. With this stuff, a program that would have exhibited undefined behaviour if compiled straightforwardly is allowed to go bananas at any time. This is compile time driven craziness, a rather different beast.
A classic example of this is signed integer overflow. You perform some arithmetic that may overflow, then check for overflow after the fact. Most architectures perform two's complement modulo arithmetic, so it works without problem. Except that in C, this is undefined behaviour. So your compiler is allowed to assume it never happens. Which means your test that `x >= x+1` is going to be "true" all the time. Which means that your error handling branch is dead code. And that could allow nasty stuff such as buffer overflows.
You could say it's the user's fault for relying on undefined behaviour (that is nevertheless well defined on x86 and ARM assembly). But you gotta admit that's not very intuitive. Yes, in some obscure platforms such an overflow could make your CPU jump in a random place. Yes, relying on this is not portable. But C is supposed to be close to the metal. Few C users would expect the compiler to optimize it all away even if it would work on their platform.
On the other hand, things like strict aliasing are explicitly about enabling compiler optimizations.
There's also the time I caused a kernel panic on CentOS 5 by running g++ on my code. It did so reliably, so I moved development to CentOS 6.
Older gdbs on CentOS 4 will segfault if you run it on a binary compiled with ghc and ask for a backtrace.
One time, I wrote a Perl module that unintentionally trashed the Perl interpreter's error tracking information. So if the module compiled, it worked fine. But if you had even a single variable mis-spelled or missing semicolon, all you'd get from perl is the error "Unknown error".
Or even something as simple as writing a negative integer literal bugs out ghc: https://ghc.haskell.org/trac/ghc/ticket/9533
Compilers are absolutely riddled with bugs.
It's all well defined. it's a social problem. it's something to be aware of.
The same can be said of any language that lacks a rigorous definition; misusing -Ofast is just an example of one. Even then, the definition can be too large or tricky for anybody to truly understand, but that's a bit of a grey area... at some point you have to draw the line and say the programmer made a mistake. The only sensible thing to do is draw the line at the definition of the language, otherwise there is chaos.
Some of us actually do need them. Soft realtime programs (games as an example) are perfect example where they are crucial. Reaching same level of performance using non optimizing compiler would balloon the development costs up multiple times.
I would understand the complaint if compilers would mess things up and you couldn't get around it.
Well, I work in the games industry and that "last bit of performance" is quite frequently the difference between having a 60fps game and something that doesn't quite hit the vsync and looks a lot worse. Seriously, this is one of the very few industries where I can say with a straight face "don't do context switching because the 5 microsecond delay is too much" and "a missed L2 cache hit costs us 50 cycles, and it a major problem here".
(a) free compilers
(b) the C standard
Crazy talk, surely! Let me explain: each of these is a Good Thing™, but together they have put us in our current predicament. It used to be that compilers were reasonably expensive products, so as a vendor you had a clear incentive not to piss off your customers. In addition (in the very olden days), there as no C standard so the definition of "correct" was very customer-based: if customer code compiled well, you were doing a good job, if not, you would probably suffer.
With compilers becoming free, the incentive structure changed, customers now matter very little. Instead, you get ahead in that pecking order by improving some SPEC benchmark by 0.2%. At the same time, the C standard, which was worded very inclusively so even compilers for odd architectures could be ANSI/ISO compliant, was used as a substitute for "my customers' code works", and people started to say silly things like "if you have undefined behaviour, demons can fly out your nose".
Well: I had a pre-ANSI compiler. ALL its behaviour was "undefined". And yet it took fewer liberties with the code than many of today's compilers do. Because the creators (a) were reasonable people and (b) had real customers.