54-line if condition in gcc's reload.c
github.com
github.com
It essentially took the place of spill placement/code/legalization. Over the years, it grew rematerialization, instruction combination, copy coalescing, stack slot sharing and all sorts of other interesting scope creep.
While i believe it has finally been replaced by LRA for some targets, there were many people who spent years of their life trying to replace reload with separate, smaller pieces of architecture. No one has yet completely succeeded.
It is essentially an interesting object lesson in what happens when you just incrementally improve architecture to achieve performance goals without a stop-loss point for requiring new design.
Spill code generation is necessary because at a program point, there may not be enough physical registers for all the live values, and code must be generated to spill some of them to the stack and reload them when needed. Legalization performs code transformations to eliminate target independent operations in the IR that can't be represented on the target. Copy coalescing eliminates the need for copies between virtual registers by assigning both to the same physical register. Rematerialization takes advantage of values that can be recomputed cheaply, and instead of holding them in a register for a long time, recomputes them where needed.
What makes this all so complex is that 1) most of these don't have polynomial time algorithms that produce optimal solutions: 2) the individual problems are mostly coupled. E.g. spill code itself uses registers, and so the interference graph (which summaries whether virtual registers are simultaneously live) might need to be recomputed or modified.
If you're interested in this sort of thing, Kieth Cooper at Rice has posted the lecture notes for his graduate compilers class online: http://www.cs.rice.edu/~keith/512/Lectures. Register allocation is lectures 26-7.
1) The goals of Reload are very complex.
http://gcc.gnu.org/wiki/reload
"Reload does everything, and probably no one exactly knows how much that is. But to give you some idea:
Spill code generation
Instruction/register constraint validation
Constant pool building
Turning non-strict RTL (Register Transfer Language, a very low level intermediate representation used in the backends of GCC) into strict RTL (doing more of the above in evil ways).
Register elimination--changing frame pointer references to stack pointer references
Reload inheritance--essentially a builtin CSE (Common Subexpression Elimination) pass on spill code"
Reload achieved them for the last 25 years(!)
2) There are more modern approaches to reach such goals, but knowing 1) it is a lot of work before the goals can be achieved by some alternative code for all platforms (I don't know how far the developers got at the moment)
3) "Local Register Allocator Project" presentation by Vladimir Makarov, working for RedHat:
http://gcc.gnu.org/wiki/cauldron2012?action=AttachFile&do=ge...
https://github.com/mirrors/gcc/blame/7057506456ba18f080679b2...
The copyright notice goes back to '87. I wonder if '92 is just where the source control kicks in.
DannyBee claims that reload was started in the late 80s.
There are commits going back to 1988, but it looks like a lot of the early commits are out of order; the repository was (at least) converted from RCS to CVS to SVN to Git.
And yes, I used RCS when working on the gcc source.
We used 860 (not sure) and 960 (definitely) too for other products.
EDIT: Apparently all previous history was lost when they implemented CVS, circa 1997. We could always ask someone in the oldest maintainers file available (https://gcc.gnu.org/viewcvs/gcc/trunk/MAINTAINERS?revision=1...) and have the real answer.
Originally? None.
Then RCS
Then it forked into EGCS and GCC, and EGCS used CVS
On remerge, we used CVS.
Then we converted to SVN.
I basically rewrote large parts of cvs2svn to make this happen (before that it took weeks to convert and ran out of memory anyway :P)
During cvs2svn conversion, the old GCC RCS versions we had data for were inserted as branches and older revisions, as appropriate
Since the original per-file version numbers were not kept, this was done by inserting it and then incrementing existing version numbers on the RCS files that made up the CVS repository, and hacking up the cvs2svn parser slightly to allow for some idiosyncracies (because now you have a billion years worth of RCS bugs to parse)
So basically, history goes back as far as there were version control systems that were being used, which is roughly 1987
Before someone asks why it got moved to SVN, at the point at which we converted to SVN,
1. git was not popular yet (or all that usable yet for that matter)
2. GCC's biggest VC problem was needing single repository version numbers for tags, and atomic commits, which SVN solved.
SVN was a huge step forward compared to CVS.
Although note that a lot of earlyish GNU codebases kept tons of older versions of source files around in the form of Emacs versioned backup files, and there were RCS-conversion scripts that would commit all of those as RCS versions...
[There would be no log messages for those older versions, of course, but the writers and creation dates of the older files were potentially available... I don't really remember whether those conversion scripts actually used that info or not, although I think they did...]
Now I feel old.
Great. Now I feel a million years old...
And that's why this will be here forever.
I.e., this style (used in this case)
&& (CONSTANT_P (SUBREG_REG (in))
|| GET_CODE (SUBREG_REG (in)) == PLUS
|| strict_low
|| (((REG_P (SUBREG_REG (in))
versus this style: (CONSTANT_P (SUBREG_REG (in)) ||
GET_CODE (SUBREG_REG (in)) == PLUS ||
strict_low ||
(((REG_P (SUBREG_REG (in)) &&
I don't have a personal preference here, just looking for any practical pros and cons I may not be aware of.More on topic, I don't see why keep such a big conditional and not move it out to its own function(s) (but this has been asked already in this thread).
SELECT
some_col
, another_col
--, and_another
, and_more
FROM blah
EDIT: too bad formatting is screwed up :/In fact, in a language like Javascript where extra trailing commas are allowed, it seems that this argument makes even more sense for commas at the end than it does for commas at the beginning. Then, there would never be a situation where you would have to edit a line besides the one you were commenting out.
Usually I only put them at the end of a line in languages that do implicit line endings, and will false-positive when you move the operator to the next line (i.e. javascript and visual basic).
http://blog.izs.me/post/2353458699/an-open-letter-to-javascr...
I would say that the most common problem is having a semicolon not being inserted if you start a line with `(` or `[`. In practice, the only time when a semicolon gets inserted where it shouldn't is when returning an object literal.
...
== NO_REGS))
#ifdef CANNOT_CHANGE_MODE_CLASS
|| (REG_P (SUBREG_REG (in))
&& REGNO (SUBREG_REG (in)) < FIRST_PSEUDO_REGISTER
&& REG_CANNOT_CHANGE_MODE_P
(REGNO (SUBREG_REG (in)), GET_MODE (SUBREG_REG (in)), inmode))
#endif
))That || needs to be inside the #ifdef, so it can't follow on the same line as the "== NO_REGS))".
And once you're doing that, might as well use them for the whole file.
When you split an expression into multiple lines, split it before an operator, not after one. Here is the right way:
if (foo_this_is_long && bar > win (x, y, z)
&& remaining_condition)
1) https://www.gnu.org/prep/standards/html_node/Formatting.htmlconstant...
or get code...
or strict low
instead of
constant... or
get code... or
strict low
<pre><code>
$moo =
[
a
, b
, d
];// insert c
</pre></code>
better than having the delimiter at the end, where you a) have to take care where you relocate the last element and b) could forget to remove it at the last element (most language parsers tolerate that though).
But, to be fair it is well commented.
x && y || !x && z
In this absence of side-effects, this basically implements a 2-input multiplexer and is identical to x ? y : z
In the 54-line condition the first obvious thing I'd factor out is SUBREG_REG(in) and GET_MODE(SUBREG_REG(in)), and then work out what else is duplicated from there. Here's my attempt at making this a little more readable. It's only 2 lines less, but this gets rid of all the repeated uppercase:The same goes for "GET_MODE_PRECISION(msrirp)" which was replaced with "msrirp."
I am frankly embarrassed by the number of bugs that have been tracked to this function.
This was in the day before register coloring.
Earlier, I had the chance to write a code generator for an implementation language targeting the 8085 (!) and even that was sufficiently hairy that the team that took it over complained about the difficulty of improving the code.
[Edit] Time sequence correction.
If they'd spent all their time needlessly fixing what ain't broke, gcc would've gone the way of GNU/Hurd.
If that corner case manifests as a bug in OpenSSL, it is by definition at least as bad as a bug in OpenSSL.
I'm not convinced that's true.
I don't trust anyone - myself included - writing code 1/10th as convoluted, age-of-product be damned. Code is not wine, old doesn't mean good. Neither do the maintainers, methinks: Looking at blame shows some refactoring.
> If they'd spent all their time needlessly refactoring, gcc would've gone the way of GNU/Hurd.
And the lack of needful refactoring may very well send it the way of COBOL - with everyone merely wishing it had gone the way of GNU/Hurd. I've seen clang and LLVM replace gcc in both of my vendor toolchains that used gcc - suggesting they've already wished and then done something about that wish.
(Either that or licensing, but code like this stomps on the scales a bit...)
It's also got prototypes that made it into production, ball of mud designs that encourage usage bugs, and C++ thrown together by that short lived intern who took a few Java classes, wrote everything assuming there was a garbage collector around, and took great care to avoid class trees with less than 3 layers of inheritance lest he be shamed for lack of 1337ness... then topped it off with a few __try/__catch blocks to deal with that one rare crash that nobody could find a repro case for.
Re-implementing bug fixes is a toll... sometimes a very worthwhile one, though.
Edit: also reading the wiki, they are fully aware how bad this code is.
LRA is by all appearances infinitely more maintainable.
Any time this if statement has been edited, it was clearly faster to simply add to the existing one than to refactor the whole thing. Sometimes it makes sense to take that extra time to refactor code when you happen to be working on it anyway... and sometimes there's something else more pressing to do.
(Amusing incongruity: "boolean" fails the spellchecker in an IT forum :o)
C doesn't support native booleans, C99 does-ish. And you can hack it in with
#typedef enum {false, true} bool;
C uses ints, all the way down, for everything, until you hit turtles.In a lot of software in the industry I've seen, I'd say he's right -- instead of an "if X condition, do X code, if U condition, do Y, otherwise do Z" you can make an object, subclass the object's base class into a default type doing Z and two other types that do X and Y instead. (If you need to share more behaviors in a more complicated setup, you make the strategies into classes as well, then have the X1 and X2 classes invoke the X strategy.) Happiness ensues! (unless you're stuck with too much ugly-looking Java boilerplate or insane C++ templatization angle-brackets)
That said...
any good advice like that should be taken with a MASSIVE grain of salt because of the exigencies of real software development, especially software like gcc. Moreover the approach is not a panacea against code complexity because sometimes the relationships between different types of conditions are just plain complicated, as appears to be the case here.
So I can totally understand where this code is coming from, and while I can imagine it being far more elegant, it's not really the prime candidate for a rant. Grandparent post should chill out.