You can’t beat a good compiler…. right?
lbrandy.com
lbrandy.com
I emailed the code to the guy who originally wrote it and his response was brief, "The compiler wasn't that smart before". But yet the code sat there for like 10 years, untouched -- largely because people don't like to touch x86 assembly code that is working.
The moral, you may get a benefit today, but you don't get any future benefits of the investment at MS, Intel, the Gcc folks, etc..., who are improving the compilers. Sometimes it is the right call to hand optimize. But when you do, I suggest having a C/C++ version alongside of it that one can test aganist it whenever the compiler revs.
I was involved in writing (Fortran) code for Crays in the 90s. For the T3D, we speeded up the code with the usual set of tricks -- loop unrolling, array padding and so on. With the later T3E, it turned out that not only could the compiler now do all that, but that code hand-optimised in this way performed worse.
Or in general - if you are considering using a tool or rolling your own, what do you know about your data that the tool author doesn't?
So in the context of real-world problems, compilers win, because none of us has the energy or attention span to write it in assembler, something the size of a real-world problem like, for example, a compiler.
We optimize small problems by hand, but problems of useful size we don't because the compiler does a better job on problems of true interest than we have the patience or time to do.
To put it another way, it is better to put energy into improving compilers than getting into a John Henry type of competition.
My experience of problems requiring massive hardware or clusters (limited, I admit, to certain classes of scientific computing) is that such assembler coding of speed-critical components is actually often done.
Rather than trying to carpet-bomb everywhere with micro-optimization hand-grenades, use laser guided precision to hit the targets that count. So instead of wasting your time manually unrolling loops in code that's just going to be refactored away before release anyway, find out exactly where your code is bogging down and concentrate on fixing that.
float temp_out = 0.0f;
for (j = 0;j<5;j++)
temp_out += filter[j] * in[i+j];
out[i] = temp_out;
solve the performance issue, and allow for out = in? (which I would think would be desirable, in cases where you don't need the original input).Also, this is way outside my area of expertise, but when I hear convolution I think FFT. I would be curious if anyone knows whether there is an algorithmic improvement to be made here.
Your second comment (about out == in), is a bit trickier. This is obviously a contrived example. In a normal gaussian filter, you'd want the filter centered so that the effective indexes of the filter are -2 .. 2, not 0 .. 5. This makes (out == in) a trickier problem to solve. I didn't really want to get into those details in the post.
An FFT wouldn't be faster in this case since the filter only has five taps. Generally there has to be a decent sized kernel for the FFT to be worth it -- the FFT is O(N log N), and this algorithm is only 5*N operations.
It was interesting to see what the compiler generated, though it won't be a big surprise to experienced C coders.
It's important to note that I never said this.
I was simply trying to separate two versions of the same rule of thumb... I said that using a compiler, a smart low-level programmer can generate the assembly code he wants. This is not "hand-made" assembly, but this is not just the magic of an optimizing compiler, either.
So, when you and the compiler are booth looking at the same code base, but you find a better local approach based on a global condition you have beaten the compiler. Even if all you did was tell it about that condition.
I'm a big fan of looking at the assembler code generated from programs - even if you are not optimising - it helps give a depth of understanding which is simply not possible without doing it.
Not to mention that the "smart" compiler is a myth - I've seen some downright stupid asm come out when the compiler is configured wrong - and it gets missed, even with profiling based optimisation targetting the area - because looking at assembler is actively discouraged.
We are spoiled with IDEs and high level languages...
EDIT: note that the off-by-one error is MY mistake - it simply isn't there. My point about "restrict" is valid, though.
I'm not seeing an off-by-one error. 'i' will be LENGTH-5 on the last iteration and we will add 'j==4' for in[LENGTH-1] not in[LENGTH].
Especially on DSPs and microcontrollers you may end up changing instruction order to avoid pipeline stalls or use on-chip RAM buffers for your often used data.
Also, if you start writing a routine in assembly from scratch it's often surprising how many different ways you can implement an algorithm since you're not as bound by the calling conventions and runtime assumptions made by the compiler.
Not that I would recommend that for anything other than the innermost loops.
By the way, are we talking about anything other than C or maybe Fortran here? I'm guessing Java or C# coders don't spend too much time looking at the IL, I know I've only done it for educational purposes.
I need to plug some holes in my skills, and this is one of them.
You could also try getting an embedded development kit, like the USB Bitwhacker, PICKit, or an AVR-based kit, and programming a microcontroller in assembly language.
It's kind of a conversational approach to learning Assembly.