Don't use assembly unless you're an expert
quetzalcoatal.blogspot.co.uk
quetzalcoatal.blogspot.co.uk
The x86 these days is effectively a poor-man's VLIW machine: you have a number of units with out-of-order execution. This means that you can look at your instruction stream as a stream of VLIW instructions for all of the CPUs units. The performance of your code might depend not just on the exact number of instructions, their execution times, their latencies, but also on the interdependencies between your instructions and scheduling constraints. Getting that right is a nightmare, and even if you do get it right on your chip, there are at least several major x86 architectures in use, each one with its own peculiarities.
As a counterexample, I recently wrote some ARM (Thumb-2, for the Cortex-M4) code. It wasn't very difficult, and the code I wrote was much better than what gcc could produce.
Better in terms of code size, speed, or both? (Hm, does code size even matter anymore? Maybe for the CPU cache.)
Cortex-M4 implies microcontroller. Quickly looking, the cheapest M4 chips have 32KB of flash for program memory.
With the move to 64-bit, code had to be recompiled anyways; hence there was a good opportunity to reshape the x86 instruction set significantly. Instead, the number of changes were relatively small; the instruction set maintained his flavour (variable length encoding, a non orthogonal operands, ...). I believe there are many reasons for that choice (made by AMD), but code density was an key factor.
64K of L1 might seem like a huge amount for a function to fit into, but that's 64K shared among every other often-used piece of code. It's for this reason that loop unrolling is no longer recommended too.
The first one is alignment of branches: the assembly code contains no instructions to align basic blocks on particular branches, whereas gcc happily emits these for some basic blocks. I mention this first as it is mere conjecture; I never made an attempt to measure the effects for myself.
And if you were, you would likely not see much on a modern x86; in fact it would probably make things slower, as it may confuse the branch predictor --- instructions are variable-length, so their addresses naturally are not aligned, and that enables e.g. lower-order bits to be used as cache tags.
And code is much smaller than data usually, it's a small code snippet operating on a lot of data, not the other way usually
I write assembly for DSPs on a near-daily basis. Up until a few months ago there didn'nt even exist a C compiler for the target architecture.
Even when you write C code for an embedded platform that does have a decent C toolchain, you cannot truly understand what you're doing without spending a lot of time looking at the generated assembly, and writing some of it yourself.
However, this kind of architecture has nothing to do with the x86. RISC, no cache, in most cases no MMU (and rarely any DMA), an extremely simple and straightforward pipeline, etc. I've written assembly for several such architectures, but compared to them I find x86 assembly intimidating.
For us, latency and determinism are the most important requirements of any processor, and it's very hard to get that on x86 without writing your own OS. But that negates the entire advantage of using the x86 platform with so much available code written on top of other kernels.
This is at least partially because a lot more work has been put into gcc's x86 backend. If the compiled code was that bad then presumably you saw quite a bit of low-hanging fruit in terms of potential improvements to the optimiser.
Run speed is objective and measurable, which is desirable for a fitness criterion.
Somehow, you'd have to specify which instructions can be moved around and which can't -- or maybe not, if you could also assess whether the program executed "successfully."
I don't really see it as practical, but it might be fun. Might uncover optimizations that no one has seen before.
Unfortunately in most modern OSes run speed on x86 is variable and profiling itself is intrusive. You might be able to get a general trendline toward a faster algorithm but it'll still be lost in statistical noise.
The second problem is that the underlying processor is a moving target. The x86 programming model doesn't directly map to hardware anymore, with out of order execution, branch predictors and register renaming. One variant of the processor with a particular caching algorithm might have totally different results than the next.
Stochastic Superoptimization (http://cs.stanford.edu/people/eschkufz/research/asplos291-sc...)
All the instruction latencies, VLIW optimizations, and so forth have a minor effect.
However, the cases where these things do have an effect, are exactly the cases where it might make sense to go down to ASM level.
The first avenue to improve performance, though, is improving memory access patterns. Compacting, aligning and rearranging data in cache windows. Even creating and maintaining redundant copies of data which is compacted or organized in a way that aids the cache can help.
Another big improvement I've witnessed is helping the prefetcher. For example, I've had code that had to iterate a few linked lists to work on all elements. When I reorganized that code to interleave the iteration of 8 linked lists at a time (rather than iterating them all serially), I went up from being memory-latency-bound, to being closer to memory-throughput-bound. The prefetcher could predict the next 8 ptrs I'll dereference, and prefetch 8 cache lines, instead of 1 at a time.
And while it's convenient to use intrinsics to generate SSE (intrinsics look like ordinary C functions and save you from having to write actual assembly), it's impossible for them to always be as efficient as a clever programmer. I'm pretty sure that generating perfectly efficient SSE from SSE intrinsics is an NP-complete problem. So a clever programmer will always be able to boost performance more than a compiler.
That said, the two biggest arguments against assembly are:
(a) performance matters much less in the modern day,
(b) you kill your portability by writing inline assembly. If you expect your code to run on Windows (and gamedevs expect their code to run on Windows, though this may be changing) you'll need to maintain two parallel, #ifdef'd hand written bits of assembly that do exactly the same thing: one for GCC, and one for MSVC. It's a Royal Pain.
Those are two damning arguments, and it's why SSE assembly crafting has become something of a lost art.
EDIT: I originally led with "It's disappointing that the author makes no mention of SSE assembly," but this was confusing due to the distinction between assembly vs intrinsics. They did briefly mention SSE intrinsics, but not that hand-written SSE assembly is sometimes superior for raw performance. Also, getting a strikethrough modifier on HN for edits would be wonderful.
I meant that I was disappointed he didn't talk about SSE assembly directly.
Right now, but I think future proofing is another issue to keep in mind alongside portability. CPUs in, say, 5 years are likely to have alternative and more powerful ways of dealing with certain things, and a good optimizing compiler should (optimistically!) be able to compile older code to run even more quickly on newer CPUs than hand coded assembly might allow.
I'm not sure how much impact this issue has in something like the Linux kernel, say, but I could imagine if significant parts of the kernel were written in i386-level x86, it would be a big issue now and demand much rewriting compared to C. (But I could be wrong as I'm not a kernel hacker, and would appreciate expert insight here ;-))
That said, for some algorithms there's a decent chance that "today's performance" is about what we can expect 20 years from now, so maybe hand-optimization is beneficial.
Being NP-complete doesn't really have anything to do with whether humans can solve a given problem better than computers or not. It just means that it (probably) takes exponential time to solve.
I believe there is no distinction whatsoever, in principle, between computers and humans (our brains are just big, complicated calculators). As such, it's only a matter of time before computers beat humans in compilation at every point, just like the situation in chess. People claimed for years that humans will always have some edge in chess, because they "understand" the game, or have "ingenuity", or something like that. Nobody claims that anymore, since computers absolutely massacre humans. And similarly to chess, computers have been getting better and better at a very fast rate, while humans have stagnated more-or-less.
As an aside, computer chess got quite boring for me, due to the homogeneity of current programs, but it did help to show me the limitations in current programming languages and compilers. So I'd like to do my part to put humans in their place in another field. My still very-much-not-even-close-to-finished contribution, that I haven't found much time to work on in a while, the Mutagen programming language: https://github.com/zwegner/mutagen
In theory, some of these shortcomings could be solved by a well-written aggressive whole-program optimiser that could deduce data usage & aliasing, but in practice they don't/can't, few people use WPO and it fails when you use library code.
In short, compilers probably don't have enough info to beat a skilled assembly programmer.
This is the main goal of the Mutagen language. See the link above for a bunch of hand-waving about compiler technology that might exist some day :)
It's not that hard to write loops in C that vectorize well. And you get the benefit that those same loops also vectorize for ARM neon.
Worse, it's almost impossible to retrofit into large existing code. Even turning on strict aliasing in the compiler might cause hideous problems. The compiler can't always warn you about these.
The 'holy grail' is still for a compiler that can infer this stuff from the code. With existing C code, that requires at least WPO and some very deep reasoning about pointer usage. The other approach is a different programming language that makes data/variable access control more explicit.
Turning on strict aliasing is a mixed bag as programs that break when it's on are explicitly breaking the C spec. But one must make amends to support badly written legacy code. The official way to handle the aliasing across types is to use unions.
In my opinion the default behaviour in C should have been what restrict does. And a keyword to allow alias and yet another to allow aliasing across types, because unions are bit tedious to use. As an example float *aliased foo. And warnings when you cast normal pointer into aliased one or cast across types if the pointer is not aliased across types.
Isn't that what LLVM tries to solve dynamically?
No. And in fact, even writing directly in LLVM IR, there are certain kinds of code (high-speed interpreters like LuaJIT) that cannot even be expressed.
LLVM is awesome, but it's a language (LLVM IR) that does not support the full range of what can be computed efficiently with a particular CPU.
Our brains seem to be primarily an I/O processor attached to a device micro controller.
This is quite unlike chess, where the rules are known up front, and are the same for both the computer and the human.
Yes, for any current programming language, humans do have some advantage by being able to utilize both low- and high-level semantic information that simply cannot be communicated to the compiler. The trump card, I believe, is coming up with new languages where this is not the case.
This is why, when I write assembly, I separate out entire functions, use an external assembler and link them into the final executable. You still have to worry about calling conventions, but in general your code is much easier to maintain than inline assembly.
Also, at least for the x86 and gcc, the inline assembly syntax is horrific, and I can't see how anyone can write software in it and still keep their sanity. You really want to use Intel syntax.
The disadvantage of writing externally-linked assembly code is that in general, x86 assemblers are crap. You don't even get register allocation, so when I program, I have those pieces of paper around with register allocation tables (1970s-style).
Would you please make a video or tutorial demonstrating you doing this? Because that's awesome.
Because it sure feels backwards to be doing that. And the guys next door who program Texas Instruments DSPs roll on the floor laughing -- they write linear assembly for their VLIW chips and the assembler composes the VLIW instructions, takes care of register allocations, and even reorders branch instructions as needed. I feel like using a chisel next to guys with jackhammers.
This will only yield performance benefits if the functions you write in assembly are large enough to actually create a performance benefit. Having a non-inlineable function call to a few instructions of assembly may yield worse performance than having a few lines of C code that may be optimized. This is especially true if you write 32 bit code where the ABI needs you to put parameters into stack.
> Also, at least for the x86 and gcc, the inline assembly syntax is horrific, and I can't see how anyone can write software in it and still keep their sanity. You really want to use Intel syntax.
This is just a matter of getting used to. It's just an assembler syntax, get over it.
You're going to have to learn to read AT&T syntax assembler anyway when looking at disassembly in the debugger or inspecting object files with binutils.
I, too, learned assembly programming with the Intel syntax and I still kinda prefer it, but for the occasional few instructions of ASM code I need here and there, it's just easier to use inline asm.
Not anymore. I use `set disassembly-flavor intel` in my .gdbinit. Modern LLVM binutils clones also allow disassembly into Intel syntax with `-x86-asm-syntax=intel`.
Modern versions of GAS can also use the `.intel_syntax` directive, which carries over into GCC inline assembly.
You should really be using conditionals in your make file to select source code files.
It is easier to test, easier to debug and you can have a portable fallback.
Plan9 uses this method, which is why porting to a new architecture is orders of magnitude simpler than other software platforms.
There is no #ifdef hell to work your way through.
For instance, plan9's lib C http://plan9.bell-labs.com/sources/plan9/sys/src/libc/ You can choose which functions you would like to hand code and test them against portable C versions with the flick of an environment variable.
tl;dr put your logic where it belongs
Please don't do this. It will make your code much harder to port. So almost every gamedev will be harmed by this.
1) create a "meta header file" malloc.h, containing the function prototypes.
2) create a "meta source file" malloc.cpp, with #ifdef switches for e.g. x86, x86_64, arm, etc, and inside the ifdefs some #include "malloc_x86.cpp"
3) inside the malloc_x.cpp source files, implement your architecture-specific asm code.
4) if you're after portable code, just do a #else #include "malloc_c.cpp" with a pure-c/cpp implementation. This way everyone can use the malloc and those with matching CPUs can benefit from acceleration.
If you start moving your programming logic (even compiletime logic) into any kind of makefile, you make life much harder for gamedevs, because it becomes harder to port your code to Windows. If you care about open source, then this is less inclusive a way to write code. You'll be excluding a segment of your programming community (gamedevs) because they're forced by the nature of the industry to compile their code on Windows.
That's changing, but it's changing slowly. In the meantime, some people don't care about that, but it would be nice to get some sympathy when it's so easy to write the code in a way that can include more groups of programmers, which is what open source is all about.
It sounded like they were talking about standard make though. Does plan9 use CMake?
My experiences from assembly code vs. intrinsics are just the opposite from yours. What I got out of some intrinsics heavy C code from GCC and Clang was way better than I could ever write.
Secondly, using inlineable functions, I could get interprocedural optimization from the compiler. I only needed to write the low level primitives (4x4 matrix multiply, matrix inverse, etc) with intrinsics and the high level code would be inlined and optimized. Adding assembly code to the mix will inhibit compiler optimization.
Additionally, the code I wrote could be retargeted for ARM NEON just by switching compilers. I was using mostly the compilers' vector extensions instead of CPU-specific intrinsics (ie. __builtin_shuffle, not __mm_shuffle_ps).
> And while it's convenient to use intrinsics to generate SSE (intrinsics look like ordinary C functions and save you from having to write actual assembly), it's impossible for them to always be as efficient as a clever programmer. I'm pretty sure that generating perfectly efficient SSE from SSE intrinsics is an NP-complete problem. So a clever programmer will always be able to boost performance more than a compiler.
Looking at the assembly code generated by GCC and Clang, I can see that the compiler does near optimal instruction scheduling and register allocation. On the other hand, the compiler was rather poor in doing instruction selection. I needed to pay a little attention to the code I wrote to avoid GCC spilling registers to the stack, Clang was a bit better.
So in my experience, a clever programmer writes poorer assembly code than a clever programmer using a clever compiler and intrinsics. You might need to look at the emitted code from the compiler and do some fine adjustments, so assembly skills are still needed.
I bet there are still cases where you can beat the compiler in an isolated problem, but when it comes to maintainability, re-targeting and the amount of effort needed to write the code, writing assembly code isn't worth it most of the time.
I wasn't arguing that people should generally do this. The tradeoffs are rarely worth it. What I was saying is that humans can be more capable than compiler magic, and when it matters, it can really matter. Sometimes by up to 600x, as the other commenter demonstrated.
This was exactly what I was doing - high throughput inner loop code with lots of matrix multiplies, etc. I wrote the primitive operations (multiply, inverse, etc) with intrinsics but my high level code was readable and maintainable C code. The compiler-emitted assembly code was at least as good as I could have written by hand, pretty close to the speed of light (ie. pretty close to CPU peak flops and/or memory bandwidth, whichever was the relevant bottleneck).
It takes an exceptional programmer and a lot of time to "wipe the floor with the compiler" with hand written assembly code. The compiler's ability to come up with near-optimal instruction scheduling and the ability to quickly change register allocation is pretty hard to match - you can do it if you spend enough time with the Intel manuals but that's not time well spent.
By analyzing your compiler output and fine-tuning your intrinsics code, you should be able to get pretty close to the peak performance with less time spent than hand writing the asm code. I stand by my point - a clever compiler and a clever programmer together can get a better job done than either of those individually.
This, however, requires a programmer with assembly skills. It's still valuable to be able to write and especially read assembly code.
What I'm saying is that you have to write an algorithm in assembly which has an inner loop. You can't have the body of your inner loop in assembly, but the looping mechanism in C. This is because you can pull out parts of the algorithm from the inside of the loop to outside of it, or you can prevent the compiler from constantly re-loading from memory into SSE registers inside the inner loop, etc.
In other words, what I think you're saying is:
for (...) {
/* inline assembly or intrinsics */
}
What I'm saying is: /* start inline assembly
preload anything that can be preloaded into SSE registers
loop:
achieve parts of the overall goal
write any finished operations back to main memory
prefetch N cachelines ahead
end inline assembly */
It may seem like a small difference, but in the former case, you're still doing a "single matrix multiply" (you're doing them one at a time). I think that might be why you're not seeing much of a difference between intrinsics vs inline assembly.Matrix multiply isn't such a good example... there are better cases like texture lookups in a software rasterizer.
Also, the prefetching step is vital. It may be you're not seeing performance gains due to lack of prefetching, which is orthogonal to the debate of SSE intrinsics vs SSE inline assembly. Something like VTune can help reveal precisely why your program is going slow. There are many pitfalls, and it's easy to accidentally optimize for the wrong thing and then trick yourself into believing that particular optimization never helps in any circumstance.
My code was something like:
__builtin_prefetch(first cache line);
for(a lot of 3d objects) {
__builtin_prefetch(next cache line);
vec4 q = read_quaternion(), p = read_position();
mat4 m = matrix_product(translate(p), quat_to_mat(q));
mat4 n = inverse_transpose(m);
mat4 p = matrix_product(camera_matrix, m);
// lots of more matrix-vector-quaternion math here
stream_store(m); stream_store(n); ....
}
In other words, my "high level" code was readable C code, only the primitives (matrix_product, etc) were intrinsics code.What the compiler did is inlined all the primitive ops, store all the values in registers at all times (no loads or stores or register spilling in the inner loop) and finally re-organize the instructions to get near optimal scheduling.
In some simpler programs, I got even more benefit from the compiler doing some loop unrolling to keep loads/stores balanced with ALU ops to get very effective latency hiding.
Writing that whole loop with assembler would have yielded next to no improvement but that would have sacrificed readability, maintainability and portability.
> You can't have the body of your inner loop in assembly, but the looping mechanism in C.
You're right in that you can't mix inline assembler and C and get the compiler to optimize it correctly, but using intrinsics you can get the best of both worlds.
This is pretty much exactly what I had, except that my "inner loop" primitives were C code+intrinsics that looked like assembly code. Because they were not assembler code, the compiler was able to optimize the whole loop and the end result is something similar to what you show as an example of good performing asm code.
If you do end up writing assembler, you have to be sure that it is worth sacrificing the compiler optimizations that would take place otherwise.
You've convinced me to leave that as an ultra-last-resort though, and just try out intrinsics first. Dumb compilers are probably less of an issue nowadays than in ye olden days of ~5 years ago. Thank you for the thorough explanation!
Yeah, at some point the compiler can't do any more magic and will start spilling. But it's not very often when a programmer intervention is required.
Quite often you can get the effect you need by looking at the compiler-emitted code, see where the spilling or other unwanted effects happen and do small tweaks to your C code. This is a bit annoying but it still beats hand writing asm code when it comes to time investment (it is not necessarily as fun, though).
> You've convinced me to leave that as an ultra-last-resort though, and just try out intrinsics first. Dumb compilers are probably less of an issue nowadays than in ye olden days of ~5 years ago. Thank you for the thorough explanation!
Yes, compilers have improved and will keep on improving. I was blown away by the quality of code I saw coming from GCC and Clang, in particular about how good the instruction scheduling was.
In any case - writing C code that looks like Assembly code (ie. plain and simple) gives very good results and doesn't require you to sacrifice compiler optimizations like using real Assembly code does. Read the output Assy code and revisit the C code if required, this way you should be able to get near-optimal code with less time investment.
Writing Assembler code is still really fun, though. Unfortunately it's not a good investment when it comes to achieving your goals in time.
I believe the term you're looking for there is "a sufficiently-clever programmer" :)
Arguably one is more likely to exist than a sufficiently-clever compiler, but it's exactly the same claim. Even the best humans will have sub-optimal edge cases just like compilers. Given perfect information, sure, I'll agree humans can do better. But it's probably impossible to have (much less make use of) perfect information.
The actual article is quite good and makes a very valid point. Don't jump to machine code whenever things get a bit slow, chances are you're introducing more problems than you're fixing, and it's very difficult to outdo the compiler. Generally you should start by looking for macro-optimizations instead. But to add my own advice, it's still an incredibly worthwhile thing to learn, even if you don't expect to ever run into the types of scenarios which will benefit from optimization at this level.
Why? Programming languages do everything they can to abstract away the machine, and abstractions do leak [1]. Even mostly air-tight abstractions over machine code [2]. Because of this, learning the entire "stack" makes debugging higher-level code much easier. When you can understand what's happening with every level right down to the metal, what were once ridiculous off-the-wall problems become recognizable and much easier to reason about.
So do use machine code if you'd like to become an expert. Just use it on your own time, and don't use it under duress.
1: Presently reads "Don't use assembly unless you're an expert."
2: http://www.joelonsoftware.com/articles/LeakyAbstractions.htm...
3: http://stackoverflow.com/questions/11227809/why-is-processin...
It's becoming very common to have a catchy title and then a disclaimer for it in the article.
When it comes to measuring the performance of code like this, averaging run times is not the way to do it.
To remove the noise caused by context switching, just run the code many times and report the single fastest run you get. This should be the value closest to running the code on an OS without preemption (i.e. you want to measure how fast the code runs on the bare metal without interruption).
Even Facebook's Folly library [0] changed their benchmarking code from using statistics to just providing the fastest run. As the comments say:
// Current state of the art: get the minimum. After some
// experimentation, it seems taking the minimum is the best.
return *min_element(begin, end);
This is explained in the docs [1]:Benchmark timings are not a regular random variable that fluctuates around an average. Instead, the real time we're looking for is one to which there's a variety of additive noise (i.e. there is no noise that could actually shorten the benchmark time below its real value). In theory, taking an infinite amount of samples and keeping the minimum is the actual time that needs measuring.
[0]: https://github.com/facebook/folly/blob/master/folly/Benchmar...
[1]: https://github.com/facebook/folly/blob/master/folly/docs/Ben...
I realize that's kind of crazy, but it'd be an awesome teaching and debugging tool. I don't look at disassembly a lot, but when I do it's not always particularly obvious why the compiler generated a thing in a certain way.
I can make things go more than 10 times faster in assembly. My main job is as manager/entrepreneur but I could read-write assembly as a result of my experience and I help-guide other people easily.
In the real world 10 times faster is nothing. You should spend the time understanding the problem in a mathematical way, and VERY IMPORTANT, documenting your work using images, text, voice and video.
This way you could make things go 100, 1000, 10000 times faster as most algorithms could be indexed, ordered in some way as to make it extremely fast, like doing log() operations instead of n squared or cubic or to the elevated to four or five(when you manage several dimensions like 3D with time or video analysis or medical tomography).
More important than that, 10 years from now it will continue working in new devices or OSs and will be something that supports the company instead of being a debt burden because the original developer is not here now(or you don't have the slightest idea of what you did so far away in the past and did not document).
The main problem is that people is not self aware that they forget things. And your brilliant idea that makes everything go 3 times faster is nuts if it makes everything way harder to understand, or if it could be forgotten even by you.
Here's a practical example: as a result of redesigning the algorithm to use fixed-point and implementing it in assembly, I got it to run 600x faster than the initial C version. Big O complexity was the same, the difference was in the constant factor. But the constant factor matters! In my case, it meant that you could get your computation done in half a day instead of a year.
Yes, it took me 3 weeks to get the algorithm implemented, instead of a single day, but even so — it was definitely worth it. And in many cases even a 3-fold improvement in speed is important, if you have long-running calculations.
Not knowing too much about processor architecture, I don't understand how fixed point can be much faster, since floating point ops are implemented in hardware.. I presume you used integer operations on your fixed point values, but could you explain a bit why it ends up being much faster than floating point?
So the speedup is not just from going to fixed point, but from managing to use the vector instructions.
That is just terrible advice.
You're probably right that in most cases you should not write assembly code to try to make some code run faster. However, it is a very valuable skill to know how to read assembly code and spot the inefficiencies.
For a low level hacker, it is a very valuable skill to be able to write and especially read assembler code. I need that skill regularly in my day job. And I would have not acquired that skill if I had not written some assembly code. And besides, writing assembly code is fun!
Sometimes you need that 10x speed improvement to be able to do what you need to. To get your game running smoothly or your video playback work. You need to know when and how to optimize for performance.
They said that about x86... 20 years ago. I have applications written in Asm that still work on the latest CPUs today. The same binaries, not even needing recompilation, now run several orders of magnitude faster. I still see a lot of potential in extracting performance from x86 and although I hesitate slightly to make this prediction, I think it'll be the dominant architecture for at least 10 more years.
That needs to be qualified as "for desktops/servers" or similar. x86 haven't been the dominant architecture for at least a decade, if ever, in terms of units shipped. It's being outsold in number of units by ARM at a 10:1 ratio, and MIPS and PPC's are shipped in higher volume as well, or at least did as of a year or two ago. Possibly even 6502 and various micro-controllers, though getting numbers is harder.
Keep in mind how many CPU's are around you. Our servers have an ARM core per harddrive, and several of our RAID controllers have multiple PPC cores, for example. We have some servers with dozens of non-x86 CPUs per x86 CPU. Even some SD cards have ARM cores on them.
Now consider your car, microwave, washing machine, dish washer, tv, set-top box, phones, camera, music player, digital radio. A lot of stuff that was semi-mechanical or employed discrete logic a few years back now have CPUs that are ridiculous overkill, but used because they're so cheap there's no reason not to.
x86 is a diminishing niche if you look at electronics as a whole.
Or my personal variant Most optimizations will end up in things being slower.
Knuth is saying that in most cases optimizing prematurely is not worth it, which is right, but just as critical is the part saying that there are parts where pays off to consider optimization early. Writing your logging function in assembly likely fits into the 97%. Writing your N-queens solver, if you're aiming for all-out speed, likely fits into the 3%.
Of course not thinking about optimization at all will lead you down a path where your software is unoptimizable, or at least very difficult to optimize. Ask anyone who's had to optimize a video game that manages all objects in a scene graph, where reorganizing a cache hostile tangled web of pointers to base classes is a herculean feat.
A persistent peeve of mine is this vapid commentary that would invariably show up in programming language related discussions
"what's the point of this language, if one needs speed, it will be written in C".
Some would replace "C" with "assembly" in that sentence.
The point is that if a piece of code is doing something non-trivial, it is extremely hard to have confidence in the correctness of handwritten code that is written entirely at a low level. Smart compilers apply drastic, and sometimes unintuitive transformation to optimize code. To do the same manually and using low levels of abstractions would require a toure de force to pull off correctly. This simply is not going to happen often. Its precisely for speed that we need a high-level, yet optimization friendly, language so that we can delegate the job of optimizing the code to the compiler. Compilers are much better at applying large correctness preserving transformations than we humans are. The "write it in assembly" works in the small, not in the large.[1] http://piumarta.com/software/maru/
[2] http://www.youtube.com/watch?v=cn7kTPbW6QQ
EDIT: There are other reasons to write machine code than performance.
There was a time when a man can eat raw meat (machine code) but he can start again today by eating bloody beef.
It's ridiculous to pretend that eating hormone-accelerated mass-processed meat is closer to nature. Sure, it's oozing blood, but that blood is tainted by an enormous industrial profit machine.
My boss also hunts animals with a bow and arrow.
Does he get a pass?
Even back in the 80s people were spoofing this attitude, with essays like "Real Men Dont Use Pascal" http://homepages.inf.ed.ac.uk/rni/papers/realprg.html
1) High level languages are easier to refactor and maintain than lower level languages (changing the algorithm in the assembly's a big job).
2) For performance, the algorithm used is almost always more important than the language you choose to implement it in.
Both of those points back up the author's assertion, it's a shame he didn't specifically discuss them.
I used to hand-code assembly (or generate it, eg "compiled bitmaps") back in the 8086 days. For many situations there were easy gains to be had that couldn't be achieved so easily in high level languages. These days I wouldn't dream of attempting it other than perhaps for vector code in an inner loop. I'd far rather spend the time profiling and optimising the algorithms and data structures because that's where the big gains are going to be. To go ahead and implement something in assembly when the algorithm clearly isn't yet optimal is perhaps fun but also kinda insane.
No one will ever know or care that a 5ms operation only takes 2ms because of how "savvy" you are. No one knows or cares what language you used. Its just a program.
If you can produce a working program in a high level language but chose to use a low level one.... whyy?
If your operation is relaying a 20ms audio packet from one side of a call to another, 50µs to 20µs is the difference between handling 200 concurrent calls and handling 500 (per CPU, minus non-linear scaling with increasing CPUs and other overhead). Unlike the cost of hardware, which every customer would have to buy separately, the development cost can be spread out over many customers.
If $50,000 of developer effort can save 50 customers $1000 on hardware, you've already broken even in a sense.
But there are still circumstances where it's necessary to write assembly, particularly in things like real-time programming. Doesn't matter how good a job the compiler did on the overall program, if you can afford 16 cycles on an inner loop and the compiler spits out 24 then you've gotta hand-tune it.
Old folks only would even understand
assembly isn't just for speed ..its for fun too.