Understanding C by learning assembly
recurse.com
recurse.com
Level 1 is what this article describes: your eyes are opened when you see how various features of C map naturally onto things that are simple/efficient in hardware. You see a pointer dereference and in your head you think "aha! load instruction." You understand that there is a real difference between "char ch[10]" and "char ch[10] = {0}" -- the latter emits extra instructions to zero out the memory.
Level 2 is when you realize that the optimizer is free to screw with all your intuitive ideas about C -> hardware mapping that you learned about at level 1. You learn that writing correct C means that you need to think in terms of the abstract machine that C defines, and not assume that any specific C statement will translate to the assembly code you expect. You can't think "pointer dereference means load instruction", you have to think "pointer deference reads an object."
There's nothing wrong with level 1, as long as you don't stop there!
I dunno where I picked it up, but my standard euphemism for writing lock-free code is “juggling knives”.
"Before I learned the art, a punch was just a punch, and a kick, just a kick. After I learned the art, a punch was no longer a punch, a kick, no longer a kick. Now that I understand the art, a punch is just a punch and a kick is just a kick." -- Bruce Lee
Unfortunately, this advantage means you can't understand all of C just by understanding assembly, even though C programs turn into just assembly at bottom.
Edit: after the responses below I've been (re)reading https://www.scss.tcd.ie/John.Waldron/3d1/arm_arm.pdf and http://www.intel.com/content/dam/www/public/us/en/documents/.... And marvelling.
Sadly it sure does. Grep through the ARM architecture reference manual for "UNPREDICTABLE" (helpfully typeset in all caps), for example…
:(
str x9, [x9], #8
Luckily it was fixed upstream. I don't like working in tablegen files. =)
https://github.com/llvm-mirror/llvm/commit/7f4f923aa57ad8d7e...
That's not quite true. For example, in my blog post on assembly [0] one of the first points I make is that on x86-64, the stack must be aligned by 16 B when calling into libc. (Note: the point I'm making is invalid if we're talking assembly with no libc interop, but c'mon). That's actually not a hard and fast rule; it has more to do with requirements of the implementation of that function for alignment. For instance, I can invoke memset from x86-64 assembly with a misaligned stack without segfaulting on _my_ machine, _but_ that is because my implementation of memset I'm linking against doesn't require alignment. That might work for me today, but let's say the next upgrade to my libc suddenly has a new implementation of memset that requires alignment for whatever reason (I think vector register access requires stack alignment, for instance). Suddenly, my assembly program that worked the day before (possibly for years) has now betrayed me. And you just have to keep track of alignment in your head. Oh, someone pushed a 64b register on the stack? Misaligned by 8B!
I'm not arguing the point about undefined behavior (though I'm sure there's probably known bugs with certain instructions on certain families of processors [1]). My point is that it's very possible for assembly to betray you, and compiler warnings from C/C++ compilers are a god send compared to the usual seg faults and bus errors experienced when writing assembly. (honestly, I find pencil & paper very useful for debugging assembly; by sketching out the state of things and trying to reason about them, you find failed assumptions and mistakes).
[0] http://nickdesaulniers.github.io/blog/2014/04/18/lets-write-... [1] http://en.wikipedia.org/wiki/Pentium_FDIV_bug
- Values of certain flags after certain instructions
- The result of bsf or bsr on 0.
- I think there are some floating point operations that can be undefined when the results fall off the edge of representability, but I don't remember off the top of my head.
These have behavior that differs from processor to processor.
This isn't to mention the various kinds of processor errata that pop up: loads being reordered with LOCK'd x86 instructions, or ARM memory reads that go backwards in time, both of which happen in real processors despite being architecturally forbidden.
In the olden days, maybe, but nowadays? No way. As the most extreme example: forget to use memory barriers when you should? Congratulations, the OOO nature of modern day CPUs means your code may not execute identically on all platforms. But even lesser things can impact you: instructions that were cheap become expensive, or suddenly registers that you know you weren't supposed to touch but did (and that happened to work) suddenly don't because a library got changed, or half a dozen other things.
With assembly on a modern CPU, congratulations: you have eliminated one level of abstraction. But there are still many, many turtles to go before you're really all the way down.
Agree.
At which point, C gives you a huge advantage. Sure, you can hand write assembly, but I bet the libc implementation you're linking against has optimized to death with intimate knowledge of timing sequences and inner state of the processor from the processor vendor. I would expect a company like Intel to contribute patches to glibc (GNU libc) based on running functions from the standard library on a very very expensive FPGA (or just expensive-software simulation of hardware).
Memory ordering operations are an advanced topic; I have a copy of C++ Concurrency in Action, currently sitting on my bookshelf, that has a chapter explaining the concepts. Should be a fascinating read!
It's almost the other way round: they have a corpus of binaries to use when making processor design decisions.
Nope. I expect _all_ early processors had undefined behaviour, because they didn't have the transistor budget to make all instructions behave predictably. The 6502 certainly had (http://www.ffd2.com/fridge/docs/6502-NMOS.extra.opcodes), as did the Z80 (http://www.z80.info/z80undoc3.txt)
The article is rather about exploring a particular C run-time at the assembly language level, and a tutorial about stepping through programs at that level.
This is something useful that would-be developers need to learn, the main reason being that programs sometimes fail in ways such that uncovering the root cause requires a trip through that level. Other reasons are that in some situations we need to verify that programs believed to be correct at the source level are actually being translated in the way we expect.
But, why do you think that understanding C by learning assembly is a terribly wrong-headed idea?
Some professors tend to disagree and prefer a bottom up approach of learning. The main advantage is : you get rid of the "magic" of computing.
Later on I went back to C since I knew I had never really done a lot with it and everything was a breeze. I understood what was going on, loved it, and on top of that was able to easily knock out loads of exploitation stack-smashing puzzles I found on the internet.
I already knew how to program pretty well but didn't really know system programming and C. I can imagine the value of, instead of writing hello world, learning assembly first and writing programs that end with some exit status instead. It would probably be way more boring, but way more informative in the long run.
Just as an anecdote, I have seen a surprising number of C programmers assume that the language's type aliasing rules are related to or somehow governed by pointer alignment, when these are entirely separate. The former exists to facilitate analysis and optimization at the compiler level -- there is no "ground" reason. It's important not to get trapped in the portable assembler mindset and lose sight of an entire level of transformations that might not have an intuitive relationship with hardware.
This is especially important in C++, where you depend on the compiler's optimization ability a lot more.
It used to be good, but it is not any longer thought of that way.
And I agree with your basic point that building one is the only to gain true understanding.
https://news.ycombinator.com/item?id=1727004
https://code.google.com/p/letsbuildacompiler/
(It looks like there are several more or less complete C versions.)
[1] http://www.plantation-productions.com/Webster/www.writegreat...
The error: "Unable to find Mach task port for process-id 13942: (os/kern) failure (0x5). (please check gdb is codesigned - see taskgated(8))"