C Internals
avabodh.com
avabodh.com
Assembly language and computer architecture using C++ and Java Book by Anthony J. Dos Reis
From book's description:
Students learn best by doing, and this book supplies much to do with various examples and projects to facilitate learning. For example, students not only use assemblers and linkers, they also write their own. Students study and use instruction sets to implement their own. The result is a book that is easy to read, engaging, and substantial.
I'm not affiliated with the author though. This book helped a lot in my career as a hardware and firmware engineer.
It talks about not only C language but C++. You'll be able to translate a C++ class into assembly language by hand. The drawback is that you'll learn a hypothetical CPU but the concepts are still the same with the real world CPU. C internals could supplement it as well.
I took a compilers course in grad school and it was a lot of fun. Especially the sort of frankenstein moment when you got it working and it actually did something, and its output would itself run.
(Another fun thing is to design an instruction set and implement a simulator and asm for it. I have a lot of fond memories).
...
void main()
...
That's where I stopped reading.Since this comment is getting downvotes, I'll explain. The "void" keyword was introduced in the 1989 ANSI C standard. That same standard specified that the valid definitions for the main functions are:
int main(void) { /* ... */ }
and int main(int argc, char *argv[]) { /* ... */ }
or equivalent (or implementations can support other forms). There has never been a version of C in which "void main()" is a valid way to define the main function.It's a small detail, and yes, many compilers will let you get away with it (the language doesn't require a diagnostic), but anyone writing about C should be aware of this, and should set a good example by writing correct code.
Maybe the site is OK other than that, but it doesn't inspire confidence.
These kinds of problems might not happen with common ABIs, but if you try to write C on the assumption that it will be compiled in a reasonable way based on your knowledge of how the plaform works, then modern compilers will punish you for your presumption.
Which is really a bug that everyone has decided to look the other way around because it wins compiler benchmarks despite coming against what C originally was meant for.
Specifically (from the C89 rationale[0], that i'm certain most people who think those benchmarks are good haven't read):
> C code can be non-portable. Although it strove to give programmers the opportunity to write truly portable programs, the Committee did not want to force programmers into writing portably, to preclude the use of C as a ``high-level assembler'': the ability to write machine-specific code is one of the strengths of C. It is this principle which largely motivates drawing the distinction between strictly conforming program and conforming program (§1.7).
Also in the same rationale it mentions how "the spirit of C" is to do operations in the way the machine would do it instead of forcing some abstract rule - yet that exactly is what happens later when everyone goes all language lawyer about C's abstract machine and how you should not rely on what you think the target machine would do.
The C standard specifies (for hosted implementations) two ways to define "main", and allows implementations to document and support more. "int main(void)" is one of them. "int main()" is not. So, strictly speaking, using "int main()" makes your program's behavior undefined.
On the other hand, as far as I know every implementation actually allows "int main()" with no problem. This was necessary to support pre-ANSI C code, which couldn't use the "void" keyword. It's still better to be explicit and use "int main(void)" rather than "int main()". (It can also affect recursive calls to main, which are legal but almost certainly a bad idea.)
However, in the specific context of stand-alone or freestanding programs, the "void main()" definition would be absolutely nonconsequential, since there is no host to return a value to.
One possible consequence I can think of: If your program is safety-critical (or even otherwise) you might be interested in running static analysis or verification tools on it. These tools implement the C standard, so they would emit a diagnostic on the use of void main.
Other than that, sure, the language police will not come and break down your door. If your compiler's docs say that it accepts this construct with the meaning you want, it is indeed an inconsequential, though also completely unnecessary, deviation from the standard.
For hosted implementations, "void main()" might be valid, but "int main(void)" is always valid.
A tutorial should not suggest "void main()" without mentioning any of this.
Reference: http://www.open-std.org/jtc1/sc22/wg14/www/docs/n1570.pdf sections 5.1.2.1 (freestanding environments) and 5.1.2.2.1 (hosted environments).
Still, I'm not aware of any implementation that explicitly supported void as the return type of main. (Many compilers would accept just about any return type without complaint.)
Somehow some authors got the idea that "void main()" was a good idea. I find it to be a good way to detect authors who don't know the language very well.
It seems like a bad idea to use void return type (presumably you want the program to always return 0), but the standard permits it.
edit Wikipedia seems less convinced that it's ok [1]
You can't just make up extra signatures for main and expect your compiler to deal with them. Of course, if your compiler vendor says that a particular non-standard signature is available, you're free to use it.
Put another way, C doesn't let you get away with the try it and see approach the way many other languages do.
Exactly, and I don't see any upside to this "tradeoff". It seems strictly worse, but maybe I'm missing something.
More generally it's because the C language standard is steered only partly by the interests of C programmers: 1) the C standard aims to easily support many different platforms and compilers, 2) it aims to remain a compact and stable (slowly changing) language, and 3) it aims for maximal performance. None of those 3 aligns with programmer convenience.
Point 1 ties in to the origins of the C standard itself, which was deliberately designed not to nail down every last detail. This was to incorporate different compilers and platforms. e.g. Where Java mandates two's complement, C doesn't. C was carefully designed so that platforms that don't use two's complement, don't have to jump through (many) hoops in order to comply with the standard. Neither should they be tempted to just break the standard for their own convenience/performance.
Point 2 means there's tremendous inertia in the C language. There are considerable upsides to the language being slow to change, though. Although mistakes in the language are slower to be fixed, they're also much less likely to make it into the standard in the first place, compared to a fast-moving language like D. It's also a tremendous advantage for compiler engineers - they can spend their time improving their compiler rather than keeping up with the standard.
(Related trivia: C++20 will break with tradition and mandate two's complement for signed integer types. I don't think C will do this though.)
And finally Point 3. C's weird and wonderful rules on undefined behaviour permit compilers to make strange optimisations and to omit runtime checks, but they require the programmer to have an eagle eye, and undefined behaviour can manifest in peculiar ways that are hard to hunt down. There are endless horror stories of bizarre undefined behaviour.
That said, I'm not sure that performance (on modern platforms) is really such a factor in C's quirks. I believe that these days (with modern compilers), C's performance is generally about the same as that of other similar languages that lack broad undefined behaviour, such as Ada.
Honourable mention: in at least one instance, undefined behaviour was deliberately introduced into the standard to permit trap-based error-reporting that was convenient on one particular hardware platform. [0] (Previously the relevant action would merely produce an indeterminate value.) I believe this unusual though even for C.
[0] http://blog.frama-c.com/index.php?post/2013/03/13/indetermin... (ctrl-f for Itanium)
FWIW, I learned "ANSI C" from a book older than 1989. I've written C programs that ran on 8-bit to 64-bit machines, big- and little-endian, on a dozen OSs (most of which no longer exist), and even no OS. I've never even heard of this being a problem.
Lets pick the "Translation of Arithmetic Operations" example and convert it to float instead of int
int a = 2;
int b= 3;
int c = 24;
a = a + b;
a = a + b * c;
http://www.avabodh.com/cin/arithmeticop.htmlAnd pack it into a function so that Golbolt can compile it.
float hn_demo(void) {
float a = 2;
float b = 3;
float c = 24;
a = a + b;
a = a + b * c;
return a;
}
And now pick a CPU that isn't that good with floating point, like AVR, using GCC 4.5.4, and we get: ldi r24,lo8(0x40000000)
ldi r25,hi8(0x40000000)
ldi r26,hlo8(0x40000000)
ldi r27,hhi8(0x40000000)
std Y+1,r24
std Y+2,r25
std Y+3,r26
std Y+4,r27
ldi r24,lo8(0x40400000)
ldi r25,hi8(0x40400000)
ldi r26,hlo8(0x40400000)
ldi r27,hhi8(0x40400000)
std Y+5,r24
std Y+6,r25
std Y+7,r26
std Y+8,r27
ldi r24,lo8(0x41c00000)
ldi r25,hi8(0x41c00000)
ldi r26,hlo8(0x41c00000)
ldi r27,hhi8(0x41c00000)
std Y+9,r24
std Y+10,r25
std Y+11,r26
std Y+12,r27
ldd r22,Y+1
ldd r23,Y+2
ldd r24,Y+3
ldd r25,Y+4
ldd r18,Y+5
ldd r19,Y+6
ldd r20,Y+7
ldd r21,Y+8
rcall __addsf3
mov r27,r25
mov r26,r24
mov r25,r23
mov r24,r22
std Y+1,r24
std Y+2,r25
std Y+3,r26
std Y+4,r27
ldd r22,Y+5
ldd r23,Y+6
ldd r24,Y+7
ldd r25,Y+8
ldd r18,Y+9
ldd r19,Y+10
ldd r20,Y+11
ldd r21,Y+12
rcall __mulsf3
mov r27,r25
mov r26,r24
mov r25,r23
mov r24,r22
mov r18,r24
mov r19,r25
mov r20,r26
mov r21,r27
ldd r22,Y+1
ldd r23,Y+2
ldd r24,Y+3
ldd r25,Y+4
rcall __addsf3
mov r27,r25
mov r26,r24
mov r25,r23
mov r24,r22
std Y+1,r24
std Y+2,r25
std Y+3,r26
std Y+4,r27
Which includes calls to a floating point emulation library and quite different from the x86 example, with numbers using multiple registers.So more of an heads up, sometimes the C translation to Assembly isn't as direct as one might think.
float hn_demo(float a, float b, float c) {
a = a + b;
a = a + b * c;
return a;
}
and let the compiler actually optimize with -02.The section on local variables assumes a downward-growing stack. This is completely fair, because the introduction specifies that the articles deal with an x86 world. What gets missed out is the fact that the direction of stack growth is determined by the processor :)
This is not really a complaint ... it just seemed to me like a missed opportunity to mention something interesting.
(I'm not saying you're wrong --- I personally know two ;-)
2. Intel 8051. 8-bit microcontroller which now is found as a core in countless SoCs long after Intel stopped making them. You probably own and use something every day which has an 8051 or 8051-core MCU in it.
(Examples of its weirdness: it only has one general-purpose register; it has three address spaces, some of which are bank-switched; every register other than the program counter exists in memory; some memory locations are bit-addressable...)
I remember the stack in ARM actually being a store that +/- the memory address it's pointing to (basically this http://infocenter.arm.com/help/index.jsp?topic=/com.arm.doc.... )
Unfortunately I don't have an arm compiler at hand to test it
I _think_ what you're describing sounds similar to the SAVE/RESTORE mechanism in SPARC, so quite possibly that's what I was being told.
Summary: original Arm gave the software the flexibility to implement an ascending stack, but in practice the stack was always descending, and more recent architecture changes assume this to a greater or lesser extent.
ARM let you choose either approach (making four different stack configurations in total!) - this flexibility is because ARM didn't have any specialised 'push' or 'pop' operations, you read/write to the stack using the normal load/store ops, which have a variety of addressing modes.
Not really having a stack ‘preference’ in the CPU isn’t uncommon for pure RISC architectures, where having an instruction that both jumps and writes a return address to memory is a no-no.
The section on local variables assumes a stack. This is completely fair, because the introduction specifies that the articles deal with an x86 world and ABI. What gets missed out is the fact that the C standard doesn’t even prescribe (edited) the existence of a stack.
(Just checked the final draft of ISO C2011. As far as I can tell, it doesn’t contain the word stack, and only uses push in the description of (w)ungetc. I also don’t think there’s wording about performance or ordering of addresses of local variables across frames that effectively enforces the use of a stack. I think it is still perfectly fine to use a linked list of environments for local variables)
This is not really a complaint ... it just seemed to me like a missed opportunity to mention something interesting. :-)
EDIT: Would any of the people downvoting this comment care to explain their affection for this historical mistake? I have never understood why anyone would choose it.
movl 8(%ebx,%eax,2), %eax
What does the 8 mean? How does the stuff in the parenthesis work? Why is there a type suffix? Here’s a general definition of AT&T syntax’s indirect form segment:offset(base,index,scale)
But the equivalent in Intel syntax (for the above two is): mov eax, [eax*2+ebx+8]
segment:[index*scale+base+offset]
Idk, but the Intel syntax is just clearer to me. eax = *(eax*2+ebx+8);
Like pseudo code almost. It honestly seems like AT&T syntax was created to facilitate easier parsing by computers, not humans.The biggest thing for me is that the parameter order doesn’t follow the Intel or AMD opcode manuals; I have to flip the operands in my head to compare them to the opcode manual.
I’m not saying people are wrong for using AT&T syntax, or that it’s not intuitive for some. Just that Intel felt more intuitive to me.
Edit: realized that gcc first target was 68k so it would make sense for gas to use right way round assembler syntax.
And it really, really bothers me when my tools do not match my documentation, for no good reason. (Just use the `-M intel` switch with x86 GNU tools, and then they will match. Or on ARM, do nothing, because by then they'd sensibly figured out not to bother with their "generic" syntax.)
Never again, after all this years I still have vague memories of how I used TASM and MASM, and trying to write x86 AT&T was such a pain.
All the GNU tools, and many of their clones, use AT&T syntax. I think I run into it more often than Intel, and I turn on the option for the latter where I can. It’s really prevalent.
As an example, scientific research papers were still being published in raw PostScript (PS) format long after PDF existed and become the defacto desktop publishing standard for 99.9% of the world outside of academia.
The use of AT&T assembler sticks out for me too, because I had learned Intel assembler back in the IBM XT days and wrote "demo" programs and all of that. And then my university used AT&T which was just so bizarre because literally 99% of the students had IBM compatible computers at home with Intel CPUs! Most of the lecturers had Intel PCs, most of the labs had Intel PCs, and it was just a handful of Solaris machines that had RISC CPUs and toolchains based on AT&T assembly.
Similarly, if you Google "Kerberos", an insane number of references pretend that this can only mean "MIT Kerberos", and is used for University lab PC authentication only. Meanwhile, in the real world, 99% of the Kerberos clients and servers out there are Microsoft Active Directory, and all configuration is done via highly available servers resolved via DNS, not static IP addresses.
Some design aspects of Linux and BSD have similar roots, and it shows. The DNS client in Linux is quite clearly designed for University campus networks. Combine this with typical University servers using hard-coded IP addresses for outbound comms, because of things like the Kerberos example above, and you get an end result that doesn't handle the requirements and failure modes of more general networks very well.
Weaponized comment flagging gone awry.
but yes, Gnu tools are ubiquitous and will continue to be so.
https://raw.githubusercontent.com/pervognsen/bitwise/master/...
https://github.com/oriansj/mescc-tools-seed/blob/master/x86/...