Why C Is Not Assembly
james-iry.blogspot.com
james-iry.blogspot.com
note: not vectorized: number of iterations cannot be computed.
note: bad loop form.
note: vectorized 0 loops in function.
However, gcc-4.5.1 and gcc-4.4.4 vectorize it when the operation is written as a standard `for` loop. Clang never vectorizes it.So I guess my perspective is that there are two different types of people using C: people who are writing systems applications and people using ANSI C. Personally, I would guess that the first group is much, much larger.
> Most real C implementations will go some distance beyond the standard(s), of course, but I have to draw the line somewhere.
You missed the guy's point, which was essentially that C code is munged and transformed by a good optimizing compiler, whereas assembly is left untouched by a good assembler.
You missed the guy's point, which was essentially that C code is munged and transformed by a good optimizing compiler, whereas assembly is left untouched by a good assembler.
That's irrelevant. They are both restricted on the basis of being functionally equivalent to the input code.
But that goes for any compiled language, unless the compiler contains a bug.
All compiled languages are "restricted on the basis of being functionally equivalent to the input code", unless you meant something by that statement that I don't understand. The reason I'm uncertain is because I wouldn't presume that you are asserting that all deterministically compiled, correct compilers are high-level assembly languages.
I would add to what he writes in the article, however (and I did in a comment there): C doesn't model the von Neumann architecture very well. It lets you create data structures just fine, but the capability to create code, dynamically, at runtime, surely the defining characteristic of a shared instruction-data architecture, is all but completely absent.
The fact that C can be relatively easily and predictably converted into reasonably efficient assembly doesn't make it an assembler.
Never said it was.
And conversely, there are many other languages that have a similar degree of predictability in the code they produce for a given architecture; Pascal, for one, which is almost isomorphic with C in most practical implementations.
You make some good points and there is one which I thought about but didn't touch on because I didn't think it was important.
Consider: most systems and kernel developers do not write C. That is, they do not simply target "the C programming language." Instead, they often target GCC or the Intel C compiler. Even more, they usually target an architecture. In theory one may have to deal with many of the complexities of different implementations, but in practice development is usually restricted (or sometimes duplicated). I think that you may be taking my position a bit too concretely: I don't support the statement that all opcode output is predictable, nor do I support the statement that no other language has elements of the same techniques which make C so useful for systems development.
Simply, view my post as a (possibly slightly exaggerated) relation of the systems C development process from one low-level kernel developer to another. From the birds' eye view (which is the feeling I get from the original post) C may look abstracted and neat, but my experience says that the opposite is in fact true and that things tend to get a lot more dirty in the details.
1. C is used in most contexts where you previously had to use assembly, and would use assembly if C didn't exist.
2. C is the language with the "lowest level" code imaginable.
3. What this means is, you can map almost any command in C into a specific command or a few specific commands in assembly.
That last line is the important part. C is optimized in the sense that every command in C will give you a deterministic amount of commands in assembly. When reading C code, someone who knows assembly can usually tell you what will happen in the compiled code. You won't see an operator which doesn't have a deterministic runtime or memory footprint. In fact, most of the questions on "Why C has this"/"Why C doesn't have that" can be answered exactly like that: you can't implement it with a deterministic set of commands.
In those senses, C is just like Assembly, at least more so than any other language around.
Is that really true? It depends very much on your compiler and on the optimizations that it uses. I'd say that the level of optimization performed is proportional to the similarity between input and output. Straightforward translation makes for crappy optimization.
For example, C won't implement the "raise to nth power" operator (in Python, 210 means 2 to the power of 10, for example), because it can't be done with one assembly instruction, only with a loop.
Even with optimizations, you can reason about your code as if every C command is one Assembly command, and you won't be far off (you can say that every C command is O(1) assembly commands). This is very different from other languages.
C has a division operator, despite the fact that plenty of CPUs (ARM, for a common example) don't have a divide instruction. Typically, the compiler generates a call to a runtime support library (e.g. libgcc) with division functions for various datatypes, and yes, those functions typically involve some looping.
Contrast to, say, Python (as an obviously exaggerated counterexample). Just calling a function in Python means performing a lookup in a hash table.
This is most definitely not true on any modern compiler. The generated assembly depends on the program as a whole and may not correspond fragment to fragment in any readily comprehensible way.
One example is a variable in C may correspond to a number of different storage locations at different points in the program graph, owing to SSA form (or individual optimizations that transform the program in a similar manner).
Another example is how pointer aliasing affects generated code: copying arguments into temporaries before arithmetic can actually result in fewer assembly instructions as the compiler can determine no aliasing is possible.
Add to this the more mundane and well understood optimizations like inlining, constant propagation, later passes of optimizations over the results of the prior and the result is that it's very hard to know exactly what assembly will correspond to any arbitrary part of a c program. Mapping memory locations to variables or assembly instructions to c program lines is not trivial.
I know this seems like nitpicking, but programmers using the mistaken mental model of c being textual macros for assembly leads to poorly performing code at best, and a variety of security vulnerabilities at worst. I think if you actually need to know what's going on at the machine level it's very important to know that the c abstract machine is most definitely NOT what the real machine on your desk is doing.
Do you have any idea how wildly inaccurate that statement is? If you have no experience at all with assembly on any processor, or perhaps only on the x86 family, then I can understand your statement. It's still wildly inaccurate, though.
Some aphorisms like "C has the efficiency and speed of assembler, with the portability and readability of assembler" come to mind, but those are from ancient history in programming years, and usually said by people proficient in both.
There are other assertions (like jwz in the 'Java Sucks' rant), that refer to C as "a PDP-11 assembler that thinks its a language", but again an exaggerated statement made by someone who knows.
I have not met anyone who primarily codes in an interpreted language who think that C is just syntactic sugar for assembly. Maybe because most of them don't program in C.
I can certainly see some amateur programming pundits (or maybe just forum dickheads) regurgitating the lines without understanding them, but there is a world of difference between the examples touched on, and the actual syntactic sugar of recent java releases (generics, unboxing).
And probably you should strip out the 'macro' bit there because in reality that's the pre-processor doing it's work and even though there isn't a C compiler without a pre-processor technically speaking it is not part of the compiler since all it does is output more C.
I disagree. Looping constructs that typically would be macros in an assembler such as 'while' and 'for' are C constructs. Also, C itself has automatic field offset computations.
A macro for subroutine entry could reserve stack space and define offset constants. A macro for subroutine returning releases the extra space.
It could be neat to use, if the processor architecture has sane addressing modes. (I might have seen/done something similar... doesn't feel like a new idea to me. :-) )
(The entry macro might need a simple preprocessor.)
Disclaimer: C and lower was another life, dimly remembered. :-)
s/FIXME/Ruby, or Python, or Perl, or any other language. Where does one draw the line that distinguishes a language from a sophisticated enough assembler?
In java any kind of minimal assignment can explode behind the scenes into a a whole pile of function calls, that would never happen in C, what you see is what you get.
C supports multiple stacks if you tweak setjmp and longjmp just right, and using co-operative multi-threading is possible without any OS support. Not very useful if you really have to do two things at the same time but a lot nicer than interleaving a bunch of code.
And an optimizing compiler is still a transformation of the input according to a given ruleset, you could not get the same level of optimization by just processing the output of the code generation stage of the compiler because you would lose a bunch of higher level information that is invaluable when optimizing the code but you could see it as just another stage in 'transforming' from one language to another without losing any functional bits along the way.
Any half-decent compiler is going to perform a non-trivial transformation of a big switch statement to make it efficient. The expected performance semantics of a switch (something better than O(n) in the number of cases) rules out simplistic iterated jumps in big cases. The programmer expects O(1), or at worst, O(log n).
But more importantly, CPUs generally have many more capabilities than are exposed by C, and this is where the "1:1 correspondence" really falls down. Assemblers generally have disassemblers that you can transparently round-trip through. That's a little harder in C.
Perhaps you meant a injective relation, rather than bijective? But that's a long way short of an assembler.
Hear, hear! Moreover, that has always been the case. As an example, try implementing multiple-word addition in C. On most architectures, you will learn that not having access to a carry bit makes that harder than it could be.
But the main point is that the difference between the generated code and the stuff you write is relatively small, when looking at the assembly that a C compiler generates I have relatively little trouble following the relationship between the two, and I can make reasonable predictions about what will pop out on the other end.
And of course processors are 'richer' than what most C compilers will use, especially when it comes to special instructions that have no equivalent in the C language.
I've worked on a 'decompiler' for the Mark Williams C compiler (yes, that's pretty long ago), and at the time the above still held true, today the boundaries are definitely fuzzier, mostly due to the increased smarts of compiler writers for the optimization stage.
Gcc is clever enough to optimize whole branches of code out of existence if you set it to be aggressive enough and the code was written naively, that's one way of dramatically losing that 1:1 correspondence.
I'm fully convinced that I could sit down with the average C++ program and accurately predict the majority of the generated assembly. It's really not that hard - the compiler doesn't have THAT many instructions, and it only uses them in a limited set of occasions!
-- Ayjay on Fedang #coding
Not a great example of your point. Here is Analog devices Assembly for the ADSP-2100 Family. Assuming Ar is loaded with 2:
sr=LSHIFT ar by 10;
Even without the assumption, no loop is needed for raising 2 to some power.
Multiplying by a power of 2 is easily done in C using the shift operator.