Massacring C Pointers (2018)
wozniak.ca
wozniak.ca
The subject of pointers was generally believed to be scary among fellow students and many of them bought pretty fat books that were dedicated solely to the topic of pointers.
However, when I reached Chapter 5 of the The C Programming Language (K&R), I found that it did a wonderful job at teaching pointers in merely 31 pages. The chapter opens with this sentence:
> A pointer is a variable that contains the address of a variable.
The exact point at which the whole topic of pointers became crystal clear was when I encountered this sentence in § 5.3 Pointers and Arrays:
> Rather more surprising, at first sight, is the fact that a reference to a[i] can also be written as * (a+i).
Indeed, it was easy to confirm that by compiling and running the following program:
#include <stdio.h>
int main() {
int a[] = {2, 3, 5, 7, 11};
printf("%d\n", *(a + 2));
printf("%d\n", a[2]);
printf("%d\n", 2[a]);
return 0;
}
The output is: 5
5
5
C was the first serious programming language I was learning back then and at that time, I don't think I could have come across a better book than K&R to learn this subject. Like many others, I too feel that this book is a model for technical writing. I wish more technical books were written like this with clear presentation and concise treatment.Doesn't it technically also contain its own address? So it contains the address of another variable and the address of itself. Pointers are hard...
Huh? Then you probably don't understand the meaning of "contains" here. Or what variables are in general. Variables are hard...
You can "get" the address of a variable (or more generally, of any l-value) with the & operator. That operator is nothing else but an "escape" character that says "Dear C, I don't want you to load the current value of that variable. Give me just the address where it lives so I can load it later".
(Same applies to storing instead of loading)
All our bytes of ram are numbered. We read a value from byte number 8 and that value is 17.
"Byte number 8" is a pain to remember or calculate so we give it a label and let the computer keep track. That's a "variable".
What does that 17 mean though? Is it a value we want to use later on? It might be. Or it could be referring to Byte 17 in our ram. Byte 17 happens to have the value 42 so if we use Byte 8 as a pointer that's the value we end up with.
But it's entirely our choice. It's just a number and we can treat it as a value or a pointer to another memory address as we like.
Given the way the C-standard is relatively carefully written, that explanation is a useful lie, but not necessarily true in every implementation of C.
But it is possible to imagine completely exotic CPU architectures, where there is a separate type of register that can only be used for pointers and it is impossible to convert to int; you could still write C compilers for that architecture and be 100% conforming to the standard.
Most processors a pointer just gets stuffed into a register and then that's used by a assembly instruction to read/write a memory location. But it's certainly possible to have systems where it's not so simple.
"Hey!"
(smacks on head)
"It's an int that points to a variable."
"Now get off the computer and take out the trash without talking about garbage collection!"
"but it could be an int32_t and sometimes..."
(smack)
In other words, the statement is incorrect.
'void' is a type, but in C values of the type void are not first class. Ie a function can return a value of type void, but you can't assign it to a variable.
Well, it is a little bit incomplete I think. A pointer contains both the address of a variable and a type. That is why pointer arithmetic works. Without type of pointer, p++ is not calculatable.
The type is implicitly stored as the pointers type. Casting a pointer can change a piece of memory to be interpreted as any type you want. Although the cast might not make logical sense or may even caused undefined behavior in C.
Pointer arithmetic works off of whatever the current type is, as the compiler needs to know how to the memory offsets.
Another way to learn pointers is via assembly. Either learn assembly or just look at the disassembled code. Look under the hood and it'll become clearer.
Of course the best way is to write your own C compiler... but that may be overkill for most.
I did not find that all that surprising, but the first time I saw that a[i] can also be written as i[a], I was very irritated.
Q: "How do you get the size of a file?"
A: "Call open, and then sizeof."
... is the only one I recall, but they all ranged from "awful" to "not even wrong".
Over my years of job hunting I'm slightly disappointed that I've never been given a question from it. :-)
This is all fun and games until you run into this kind of code in production. I once had to do an emergency rewrite of some firmware by someone who did not appreciate the benefit of functions. That one was so bad that I stopped looking at the original code for any reason because I was afraid I'd accidentally adopt one of the author's many, many misinterpretations of the hardware, or of base reality.
These days you can outsource technical tests to online businesses that specialise in that kind of stuff but that wasn’t always an option, particularly if you travel far back enough in time, and hiring contractors to support you during the recruitment process is often seen as too expensive for some businesses.
https://en.m.wikipedia.org/wiki/Header-only
Maybe the author was cargo-culting that?
It's stupid. It messes up `using`. It doesn't do incremental rebuilds. But it does clean builds fast for the release script, and it saves me from setting up cross-compilation on a faster computer.
It's also a really great idea, for the reasons you give. This is how the Firefox tree is built, and it gives massive compile and link time speedups. (We call it "unified builds".)
Some runtime speedups too. I guess that's from better cross-file inlining?
Most of the manufacturers eventually just wrote gcc backends for their architectures, but the era of crappy proprietary compilers went on far too long in many cases.
Another highlight, no main loop. Everything was in an interrupt. This allowed the code to run despite a very broken i2c driver because it was given a low enough priority that other tasks could preempt it while it was hanging indefinitely.
Come to think of it, I refactored some code from a guy who didn't understand procs. He just cut & pasted code 'snippets' (a snippet being about a screenful) a couple of hundred times.
I find it difficult to understand such people.
Everything was done through a small terminal emulator for AS400, 24x80 characters. Every programs were copy pasted from other places.
I eventually found a way to transfer the file out via FTP then upload it back that way I could at least use vim and see what I was writing but I was alone doing that.
I’ve been reprimanded after documenting my code via comments because I was using some vertical space that should have been instead reserved to code (when you only have 24 lines at a time, people become really petty about things).
That was around 2009 :p
Hopefully adoption of python based tooling in many fields could help alleviate this in coming years..
They even introduced first class functions more than a decade ago.
> "A pointer to a function serves to hide the name and source code of that function. Muddying the waters is not normally a purposeful routine in C programming, but with the security placed on software these days, there is an element of misdirection that seems to be growing." (p. 109)
It’s like a book by Calvin’s dad.
If, say, the publisher was Addison-Wesley or Prentice Hall, would they make it into print.
I am sure that this is not the only programming book that made it to publication with a lack of meaningful editorial oversight.
"Prior to publication of the first edition, the manuscript was reviewed by a professional C programmer hired by the publisher. This individual expressed a firm opinion that the book should not be published because “it offers nothing new -- nothing the C programmer cannot obtain from the documentation provided with C compilers by the software companies.”"
The author's claim is that there was a pressing need for a book that just explains C pointers rather than covering the whole language. (It would be interesting to see the rest of what the reviewer wrote.)
"Professional C programmers hate him! A BASIC programmer discovers a clever way to use pointers..."
There is a line between crackpotness and groundbreakness which sometimes is not visible even for deep experts in the subject (though to be fair most of the time you only need high-school science to debunk most stuff).
The reviewer should have been more explicit "this guy doesn't know what he's talking about" rather than "don't publish this book"
More like a book by someone who fell for his tall tales!
This is one of those things that works by accident for some compilers for some platforms, but because it works for the developer, they think it's a brilliant idea.
From what I remember, and I'm taxing my memory here, some of the early Turbo compilers did lay out arguments in this fashion, for DOS, when all optimisations were disabled. Most of the time.
I do recall seeing similar layouts in some of the programs I played around with at the time, and thinking it was a genius idea. For context, I was about seven or eight years old.
Which is not to excuse the author. His description of how this works and why you would want to do it is just wrong.
<varargs.h> was a pre-Standard header that provided a consistent interface for accessing variable arguments. It was superseded by <stdarg.h> in the 1989 ANSI C standard.
gcc in particular dropped support for <varargs.h> a number of years ago.
One of which only seems to praise the delivery time and condition of the book.
(“This is a terrible book ...”)
> Expressions “return a value.”
> “A pointer is a special variable that returns the address of a memory location.”
I don't see how these are clearly wrong? I guess the 'returning' is problematic?
Somehow I feel this was written with an air of superiority, the author doesn't expand on anything. I get that feeling a lot with articles about c.
The second one reads really strange to me, and variables don’t return things.
Your examples are just wrong. "Returning" has a very specific meaning in programming. An expression does not "return a value". A pointer very very much does not "return the address of a memory location".
The code example given was utterly awful. If you cannot see why that code example is almost guaranteed to cause problems, you might not want to continue in this discussion...
Why not? Wouldn't it be especially important to educate these people? Overly ignorant people sure can be annoying, but at least give people some information on the right concepts instead of only derision.
Yes, "returning" is problematic, because that has a very specific meaning in the context of a function call. But even if we forgive the author the inaccurate terminology, it is also an incorrect statement. You can evaluate an expression. But in C in particular, an expression does not necessarily evaluate to a value.
A somewhat better statement might be:
"Expressions can be evaluated. Some expressions produce a result value."
This statement can probably further improved upon; it definitely still does not capture everything that expressions are in C: such as being able to cause side-effects! That might not be evident, but can certainly be important. But then again, I am not writing a book about C.
> “A pointer is a special variable that returns the address of a memory location.”
Indeed "returning" is problematic here. A pointer does not return anything, it points to something.
> Somehow I feel this was written with an air of superiority, the author doesn't expand on anything. I get that feeling a lot with articles about c
If you're going to write a book about C in order to teach programmers how to work with pointers, better explain things with the correct terminology.
If this were a journal, then of course you can forgive the inaccuracies. But it's a published book that supposedly teaches certain concepts to readers. The author of the blog post says as much in their notes:
"This provides some insight into the mind of the author: he's just picking up concepts and terms as he learns about them and tossing them in without any regard for the reader. This book is pretty much his journal — that somehow became a book with two editions."
> the author doesn't expand on anything
The author expands on a lot more than I would have the patience for; see the notes [0]. Why certain things are incorrect, misleading or dangerous does require some C knowledge. The article + notes won't be a place to learn more about C; but neither is the book that's being reviewed.
For example. If they come from a 'toy' programming background where they just call functions that return values, set variables, jump to labels, and do some math... Even if they're experts at that task, the words "evaluating expressions" might likely have no meaning to them at all. You'll have to explain first with concepts they already understand.
In this case, it seems the writer might've come from such a background. (And assumed his contemporaries were in a similar mindset.) :)
If you show the wrong way to do things, it's always good to give a few pointers to a better path so part of your reader base doesn't feel excluded.
Colloquially, an expression is something on the right hand side of `=` or in parentheses, and that returns a value. Most importantly, "returns" implies that some kind of computation may go on, maybe with side effects. I'm aware that the C standard uses the terms differently, but there is usually a difference between the mental model and the C standard, or the parser.
char *combine(s, t)
char *s, *t;
{Here is the "modern" equivalent:
char combine(char s, char* t) {
As an aside it's not always more verbose because you can group parameters by type, eg:
int foo (c1, i1, c2, i2)
char c1, c2;
register int i1, i2;
{
... char *combine(char *s, char *t) char* combine(char* s, char* t) { ... }
which means combine takes in two pointers-to-char and returns a pointer-to-char.Opinions differ on whether the * should be next to the type or next to the identifier. I prefer putting it next to the type.
I mean, short of enabling declarations like:
char *a, *b;
But I have long since found the tradeoff to be worth the syntactic clarity. int const * const i;
This is nice because you can naturally read the signature from right to left. "i" is a constant pointer to a constant integer. It's a little unconventional, but I think it's a really clear way to convey the types. char *
. Some people (me included) find it clearer to write the type, followed by the variable name. So just as you’d write int a
, you’d write char* a
.The fly in this ointment is that C type syntax doesn’t want to work that way. It’s designed to make the definition of the variable look like the use of the variable. A clever idea, but unlike nearly every other language, which BTW is why I think you should really use typedefs for any type at all complicated in C.
For example, the type-then-variable style falls down if you need to declare an array
int foo[4]
or a pointer to a function returning a pointer to an int int *(*a)(void)
(...right?).So I’m perfectly willing to do it the “C way”, I just find out more readable to do it the other way unless it just won’t work (and then prefer to use typedefs to make it work anyway).
Note that this was rethought for Go syntax.
C does this very cute (read: horrifyingly unintuitive) thing where the type reads how it's used. So "char ⋆a" is written so because "⋆a" is a "char", i.e. pointer declarations mimic dereferencing, and similarly, function pointer declarations mimic function application, and array declarations mimic indexing into the array.
It helped than I've learned by K&R C book. Windows API and code examples are horrific, like another language.
https://docs.microsoft.com/en-us/windows/win32/learnwin32/wi...
https://docs.microsoft.com/en-us/windows/win32/learnwin32/ma...
char *a => *a is a char
char a[3] => a[i] is a char
char f(char) => f(c) is a char
char (*f)(char) => (*f)(c) is a char (short form: f(c))sizeof(a[3]) is not evaluating a[3], so it also isn't UB.
char* a, b;
Now a has type char * but b is just char. It’s probably not what the author meant and it’s definitely ambiguous even if it was intentional. Better to write: char b, *a;
Or, if you meant it this way: char *a, *b;
“Well, don’t declare multiple variables on the same line,” you respond. Sure, that’s good advice too. But in mixed, undisciplined, or just old code, it’s an easy trap to fall into. char* a, b;
apply the char* type to both? (That is, why didn't they design it that way?)I assume there was some reason originally, but it's made everything a bit more confusing ever since for a lot of people. :/
Edit: Apparently it's so declaration mirrors use. Not a good enough reason IMO. But plenty of languages have warts and bad choices that get brought forth. I'm a Perl dev, so I speak from experience (even if I think it's not nearly as bad as most people make out).
In the above, you're declaring the types of * a and * b to be char, making a and b pointers to char.
(EDIT: how do you escape * properly inline with other text?)
The difference is the function prototype `void newprint(char *, int);` at the start, which is missing in the second example. With the forward declaration, the compiler knows what arguments newprint takes and errors out if you pass something else. C is compiled top to bottom so in the older version of the example the compiler has no way of knowing what number of arguments the function takes at the point whree it is called. In (not so) old versions of C that implicitly declared a function taking whatever you passed it.
main(int argc, char **argv)
These all produce the same result main(int argc, char* *argv)
main(int argc, char** argv)
main(int argc, char* argv[]) char * a;
declares a variable, `a` that, when dereferenced using the `*` operator, will yield an int.In C++, the same line declares a variable, `a`, of type `pointer-to-int`.
C cuddles the asterisk up to the variable name to reflect use. C++ cuddles it up to the type because it's a part of the type. Opinions don't really differ on whether C-style or C++-style is better, but a lot of cargo-cult programmers don't bother adjusting the style of the code snippet they paste out of Stack Exchange so you see a lot of mixtures.
modern syntax is:
char *Combine(char *S, char *T) {
...
}
which means the Combine function returns a string (char pointer), and takes as arguments two strings, S and T.As of the 2011 standard, that's still the case. I think that C2X will finally remove them.
Because the parameter passing was uniform, you didn't need to inspect anything at the call site. All functions get called the same, so just push your params and call it. Types were for the callee and were optional. This is what powered printf, surely the highest expression of K&R C.
In modern C-lineage style, we enumerate our formal parameters, and variadic functions are awkward and rare. But LISP and JavaScript embrace them; perhaps C could have gone down a different path.
The interesting thing is that with K&R C a function declaration/prototype is optional. That means you can call a function that the compiler has not even seen. Mismatches in parameter/return types (which are optional and default to int in declarations as well) are normally not a problem, because of the aforementioned promotion. If you have the declaration, then the compiler will at least let you know about wrong number of arguments.
But I think it's good to know other pitfalls:
- The "default" return type of a function is int
- A function that does not use void to tell that it takes no argument just have an undefined number of parameters
See this example: https://ideone.com/GfwS4O
You could in theory use the address of a local variable (allocated on the stack) to access the arguments passed to the function directly on the stack.
But this is just madness... Or is it ? Isn't C just assembly with a "nicer" syntax ?
A blast from the comp.lang.c past.
Yikes, that could kill somebody if he takes the same approach... otoh it looks like he was a scuba instructor?
https://trackbill.com/bill/virginia-house-resolution-92-comm...
Bob actually sounds like a guy who had many talents. Writing books perhaps wasn’t one of them. https://www.findagrave.com/memorial/29007415/robert-joseph-t...
[0] https://www.thriftbooks.com/w/complete-reloading-guide_john-...
Honestly, I feel like I am witnessing a class room teasing act with all the comments just resonating. Do we really feel such an urge that if we don't tease this old author (not just the book or code) loud, it will lead a new generations of young programmers astray?
Reminds me of https://m.xkcd.com/1096/
https://b-ok.xyz/book/2368114/60fa20 https://b-ok.xyz/book/2368115/f2ecc8
That's the question of how the author should feel. I'm sure that if he found out about Woz' review, he would dismiss it by saying "who is this Wozniak guy anyways, he probably knows nothing"
> With BASIC, the key thing to know about most implementations at the time is that there were no functions and no scope aside from the global scope.[2]
Take a look at DEF FN in the Applesoft Basic Manual:
https://www.landsnail.com/a2ref.htm
(To be fair, it wasn't used that much, but it was there.)
Also, making suggestive remarks is not very helpful.