Massacring C Pointers
wozniak.ca
wozniak.ca
It seemed like a good way to really try to understand where someone's comprehension broke down. It felt like it added some legitimacy to the students who got the answer wrong, if others could look at it and go "oh, I see why you thought that!" instead of just "wow try to get it right next time". I believe it's part of what good teachers (not limited to school educators, mind you) are doing all the time: looking for the student's gaps and trying to correct those, instead of just repeating the lesson that has already failed to stick.
I guess this just makes me think of that teacher, trying to work out what her pupils were misunderstanding, by looking at their answers.
The amount of students prevents us from real selection, instead we take exercises which failed tests (and that are not empty) and explain what was wrong, identify common mistakes or bad style and try to make it correct.
We had a very good feedback on this. And we saw an evolution in students' code. It was way better than trying to do collection of bad practice ...
I remember trying to help fellow students with their C and C++ assignments back in undergrad, and reading some of their code I got the distinct feeling that most of it was “random-walked” code. Like they tried writing things that looked like C, making random, iterative changes until it compiled. Then once it compiled, made random, iterative changes until the output approached what the expected output would be. Sometimes the functionality ended up right, but with bizarre stuff like: unused variables all over the place, variables assigned then reassigned without using the first value, uses of pointer indirection that look like the student kept adding or removing asterisks until it didn’t crash, elaborate class hierarchies that went unused, etc. It’s like the Infinite Monkeys Eventually Producing Shakespeare thing but with programming assignments.
Oh c'mon, that's practically a right of passage when learning C or C++ for the first time, especially if you're new to programming as a whole :-)
The tutorials about these exercises were really cool, because we'd spend a short time wrapping up mistakes in the notation. But then he usually classified and grouped the different solutions according to the design decisions made and it would evolve into large discussions about tradeoffs of choices.
Someone might go "Well I figured the FSM would be smaller if I fail on the first wrong number, and it'd be user friendly because it fails early", and someone objects "But that's insecure!" and someone else figures "And also it'd be much easier to write up a control circuit that always expects three digits" and eventually everyone is confused if we're looking at the user input as a stream of tuples of three digits, instead of looking at the last three digits in a stream of digits.
It was a very instructive class.
I certainly didn't feel like I was on my own.
Covered everything from the programming to the multivariable calculus and data science (much harder to find students confident enough to teach others for those last two though)
Eh... no. It's not 1975 anymore. This knowledge is widely distributed.
On one of the early homework assignments I realized my professor misunderstood how pointers work - he seemed to believe that (re)assigning pointers created chains rather than changing what the pointer points at. I.e. given "int x,y; int_pointer a,b,c; a=b=c=&x;" he seemed to believe that then executing "c=&y;" would also cause "a" and "b" to point at "y".
I spent the first page of my turned-in assignment excoriating his lack of understanding. Then I presented a class template that implemented a special smart pointer which did behave in the unusual way he seemed to think C pointers work, so that I could write the code exactly as he presented in the assignment and make it actually work.
In retrospect I could have been nicer and considered other pedagogical factors beside technical correctness. I think he took it with more grace than I deserved.
Lastly, and this is taught by another professor, is my class for Linux. It includes bash and C++ programming and focuses a bit on the POSIX API. Our exam was last week and we had 3 parts. The first part was a theory exam by the professor that taught the class. It was about bash/POSIX commands and very little Linux specific stuff. He expected us to know all of the options from all of the commands you can think of: cut, split... You had to know all of them and off the top of your head. It's ridiculous. I'm pretty sure I failed that one, but I absolutely nailed the next part which was about bash scripting (where you could use if statements and the like, the first part did not allow it, only redirection and piping). The man is mad (hah man, that's what I needed during that first exam part).
You are not wrong. This completely breaks my brain. I really would love to hear the professor's explanation for this.
It's like people see inheritance and then forget that containment is a thing.
That happens when you learn OOP from a bad teacher or text book. Far too many books focused entirely on class hierarchies, as if an inheritance diagram was an almost completed program. That leads to a weird kind of brain damage where you think all problem can and must be solved by inheriting from something. It was the STL that taught people that algorithms matter more than taxonomies, but clearly some professors never got the memo.
I will never in a million years understand why some people believe this.
You can use a goto then... (yeah, even worst practice, but as long as it's allowed)
That kind of righteous fury is very common in high school and undergrads. Teachers have to perform in front of a tough crowd.
Certainly true. OTOH, there's hardly anything worse than a teacher with Dunning-Krunger syndrome when it comes to conveying the essentials of low-level programming to newcomers, so I would say that such a reaction is justifiable to a certain extent (as long as the tone wasn't too aggressive).
Of course, it doesn't necessarily have to be the teacher's problem. I remember my first programming teacher in high school explaining in the very first lesson that honestly, he didn't have much programming experience (being a math professor) and had just dabbled a bit with pascal, but [sic] "the administration desperately were looking for someone to teach this course" and he didn't really have a say in it.
You mean he gathered all of his students, colleagues, and superiors; officially admitted to faking knowledge where there was none, and solemnly apologized, promising never again to teach things he doesn't understand?
I somehow doubt it, even though teaching nonsense to eager youth from the position of authority is an offense damn close to sexual molestation of minors. Dealing with subtly (or not so subtly) twisted minds, and holes (and lies) in knowledge, of the corrupted students happens years later and is someone else's problem, but it's all caused by teachers like this. I don't see anything graceful in it, at all.
You really think this?
Wow, that's remarkably offensive.
I just can't understand why teachers are not judged by the same standards that lawyers and doctors are. The damage bad-educators do is very real. It may not look bad on average, but in specific cases, it can be really devastating to students' minds (vulnerable as they are).
Is it because people being taught don't (in general) vote? Or for some other reason? I don't know, but I don't think a bad educator should have an easier time than a bad doctor, who (at least in theory) would be removed from the profession outright if found out. Instead, bad teachers are left alone or sometimes transferred, and that's it. It looks incredibly similar to how some churches handle offenses of their priests.
If you're a teacher, please realize that you're partially responsible for the future life of many of your students, and start acting like it. Removing people who are not qualified (to teach) should be a priority in your school like it is in courts and hospitals. Please, stay cautious and vigilant, and don't let your colleagues tarnish the profession's reputation by teaching things that are provably (and, sometimes, obviously) wrong.
If you insist on conflating it with sexual abuse of children, you are an idiot.
Basically there was a monolith in the middle of the course, a crap textbook, and monkeys were praying to it as the authority on everything.
I spent a couple of months "hacking", just familiarizing myself with the microcontroller and its instruction set, and the tooling. (In fact, I had no source code to look at during this time, as he would not surrender it!) When I finally got his source, I rewrote it to a 100% functional equivalent in two weeks, 1/5 the size without all the unnecessary control path duplication and without pointless register moves. At the time I found the way he had used register moves particularly puzzling because it resembled the way an un-optimized compiler might work.
Months later after I was no longer on that project, I got a call from him. He was very flustered and wanted to know where I "got" a particular sequence of instructions from. I was like, come again? He said, "It's not in the book. You used a sequence of instructions that's not in the Book." (The Microchip programming manual.) He asked about another block of 3 or 4 instructions - also not in the book (in the combination I used them in).
Slowly it dawned on me - I'm quite certain he didn't understand what any individual instruction "did". That whole level level of abstraction didn't exist for him. He programmed in assembly, yes, but only using blocks of example instructions from the Book. Suddenly the pointless register moves made sense - he was acting as a human compiler, without an optimization step.
Years later I realized I should have asked him, "and how do you think the example code in the book was written?"
Great story though, even though it makes me shudder.
Ironically, I think this is one of the best lessons anyone can learn . . .
He did his best but I have taught my self faster than he did.
Though because I was on the CSE stream even though I got a grade 1 they wouldn't let me do computing at A level - as CSE kids where supposed to leave at 16.
What is CSE? Google seems to think it's Certificate of Secondary Education. So, middle school -> early high school equivalent?
Grade 1? A level? Leave where, the school?
(I'm in the US, so that's probably the disconnect)
The O levels where for Kids in grammar and private schools and the CSE's where for the kids who went to secondary modern schools who left at 15/16.
As I was dyslexic I was put in the CSE stream though I did get a grade 1 in maths and computer studies which is the same as a pass at O level I didn't get to do A level computing.
CS courses generally don't have this issue but they do have a large amount of mediocre professors. Reason being anyone talented in CS can make 2-3x the salary out in the real world and so CS grads rarely stick around to teach.
The saying "those who can't do, teach." is as relevant as ever.
> Reason being anyone talented in CS can make 2-3x the salary out in the real world and so CS grads rarely stick around to teach.
I would call bullshit on that, money is not the only thing that motivates people to do things.
It was quite an experience.
Students who wish to explore topics and solve problems at the limits of current human knowledge generally enroll at the University. At these institutions, undergraduates have an opportunity to work in research laboratories and graduate students, postdocs and faculty have the responsibility to obtain grant funding to keep their labs operating. As a courtesy, occasional classes are taught (to get students up to speed on basic theory) but there is an expectation that students are self-driven and can figure things out on their own (hence its ok if the TA is not a native English speaker, etc).
TLDR: If you want to learn a trade, go to vocational school. If you want an education, go to University.
https://www.findagrave.com/memorial/29007415/robert-joseph-t...
> ... The Hacker News army is here.
"Screw You Bob" you do not speak for me. Hopefully, you don't speak for many people on HN.
Words fail me. Can't we let the dead rest in peace?
Maybe I'm just sensitive because I have gotten excited about having learned something and rushed to share my knowledge, only to be cut off at the knees by the more august folks who told me I was in fact doing it wrong. :-)
But also, we've all written bad code, right? Have you ever written code that you've held your nose while writing, working under annoying idiosyncrasies of the problem domain / business requirements, and you just know at some point another developer is going to come along and read that code without having the full context of the situation, and think you're an idiot, and you won't be there to defend yourself?
While it's not quite the same situation that's the kind of empathy I'm feeling here. Who knows the circumstances that led to this book. The author of this blog post admits there was a dearth of material on learning C pointers at the time so I think Traister's heart was in the right place in trying to fill that gap, however misguided.
I wonder if these people ever got to know a bit about the things they didn't understand at the time. I also wonder if I'll ever get to know about the unknown unknowns in my life :)
Rest in peace Mr. Traister, we love you even though C kinda thorned ya. :-)
> The Keil C51 C Compiler works with the LX51 Linker to store function arguments and local variables in fixed memory locations using well-defined names
[0] http://www.keil.com/support/man/docs/bl51/bl51_overlaying.ht...
This is similar to what actually goes on under the hood of the dynamic loader. http://tldp.org/HOWTO/Program-Library-HOWTO/dl-libraries.htm...
The other example was data, with the compiler assisting by changing local variables to have fixed addresses that get carefully reused for different variables at different times.
That compiler uses a run-time call stack.
Can someone please elaborate why it is bad? Are their any good resources to fill gaps in my knowledge?
Thanks in advance.
Edit: Thank you guys for pointing out so many problems. It seems that I have a lot to learn. :)
It copies s and then t to a fixed size buffer, without any checks. That will write to invalid memory (probably smashing the stack) if len(s), len(t) or len(s) + len(t) > 100.
It returns a stack allocated buffer (r) pointer to the caller. The array will be invalid when the function returns, as the automatic variables only live in the function scope (during the call), they are deallocated when the function returns.
To do this right you have various strategies.
1. Allocate a buffer of len(s) + len(t) + 1 with malloc, copy the strings and return it. Have the caller free it when it's done. It can be inapropiate because of the dynamic allocation.
2. Have the caller pass a destination buffer and its size. If you know the char *'s are zero terminated, check if you have space for them in the dest buffer. If not, truncate or error out. Most of the times, this is the prefered solution.
3. Use a static local buffer, and return it to the caller. You may need to truncate the copy too. Not recomended. This is not a good solution as the function also will not be reentrant (unsafe with multiple threads).
You can use libc functions like strncpy (C89+), snprintf (C99+) etc to make a "size-checked" copy with various automatic truncation semantics. You can refer to their man pages for details.
Edit: to the downvoters, please point out what is wrong in the comment.
You can argue that truncation is an error condition but then it ought to notify it somehow, for instance by returning NULL in such a case. And even then it's incoherent with snprintf which doesn't have the same behaviour and does always terminate with '\0' even in case of truncation (assuming non-0 buffer length, of course).
It's just an unnecessary footgun that serves no practical purpose. It would be like a date function that gives you the today's date except on the 4th of December where it replies that it's the 31st of February. Not hard to work around but still broken.
This is because it is not intended to work with the same kind of string that the other str* functions work with (ie. an ordinary null terminated string).
Instead it's supposed to work with fixed-width string fields that pad out values shorter than the field width with nulls. This is how original UNIX directory entries were stored.
See how the name is copied into u.u_dbuf here: https://github.com/hephaex/unix-v6/blob/daa355109625a50e6b10...
The strncpy function writes to the entire buffer. This is important if you will be passing the buffer across a security boundary, for example in a network packet or as a struct copied into a publicly visible file. If trailing bytes are not cleared, then secret data (which happened to be sitting in memory) can get leaked.
I hate this mentality in a lot of C circles that boils down to "there's no bad language, just bad programmers, man up pussy". I like C, I use it a lot, it's one of the first languages I learned and it's been my main "professional" language for more than a decade. Yet I can also see that it has many unnecessary sore points. Having switch not break by default, gets(), array shenanigans, some aliasing rules, the hundreds of completely different meanings for "static", macro hygiene and I could go on... You can say "it's not a big deal and it's not going to change at that point anyway" and sure, I'm not arguing for a revolution, but let's not act that it's an absolutely perfect language and I'm an idiot that doesn't get it for pointing out these issues.
The linux manual for strncpy makes few good notes about its usage and that has a meaning, since it is official and will not change due someone ranting about something again.
Principle of least surprise. Idiot proof. Semantics that match common usage. Call it what you like.
I think the API designer was trying to be parsimonious and not assume too much about what you want to do with the data in your destination buffer. The requirement was to provide buffer overrun protection when the source string is too large, and this function provides that. Beyond that, the requirement that the character array be null-terminated is your decision.
In fact, if you determine that it isn't null-terminated you can do other things besides null terminate it yourself. You might want to provide a helpful error message to consumers of your API rather than truncating their input data silently. Or you may try reallocating your buffer until the string fits. There's actual error handling that you can do with this function. In contrast, automatically null terminating the string makes that more difficult.
The other issues that you have with C seem like preferences. There's nothing wrong with not breaking by default in a switch statement as long as you know that that's what happens.
Can you come up with an example of an API which you would consider to be effectively broken and yet is not actually broken? Presumably, it would be an API that's easier to misuse than strncpy.
I've written about strncpy: http://the-flat-trantor-society.blogspot.com/2012/03/no-strn...
Method 1 is a contract you often make as part of defining the interface. For an example, such things are out of scope so this is a perfectly reasonable point.
Method 2 is common. Most of the strn*() and snprintf() etc, do this. This is the preferred method for some of my colleagues.
Method 3 is used in the standard library, though as mentioned is usually avoided.
It's been years since I have written C professionally (and I only did it for two years), but I felt that snprintf was the single biggest improvement that C99 brought to the table, at least for people that had to deal with strings. ;-)
1. Returns pointer to stack-allocated data, which immediately becomes invalid. Instead, it should be using some sort of allocation (e.g. 'malloc'), or taking in a destination pointer.
2. 'r' is arbitrarily set with length 100. Smaller strings don't need all that space, and larger strings definitely will overrun.
3. The function signature is really awkward. Without any of the surrounding textbook content, I'm not sure what behavior is supposed to happen. At first, I expected something like 'strcat', which takes two char* and appends the second one to the first one. But that isn't happening here and instead it seems to require dynamic allocation. (Hiding allocations inside a function is generally kind of weird. Usually the caller should be responsible for passing in a handle to the destination.)
4. There's no sensical limit on the loop iteration. If the input 't' doesn't have a null terminator, this is going to throw a ton of garbage into the stack space (because 'r' is stack-allocated to a fixed size). And also maybe run for a really long time.
5. 'strcpy' should usually be replaced by 'strncpy', which performs the same function but also requires you to provide a limit ("copy this string, but at most 'n' bytes"). That prevents a class of exploitable errors known as "buffer overruns". I don't know when the 'n' string functions were added to C or became popular, though.
This is a teaching exercise, so the fact that this is implemented as a separate function instead of calling 'strcat' from <string.h> doesn't seem like a big problem.
Sorry to butt in, but this is a bit of a trigger for me: I’ve had to fix a number of programs infected with this idea.
The main problems with strncpy are:
When the source string is shorter than n, strncpy will pad the target to n bytes, filling with zeros. This is bad for performance.
When the source string is longer than n, strncpy will copy n bytes but _not_ nul-terminate the target. So you need extra schenanigans every time you use it to cover this case.
So strncpy is hardly ever a good idea. Sadly there is no standard replacement that is widely accepted. More details at https://en.wikipedia.org/wiki/C_string_handling#Replacements
Before writing to the buffer you should've ensured that it's big enough, and decided what to do if it's not, long before actually doing it. In other words, what happens if it's not big enough? These "always use $length_checking_function" proponents miss that point. Yes, you've avoided an overflow here, but chances are something was already too small long before the flow reached here, and the fix is not to replace an overflow with truncate/not copy/etc. here, but fix the check/sizing that came before elsewhere.
If you planned all this out, you're still making an assertion as to the length. The contract is "give me a string of this length" and if that's not enforced by the compiler, it ought to be enforced at runtime so that the error is detected and dealt with as soon as possible.
So maybe "safe string functions" should really be "fail fast string functions."
True, but practically: when having access to BSD extensions, strl* are used. If only C standard is available, snprintf is preferred. I have seen C libs that will check for strl* availability, and if not, reimplement them using snprintf.
So for portability, snprintf is the way to go. For correctness, and pushing for their extended use, strl* is nice.
#define strlcpy(d, s, n) snprintf(d, n, "%s", s)
Not quite the same (different return type) but close.To be honest, strncpy is barely better in this respect (as a security improvement) - truncating against arbitrary size limit in this day and age of text-only protocols... I wonder if outright crashing at the testing stage would be preferable rather than subtle misbehavior creeping into the release.
Both are bad IMO, the actual required buffer size should be known in advance.
Absolutely agree. Scanning the memory until "we find it", potentially crossing boundaries between segments of memory with different characteristics (caching etc.) just doesn't seem right in general, and if I recall, some CPUs even used to have published errata related to that.
> std::string` and the bafflingly just-introduced `std::string_view` are the right way to handle strings.
I'd even go straight to custom implementation of Hollerith strings. Literals have lengths known at compile time, protocols would either carry the lengths alongside the strings, or be trusted (to have good strlen behavior) until they do, composite strings would compute the length out of the components, etc. This doesn't seem too complex to do, looking from from my bell tower, but I know many people here would frown upon mentioning C++ in the context of embedded development (my area).
Even if it were, there's safety and there's safety. A function (like strncat, for example) that quietly truncates your data if it's too long isn't necessarily better than one that quietly ignores array overruns. Consider what happens if "rm -rf $HOME/tmpdir" is quietly truncated to "rm -rf $HOME/"
Then it's not a string.
It returns a pointer to memory on the stack; but the stack will shrink when the function returns, and that piece of memory will get re-used. Which is consistent with the article's later claim that "I don’t think he understands the call stack."
Ok, so it's reasonable to think that this is supposed to be a teaching example and that considering these concepts might be a bit too early in the process. This leads into the second problem and one that is more subjective: this code is incredibly dense and relies on enough quirks of C that it's almost never going to be clear to a beginner reader what they are supposed to take away from it. It's maybe useful as a quiz question on C syntax and semantics, but there are enough barriers to understanding what the code is supposed to do that the amount of explaining the text would need to do to describe what the code is doing is most likely prohibitively long. Instructive examples should be unambiguous in what they are trying to show, otherwise students will be confused and potentially conflate issues in a way that is difficult to untangle later.
Edit: Ha! I spent so long looking at the first half of the function I totally missed that it was returning r! So, not only does this code have minor issues here and there from its careless implementation, it has a fundamental flaw that, if it were to work, would do so only by accident. I can only imagine that a student might walk away from this example thinking that C functions can return arrays and possibly misunderstand scoping in C.
> Not only does it not check, the design of the function means that there is no possible way for it to check safely. s and t are both pointers to characters? How long are the strings they might represent supposed to be?
strlen
> This leads into the second problem and one that is more subjective: this code is incredibly dense and relies on enough quirks of C that it's almost never going to be clear to a beginner reader what they are supposed to take away from it.
Most "nice" C functions are much terser than this.
Most "nice" C functions are also not used as teaching examples. There's a difference in how one writes C code for production use and instructive use.
void strcpy(char *s, char *t) {
while ((*s++ = *t++) != '\0') ;
} _String == 0 ? 0 : strnlen(_String, _MaxCount) nullCharPtr - str
? Is there something unsafe with that?Sure you can say that good programmers will always ensure a C-string is a C-string but there's decades of programming history that shows that's not true in practice.
* What happens if s or t are longer than 100 characters?
* What happens if s and t are both longer than 50 characters?
* What happens if no element of s == '\0'? How about t?
* r is allocated on the stack. What happens to the memory pointed at by r when you call another function after calling combine? Say you wanted to combine three strings; could you call combine twice to build up the result?
For the third point, you have two main ways of storing strings in general. The first was called pascal style, where the length is stored first in either a byte or two and then the string data with no (or an optional) terminating null character (the '\0'). The second is referred to as c style, and is done by storing the string in a memory and denoting the end of the buffer by a null character.
The c-style enables certain nicer looking c code with loops and such, but it is more dangerous and potentially expensive if you recompute the length all the time. You can always mix and match the way you store the string, such as in C++ where the string class stores a length as well as a null terminated string in a buffer.
For your forth point, yes, you end up corrupting the string leading to more crashes and/or it can be used as an entry point to screw around with your program's stack.
How so?
> I would have used strcpy 2 times
You can't do this, because strcpy doesn't give you the length of the string you copied, which is necessary to put the trailing null byte.
> The strcpy() function copies the string pointed to by src, including the terminating null byte ('\0'), to the buffer pointed to by dest.
But,yeah... It's a massacare.
This is wrong. Postfix increment has a higher precedence than the dereference operator.
Note that postfix increment is at the highest precedence and the dereference operator is at the second highest precedence.
The postfix operators behave in a bit of an interesting way; the expression "t++" evaluates to just "t" but increments t as a side effect. Consequently, it appears to update t after the expression it is part of, even though it is one of the first things evaluated.
If you actually test the function, you will find that it works provided that you make r static or global instead of a local stack variable -- and, of course, carefully mind the restriction on the length of the input strings to avoid overflowing the buffer.
How so?
In one sentence: this guy managed to reinvent a variant of strcat and get it completely wrong.
Here's a somewhat better version:
char *combine(char *s, char *t) {
size_t m = strlen(s), n = strlen(t);
char *ret = malloc(m + n + 1);
if(ret) {
memcpy(ret, s, m);
strcpy(ret + m, t);
}
return ret;
}If s and t are very long and overlapping, m + n could wrap around.
Furthermore, r is a fixed 100 bytes. There is no overflow checking whatsoever.
If you're actually learning C and writing programs in it then the two best things to do would be turn compiler warnings up to 11 and run valgrind often. So you want -Wall and -Wpedantic when you compile and if possible run with valgrind as part of your build script or run it with your tests, just run it often (the longer between runs the harder it is to track back to which change you made).
1 char *combine(s, t)
2 char *s, *t;
3 {
4
5 int x, y;
6 char r[100];
7
8 strcpy(r, s);
9 y = strlen(r);
10 for (x = y; *t != '\0'; ++x)
11 r[x] = *t++;
12
13 r[x] = '\0';
14
15 return(r);
16
17 }
There are several critical memory errors here.1) The function is returning the address of a local variable. This alone makes this function rubbish.
2) The pointers s and t are unknown length, we have no guarantees that concatenating them will fit in a 100 character array. Also, he probably should be using malloc to dynamically allocate the space.
3) Use of the function strcpy rather than strncpy. He should have measured the length of s first, and then use strncpy if s was longer than 99 characters (don't forget the null terminator!). Then, rather than using a for loop, call strncpy again to copy the rest into a safe buffer size (the for loop is rather silly). The reason for this is that strcpy will cheerfully start copying past the array 'boundaries', and in this case, since he's copying into a local variable on the stack, is setting himself up for a Remote Code Execution attack if this ever gets untrusted input.
So those are the critical errors. These tie directly into why Geoff argues that the author doesn't understand the stack.
So let's, for educational purposes, go into this. We're going to go a bit into the weeds here. Sorry about that. This would be easier with a whiteboard. :)
When you fire up a program, the programs machine instructions get copied into memory, let's pretend at the memory location 0x1000. Far away from that code, at the highest memory values (more complicated on modern virtual systems, but hey, let's go back in time here :) ), the computer keeps track of a location called the stack pointer.
I'm going to put forward 3 diagrams now. Please forgive any off by one errors.
(Diagram a)
Registers
A 0
B 0
C 0
SP 0xffff
PC 0x1000
Address Mnemonic DATA
PC0x1000 MOV 1, A 0x00 0x01 0x01 0x01
0x1004 MOV A, C 0x00 0x01 0x03 0x01
0x1008 PUSH 3 0x01 0x00 0x00 0x03
... ...
... ...
... ...
0xfffc XXXXXXXX 0x00 0x00 0x00 0x00 <-- Stack Pointer is here
The program starts at 0x1000, then after executing the first two move (MOV) instructions, the state of the world becomes as follows (Diagram b)
Registers
A 1
B 0
C 1
SP 0xfffe
PC 0x1000
Address Mnemonic DATA
0x1000 MOV 1, A 0x00 0x01 0x01 0x01
0x1004 MOV A, C 0x00 0x01 0x03 0x01
PC0x1008 PUSH 3 0x01 0x00 0x00 0x03
... ...
... ...
... ... v------\
0xfffc XXXXXXXX 0x00 0x00 0x00 0x03 ^-- Stack Pointer is here
When you have code like void function() {
int a = 5;
int b = 2;
return a;
}
void main() {
return function();
}
It'll get turned into something like (I've set a 'break point' at 0x2010) (Diagram c)
Registers
A 5
B 0
C 0
SP 0xfff8
PC 0x2010
Address Mnemonic DATA
# Main starts here
# (Note, in C, there is actually code that gets executed before this)
PC0x1000 PUSH 0x1008 # We want to remember where to return to, so we push it to the stack.
0x1004 JMP 0x2000
0x1008 EXIT A # In this implementation of C, the A register will propagate return values
... ...
# Function 'function' is here
0x2000 PUSH 5 # Local variables go on the stack.
0x2004 PUSH 2
0x2008 MOV [SP+2], A # Locally, we refer to local variables by
# offsets to the stack pointer, so if this function
# were to call itself, the stack would keep growing down
# but these values would be good.
0x200c MOV SP+2, SP # Reset the stack before returning
PC0x2010 JMP #SP # Made up notation. Look at the value of the stack pointer, pop it, and jump to it.
# in x86, this is kinda what RET does.
... ...
... ... ...
... ... ...
0xfff8 XXXXXXXX 0x00 0x00 0x00 0x00SP
0xfffc XXXXXXXX 0x02 0x05 0x10 0x08
Okay! So with the above diagrams in mind, let's recap what goes on the stack. Local variables and return addresses. Each time a function gets called it moves the stack pointer down[1] (to lower memory addresses) to make room for local variables. So after you return from that function, and then call another function (or heck, the same one) that pointer you have that was supposed to be the concatenated string is now going to have it's values overwritten.Furthermore, if the input strings are longer than expected, than they can overwrite values on the stack itself, including the return addres, causing your program to jump to some (if you're lucky) random location in memory.
Honestly, some of the best ways to get intuition for how the stack works, and the things that can go wrong, are CTFS at overthewire.org.
Also, https://microcorruption.com/
http://overthewire.org/wargames/bandit/
http://overthewire.org/wargames/leviathan/
[1] Sorry, 'down' means lower memory addresses, even though the displays of memory layouts always have lower memory addresses "up". :(
The function in question:
char *combine(s, t)
char *s, *t;
{
int x, y;
char r[100];
strcpy(r, s);
y = strlen(r);
for (x = y; *t != '\0'; ++x)
r[x] = *t++;
r[x] = '\0';
return(r);
}
1. The array 'r' is allocated on the stack, and returned from the function. This is bad because 'r' goes out of scope as soon as the function is returned. This function returns a pointer to memory with essentially unknown contents.2. The array 'r' is allocated at a fixed size. This is okay if you know ahead of time that you know this size, however for this function we don't know the lengths of s and t, so the odds that we are allocating the right amount of memory is slim.
3. As a result of points 1 and 2, the strcpy and loop may cause a buffer overflow. This is a class of bug where it is possible to overwrite memory that should be unavailable to us. In this case, if the combined lengths of s and t happen to be equal or greater than 100 characters, we will be overwriting memory that does not belong to r, corrupting it, and potentially crashing the program. This may additionally be as security risk, as buffer overflows can be exploited to execute malicious code.
4. There are no checks to see whether s and t are valid pointers. If they are NULL then the function would generate a segmentation fault. This is a check that is often ignored in cases where it is deemed to potentially hinder performance if the function is used frequently.
saulrh also mentions the case where s or t are not NUL-terminated. This is often considered to be a pre-condition of the function in C, suggesting that providing the function strings that aren't NUL terminated is an issue for the user.
For some comparison the following would be my first cut at the same function. Note I the comments are only for illustrative purposes. I'd omit them in actual code.
char* combine(const char *s, const char *t)
{
size_t slen, tlen;
char *str;
// optional checks for validity (point 4)
if (NULL == s || NULL == t)
{
return NULL;
}
// get lengths of s and t, to calculate allocation size
// also save them to use with memcpy later
slen = strlen(s);
tlen = strlen(t);
// allocate to heap (point 1)
// allocate correct size (point 2)
str = malloc(slen + tlen + 1);
if (NULL == str)
{
return NULL;
}
// use memcpy since we already know the size
memcpy(str, s, slen);
memcpy(str + slen, t, tlen);
str[slen + tlen] = '\0';
return str;
} if (NULL == s)
Ouch.I haven't seen a single case where this abomination actually helped catching the fearsome 'if (s = NULL)' typo. One needs to be a sloppy typist, not paying attention to what they write, not proof-reading the code before committing and ignoring compiler warnings for this disaster of a notation to be even remotely justified.
if (!s)Your statement there just does not read correctly, always have the condition explicit.
The statement '!x' tells me nothing about x, just that I'm expecting a given logical value. The semantic meaning of that value however, I have no idea. Odds are it is probably a NULL or 0, but that doesn't help much because (at least in the linux world) they can have different meanings despite having the same logical value. I cannot differentiate between success (!(x == NULL)) or failure (!(x == 0)) without more context.
The statement 'x == NULL' immediately tells me at a glance that I am dealing with a pointer, and therefore I should pay more attention to how it is used. I know I should now be looking for other patterns of safe/unsafe pointer management that are immediately relevant to the function I am reading right now. There is no need for me to look elsewhere and go on some tangent to find out, then to have to recall what I was doing when I come back later. I can immediately start considering the likelihood of segmentation faults, or other memory management issues. x == NULL also tells me that here I am expecting failure of some kind. The context around that condition should then tell me which kind, eg NULL == alloc_thing() vs NULL == find_thing().
Similarly, the statement 'x == -1' tells me I am working with an integer. This can tell me immediately whether I am expecting success (x == 0), failure (x == -1), or that I should make a mental note of more complex possibilities (THING_WORKED == do_thing()).
Code is a narrative. It is documenting the answers to the questions you are asking while writing it, and it should answer the questions that someone else is asking while reading it. When somebody asks you to explain something, the least helpful thing you can do is answer a plain "yes" or "no". It is better to give context and answer further related questions before they need to be asked. Readability and comprehension of code are no different - any given programmer is asking questions of the code. The more questions they have to ask, the longer it will take them to discover the answers in order to understand the code. Context helps. Shorter is not always better.
If you have to write “== true” everytime that’s your own idiosyncratic hangup. It has nothing to do with explicitness, it’s just redundant. The true test was explicit the moment you wrote if (
I actually recently saw an 'if (x = 0)' or similar get caught in review recently. Time pressure increases, tests are rushed, and authors proof-read the code in their head, and not what's on the screen. These things do happen, and if your personal style preferences get in the way of using a simple trick to save your own time at best case, or multiple other people's time at worst, then you might wish to reconsider your preferences - even if it is at 1 in 1000 odds.
I've seen weirder stuff get through the reviews and compile cleanly, something like "f,()". Typos happen, but (a) it's not a good enough reason to make the code less readable (b) if the code is prone to this sort of errors, just pay closer attention to them during the review phase. Hedging against a single exotic type of mistake that virtually never happens at the expense of code readability is unacceptable.
But, my biggest memory is the book I didn't buy. I once worked with a programmer who wasn't very good, and then I heard he wrote a book. I typed his name into Amazon, and there was his book. It was all about the half-baked concepts he was trying to put into our failing project. (Ultimately canceled because we couldn't ship a very simple product. We couldn't ship it because everyone just wanted to add code generators and additional layers around a database... Instead of learning how to use a database.)
I couldn't get out of that job fast enough.
I can't warn you because I put it to the trash bin long ago.
Hmm, I'm starting to see a pattern. Is it possible BASIC, plus lack of internet back in the day, plus attrocious books are the reasons for truning poeple into terrible programmers? I happen to know only a couple of seniors but without exception their code, no matter what language written in today, is horrible on all fronts. I used to think it was a lack of attention to detail, their lack of wanting to strive for even the tiniest bit of more than just 'good enough for tady'. Possibly stemming from lack of education and lack of continuous self-education. But maybe there's more to it. Maybe they were influenced by a bad book. And/or by a not-so-optimal language like BASIC.
Just as poor programmers came later from Visual Basic, or MSVC++ (somehow we went through a phase were everyone coming for interview with MSVC++ actually had C with classes, and for a while it was a warning sign and standing joke at the place I then worked). Getting people who claimed C++ who actually knew it was pretty challenging around the millennium.
In the 80s just about everyone who had an 8 bit started with BASIC, and an awful lot of them managed to go on to be perfectly acceptable programmers when they moved on to C, C++, assembler, and more recent languages, or even turned out a decent game in something like Blitz Basic on the Amiga. Then again there were some who couldn't move beyond BASIC, because it was simple enough almost everyone could piece something together with it - so even back then it was a warning like PHP can be today.
It's not age, and it's not BASIC - there's plenty of younger folks turning out abysmal code who've never been near it.
I wonder how I'd have turned out if I'd learnt C from one of this guy's books instead of K&R and having a couple of experts handy.
[1] https://github.com/ValveSoftware/source-sdk-2013 (originally released in 2004)
C wasn't much used in MS-DOS, as it was yet another systems language trying to gain the place of Assembly for high performance applications.
In some countries Pascal dialects (mostly TP compatible), Modula-2 reigned, while in others C and C++ were other contenders.
By the time Windows (written in C) became mature enough people started caring about it (3.x), C++ was already having its place via OWL (later VCL) on the Borland side, and Microsoft eventually came up with MFC.
However these were the days when C++ still didn't had a standard (which came in 1998), beyond the C++ARM book, so either you would stick with the compiler framework, or try to minimize language features for better portability.
Additionally on Windows everyone was learning it via Petzold's book, where he takes the approach C compiled with C++, not even "C with Classes".
Also although OWL and VCL were great OOP libraries with nice abstractions, MFC was pretty much a Win32 thin wrapper as its initial internal implementation (AFX) wasn't well received by internal MS employees as not being low level enough over Win32.
So there were lots of issues going on that lead to such cases.
Course back then STL (mostly now the standard library) was still sgi STL, and there was also RogueWave, both still fairly new and on the up.
> It is practically impossible to teach good programming to students that have had a prior exposure to BASIC: as potential programmers they are mentally mutilated beyond hope of regeneration.
--Edsger W. Dijkstra
Though, to be fair, he would likely say that about a lot of mainstream languages today. Iirc he was very fond of Miranda, which in many ways was a precursor to Haskell.FORTRAN —"the infantile disorder"—, by now nearly 20 years old, is hopelessly inadequate for whatever computer application you have in mind today: it is now too clumsy, too risky, and too expensive to use.
PL/I —"the fatal disease"— belongs more to the problem set than to the solution set.
It is practically impossible to teach good programming to students that have had a prior exposure to BASIC: as potential programmers they are mentally mutilated beyond hope of regeneration.
The use of COBOL cripples the mind; its teaching should, therefore, be regarded as a criminal offence.
APL is a mistake, carried through to perfection. It is the language of the future for the programming techniques of the past: it creates a new generation of coding bums.
At this point I'm ready to say "Dijkstra considered harmful" or "the use of Dijkstra quotes cripples the mind, their use in discussions should, therefore, be regarded as a criminal offence."
Agreed. Most of the time, it seems the pithy quotes are the only thing that the quoter knows about Dijkstra.
> You end up having to explain why blanket bans on goto don't make sense
For Dijkstra's GOTO quote specifically, there's a wonderful document by David Tribble named "Go To Statement Considered Harmful: A Retrospective" that does a line-by-line analysis of Dijkstra's paper and explains what Dijkstra's meant in the context of his time and examines where usage of GOTO still makes sense today.
What I find disturbing is that author thought every language was the same save a few keywords, and then publisher was happy to put that crap to print.
I think your conjecture is based on confirmation bias and a bad experience you’ve had with some colleagues.
For people of certain of a certain psychological orientation, there is the additional challenge that having put away your first language, you now think you are a "Programmer (TM)", and learning that second language and learning that you have a number of misconceptions can strike at your very identity. People can get psychologically attached to their misconceptions if it means retaining the illusion that they have mastery.
Nowadays the easiest way to screw this up is to go to a computer science/engineering program that uses just one language. As tempting as it may be from a curriculum simplicity perspective, it's a big mistake. I've interviewed a number of people who think that Java === computing. Not even the "JVM", mind you, but Java, the language, itself. I don't blame Java for this, it's the education. Java itself is not a great lens to understand computer capabilities through, and it's a miserable language to be your lens to understand the general capabilities of programming through, especially 10 years ago. (It's slowly getting better, with easy closures and such, but it's still stuff bolted on the side 20 year later.)
Looking at it from that perspective you can see why 8-bit-era BASIC was even worse than that. It offers a very impoverished view of the computer's capabilities and a very impoverished view of the possibilities of computing. (It was possible to rehabilitate BASIC into at least a passable language; I'm glad I don't have to use Visual Basic to do my job, but it's still light years ahead of the BASICs that still used line numbers, and I've done Real Work (TM) in it, albeit a long time ago.) A 21st-century Java-only programmer is substantially better equipped than a 20th-century 8-bit-era BASIC-only programmer.
(By "8-bit-era", I mean the timeframe, not necessarily the CPU. I'm fairly sure there were BASIC implementations with line numbers and such on non-8-bit-machines, and they'd still be dangerous. But as computers got into the 16- and especially the 32-bit era, even BASIC had to grow up.)
http://www.geocities.ws/slaszcz/mup/prgnames.html
https://github.com/ggnkua/Atari_ST_Sources/blob/master/GFA%2...
Could you elaborate on this? What exactly made you realize that was how/what they thought?
And let me say again that it's not specifically Java. I've seen a couple of people that way with C, for instance, though not in an interview situation.
Using the Windows 'ZeroMemory' macro to assign bitwise zero over a newly declared object, rather than using a constructor like god intended.
In C++, null isn't required to be bitwise zero, so I'm fairly sure nasal demons are possible here. (Do we still have 'effective type' in C++?)
> "GIGO (garbage in, garbage out) is a term coined to describe computer output based on erroneous input. The same applies to a human being." (p. 152) — ???
(like the readers of this book?)
Priceless.
https://www.amazon.com/Conquering-Pointers-Robert-J-Traister...
> I believe that the author thinks that integer constants are stored somewhere in memory. The reason I think this is that earlier there was a strange thing about a "constant being written directly into the program." Later on page 44 there is talk about string constants and "setting aside memory for constants." I'm wondering now…
Yes, most of the book is wrong. In this example, the author probably also presented this idea in a wrong way.
But the author is correct for having the idea that "constant being written directly into the program" (by the compiler!) and "integer constants are stored somewhere in memory", they are correct and make perfect sense. Of course the integer constants and string constants are all allocated and stored somewhere in memory (or somewhere that can be mapped as memory). They are usually known the text segment and data segment.
> …(remember, the array name becomes a pointer when used without the subscripting brackets)"
> "…while a pointer, as always, is a special variable that holds the address of a memory location." (p. 57) — Still wrong, but slightly less wrong.
Good enough IMHO. It is true that an array "name" is a pointer to its base address.
No, it's not. It is true that an expression of array type, when it is not the subject of either the unary-& or sizeof operators, evaluates to a pointer to the array's first element.
sizeof array
gives the size of the whole array, not the size of a pointer. &array
gives the address of the whole array, not the address of a pointer.> array[x] and *(ptr+x) is completely equivalent, so array and ptr is equivalent.
until now. Thanks for the clarification.
array[x] and *(array+x) are indeed equivalent for any identifiers 'array' and 'x' (assuming one of those evaluates as a pointer value, and the other as an integer value; otherwise the code is incorrect). In fact, in this context an actual array is not subject to either unary-& or sizeof operators, so it evaluates to a pointer value, fulfilling the precondition.
This is why "array subscription" also directly works with pointers (i.e. "ptr[x]"), and from the equivalence above follows one of the common useless facts that you can swap the identifiers (i.e. "x[array]").
(This comment is probably confusing enough without saying that "(&array)[x]" is valid code too, but isn't the same thing as those before.)
If "program" refers to the object code unambiguously, I don't think this expression is problematic per se.
Additionally another common domain of clueless writing is computer graphics and the related math. There are so many articles written by enthusiastic people (no doubt) where the information is just adding noise. Finding trustworthy good quality information requires that you know a considerable amount already so you know what is good and what is not (talking about online content here) :)
I think it was this book (https://www.amazon.com/Flights-Fantasy-Programming-Video-Gam...) that taught me 3D programming better than anything else. The code was readable, the maths was well explained and it included sections on how to do things without those newfangled maths co-processors. I'd love to buy a copy now just to see if it really was a good book or if it led me astray.
I see beginners writing code like that all the time, which makes me sad.
There hasn't been a lot of activity on it recently so it's probably still wrong.
[1]: https://developer.mozilla.org/en-US/docs/Web/JavaScript/Guid...
Not that pointers are particularly difficult in the scheme of things, but if someone who doesn't understand them tries to teach them, pointers inevitably do become "difficult"
I do too. Many times I find it much more interesting understanding how the sausage is made than what the actual sausage tastes like.
This is a great meta review. As somebody who's just published a book for tech teams, I need to be acutely aware of my own work to make sure I'm not falling down the same hole: taking a little bit of knowledge and and trying to fluff it out to appear to be a comprehensive body of work.
It's not just the 80s and old coding books. There has been quite a trend over the last decade or so of people over-publishing (I guess that's the term). Promising all sorts of things while delivering on little.
In some ways I think this is okay. The received wisdom is that you don't have to know everything, you just have to know more than the reader and be able to explain how to move them a bit forward. Perhaps the key is attitude. Woz says "..'Ive always found the authors come from a position of earnestness, attempting to draw the best conclusions based on decent principles and what they knew at the time they wrote it..."
A little humility, careful scoping, and honesty in a tech book can go a long ways. (Insert long discussion here about whether people would buy such a book, and how people are much more naturally attracted to books with a strong emotional impact "Make money with C Now!" than they are books that simply try to helpfully explain something without all the glitz) There is a natural tension at work here.
Now I know the even darker truth.
I'd like to read that book.
Even more concerning is the book seems to have some positive reviews on Amazon(!), and just one shredding it.
It did not even contain working example programs, let alone exercises. It was so bad, as the saying goes, it was not even wrong. I still have that book on my shelf as a reminder to not blindly buy the first/cheapest textbook I can find.
The modern equivalent would probably be Arduino experience. I wonder if there are similar examples in books out there about C++ written by someone with only that...
I briefly tried to learn C++ an C in the 90s. I'm somewhat glad I didn't find this book in the library. I think it would have made attempting to learn harder, or given me some bad and dangerous habits.
Made me wonder if the quality of the content in the other 11 books is anything like the one the Woz took apart.
> I’m Geoff Wozniak, just one of those persons on the Internet. My blog is hosted here, but not much else at the moment.
"Regrettably, this insight was obscured by a regrettable poor choice of terminology ("bad" and "good"). The comments, azernik, has enough HN karma to suggest that this error ought be assigned to the casual, off-the-cuff nature of internet commenting, and that this comment is not up to the usual work (i.e. comments) of the author.
y = strlen(r);Can you please add some xBase book review to the mix for more outdated fun and facepalming from the heightened point of view on a hill of three decades next? /s
OK, ignoring the s/Brian/Dennis/ snafu... WTF is a "code drivel"? What are you trying to say here?
Second, your point is irrelevant: It's pointless to write a book about driving cars and fill it with a long rant about how riding motorcycles is so much better. It's a non sequitur, and false advertising. Much like how this book is presented as a good resource for learning C and is, in fact, a horrible example of precisely how little the author understood C.
> Then we can argue that back in those times it was something progressive etc.
Aside from your infelicitous attempt at English, this is wrong: This book was never good, and claiming it was insults the past.
i dont know, awk has its faults -- but blaming it for C?
I read plenty of programming books around that time. They were almost always specific to some platform, but they always said so. I never saw, say, a “Pascal” book that turned out to be completely specific to ORCA/Pascal. It would say what it was for.
Do you often make fun of ridiculous movies from 30s?
We could begin from your post starting this thread, where your implication that its errors are just optimizations suggests that you did not understand what's actually, indisputably wrong with the code from the book. Your comment here about current knowledge strongly suggests that you still do not. There are several posts here that explain exactly why it was simply wrong the instant it was written.
Actually, reading that code was fun for me. I wish I could go so low level these days in the mainstream...