Printable True Bugs Wait Posters
natashenka.ca
natashenka.ca
gcc -ansi -pedantic -Wall
At university I often had people coming to me and asking, "my program compiles correctly, but it keeps crashing / doing weird things; what did I do wrong?". I always said, "add '-ansi -pedantic -Wall' to the compile flags and come back to me when you cleared that wall of warnings down to nothing". Every single time the problem was directly or indirectly caused by the things that were in the warning messages.It's a shame that the default choice of GCC flags make this compiler pretty much useless for work.
The warnings are there to help you. You can choose to ignore them and that's fine, as long as you do it deliberately and with a damn good reason, not because you don't even know they're there.
(I say this because these days, I hear "C89" or "C90", "C99", or "C11". People who say "ANSI" get asked for clarification.)
-Wall is enough to catch most relevant things, throw in -Wextra if you want overkill, not -ansi -pedantic.
I disagree (not "strongly" because your clarification in the next paragraph). That's only true if you don't know what are you doing.
That said, I am a big fan of gcc's -Wall since the late 90s, and I can tell you gcc has improved a lot since then.
gcc -std=c99 -Wall -Wextra -Werror
I'm not sure if -pedantic is still beneficial in that case. The gcc documentation had a list of every warning category, and which parameters enable each warning. This SO question was also interesting: http://stackoverflow.com/questions/11714827/how-to-turn-on-l...Also —analyze to run the static analyzer: http://clang-analyzer.llvm.org
"Your compiler is trying to help you. Listen to it."
warning: C++ style comments are not allowed in ISO C90
warning: ISO C90 forbids mixed declarations and code
-Wall -Wextra gives me: warning: unused parameter ‘type’
For me, -Wextra is much more useful for finding messy code.[1] http://gcc.gnu.org/onlinedocs/gcc-4.3.3/gcc/Diagnostic-Pragm...
Remember, many more of us use the compiler to build packages in the wild, we need it to be fairly permissive to hope to get things to build.
Perhaps, just like condoms, memcpy, strcpy, strcat are dangerous or ineffective 10% of the time. But that other 90% of the time, when used correctly, they are perfectly safe, fun and essential to learning and growing. Avoiding them completely does not create a healthy relationship between a language and its programmer. I realize that not everyone in this thread is a C programmer, but this prevailing attitude of saying "this should never be done/used" just because you're heard that from someone else, is a really petty, annoying, and persistent aspect of programmer culture.
You might be right about the sarcastic message 'wrapping', but these posters only work because the messages they're wrapped around are true.
I certainly didn't come away from those posters going, 'dang, those guys... so over the top; there's nothing wrong with using strcpy, you're making a fuss about nothing'.
I've found bugs in libraries where sprintf was used instead of snprintf and it caused crashes. So I know from experience of cases where snprintf was a better choice than sprintf. And I've certainly coded a few off-by-one errors in my time so I can sure understand it's easy to get the destination buffer size wrong!
I'm not doing anything performance-critical with strings. Under what circumstances should I use sprintf instead of snprintf and what benefits will I see?
As for an example of when you might want to use sprintf. If you compile to C89 (ANSI) snprintf isn't included in the standard so you're left with sprintf if you want your code to be portable.
A better example is strcpy though. A function which is completely normal and safe when used correctly, but someone with the wrong idea might use strncpy as a "safer" version and cause themselves more trouble for not knowing what is going on. As if to prove my point you can see someone advocating this exact thing in another comment tree in this thread. They've probably heard from someone that the "n" version of functions are safe. If they'd actually looked into it they'd realize strncpy doesn't do what you might think.
It's too much trouble for too little gain.
I think the areas where C is a good and effective solution are shrinking and safer languages are becoming more common and faster (even C++)
The C syntax, not its library doesn't allow for good string handling. String handling should be built deeper into the language, and, yes, even though it is possible to have safe C code, it's very hard. So try when it's worth it.
The point is some languages encourage you to do things correctly, others, make it an uphill struggle.
(but your point is totally valid with regard to C in general~)
The reason is it is resiliant to length overflows is that
bstring lengths are bounded above by INT_MAX, instead of
~(size_t)0. So length addition overflows cause a wrap
around of the integer value making them negative causing
balloc() to fail before an erroneous operation can occurr.
I wouldn't touch this library.One can also use the Boehm GC if they feel it necessary.
It's a very conservative garbage collector at the end of the day. It can be tweaked so as to be completely bare bones, and the performance impact is very benign. Rather, memory consumption is its weakness.
There is no stable correspondence between number of malloc calls and number of free calls. I might allocate things in 10 places that get freed in one, or vice-versa. A simple example would be "parse a packet and send the built packet (through a message queue) to handling code."
Naturally there's a tradeoff; if redesigning your code to allow one malloc() to one free() would introduce more bugs in logic than it would solve in memory issues, then it's not worth it.
> though it is possible to have safe C code, it's very hard
A dull knife is pretty safe for most people. It's still possible to shove it into your eye and blind yourself, but other than those extreme cases it won't cause much injury when used in the regular manner. However, it is also extremely inefficient at the purpose it was designed for: cutting things.
There are, of course, safer alternatives to a knife. EMTs use special tools designed to fit a seatbelt or cloth into a small slot and slice through without any risk to a person; they also have specially designed shears which make it difficult to cut flesh, but easily cut through nylon and leather. Utility knives have retractable blades to reduce injury, and other tools are designed to fit specific materials into slots and make cutting people impossible.
All of those are purpose-driven and application-specific solutions, however. For the most high performance and general purpose application, a really sharp fixed-blade knife is still the most precise and efficient tool for the job. When wielded correctly it is still safe and efficient. But the practitioner is not protected from harming themselves; it's expected that they know what they're doing. And really, it's not that hard to learn how to use it correctly.
But I totally get that it's easier to use a dull knife or scissors than learn all about knives, and it gets the job done.
I wouldn't say a hatchet or a machete is a "crappier version" of a knife. Surely a hatchet is much better suited for chopping down a tree, for example. On the other hand, a knife would be very inefficient at the job, even if it could get the job done eventually. However, a hatchet is arguably less safe than a knife for many kinds of jobs, and a knife for hatchet-jobs, etc.
So really I guess my point was the idea of a "safer" language is dumb, because not every tool is "safe" for every job, and not every job is suitable for a "safe" tool.
Patronizing and 100% wrong.
Or do you think there are only idiots developing C code?
How does things like OpenBSD get so secure? With a lot of code revisions by people that are good at catching problems. And even they have problems sometimes.
Your comparison with a knife is false, I can have multiple languages in my development machine without a big burden, as opposed to carry a lot of specialized cutting equipment.
So is Java/Python/Go a dull knife? Let's try something then, create a Web application in C as fast as it's doable with these languages and as safe as them.
... So you're saying that C was designed for string handling?
> When wielded correctly it is still safe and efficient
You do realize that even the most skilled chefs have 'battle-scars', right? Your analogy falls on its face in this respect.
While your point is true, I want to object to your metaphor. I object for safety purposes. In fact, a dull knife is more dangerous than a sharp knife for most anyone who needs to cut things.
You need to press much harder with a dull knife and sometimes even to saw down into the object. You may even need to get a firmer grip on the object you're cutting to counter all that force. Those are very dangerous behaviors. A sharp knife that cuts easily is much, much safer.
The point of the parent is use the correct tool. Most of time people don't need to cut things, they just need to spread some butter.
while(*dest++ = *src++);There's nothing here to prevent a buffer overrun. If your src doesn't end in 0, or your dest is too small, you'll be reading memory you shouldn't and/or obliterating memory past your dest buffer. Your method is pretty much on par with an inlined strcpy.
warning: strcpy() is almost always misused, please use strlcpy()What should I use instead of these things if they're so dangerous?
strcpy -> strncpy (or strlcpy on BSD)
strcat -> strncat (or strlcat on BSD)
sprintf -> snprintf (but still watch out for printf format attacks)
gets -> something else entirely. The man page says "programs should NEVER use gets()".
They told me that... ˚sob˚ but they were wrong.
Programs (often) handle text. Apparently, that pretty much means you're fucked. So what is a reasonable way to write such programs?
Edited to add:
Apparently there is this: http://www.hpl.hp.com/personal/Hans_Boehm/gc/gc_source/cordh...
The problem isn't necessarily the functions themselves, it's coders who make assumptions that don't pan out to be true.
struct string {
char *str;
unsigned length;
} "%.*s"
or possibly even "%*.*s"
if you want it space padded. In principle you can bound your space usage and avoid an snprintf with such constructs; in practice, it's probably better to still use snprintf (if you're using standard-library string functions at all).This is useful for copying a string into a fixed size buffer that is sent over the network, to give an example. It's not what the programmer generally means when using strncpy.
Many programmers are surprised when they learn that strncpy() really writes 1M-strlen("abc") zeros in the 1M char array every time it's called...
This is false.
From the man page: "The strncpy() function is similar, except that at most n bytes of src are copied. Warning: If there is no null byte among the first n bytes of src, the string placed in dest will not be null-terminated."
And the following code prints foo:ar on my system.
char buffer[10];
strcpy(buffer, "foobar");
strcpy(buffer, "foo");
printf("%s:%s\n", buffer, buffer+4);
Edited to add: huh, scratch that. Obvious error in above test :-P. Testing it with strncat like I had meant to, it seems it is in fact padded, not just (possibly) terminated. Interesting, and very worth knowing if you are trying to move a probably small string to a large buffer under time pressure.Not sure what you want to say with the example code. Maybe swap it for strncpy and strlcpy and see whether that matches your expectations?
However, despite the n functions being generally safer, you're still propagating misinformation by touting them as secure alternatives.
The original use of the n functions was to manipulate strings in matters of fixed size arrays. If you don't know what you're doing and just blindly use strncpy() as a strcpy() replacement, you could end up truncating your strings.
OpenBSD's l functions, on the other hand, were specifically designed with security in mind.
* http://natashenka.ca/strcpy/
* http://natashenka.ca/strcat/
* http://natashenka.ca/sprintf/
There are many, but here's one such example from last year: http://osvdb.org/show/osvdb/94336
The people who brought us the [in]famous "Why Pascal is not my favorite language" article would have done well to look at their own glass house.
OTOH, C does make a great portable assembler if you are using it to implement another language (which is exactly what we did at one of my jobs in the early 90s)
... Had to explain to my daughter this morning why I was laughing at the "abstinence" posters ...
For you younger people reading this who have never been exposed to pascel, go dig up that article and scroll down to section 2.1 and just think about that for a minute. Ask yourself 'is this the kind of computing environment I want to work in?'
2.4 - separate compilation was added in Turbo Pascal 4. (one could argue that the result is Modula, rather than Pascal - so be it)
2.2 - initialization of module data was added in Turbo Pascal 5. Yes, "static" data has to be at the module level instead of hidden within individual routines. Bug or feature? (let the jihad/crusade commence...)
As the commentary on http://c2.com/cgi/wiki?WhyPascalIsNotMyFavoriteProgrammingLa... points out, the critique was against the 1981 academic version of Pascal.
I had no clue printf was problematic.
I'm also having a super fun time with the microcorruption challenges. Highly recommend.
Basically, don't let the user control the very first argument (which would allow them to add format specifiers like %d or %x) and you're safe.
You can always say "dont do X and you'll be fine." But that's kind of like saying "don't point a gun at someone" and the gun will be completely harmless. That's the trick, isn' it? That seems to be the point of the "True bugs wait" site at least. It only takes one mistake.
Or are liberals going to start teaching an honest version of their view: "A combination of condom use and abortion when condoms fail, is 100% effective"?
The phrase "only abstinence is 100% effective" is a slogan for teens, not a claim that only abstinence only sex education is 100% effective (which is clearly false).
>There is room to mock it on a personal level, by pointing out that we are willing to engage in far more dangerous activities without 100% guarantees on safety. 'Seatbelts fail, only not driving is 100% effective'.
It depends on how much a person wants to avoid having an abortion or unplanned pregnancy, and how much value they get out of sex.
You're misunderstanding the point -- "only abstinence is 100% effective" is mocked precisely because it is a remarkably bad way to get kids not to have premarital sex.
It's a slogan associated with a moralistic, backwards, prude, stupid and failed sex-ed program. It deserves to be mocked.
>> It depends on how much a person wants to avoid having an abortion or unplanned pregnancy, and how much value they get out of sex.
Hmm, one of the more important drives, if not the primary hard-wired drive in a human animal, gee, I wonder how much value people get out of it?
You don't have to claim that a stupid course of action is perfect to deserve mockery. Following the stupid course of action will suffice.
> It depends on how much a person...
No, the success of abstincence-only education as public policy does not depend on a single person's wants, desires, and incentives. It depends on the wants, desires, and incentives of an existing imperfect population.
Claiming that "our policy would work if only people were moral and rational in X, Y, and Z ways" is irrelevant if people are known to not be moral and rational in X, Y, and Z ways, which is indeed the case with respect to sex-ed. At the end of the day it either works to reduce teen pregnancy or it doesn't, and in this sense abstinence-only education doesn't work.
Effectiveness of any form of birth control is usually quoted as two numbers: an perfect-use rate, and a typical-use rate.
The perfect rate assumes that you are able to follow the instructions exactly every time. For example, perfect use of a hormonal pill means that you take it every single day, at the same time, without ever forgetting; perfect use of condoms means that you use a condom every time you have sex, before beginning sex, that you put it on correctly and that you stop immediately if it breaks; et cetera. Even the much-lambasted "withdrawal" and rhythm methods have really quite good efficacy, assuming you implement them perfectly.[0]
Of course, people aren't robots, especially when it comes to sex, and that's why we have typical-use statistics that reflect the reality of the situation: People forget to take the pill, or take it at the wrong time. People skip using the condom, just this once-- and forget about "pulling out". And people who were abstaining, well, don't.
It may be the case, if we ignore certain unpleasant factors, that abstinence, done perfectly, has 100% efficacy for preventing pregnancy. But we need only to glance at the statistics to see that, in the real world, held to the standard we hold any other procedure to, it is the very worst form of birth control anyone has come up with yet. And that makes your "true statement" nothing more than a bald-faced lie.
[0] Look it up. About 5% failure rates, only twice as bad or so as condoms.
Needless to say, these policies have proven to be counterproductive when it comes to actual teen pregnancy and STI rates.
To the unfortunate consequence of many teenagers.
I wish it wasn't so easy to be a misanthrope.
Evil, but secular.
There's also a string module in the standard library that includes a lot of convenience methods, including ones for interoperating with C and for working with Unicode[1].
For someone like me with a background in dynamic languages like Python and Lua, as well as some background in C, D was a great fit. Unlike when I started learning C++, D felt like a very natural extension of C to include GC and lots of modern language features.
Lamentably, not many people are interested in D, and so aren't yet a lot of third-party libraries, and many of those that are abandoned. But the core language and standard library are great, and there's a nicely growing and incredibly fast web framework (vibe.d). And Facebook has started supporting the language. So hopefully it will start seeing some growth.
As another respondent mentioned, there are other languages that fill that need (C++, D, Rust, Nimrod), basically all of the systems languages that are designed to do C's job easier and safer (but probably not faster).
Also, what's pickle? PS: don't type "man pickle" into the google, you won't like it.
>>> import marshal, cPickle
>>> x = object()
>>> marshal.dumps(x)
Traceback (most recent call last):
File "<stdin>", line 1, in <module>
ValueError: unmarshallable object
>>> cPickle.dumps(x)
'ccopy_reg\n_reconstructor\np1\n(c__builtin__\nobject\np2\ng2\nNtRp3\n.'
That's a bit tongue-in-cheek, but it shows that the marshal serialization only handles a handful (I count 9) of built-in object types.What's "special" about pickle is it's the default recommended serialization method. Quoting http://docs.python.org/2/library/marshal.html:
> "If you’re serializing and de-serializing Python objects, use the pickle module instead – the performance is comparable, version independence is guaranteed, and pickle supports a substantially wider range of objects than marshal."
The resulting photo reference, in addition to being rather NSFW, was not useful for the particularly game we were making.
If you mean the buggy software that uses the unsafe C APIs (and is not careful enough), then, that's exactly what those posters are about.
The 'safe' versions of the functions have their own risks too. No amount of use-this-thing is a replacement for conscientious programming and review.
I wonder what the next pass of my loop will do now, or what will happen when this function returns???
If you allocate a 10-element character array, it can hold strings of up to 9 characters (plus the null terminator). But, the library functions ("runtime") don't know how big it is.
So if you try to copy a 20-character string into that 10-element buffer, it will copy 21 characters. The first 10 will go into the array, and the next 11 characters will go... somewhere else. Depending on where in memory your array and other things are, something important will probably get overwritten with garbage.