Personally I much prefer APIs that will fail safe until you know you need the performance of a less safe one, and can spend the time auditing the code.
Personally I much prefer APIs that will fail safe until you know you need the performance of a less safe one, and can spend the time auditing the code.
> These words make sense. The problem with strlcat and strlcpy is that they assume that it's okay to arbitrarily discard data for the sake of preventing a buffer overflow. The buffer overflow may be prevented, but because data may have been discarded, the program is still incorrect. This is roughly analogous to clamping floating point overflow to DBL_MAX and merrily continuing in the calculation. ;)
This is still a bug in your program, and it can still specifically be a security risk; if you want to cap a buffer, you need to fail the operation, not just silently discard some of the data.
Further down-thread, here is Ulrich's (I will argue 100% correct) rant:
> Dammit, it is not safe. It hides bugs in programs. If a string is too long for an allocated memory block the copying must not simply silently stop. Instead the program must reallocate or signal an error. I can construct you cases where the use of these stupid functions is creating new security problem.
I realize it is fun (and sadly part of the zeitgeist at this point) to rag on Ulrich, but a lot of the things he's said over the years have been quite insightful, and you should at least read the full conversation before judging him.
*((char *) mempcpy (dst, src, n)) = '\0';
is a good solution then?Also, I don't know where this 'srlcpy silently fails' meme comes from. It's your responsibility to check the return value. It only 'silently fails' if you don't check. You can argue about whether it is a good idea or not to make client code responsible for that kind of thing but it is the C philosophy. If you think string copy functions should call abort on failure then, like Theo says in the OP, implement that and submit patches to all those projects using strcpy.
For more on this (and a specific ludicrous example from OpenSSH where someone spent the time to write a comment above the strlcpy about truncation and yet still didn't actually bother to check the return value of the function), please see this response I made to someone else:
For chain copying? Why not.
Define "fail safe". They don't write more than what you want it to write, which could be exactly what you're looking for. Does it hide bugs? Maybe (depends on how you use it), but it's easier to figure out a bug regarding a short string, than trying to figure out why the program is crashing because you have a buffer overflow somewhere in your program.
> Instead the program must reallocate or signal an error
Who is he, to tell us what is and what isn't acceptable behaviour in our programs? There are numerous cases where discarding some parts of a string is perfectly acceptable.
> I can construct you cases where the use of these stupid functions is creating new security problem.
Sure, I can construct such cases for almost every single function in the C std library. This is C, that's the tradeof we make for speed and efficiency.
Also, I disagree C is a tradeoff for speed an efficiency. Runtime speed and efficiency? Development time perhaps? I can easily write parallel code in scala that defeats idiomatic C on both. I use C when I want something portable. I can't use the scala code on a tiny mips device, but I can port the C code if it's reasonably well written. Same goes for any other target, there's nothing so pervasive as C.
I was mostly thinking about about strlcpy and strcpy. The latter indicates no error and can cause a buffer overflow, which IMHO is harder to debug than erronous logic due to a stringer that is shorter than it is supposed to be.
> I get a nice stack trace and core dump to gaze at
I keep forgetting that C actually has stack traces on normal platforms. We lose all information when the program crashes, stupid custom proprietary platform.
> Also, I disagree C is a tradeoff for speed an efficiency
In C, you sacrifice safety (undefined behaviour etc.) for speed. Rust gives you safety, but (acording to themselves) you can't expect Rust to beat C/C++ in performance for certain operations.
> that defeats idiomatic C on both
But has a minimum of three times the memory overhead. I could write a paralell Clojure application that is way easier to reason about in less time than I could write said application in Scala, but I sacrifice memory and speed efficiency. Every language has tradeofs. C sacrifices safety for speed and memory efficiency.
> I can't use the scala code on a tiny mips device
Why not? There is a JVM for mips isn't there? Or don't you have enough memory? JVM languages uses alot more memory than a native program, that doesn't necessarily make JVM-based languages any less portable.
The same .jar file can be executed on Windows, Linux, Mac and more operating systems without a recompile. That's portability that's hard to beat. Considering you can also compile Java applications for iOS now, and of course Android as well, JVM-based languages are remarkably portable, and I'd argue that they are more portable than C.
The reasons for using C IMHO, is that it's fast, memory efficient, doesn't do any magic (like GC and JIT compiling and whatnot), and that it doesn't require warmup to be fast. It is however slower to develop in, and the language is inherently unsafe.
Last but not least, it's important to use the right tool for the right job. I would never write a VM in a JVM-based language (I'm writing one in C now, fun fun fun!). On the flip-side I would never write a GUI application in C.
Ah, what you wrote earlier makes sense in light of this :)
> In C, you sacrifice safety (undefined behaviour etc.) for speed.[...]
You can write slow code in any language, a program isn't necessarily fast because it's been written in C. It's fast because it's been written with runtime efficiency in mind. Undefined behaviour comes from not making unnecessary assumptions on the underlying platform, which aids portability.
> But has a minimum of three times the memory overhead.[...]C sacrifices safety for speed and memory efficiency
Memory efficiency, yes. I, for one, prefer not to rely on a garbage collector and verify correctness with valgrind. You do have to diligently manage memory, or any language that assumes GC may win in the long run. Speed, however, was never a goal of C. You could argue the same thing with assembly, if you do all of the hard work yourself the result can be faster and more efficient than the same application in C. Neither was meant to be fast or efficient on its own, but may be considered as such when compared to higher level languages.
> There is a JVM for mips isn't there?
It exists, but it's proprietary.
> Or don't you have enough memory?
Not by a long shot, usually.
> JVM languages uses alot more memory than a native program, that doesn't necessarily make JVM-based languages any less portable.
Not on principle, but in practise it does.
> The same .jar file can be executed on Windows, Linux, Mac and more operating systems without a recompile. [...] I'd argue that they are more portable than C.
You simply moved the problem from the program itself to the underlying platform. By that logic, I can prove a PPC binary is portable by running it in qemu on x86. As long as there's a PPC or Java bytecode implementation, the hypothetical programs will work. Fact remains that given Java or C source code, the C code will end up being more portable.
> Last but not least, it's important to use the right tool for the right job. I would never write a VM in a JVM-based language [...] never write a GUI application in C.
I absolutely agree with using the right tool for the job, so never say never. Both examples can be perfectly reasonable ways to get a certain job done, we just don't see them very often.
I do disagree regarding comparing the JVM to Quemu though. A PPC binary can't be said to be portable when you're emulating the CPU that it was meant to run on. I do understand the point you're making though, we just happen to disagree.
Scenario 1: a bug in your program leads to incorrect program behavior because strlcpy just did what it was told and the client code didn't detect the error.
Scenario 2: a bug in your program leads to a buffer overflow vulnerability which can potentially lead to remote code execution.
I think I know which set of problems I'd rather have.
Personally, I'm making the assumption that scenario two will lead to a probably exploitable buffer overflow whereas the scenario one might lead to another class of exploitable vulnerability.
By significantly reducing how much time I have to spend focusing specifically on the buffer overflows of scenario two, I have more time to spend on the other logical errors that may lead to vulnerabilities in scenario one. It's not fail safe -- nothing is in the most extreme interpretations of the phrase -- but it is fail safe against one specific failure case, which is valuable.
FWIW, if people want to say "this API still sucks and users still have to be really careful when using it, but it slightly decreases a specific class of error" that's one thing, but that doesn't mean it "fails safe", nor is the panacea that people make it out to be: it is seriously dangerous for people to keep proselytizing that strlcpy is so fundamentally more safe that it can be used with reckless abandon (which, as I point out in the other comments I've made in this thread, is precisely what has happened due to the dogma that "strlcpy is secure and strcpy is not").
Sure. Hell, I apply this to entire languages and prefer C# where you generally have to opt into rather than opt out of corrupting buffer overflows, and where finding the length of a native string literal is an O(1) operation rather than an O(N) operation.
But 3rd party APIs mean you will at some point have to deal with fixed length character buffers both for reading and for writing. Even in C#, if you ever find yourself needing to do native interop, you'll almost certainly eventually run into them.
> nor is the panacea that people make it out to be: it is seriously dangerous for people to keep proselytizing that strlcpy is so fundamentally more safe that it can be used with reckless abandon
I think the linked article shows we have a much bigger problem: people simply do not care about security and use strcpy with reckless abandon. I won't say "don't add strlcpy to your banned.h". I won't say, "don't annotate the documentation to point programmers to better alternatives". I won't say, "make use of the function generate warnings or errors in the default settings of your project, standard library, or compiler configuration".
But I'm unconvinced that strlcpy is ineffective at harm reduction, and thus that Christoph Hellwig's stance is the correct one.
Until about 2 years ago there wasn't really a better solution around...
Thanks.
We then also have to take into consideration that actual stack overflows are nowadays rare with strcpy due to compilers being able to automatically swap in __strcpy_chk whenever sizeof the stack buffer is known at compile-time. Humorously, swapping strcpy out for strlcpy in these cases will take a situation where the program aborts when the buffer overflows (which isn't perfect, of course, as you might be able to use that to DoS part of a system that you need to temporarily go away while you attack something else) and replaces it with strings that randomly truncate their values at different points in your program's logic.
Heap-allocated buffers are still problems, but are also much more difficult to deterministically exploit. However, randomly truncated filenames, usernames, or keynames offer tons of opportunities to mess with program logic: get a system to verify properties of one path (which can be made arbitrarily long using a cyclic symlink) and then access another one, for example (which is incredibly common due to everyone's misconception that you can use a buffer of size MAXPATHLEN to store a complete path...).
These functions do solve a real issue. Of course, people not using the functions correctly should be made aware of that fact, but still, avoiding buffer overflows > bugs due to short strings.
strlcpy(buf, s, sizeof(buf)) >= sizeof(buf)
Rocket science.
Ulrichs reputation is well deserved, he worked very hard on it.
The thing is, though: no one actually checks that return value, as this idea that using this API is what makes things secure has become zeitgeist. This post from OpenBSD is even making this mistake: they care only that people are using the API, not that they are using it correctly (despite pointing out the truncation issue at the end, they are just analyzing the usage, not the correctness). Just using this API is not a magic bullet: if you are using the API incorrectly, your program is still incorrect, and if you are using it correctly you can just use strcpy correctly.
If you don't believe me, this is a really depressing search to perform:
site:opensource.apple.com "strlcpy" -"strlcpy.c" -"strlcpy.s" -"strlcpy.3" -"strlcpy.h"
OpenSSH doesn't even check strlcpy correctly...
317 /* No MIN_SIZEOF here - we absolutely *must not* truncate the
318 * username (XXX - so check for trunc!) */
319 strlcpy(li->username, pw->pw_name, sizeof(li->username));
Yeah, that's great: we "absolutely must not" do something, but it's marked as a XXX?!?!?Yet somehow, even you admit that they have some use because "almost no value" is infinitely more value than "no value". The benefit you describe is exactly what's beneficial. strlcpy does the length calculation, and prevents overflowing the buffer, which is a reasonable improvement over calling strcpy and trusting that previous checks (if they exist at all) are correct. Truncation (unless intended) can never occur when the program is correct, it merely exchanges one vulnerability for another one, which rarely can be exploited.
That noone checks for the return value is trivially false. That Theo only cares about usage and not about correct usage is also false. Just because something isn't in a mail doesn't mean you can come to the reverse conclusion and claim the person doesn't care about a topic not mentioned. Most people tackle bigger problems first, before getting into discussions about writing "perfect" code.
Given you are a perfect programmer (I assume), why don't you check the openssh code you quoted and let us know how you would've handled it? I looked at it, it's part of (portable) utmp logging and will truncate the username from passwd(5) to fit into the utmp structure. Nothing much can be done about it, because the two structures do not agree on common size limits.
Demanding perfection while being unwilling to improve the status quo is hypocritical.
Meanwhile, how many overflows have even the naive conversions of strc=>strl saved?
We should ask you.
I like C, I use it, I use it in my work, but I simply cannot understand why we don't pay more attention to its huge problems to give compatibility with an ancient standard library.
If you don't know what your source is, how can you pass its length to the function? On the other hand, if you know its length, then you should know that it is NUL-terminated (i.e. it is a string), or you should be able to NUL-terminate it.
Perhaps there is some scenario where a helper function could also be used to force NUL-termination on a fixed-size source buffer (say, after fetching input with read()). But I cannot think of one that makes a lot of sense. For instance, having to forcibly NUL-terminate a bufferful of data from read() would imply data loss: you've read too much and now one byte was replaced with a zero. How would your safe function signal this error condition, and how would you react? In the end you'd have to fix the broken code so it leaves one space for the NUL.
By contrast, strcat/cpy generally speaking do not catch broken-code errors (that is difficult anyway); they only implement some correct code.
Also, reading past the end of the input buffer is much less likely to lead to an exploitable hole (unless you count DoS, but then how would your "safe" function react anyway?).
EDIT: One of the goals in OBSD helper functions is simplicity. The rationale being that if doing the right thing is too complicated, the right thing won't be done (lazy programmers!). And, of course, all complications make human error more likely. This is why strlcat and strlcpy are not the solution to any and all string handling ever; sometimes you need something smarter, e.g. for performance. But these functions are great when they're good enough (which is plenty often, if you trust any of the code that actually uses these!).
I think it's actually quite appropriate that the maintainer of the software that (in one form or the other) is present on hundreds of millions of devices is a bit conservative.
While superficially the reasoning may seem to be purely technical, freedom from Drepper is not an unacknowledged feature of these alternatives.
For example he's "famous" for refusing to patch ldd to work with any posix shell instead of bash: https://www.sourceware.org/bugzilla/show_bug.cgi?id=3266
It was recently fixed, but only after more considerate guys took over maintainership...