On the matter of strlcpy/strlcat acceptance by industry
marc.info
marc.info
Personally I much prefer APIs that will fail safe until you know you need the performance of a less safe one, and can spend the time auditing the code.
> These words make sense. The problem with strlcat and strlcpy is that they assume that it's okay to arbitrarily discard data for the sake of preventing a buffer overflow. The buffer overflow may be prevented, but because data may have been discarded, the program is still incorrect. This is roughly analogous to clamping floating point overflow to DBL_MAX and merrily continuing in the calculation. ;)
This is still a bug in your program, and it can still specifically be a security risk; if you want to cap a buffer, you need to fail the operation, not just silently discard some of the data.
Further down-thread, here is Ulrich's (I will argue 100% correct) rant:
> Dammit, it is not safe. It hides bugs in programs. If a string is too long for an allocated memory block the copying must not simply silently stop. Instead the program must reallocate or signal an error. I can construct you cases where the use of these stupid functions is creating new security problem.
I realize it is fun (and sadly part of the zeitgeist at this point) to rag on Ulrich, but a lot of the things he's said over the years have been quite insightful, and you should at least read the full conversation before judging him.
*((char *) mempcpy (dst, src, n)) = '\0';
is a good solution then?Also, I don't know where this 'srlcpy silently fails' meme comes from. It's your responsibility to check the return value. It only 'silently fails' if you don't check. You can argue about whether it is a good idea or not to make client code responsible for that kind of thing but it is the C philosophy. If you think string copy functions should call abort on failure then, like Theo says in the OP, implement that and submit patches to all those projects using strcpy.
For chain copying? Why not.
For more on this (and a specific ludicrous example from OpenSSH where someone spent the time to write a comment above the strlcpy about truncation and yet still didn't actually bother to check the return value of the function), please see this response I made to someone else:
strlcpy(buf, s, sizeof(buf)) >= sizeof(buf)
Rocket science.
Ulrichs reputation is well deserved, he worked very hard on it.
The thing is, though: no one actually checks that return value, as this idea that using this API is what makes things secure has become zeitgeist. This post from OpenBSD is even making this mistake: they care only that people are using the API, not that they are using it correctly (despite pointing out the truncation issue at the end, they are just analyzing the usage, not the correctness). Just using this API is not a magic bullet: if you are using the API incorrectly, your program is still incorrect, and if you are using it correctly you can just use strcpy correctly.
If you don't believe me, this is a really depressing search to perform:
site:opensource.apple.com "strlcpy" -"strlcpy.c" -"strlcpy.s" -"strlcpy.3" -"strlcpy.h"
OpenSSH doesn't even check strlcpy correctly...
317 /* No MIN_SIZEOF here - we absolutely *must not* truncate the
318 * username (XXX - so check for trunc!) */
319 strlcpy(li->username, pw->pw_name, sizeof(li->username));
Yeah, that's great: we "absolutely must not" do something, but it's marked as a XXX?!?!?Yet somehow, even you admit that they have some use because "almost no value" is infinitely more value than "no value". The benefit you describe is exactly what's beneficial. strlcpy does the length calculation, and prevents overflowing the buffer, which is a reasonable improvement over calling strcpy and trusting that previous checks (if they exist at all) are correct. Truncation (unless intended) can never occur when the program is correct, it merely exchanges one vulnerability for another one, which rarely can be exploited.
That noone checks for the return value is trivially false. That Theo only cares about usage and not about correct usage is also false. Just because something isn't in a mail doesn't mean you can come to the reverse conclusion and claim the person doesn't care about a topic not mentioned. Most people tackle bigger problems first, before getting into discussions about writing "perfect" code.
Given you are a perfect programmer (I assume), why don't you check the openssh code you quoted and let us know how you would've handled it? I looked at it, it's part of (portable) utmp logging and will truncate the username from passwd(5) to fit into the utmp structure. Nothing much can be done about it, because the two structures do not agree on common size limits.
Demanding perfection while being unwilling to improve the status quo is hypocritical.
Meanwhile, how many overflows have even the naive conversions of strc=>strl saved?
We should ask you.
Scenario 1: a bug in your program leads to incorrect program behavior because strlcpy just did what it was told and the client code didn't detect the error.
Scenario 2: a bug in your program leads to a buffer overflow vulnerability which can potentially lead to remote code execution.
I think I know which set of problems I'd rather have.
Until about 2 years ago there wasn't really a better solution around...
Personally, I'm making the assumption that scenario two will lead to a probably exploitable buffer overflow whereas the scenario one might lead to another class of exploitable vulnerability.
By significantly reducing how much time I have to spend focusing specifically on the buffer overflows of scenario two, I have more time to spend on the other logical errors that may lead to vulnerabilities in scenario one. It's not fail safe -- nothing is in the most extreme interpretations of the phrase -- but it is fail safe against one specific failure case, which is valuable.
FWIW, if people want to say "this API still sucks and users still have to be really careful when using it, but it slightly decreases a specific class of error" that's one thing, but that doesn't mean it "fails safe", nor is the panacea that people make it out to be: it is seriously dangerous for people to keep proselytizing that strlcpy is so fundamentally more safe that it can be used with reckless abandon (which, as I point out in the other comments I've made in this thread, is precisely what has happened due to the dogma that "strlcpy is secure and strcpy is not").
Sure. Hell, I apply this to entire languages and prefer C# where you generally have to opt into rather than opt out of corrupting buffer overflows, and where finding the length of a native string literal is an O(1) operation rather than an O(N) operation.
But 3rd party APIs mean you will at some point have to deal with fixed length character buffers both for reading and for writing. Even in C#, if you ever find yourself needing to do native interop, you'll almost certainly eventually run into them.
> nor is the panacea that people make it out to be: it is seriously dangerous for people to keep proselytizing that strlcpy is so fundamentally more safe that it can be used with reckless abandon
I think the linked article shows we have a much bigger problem: people simply do not care about security and use strcpy with reckless abandon. I won't say "don't add strlcpy to your banned.h". I won't say, "don't annotate the documentation to point programmers to better alternatives". I won't say, "make use of the function generate warnings or errors in the default settings of your project, standard library, or compiler configuration".
But I'm unconvinced that strlcpy is ineffective at harm reduction, and thus that Christoph Hellwig's stance is the correct one.
Thanks.
We then also have to take into consideration that actual stack overflows are nowadays rare with strcpy due to compilers being able to automatically swap in __strcpy_chk whenever sizeof the stack buffer is known at compile-time. Humorously, swapping strcpy out for strlcpy in these cases will take a situation where the program aborts when the buffer overflows (which isn't perfect, of course, as you might be able to use that to DoS part of a system that you need to temporarily go away while you attack something else) and replaces it with strings that randomly truncate their values at different points in your program's logic.
Heap-allocated buffers are still problems, but are also much more difficult to deterministically exploit. However, randomly truncated filenames, usernames, or keynames offer tons of opportunities to mess with program logic: get a system to verify properties of one path (which can be made arbitrarily long using a cyclic symlink) and then access another one, for example (which is incredibly common due to everyone's misconception that you can use a buffer of size MAXPATHLEN to store a complete path...).
These functions do solve a real issue. Of course, people not using the functions correctly should be made aware of that fact, but still, avoiding buffer overflows > bugs due to short strings.
Define "fail safe". They don't write more than what you want it to write, which could be exactly what you're looking for. Does it hide bugs? Maybe (depends on how you use it), but it's easier to figure out a bug regarding a short string, than trying to figure out why the program is crashing because you have a buffer overflow somewhere in your program.
> Instead the program must reallocate or signal an error
Who is he, to tell us what is and what isn't acceptable behaviour in our programs? There are numerous cases where discarding some parts of a string is perfectly acceptable.
> I can construct you cases where the use of these stupid functions is creating new security problem.
Sure, I can construct such cases for almost every single function in the C std library. This is C, that's the tradeof we make for speed and efficiency.
Also, I disagree C is a tradeoff for speed an efficiency. Runtime speed and efficiency? Development time perhaps? I can easily write parallel code in scala that defeats idiomatic C on both. I use C when I want something portable. I can't use the scala code on a tiny mips device, but I can port the C code if it's reasonably well written. Same goes for any other target, there's nothing so pervasive as C.
I was mostly thinking about about strlcpy and strcpy. The latter indicates no error and can cause a buffer overflow, which IMHO is harder to debug than erronous logic due to a stringer that is shorter than it is supposed to be.
> I get a nice stack trace and core dump to gaze at
I keep forgetting that C actually has stack traces on normal platforms. We lose all information when the program crashes, stupid custom proprietary platform.
> Also, I disagree C is a tradeoff for speed an efficiency
In C, you sacrifice safety (undefined behaviour etc.) for speed. Rust gives you safety, but (acording to themselves) you can't expect Rust to beat C/C++ in performance for certain operations.
> that defeats idiomatic C on both
But has a minimum of three times the memory overhead. I could write a paralell Clojure application that is way easier to reason about in less time than I could write said application in Scala, but I sacrifice memory and speed efficiency. Every language has tradeofs. C sacrifices safety for speed and memory efficiency.
> I can't use the scala code on a tiny mips device
Why not? There is a JVM for mips isn't there? Or don't you have enough memory? JVM languages uses alot more memory than a native program, that doesn't necessarily make JVM-based languages any less portable.
The same .jar file can be executed on Windows, Linux, Mac and more operating systems without a recompile. That's portability that's hard to beat. Considering you can also compile Java applications for iOS now, and of course Android as well, JVM-based languages are remarkably portable, and I'd argue that they are more portable than C.
The reasons for using C IMHO, is that it's fast, memory efficient, doesn't do any magic (like GC and JIT compiling and whatnot), and that it doesn't require warmup to be fast. It is however slower to develop in, and the language is inherently unsafe.
Last but not least, it's important to use the right tool for the right job. I would never write a VM in a JVM-based language (I'm writing one in C now, fun fun fun!). On the flip-side I would never write a GUI application in C.
Ah, what you wrote earlier makes sense in light of this :)
> In C, you sacrifice safety (undefined behaviour etc.) for speed.[...]
You can write slow code in any language, a program isn't necessarily fast because it's been written in C. It's fast because it's been written with runtime efficiency in mind. Undefined behaviour comes from not making unnecessary assumptions on the underlying platform, which aids portability.
> But has a minimum of three times the memory overhead.[...]C sacrifices safety for speed and memory efficiency
Memory efficiency, yes. I, for one, prefer not to rely on a garbage collector and verify correctness with valgrind. You do have to diligently manage memory, or any language that assumes GC may win in the long run. Speed, however, was never a goal of C. You could argue the same thing with assembly, if you do all of the hard work yourself the result can be faster and more efficient than the same application in C. Neither was meant to be fast or efficient on its own, but may be considered as such when compared to higher level languages.
> There is a JVM for mips isn't there?
It exists, but it's proprietary.
> Or don't you have enough memory?
Not by a long shot, usually.
> JVM languages uses alot more memory than a native program, that doesn't necessarily make JVM-based languages any less portable.
Not on principle, but in practise it does.
> The same .jar file can be executed on Windows, Linux, Mac and more operating systems without a recompile. [...] I'd argue that they are more portable than C.
You simply moved the problem from the program itself to the underlying platform. By that logic, I can prove a PPC binary is portable by running it in qemu on x86. As long as there's a PPC or Java bytecode implementation, the hypothetical programs will work. Fact remains that given Java or C source code, the C code will end up being more portable.
> Last but not least, it's important to use the right tool for the right job. I would never write a VM in a JVM-based language [...] never write a GUI application in C.
I absolutely agree with using the right tool for the job, so never say never. Both examples can be perfectly reasonable ways to get a certain job done, we just don't see them very often.
I do disagree regarding comparing the JVM to Quemu though. A PPC binary can't be said to be portable when you're emulating the CPU that it was meant to run on. I do understand the point you're making though, we just happen to disagree.
I think it's actually quite appropriate that the maintainer of the software that (in one form or the other) is present on hundreds of millions of devices is a bit conservative.
While superficially the reasoning may seem to be purely technical, freedom from Drepper is not an unacknowledged feature of these alternatives.
For example he's "famous" for refusing to patch ldd to work with any posix shell instead of bash: https://www.sourceware.org/bugzilla/show_bug.cgi?id=3266
It was recently fixed, but only after more considerate guys took over maintainership...
I like C, I use it, I use it in my work, but I simply cannot understand why we don't pay more attention to its huge problems to give compatibility with an ancient standard library.
If you don't know what your source is, how can you pass its length to the function? On the other hand, if you know its length, then you should know that it is NUL-terminated (i.e. it is a string), or you should be able to NUL-terminate it.
Perhaps there is some scenario where a helper function could also be used to force NUL-termination on a fixed-size source buffer (say, after fetching input with read()). But I cannot think of one that makes a lot of sense. For instance, having to forcibly NUL-terminate a bufferful of data from read() would imply data loss: you've read too much and now one byte was replaced with a zero. How would your safe function signal this error condition, and how would you react? In the end you'd have to fix the broken code so it leaves one space for the NUL.
By contrast, strcat/cpy generally speaking do not catch broken-code errors (that is difficult anyway); they only implement some correct code.
Also, reading past the end of the input buffer is much less likely to lead to an exploitable hole (unless you count DoS, but then how would your "safe" function react anyway?).
EDIT: One of the goals in OBSD helper functions is simplicity. The rationale being that if doing the right thing is too complicated, the right thing won't be done (lazy programmers!). And, of course, all complications make human error more likely. This is why strlcat and strlcpy are not the solution to any and all string handling ever; sometimes you need something smarter, e.g. for performance. But these functions are great when they're good enough (which is plenty often, if you trust any of the code that actually uses these!).
strlcat() has the same overhead issue, and its return value is downright cryptic: it is strlen(source) + min(length, strlen(target)).
When it tries to compute strlen(), ignoring your requested length parameter, this means it will access out-of-bounds memory if that string is not properly null-terminated.
And if you're absolutely paranoid about security, these functions not having both source size and target size parameters gives you a second way to go out-of-bounds.
Overall, they're okay functions, but you can easily do a lot better. And since these aren't universally available, if you're going to declare your own cpy/cat functions, you may as well use a better implementation than strl.
Very true; people have caused serious issues by blindly doing things (supposedly making things secure or else), without bothering to understand what they're doing. However, for the specific (and reasonably common) scenario of dealing with proper C strings, the strl* functions are very simple and easy to use correctly, and their use is easy to audit.
> So say you have a big 20MB XML document, and as you are scanning along, you use strlcpy() to store a node name or value into another string. [..] If you keep doing this over the entire document, your parser will end up with quadratic overhead.
Right. It is going to be ridiculously slow (and that should be obvious). But this isn't the common scenario that strl* family of functions were designed for (and which a lot of people get wrong with standard C string functions). But chances are in code like this C strings' shortcomings would be problematic anyway.
> strlcat() [..] its return value is downright cryptic: it is strlen(source) + min(length, strlen(target)).
Sorry but that is just wrong.
> When it tries to compute strlen(), ignoring your requested length parameter, this means it will access out-of-bounds memory if that string is not properly null-terminated.
It's not "requested length", it's "this many bytes fit in my buffer". These functions are not supposed to validate your input any more than are they supposed to check the validity of your pointers, intent, etc. If you're passing a pointer to something that is not a string (if you don't have a NUL terminator, your array by definition does not contain a string), you have a problem that you shouldn't hope string handling functions to magically cope with.
> if you're absolutely paranoid about security, these functions not having both source size and target size parameters gives you a second way to go out-of-bounds.
Target size ought to be enough for simple code dealing with actual strings. But if what you're dealing with isn't a string (and you need extra reassurances because you're not sure what it is), then how are you going to compute the source size? But if you can just compute the source size, then you can make sure what you have is a real string.
I will give you that. They definitely weren't designed for this scenario. But in practice, people tend to ignore the return value (which is using the functions incorrectly), and if you were used to strncpy (which has its own performance problems), you might wrongly assume that you could use these to pull substrings out as well.
The regrettable part is that it's easy to combine overflow detection with substring copy functionality at the same time.
> Sorry but that is just wrong.
Sorry, but it is not wrong.
See the actual implementation source: http://www.openbsd.org/cgi-bin/cvsweb/src/lib/libc/string/st...
"Returns strlen(src) + MIN(siz, strlen(initial dst))."
Even the official man page says the return value is confusing: "While this may seem somewhat confusing, it was done to make truncation detection simple."
I mean no disrespect, but you're proving my point. Even most proponents of these functions do not understand all of the nuances in how they work. That's a real problem, as if you don't understand strlcat's return value, you can't use it safely. The functionality is too specialized to be general purpose functions.
> It's not "requested length", it's "this many bytes fit in my buffer"
Yes. Sorry if my wording was poor, I'm not really good at explaining things.
> These functions are not supposed to validate your input any more than are they supposed to check the validity of your pointers, intent, etc.
Exactly. They're useful functions for certain purposes. My only point is that they are by no means perfect. There are many use cases where they are not appropriate. Unfortunately people tend to either praise this family as the solution to C memory safety problems, or they dismiss them entirely.
Absolutely use them when they're appropriate, but don't simply use them everywhere and think you're making your program safer. There's performance implications, out-of-bounds memory violations, security issues raised by truncating values without catching it, etc. These are all the programmer's fault for misusing the functions, of course, but safer and/or more versatile functions can be made.
That might be true. It would be very interesting to see some statistics. However, with these functions, it is very easy to check whether the truncation detection is being done or not. Grep is mostly enough, but it can even be automated for most cases.
Also, that is hardly an argument against these functions: any string utility function will inevitably have to deal with errors somehow (or worse, ignore them and continue!).
> if you were used to strncpy (which has its own performance problems), you might wrongly assume that you could use these to pull substrings out as well.
Yes, and sadly I've seen something like that (and worse) in the wild. These are easy to spot though.
> Sorry, but it is not wrong. > > See the actual implementation source: http://www.openbsd.org/cgi-bin/cvsweb/src/lib/libc/string/st.... > > "Returns strlen(src) + MIN(siz, strlen(initial dst))."
Interesting. You're right about the code and its return value. However, the current manual documents it a little bit differently; the weird return value for the case when the initial destination string is not a string is never mentioned (but the fact that NUL-termination won't occur in this case is there). So in practice you can forget the MIN() and say it is strlen(dst) + strlen(src). From my perspective it makes all the sense, because, going back to what I've argued before, a string handling function should at least be allowed to assume you're giving it strings to work with; it's not there to validate your input or check pointers, etc.
> Even the official man page says the return value is confusing: "While this may seem somewhat confusing, it was done to make truncation detection simple."
You quoted a line that is not describing the corner-case return value mentioned above. It is describing the fact that the return value is (assuming actual strings are processed!) the total length of the string that would result from the operation if the destination buffer was large enough to hold it without truncation. There is nothing confusing about that. Perhaps the author initially thought it would be confusing for people who expect it to resemble functions such as read() that return the number of characters actually given to you. But there is nothing to be confused for, and the official man page has been simplified to reflect that:
commit: http://marc.info/?l=openbsd-cvs&m=133338810915771&w=2
current version: http://mdoc.su/o/strlcat
> I mean no disrespect, but you're proving my point. Even most proponents of these functions do not understand all of the nuances in how they work. That's a real problem, as if you don't understand strlcat's return value, you can't use it safely.The correct way to detect truncation with strlcpy and strlcat has always been exactly the same, and documented:
if (strlcat(dst, src, dstsize) >= dstsize) { /* truncated */ }
The fact that I wasn't aware of a corner-case scenario where dst does not contain a string to begin with has in no way impeded my ability to type the above line correctly. The branch will still be taken if the string would be too large for the buffer, but your code is broken and strlcat cannot fix it for you.> The functionality is too specialized to be general purpose functions
Yes, they perform a very specific function; they are not Perl. Yet these are "general purpose" enough that the less secure variants (and not only these as of C11) were introduced in the C standard, and are used everywhere. These functions are convenient, available everywhere, and recognized by any C programmer, and that is why they are used. The OpenBSD functions came to be because of the widespread (mis)use of the C functions, to make it easier to do it right.
> Exactly. They're useful functions for certain purposes. My only point is that they are by no means perfect. There are many use cases where they are not appropriate. Unfortunately people tend to either praise this family as the solution to C memory safety problems, or they dismiss them entirely.
Good to know we are in agreement. :-) Perhaps I interpreted you too strongly when you said one may as well use another implementation; to me that sounded like a suggestion to dismiss these entirely.
> These are all the programmer's fault for misusing the functions, of course, but safer and/or more versatile functions can be made.
I would be interested in a concrete example. More versatile? Yes, that is easy. Safer? Well I'm not so sure about that. Perhaps by trading some correctness. Perhaps in some non-critical user-land application where unbounded allocations are OK and aborting on memory exhaustion is "safe". Simpler and safer? Almost definitely not, but please show me I'm wrong :-)
True. And again I was hesitant to say that, so I hope you didn't take any offense. These functions surprised me as well at first (I was personally bitten by that strlcpy() running full strlen(source) issue in my own XML parser.)
> I would be interested in a concrete example.
Well I don't really have a super-secure version. Honestly 99% of string stuff I do is using a C++ library that keeps track of the allocation size (and also has small-string optimization.)
What I came up with in place of strlcpy/cat, which worked better for my own purposes, is below, which I call strmcpy/cat:
//return = strlen(target)
unsigned strmcpy(char *target, const char *source, unsigned length) {
const char *origin = target;
if(length) { while(*source && --length) *target++ = *source++; *target = 0; }
return target - origin;
}
//return = strlen(target)
unsigned strmcat(char *target, const char *source, unsigned length) {
const char *origin = target;
while(*target && length) target++, length--;
return (target - origin) + strmcpy(target, source, length);
}
Truncation is detected by seeing if source[strmcpy/cat(...)] != 0, and very easy to wrap with another function. And in general, I find strcat functions kind of unpleasant (eg Schlemel the painter), so for that a strpcpy() that writes to a pointer will let you chain those, eg: void strpcpy(char *&target, const char *source, unsigned &length) {
unsigned offset = strmcpy(target, source, length);
target += offset, length -= offset;
}
Now I am sure you can point out a myriad of serious problems with these functions, just as I have for strlcpy/cat.Hell they're probably much worse than strl, but they work for me at least. If nothing else, they are at least easier to read =)
Full info on these functions is at the bottom of this dead article: http://web.archive.org/web/20130121234920/http://byuu.org/ar...
Strcpy_s and strcat_s for the win :) They have much better error semantics, do the right thing upon failure (set dst[0] to '\0' if dst is not NULL and return an error code), and if you are working with C++ in Visual Studio, come with template specializations that can infer the length of the destination string. Which is really handy when the compiler substitutes strcpy and strcat calls with calls to strcpy_s and strcat_s with the appropriate implied length :)
This should be embarassing when 'script kiddie' languages like Python, Ruby, JavaScript and co. do a fantastically better job. It has nothing to do with the level of abstraction and everything to do with poor implementation of the standard library on /every platform/.
And I am supposed to trust 3rd party libraries when these giants make mistakes with both trivial, yet highly important functionality? It makes me cynical that they want to force me to use their NSStrings, Java string or Platform::String^ crap, which I really don't like to trust because they are bad in their own ways too... how can I trust these people with my strings?
C and C++ are ridiculously cross platform and utilitarian, but this area is definitely one of the weakest imo :/
Still, I don't advocate rolling your own, but making things work safely and nicely across platforms is a pain in the arse
i don't like introducing any dependency when I shouldn't have to - using Java strings on Android is not something I've yet had to do but I suspect I will if I want good working functionality because the wcs* functions in the C stdlib in the NDK are crap.
there are a number of performance and aesthetic reasons for not wanting to use java strings as well but they are largely irrelevant imo...
(http://blog.liw.fi/posts/strncpy/ is my rant about this. Others have made the point better, of course.)
std::stringThere are no excuses for not including strl* functions in all standard C libraries, it's just NIH that causes the status quo.
strcpy and friends are broken because they are C string functions that can generate invalid C strings and fail the Principle of Least Astonishment. The strl* variants are safer, easier to debug (it's easier to catch a bogus truncation than a missing \0 that might still execute properly by chance on certain systems/builds but not others) and only requires a couple instructions on top of the existing string copy functions. I have yet to hear a reasonable argument about why they shouldn't be included (no, "horribly inefficient BSD crap" is not a reasonable argument).
In some case this could be even useful, but it's for legacy and sloppy programs that no one bother to really fix.
For anything that requires some minimal level of quality the best option is to make the program to intentionally crash if a buffer overflow condition is detect. Exactly like a "Division By Zero".
EDIT: In your world the kernel could panic because someone tried to open() too long a filename. Or should the application trying to open the file crash instead? And an IRC server would crash because my client sent a message with 513 characters.
I would rather detect these conditions and do the right thing.
It doesn't matter what function you use to copy strings. Hell, you can use strcpy in a way that avoids buffer overflows! If your application is poorly designed, it will be insecure. In this case what we're talking is not doing bounds checking on input before copying data, which will always end up fucking you in one way or another, regardless of what function you use.
Going on and on over this function vs that function is a distraction from the bigger problem, which is the overall poor design of many applications.
Insecure software is broken software; it doesn't work correctly (unless insecurity is an intended feature, such as in a backdoor).
> It doesn't matter what function you use to copy strings.
Technically you're right, you can even roll your own loop. However, the purpose of library functions is to make your (and your users' & co-developers') life easier. The strl* functions provide a simple idiom that just works, and it takes less effort than doing arithmetic prior to strcpy/cat calls. It makes it easy for everyone to check that it does indeed work.
> In this case what we're talking is not doing bounds checking on input before copying data, which will always end up fucking you in one way or another, regardless of what function you use.
The strl* functions are all about doing bounds checking in a simple and compact way while copying data, and they inform you in case the input was too large. They're all about replacing broken (or entirely missing!) hand-rolled bounds checking that you'd write before every strcpy/strcat.
Why don't you try audit code that uses strlcat & strlcpy? Grep for these and check that every invocation is correct. Now do the same for code using strcat, strncpy, and strcpy. Check all the arithmetic (which may take up goodness knows how many lines, spread all over the place, in loops, etc.). After a few hundred of these, you should begin to understand why strlcpy and strlcat exist.
People have done such massive audits before. And misuse of the standard string functions has indeed caused many many bugs. The secure alternatives weren't invented just for the heck of it; there's actual history behind it.
I realize that many out here don't care about auditability, and never spend hours systematically going through hundreds or thousands of invocations in other people's code, checking for correctness of exactly the kind of code that historically has caused lots of easy-to-exploit security problems. I realize that most people here would probably trust themselves to be able to use strcpy, strncpy, and strcat correctly; I would trust myself too. But that is not the point. With the right idioms, I make it easier for other people to verify the correctness of my code, so they don't need to depend on blind trust.
And to be honest, it doesn't take a whole lot of "design" to use strlcat correctly; it is simple. Yes you can misuse it, but I would hold it very likely that most people who actually learn about it, check the manual, and decide to use it, are going to use it correctly.
You don't think that's dangerous? If your method of auditing code is skipping all the strlcpy's because you assume it's being used correctly, you're going to miss the one or two times it's not, and end up with vulnerabilities.
I haven't needed to audit code professionally, but when I do it personally, I don't look for common pitfalls in function use - I look for general scenarios that are dangerous, like working with potentially tainted data. So when I talk about design security, I don't mean what function you use, I mean how it fits into the broader scheme of what the application is trying to do.
That is not at all what I have said, how did you come up with that? I've been arguing that the idiomatic use of strlcpy is easier to audit than any use of strcpy.
> I haven't needed to audit code professionally, but when I do it personally, I don't look for common pitfalls in function use - I look for general scenarios that are dangerous, like working with potentially tainted data. So when I talk about design security, I don't mean what function you use, I mean how it fits into the broader scheme of what the application is trying to do.
What you're doing isn't wrong, but looking for common pitfalls is a very easy way to get a baseline assesment of code quality. You'd be surprised at the amount of problems you can find just by looking for arithmetic overflows near memory allocation, misuse of realloc, and the string handling functions. And if the search turns out lots of such problems, trying to understand the design and "broader scheme" is not very useful (at least for me, because at this point I would've decided the software in its current state is simply not good enough for me).
There's another reason to look for common pitfalls: you have something very concrete to focus on. You can work mechanically. You can work with grep. It really does make a huge difference. And let's not forget: these common pitfalls are common. And commonly, they are exploitable. And they are the things many attackers would also initially look for.
If you "just look" trying to find potential security issues in design or otherwise, but are not looking for anything very specific, you can't be anywhere near as focused, and it will take more time to usefully process the code. If it takes a lot of time, focus is going to get worse with the degrading attention span.. it'll be harder (at least for me) to be certain that the review has actually been thorough and has managed to flag things that need attention.
Oh I do look at the broader design of software, but only after getting the low hanging fruit out of the way.
First of all, all of those software projects rolling their own implementation should consider libbsd (installed on my system as a dependency of the Samba client libs, so it should be on most desktop Linux distros already)
http://libbsd.freedesktop.org/wiki/
Secondly, it's fallacious to point out that lots of people already use these functions and therefore it must be sane and should be encouraged. People used gets() for decades.
https://en.wikipedia.org/wiki/Argumentum_ad_populum
Thirdly, strlcat and srlcpy are awful. The fact that people have already moved from awful to maybe slightly less awful variants actually means they have some will to improve their code, which should be encouraged. strl* should not be encouraged. You should be strung up (pun intended) if you're not using a modern string handling library in 2014.
In fact the Linux kernel uses both strcpy and strlcpy -- strcpy is used more often. Sometimes strlcpy makes more sense than strcpy, other times not. It's not a popularity contest.