Strlcpy and Strlcat – Consistent, Safe, String Copy and Concatenation (1999) [pdf]
openbsd.org
openbsd.org
The tiny performance hit is usually worth the safety and simplicity. Only one error condition to check, covering both concatenation and copies, sensible 'size' argument behaviour (size includes null terminating byte, null terminating byte is always written)...
Just use snprintf().
Especially since these functions' return values make it hard to use them to step forward in the buffer, so subsequent calls need to walk over the already-inserted characters again and again. That kind of redundant work makes me grit my teeth, which I guess is a sign that I'm a C programmer at heart. :)
Compare:
char buf[1024];
snprintf(buf, sizeof buf, "%s%s", "foo", "bar");
and char buf[1024];
strcpy(buf, "foo");
strlcat(buf, "bar", sizeof buf);
The latter is given a pointer to the base buffer, so it has to walk past the existing characters in order to find the end where the concatenation should start.I also find the "size" arguments confusing and weird. I think it's idiomatic in C to (tightly) associate a size with a buffer pointer, and speaking in terms of "there are n bytes available at this location" makes sense to me. So the prototype really should have been
size_t strlcat_unwind(char *buf, size_t buf_size, const char *src); int xasprintf(char**p,const char *fmt,..) {
int n=0;
char *p2=nullptr;
if(fmt) {
va_list v;
va_start(v);
n=asprintf(&p2,fmt,v);
va_end(v);
}
if(n>=0) {
free(*p);
*p=p2;
}
return n;
}
If you thought asprintf was inefficient, this is worse. But it's very convenient!The snprintf() and vsnprintf() functions will write at most size-1 of the characters printed into the output string (the size'th character then gets the terminating `\0'); if the return value is greater than or equal to the size argument, the string was too short and some of the printed characters were discarded. The output is always null-terminated, unless size is 0.
On a more anecdotal front, I did a quick test on Ideone [2] that seems to behave exactly how I would expect it to (and contradicts your description).
In short: always call snprintf() with the exact size of the buffer. No Obi-Wans.
strcpy(buf, "foo");
sprintf(buf, "%s%s", buf, "bar"); //buf contains "foobar"
strcpy(buf, "foo");
snprintf(buf, sizeof buf, "%s%s", buf, "bar"); //buf contains "bar"Some programs imprudently rely on code such as the following
sprintf(buf, "%s some further text", buf);
to append text to buf. However, the standards explicitly note that the results
are undefined if source and destination buffers overlap when calling sprintf(),
snprintf(), vsprintf(), and vsnprintf(). Depend‐ ing on the version of gcc(1)
used, and the compiler options employed, calls such as the above will not
produce the expected results.The glibc implementation of the functions snprintf() and vsnprintf() conforms to the C99 standard, that is, behaves as described above, since glibc 2.1. Until glibc 2.0.6, they would return -1 when the output was truncated.
- printf(3), Linux man-pages 6.04, 2023-04-01
It need not even be a performance hit. The compiler could recognize the idiom of code such as
snprintf(buffer, bufLen, "%s%s%s", s, t, u);
and optimize it into (example possibly buggy; my C is rusty) b = buffer;
e = b + bufLen:
while((b < e) && (*b++ = *s++));
while((b < e) && (*b++ = *t++));
while((b < e) && (*b++ = *u++));
*b = 0;https://www.pixelbeat.org/programming/gcc/string_buffers.htm...
What next? merncopy?
it should be
char *strxcpy(char *restrict dst, const char *restrict src, size_t len)
{
char *end = memccpy(dst, src, '\0', len);
if (!end && len > 0) { dst[len - 1] = '\0';}
return end;
}
Not the end of the world but just another subtly bugged implementation..This illustrates the issue..
Notice in the code below how it wipes out the dest string at char 0 when we supply buf[1]
if we didn't supply buf[1] the zero gets written at buf[size_t_max]
#include <stdio.h>
#include "string.h"
char *strxcpy(char *restrict dst, const char *restrict src, size_t len) {
char *end = memccpy(dst, src, '\0', len);
if (!end) { dst[len - 1] = '\0';}
return end;
}
int main()
{
char buf[3] = "??";
printf("Hello World%s", buf);
strxcpy(&buf[1], "test", 0);
printf("Hello World%s", buf);
return 0;
}(I mean, of course it doesn't null terminate given its intended usage, but that's what it would need for that claim.)
Those are not valid C-strings! String functions are(have) "undefined behavior" for non string inputs!
Which of those two error messages would you rather see?
sprintf: output buffer is to small
or fopen: cannot open file /usr/bin/superlong-dist-path/208277409874874/fiThe first question is: do I need a copy? Obviously, often, you do, but maybe you can avoid it. Copying data can be costly in terms of performance, and of course, without copying, no need to allocate memory or deal with buffers sizes.
Now, that you are sure that you indeed need a copy, maybe you should consider using memcpy() instead. If you already know the sizes of your strings, then it is a no-brainer. If you don't, maybe doing strlen()/strnlen() first on your input data is a good idea, and if it is too big, handle the error appropriately, and if it is not, then do a memcpy() now that you have the size.
The default behavior of "safe" variants of strcpy() it to truncate the string to the maximum size. It is certainly better than a buffer overflow, but it may not be what you want, just ask your ass
isstant.
If I'm dealing with lots of string manipulations I usually create a struct with a count and an alloc size and a char* pointer and write init, cleanup, extend, and append functions specific for it and use that as a string type. It's less of a mental load for me to deal with reallocating an array and keeping count and alloc updated correctly once than it is knowing that character arrays are being recounted all the time and worrying about exactly where the '\0' is going to end up and if my buffer is large enough for every string I'm dealing with.
I wonder what an alternative world would look like where C had decided to use pascal-strings instead (prefixing the string with its length). We would undoubtedly have argued about the length of the size prefix, but a mistake there is trivial to spot compared to wrong application of str[nls]*cpy(_s)?.
First, you cannot really "argue" about length of the size prefix - it is a part of ABI/calling convention, and unless you have template-like mechanism and overloading/conventions, you will have to use whatever stdlib was compiled with.
And Borland Pascal's stdlib was compiled with 1 byte prefix, making max string length of 255 bytes. This made sense when total RAM was measured in hundreds of kilobytes, but generated so many bugs when a normally short string had to be just a bit longer.
I was somewhat envious of my C-using friends who could put entire help message into a single string. But not envious enough to switch to much slower C tooltchain :)
There were other reasons, too, like using less registers in assembly-level loops (old machines had much fewer registers than modern ones). Or not having to worry about validating strings when reading binary data from disk.
The above is, of course, speculation on an alternate timeline -- but it's certainly consistent and believable. What's interesting is what impact such a path would have elsewhere. I'd expect two other consequences: First, our existing uses of varints (e.g. protobufs) would become faster and more efficient, and would probably appear earlier in the timeline. Second, UTF8 would likely become varint-representation-compatible, giving up the (valuable) ability to tell which byte of a codepoint is being decoded in isolation, in exchange for hardware acceleration of encoded<->codepoint conversion; and likely giving us 2^28 codepoints based on a 4-byte maximum length, rather than arbitrarily choosing a maximum length of three bytes to give 2^21 codepoints.
What else would be different? Hard to say… but I do think that size prefixes could have been an alternate route.
Varints can be useful on disk/network streram, but I've never seen them used for actively modified in-memory storage.
I totally agree that as a general tool, varints have failures. But for (a) encoding pascal-style string lengths; (b) encoding often-small numbers on the wire; and (c) encoding codepoints they seem to apply fine.
Yes, they make it harder to get memory corruption. But that's a very narrow definition of "safe".
If you tell the computer to copy a string, or append to a string, it's wrong, and not safe, to only do a partial copy or concatenation.
It's a huge footgun. Smaller than strcpy(), sure, but "safe" it is not.
I don't really have a better answer for C.
Unfortunely WG14 has proven in 30 years of existence, that it isn't something that they care to fix, and while 3rd party solutions exist, without vocabulary types on the stardard library adoption will never take off.
The set of projects in 2023 that should be done in C are not that big. Even the Linux kernel is getting Rust.
C is still good for embedded stuff, but there tends to be much less string manipulation there than in most other code.
Embedded has enough memory buffer manipulations to be a concern as possible CVE root causes, specially in the IoT age where the S stands for security.
My point is that adding a higher level abstraction for e.g. strings in C runs a strong risk of not being fit for purpose for the people who reach for C.
People code in C because they have very specific requirements. If they could just pick up an opinionated library then it's likely that they don't need to use C in the first place.
That's exactly what I said.
> exactly because of that, C security model must be improved.
I'm all ears. For 75 years we've been trying to perfect the need for control with higher level save abstractions.
Most coding nowadays will gladly sacrifice this control, because the code will run in an operating system, and on a fast CPU. And rightly so.
A lot can be gained (without losing control) just by allowing some features of C++. But still, it's not like std::string is a simple fix to this problem. Some C environments cannot allocate memory, or if they can then they need to have an upper bound in how much memory they will allocate.
It's possible to have a mystring type that satisfies the environments like the ones I (and then you) mentioned, but by their nature (because they chose C) they have specific requirements that it may not be possible to solve the general case for.
That said, is strlcpy() a safety improvement over strcpy() and strncpy()? Well... it's not worse. But I'm not sure "if (s >= strlcpy(dst, src, s)) {...}" is any safer than "if (strlen(src)>=s) {...} strcpy(dst,src)". Both are (unless I made a typo) correct, and both are very easy to make a typo with.
And you need to do this at every call site, or you have a bug. Yes, the strcpy() version is much more likely to be exploitable, but the strlcpy() version at the very least produces incorrect output.
And I would not consider something that produces incorrect output "safe". "So just call it correctly" also applies to strcpy(), so same thing.
But yeah, strncpy() was clearly misdesigned, and probably most people would incorrectly guess that its semantics are that of strlcpy(), the latter semantics making sense.
We are decades away from OSes using memory safe languages across all layers.
In the meantime, the security story of C, C++ and Objective-C also needs continuous improvement, of which bare bones C is the worse.