Ushering out strlcpy()
lwn.net
lwn.net
struct bytes { size_t size; unsigned char *pointer; };
Seems simple enough... A little structure like that has always been useful in my projects.I find myself agreeing with the glibc maintainer in the extended discussion linked from the article:
https://lwn.net/Articles/612244/
> Correct string handling means that you always know how long your strings are and therefore you can you memcpy (instead of strcpy).
I think he's absolutely right. The str* functions are just worse versions of the mem* functions.
For example, calls to printf can be transformed into calls to puts under certain conditions. The compiler can check + find those, even after optimizations. There are many of these tricks for the str* functions that assume NUL terminated C strings.
Though perhaps if the mem* functions were used to implement a fat strings implementation, many of those might still apply.
Point being; the compiler can help when your strings representation is first class by the language spec.
The compiler can also help with FORTIFY; it can insert compile time length checks in certain cases that can make code safer. This avoids treewide rewrites which are relatively painful to do for the Linux kernel due to its development model, but not impossible. That's another barrier to a new string representation and a library of routines for these.
That said, strscpy is not part of the language spec, so I'm guessing unless it's implemented in terms of language defined functions, it gets neither libcall optimizations nor fortified.
(Also the structure can be passed by value (without a pointer to the structure), which also allow to easily use in other function, e.g. to convert a null-terminated string or Pascal string into this structure, without allocating additional memory.)
No, just calling a conversion function before (that would need heap allocation) would never be accepted by C programmers (for the overhead).
Then there is inertia - how many really would want to port their application to a different string type? Not to mention, all the libraries you're using would also have to been converted.
Of course that also generates a giant foot gun because you might manage to get the old and new strlen to disagree on the size, because one reads the length field and the other searches the \0.
As it is, it will never happen.
Surely they could add better designed structures and functions for dealing with memory.
I'm not saying don't use length fields, I'm saying use nul terminators where possible and use length fields where needed. And, they are not mutually exclusive.
And C doesn't need a standardized length delineated string structure in my opinion. Nul terminators serve the job fine for most Standard APIs (which take only short strings), and can receive length fields as separate function parameters where required.
• On the most widely used architectures, reading a string is much easier if the string is a known length. x86 has its string instructions, ARM has its Load Multiple instructions.
• Even with length-prefixed strings, many uses of short strings are with string literals and so the length does not need to be stored anywhere.
It's inherent to the language. Writing "string" gives you a NUL-terminated string; converting that to another format takes effort.
Interestingly, some Mac OS compilers would let you write "\pABC" to get a structure containing the bytes {3, 'A', 'B', 'C'}. (The "p" stands for Pascal.)
Couldn't we just leave the NUL byte in there and pretend it doesn't exist?
const char *literal = "string literal with NUL byte";
struct bytes text;
text.size = strlen(literal); // strlen doesn't count the NUL terminator
text.pointer = literal;
// I know I discarded the const qualifier up there, but it's just to illustrate
Then a copy(text, other) function would conveniently ignore the entire NUL issue. The copy would not even have the terminating NUL.> Interestingly, some Mac OS compilers would let you write "\pABC" to get a structure containing the bytes {3, 'A', 'B', 'C'}. (The "p" stands for Pascal.)
Pascal strings are nice but their sizes are too limited. Same idea as my structure above but with a uint8_t for length instead of uint64_t = size_t.
Say two highest bits of the counter set the size of the counter field. 00 = 6 remaining bits, 01 = 14 bits (2 bytes), 10 = 30 bits, 11 = 62 bits (8 bytes).
A simple `counter* & 0x3f` would remove the width-setting bits, without any shifts, additions, etc.
This allows small strings to use only 1 byte for the counter, while allowing huge strings that span the entire RAM.
The kernel has perfectly reasonable I/O interfaces.
ssize_t bytes_read = read(file_descriptor, buffer, size);
ssize_t bytes_written = write(file_descriptor, buffer, size);
You always know the length.Well... At least you would always know the length if the standard C library didn't abstract that perfectly good interface away behind stdio just so it could do buffering and return NUL-terminated strings.
It's just like errno. The kernel simply returns a negated error constant on failure. The C standard library takes that sane interface and turns it into a thread local global variable.
We can also simply get rid of all that bloat and just use the system calls directly.
> you are not supposed to use the raw API everywhere
Linux system calls are a stable interface and the entry points are even programming language agnostic. It's okay to use them directly.
> e.g. you if don't quickly stop doing that with the BSD sockets API, you are part of the problem
Yeah it's not a good idea on other operating systems since the system call interfaces are unstable. We have to use their C libraries on those platforms.
If I were to survey the state of modern software development and try to characterize the skills lost compared to decades past, "not enough wrappers" would be nowhere on my list.
Indeed, it is just an input check/"sanitization" issue - just like one carefully checks that a JSON or XML input is well formed, if a protocol spec says that some part is an ASCIIZ string, one has to check that there's indeed a zero byte before the end of the data packet.
This isn't unique to my example though. Traditional C strings have the exact same problem and they do get copied all the time.
Stock GCC doesn't have such option though. I guess it's implemented with a macOS-specific patch?
Of course this was before the days when everyone knows only C can be used to write OSes. /s
At least for string literals, no real effort required :)
I think they aren't "worse versions of the mem* functions"; they are different functions. The "str" functions deal with null-terminated data, "mem" functions deal with data of a specified length, and "strn" deals with whichever is shorter.
Some "str" functions do not have "mem" variants (and vice-versa). For example, there is no "memdup" function.
Well, yes. I say they're worse because the only reason they exist is to deal with this NUL terminator nonsense. The str* functions all reduce to the mem* functions after the string length is computed. To me it's like this:
str_function = mem_function(string, strlen(string) + 1)
There is no need for these functions to exist if we get rid of this NUL terminator business.> Some "str" functions do not have "mem" variants (and vice-versa). For example, there is no "memdup" function.
There could easily be. For example, strdup is essentially strcpy(malloc(strlen(string) + 1), string). A memdup function would be even more efficient because the length is already known: memcpy(malloc(length), source, length).
In the early 1990s, there was a popular shareware string library for C in circulation that added a whole raft of extra strXXX() functions, such as strrtrim() and strend(), to a C runtime library. All of these were useful. Indeed, they were fundamental in some other languages, e.g. some dialects of BASIC with their various string functions like MID$ and RIGHT$. But the standard C library never gained them.
The strXXX() set in the C standard library is, rather, in large part an ad hoc set of useful wrappers around stuff that could be done with assembly language idioms, like REPNE SCASB on x86 instruction sets, that had grown up by 1987. The functions weren't intended to be reducible to memXXX(). They were intended to be reducible to assembly language, or even to compiler intrinsics.
The sad part is that the context here is kernel code, in particular Linux kernel code, where human-readable text interfaces (as opposed to machine-parseable interfaces) are the norm. Whitespace-terminated or LF-terminated strings are the norm, in things like procfs for example, and it is a double irony that the C standard library addresses NUL termination more readily than it does those, and even then provides only an ad hoc collection of NUL-terminated string functions with long-since well-known glaring holes unsuitable for kernel code.
Getting rid of the problem in the way that you suggest would necessitate redesigning a lot of kernel APIs to not be human-readable text that operates in terms of variable length strings terminated by special character values with no explicit length counters. No more redirecting the output of the "echo" command to /proc/something . This is exceedingly unlikely to happen.
struct bytes { size_t size; unsigned char *data; };
now you've got two allocations (or at least two separate memory regions, or at least a pointer wasted assuming it's constant) per dynamically-allocated thing.On the other hand, if you take a more direct mirror of a Pascal string:
struct bytes { size_t size; unsigned char data[]; };
You're back to one memory span but can't reslice it.And of course the worst codebase is when someone uses the first one because they want to keep slicing and someone else uses the second one because they need to save memory / indirections. To support both you end up writing functions that take a separate size and data pointer, and... well, then what's the point?
size_t length = strlen("some string");
It's so common. Might as well memoize it so it's always available with no need to constantly loop through strings which is an O(N) algorithm. So many string algorithms call strlen, often multiple times on the same string. I remember GTA V took 6 minutes to parse a goddamn JSON file because of stuff like this and part of the fix was to store the string lengths.https://nee.lv/2021/02/28/How-I-cut-GTA-Online-loading-times...
So is a length variable really such a big deal? It even fits in a register.
I understand and agree with your one memory span point. Ideally they should be located as close as possible in memory.
struct bytes { size_t size; unsigned char data[]; };
const char *literal = "some text";
size_t length = strlen(literal);
struct bytes *text = malloc(sizeof(*text) + length);
text->size = length;
memcpy(text->data, literal, length); struct bytes { size_t size; unsigned char *data; };
struct bytes { size_t size; unsigned char data[]; };
you want depends heavily on what you're doing with the string(s) - plus other common variations like len+cap instead of just size, SSO, etc.So if you can't standardize the data structure, what's the common interface? A function that takes a pointer and a length - which is what we already have. So everyone in this thread appealing to the C standardization process or stdlib to do something wants instead - what, exactly?
> then what's the point?
There's no point. Moreover, passing the length separately is annoying and error-prone.You can have both your models in a single type. That's what my buffer lib achieves : SSO + slicing.
Also some major impl are 24 or even 32 bytes. With a generous SSO you catch a lot of strings w/o overhead.
MSVC 32
GCC 32
Clang 24
SDS is not typesafe, no SSO, no slices.In order to allow for the pointer to live on stack and minimize copying, the data is also reference counted and the compiler takes care of inserting the necessary reference counting calls where needed.
Overall it's pretty flexible, but the reference counting means it's not ideal to use shared strings in heavily threaded code. Of course the second a thread modifies a string, a new string is allocated and that thread can happily work on its "own" string.
Anyway, just yet another way of implementing strings.
Having a size_t on the stack is hardly an issue, it's what every low-level modern language does. It's fast, convenient, and pretty efficient. It also doesn't require deref'ing to get the length, which is a pretty common use case (e.g. checking if a string is empty, or too big, or something along those lines).
I would be also worried about any difficulties separating the length and data causes for prefetching/cache lines though.
The second form is also more amenable to SSO; I'm not sure how often that would come up in the kernel but it's saved me a decent chunk of memory in at least one past project. (Still today I'm sometimes frustated by Go `string` porting from Java `String`, like great now I don't have to pay a boxing overhead but if it's often absent/empty my base size is now 2x what it would otherwise be...)
Per Todd C. Miller (creator of strlcpy and maintainer of sudo) and Theo de Raadt (OpenBSD) in these 1999 (PostScript) slides, the simplest implementation of strlcpy() uses (can use) memcpy():
size_t strlcpy(char *dst, const char *src, size_t siz)
{
size_t n;
size_t slen = strlen(src);
if (siz) {
if ((n = MIN(slen, siz - 1)))
memcpy(dst, src, n);
dst[n] = ’\0’;
}
return(slen);
}
* https://www.openbsd.org/papers/strlcpy-slides.psThe current OpenBSD is currently different:
* https://github.com/openbsd/src/blob/master/lib/libc/string/s...
* https://cvsweb.openbsd.org/src/lib/libc/string/strlcpy.c
One problem with memcpy() is that it returns (void *), so there is no way to know if you've truncated things. From the USENIX paper on strlcpy():
> The strlcpy() and strlcat() functions return the total length of the string they tried to create. For strlcpy() that is simply the length of the source; for strlcat() that means the length of the destination (before concatenation) plus the length of the source. To check for truncation, the programmer need only verify that the return value is less than the size parameter. Thus, if truncation has occurred, the number of bytes needed to store the entire string is now known and the programmer may allocate more space and re-copy the strings if he or she wishes.
* https://www.usenix.org/legacy/events/usenix99/full_papers/mi...
I understand the topic is strXxx() funcs which are ascii only, but it does need to be said that size!=len for wide and multi char sets.
Honestly "string" is a very harmful word that we've all grown used to. As an abstraction it sits somewhere between raw bytes and properly encoded text with proper unicode functions such as those provided by ICU. Python 3 finally forced people to start thinking about this stuff and nobody liked it.
Bytes are relevant when I have to allocate memory otherwise some definition of "character" is often more relevant. Even if I trim text to fit in a buffer I don't want to trim inside a "character" but get the most number of fitting "characters" Now "characters" are of course complicated as grapheme clusters are what is useful the most for human interaction ... but those are quite out of scope for a "simple" string library ...
The Linux kernel has data structure libraries for things like linked lists and maps. Does it really not have a higher-level string abstraction?
https://stackoverflow.com/questions/20935769/sse42-sttni-pcm...
Sometimes you just need to do the boring work so it's done. The Linux kernel is one of the most important pieces of software on the planet - limiting it's performance and safety due to C's string handling legacy is madness.
* https://github.com/openbsd/src/blob/master/lib/libc/string/s...
(For a more detailed discussion, I wrote a whole blog post about why strncpy is good and it got posted here a while back: https://news.ycombinator.com/item?id=27537900)
TBF that was definitely incorrect, as strncpy was never intended to work with C strings. Most (though not all, that would be too easy) of the n functions work with fixed-size strings. The behaviour of strncpy makes perfect sense then.
memcpy is more useful if your fixed-size strings are e.g. space-padded. Obviously it also works if you absolutely know your input is already a fixed-size string, but strncpy will work on both fixed-size and null-terminated inputs.
An "anti-pattern" I seem to be seeing increasingly often in newer code is that of allocating a string (or worse, a dynamically expanding buffer), copying several other strings to it, then only calling another function with that concatenated copy before freeing it, when the other function could've simply been called several times with the individual parts successively.
It should have just returned the number of bytes copied to dst if the string was successfully copied and a -1 with an errno of E2BIG if the dst was too small for the src. It would still do the copy and termination of course, and the programmer will know the length in this case because they specified it in the function call. Of course this is what strscpy does.
If you aren't sure about the length of the src buffer strlcpy can be a ticking time bomb. If you aren't sure if the src string is NULL terminated it's also a problem, which is especially bad since one of the big reasons to use strlcpy is to avoid buffer overruns. GTA Online suffered from outrageous load times due to this very same API quirk. This would also make it easy for the function to return -1 and set errno to EINVAL if you specify NULL for the src or dst.
Like all things, never say "for all" or "always" but almost all strings can be buffers packed with a size in a struct like in the top comment.
The 4KB of RAM in my microcontroller would appreciate those extra 4 bytes (or 3) used for every string.
I really hope people aren't seriously doing string parsing on microcontrollers.
String parsing is sometimes necessary, and it could be as simple as formatting and sending logs lines through an UART.
One byte is the actual overhead for strings.
Also, it is unclear what side effects the function has if it returns ERR2BIG. Is dest null terminated or not in that case?
I could figure these things out, but my point is that all these str*cpy functions are fundamentally error prone because other people can't keep those details straight either, apparently.
I've never had an issue using 1+strlen() to figure out how many bytes to copy, checking the destination buffer size and then invoking memcpy. It's a bit inefficient (two passes, though both are usually using vectorized instructions) but at least it is really clear to the reader.
The main problem is that C lets you typo this:
strlen(s+1)
When you mean this:
strlen(s)+1
Also, of course, for untrusted buffers, you can't use strlen().
There are tools to find that, like:
https://clang.llvm.org/extra/clang-tidy/checks/bugprone/misp...
https://manpages.debian.org/testing/linux-manual-4.8/strscpy...
These APIs are too complicated/subtle. Honestly, at this point, I just use (ptr, len) pairs that point to non-null terminated strings whenever possible.
No, seriously, shut up. After all that work you don’t get to go “uh, this is actually unsafe because what if the original string is not null terminated and the buffer is not actually the size you said it is”. That’s not how strings in C work. Strings in C work in exactly one way and you don’t get to change that, because that’s just how they work, and there’s 50 years of code that uses string copying routines that work with this that would very much like to have the holy grail of string copies and doesn’t need you muddying the waters and saying that this cannot be invoked safely. It can. The invariants are challenging to uphold but this implementation provides every accommodation that the C standard can provide to this API.
Every single time string copying comes up on Hacker News some enlightened commenter is like “yeah why even bother with this I’m just going to use pointer+length pairs and solve everything”. No. You don’t get to drag this discussion about a very real problem into your half-baked “solution” for C. You especially do not get to reply to me explaining exactly why this API improves upon the state of the art with a lazy dismissal where you retreat to something completely different and go “yeah this is way safer lol”. This would be like if I just responded “yeah I’d just use Rust here I don’t get what the deal is” to you. It doesn’t help.
Look, I’d love to have a conversation about fat pointers and alternative strings in C. I’d love to talk about how we can migrate all this code to safer languages. But this isn’t the time or place for that, and I’m sick of this discussion repeating every single time we talk about, for the love of god, projects that are using normal C strings and safer ways to work with them. It’s not just you but you responding to my comment with something I perceive as fundamentally uninspired set me off.
Count is necessarily the target buffer’s size since it’s what you don’t want overrun.
Not BSD, it was OpenBSD specifically.
As explained in the parent article, using something equivalent with strlen is inefficient when the source is much longer than the destination, because you do not need the length of the source, you just need to know that it is longer than the destination.
Instead of defining yet another strcpy-like function, I would have preferred an alternative to strlen, taking 2 arguments, the size of the destination and a pointer to the source, and returning either the length of the source or an error when the source is longer than the provided size.
Then every strcpy can be replaced with that strlen-replacement function followed by memcpy, and the same strlen replacement can be used with all other string functions, e.g. with strcat.
POSIX 2008 has added the strnlen function ("size_t strnlen(const char *s, size_t maxlen);"), which does not have the inefficiency of strlen, by returning at most the size of the destination, but it does not signal an error to indicate truncation.
You can still use strnlen, if you give it a size larger by 1 and you check the return value to detect an oversize source, but it is slightly less convenient than if the strnlen return value would have been defined like the return value of strscpy.
The length argument of strnlen is under the control of the kernel writer, not under the control of whoever provides the source string.
The new strscpy function can also read anything, if the kernel writer allocates a destination buffer long enough, there is no difference between strscpy and strnlen, from the security POV. What they read is determined by the length of the destination buffer.
Because the destination buffer must have an extra byte for the null terminating byte, strnlen must be called with the size of the destination buffer as argument, but if it returns that size, that signals an error, and then a number of bytes that is one less than the return value must be copied from the source (and a null byte must always be written after whatever is copied by memcpy).
The kernel function strscpy is exactly equivalent with using correctly strnlen and memcpy. While strscpy encapsulates that behavior, so it is more convenient, anyone who wants to use strscpy in a user C program must define a strscpy macro or function, using strnlen and memcpy, which are standard functions.
memchr is indeed an alternative for someone who would need to use some old libc, without the strnlen function.
The strscpy function is available only inside the Linux kernel, but strnlen should be available in any environment that conforms to POSIX 2008, so strnlen + memcpy is what I would recommend for copying null-terminated strings in any C program.
So that would mean either you abolish string literals in freestanding (making C worse) or you have a separate type for these string literals from the type for a expandable string.
Now, Rust pulls off the latter in my view fairly elegantly, but that takes a lot of sophisticated type features in the language including a DeRef and several AsRef implementations. In Rust these are general features open to other types, C could special case them instead, but that's still a lot of engineering for a small feature.
The result is definitely not the elegant language which can be described in a slim book. Maybe it's better anyway, but it's not C.
You can easily implement a stack allocated C++ string:
https://forums.4fips.com/viewtopic.php?f=3&t=1075
You can also wrap a const char* in a string_view which has a length.
Still. Why isn't there a good quality library or set of libraries that would cover these features is my question.
Given how common of a scenario this is for Linux, I imagine this sort of a change would be more trouble than it's worth...
struct string {
size_t len;
union {
char* ptr_to_malloced_string;
char short_string[sizeof(char*)];
}
}
So if the length of the string is shorter than the size of a pointer, the string is stored inline in the struct, and if the length is longer it's a separate allocated object.Can't you just copy a byte to a buffer checking for null, counting down the length.
At the assembly level you have to copy a byte at a time anyway.
I don't see why you'd want to check the entire length of the input string for strlcpy either. It seems less than optimal.
I'm guessing the people involved know more than me. So...?
char big_enough[512];
strncpy(big_enough, sizeof 512, small_string];
strncpy now fills the rest of big_enough with nuls, a lot of waste if small_string is 4 characters long.That’s like saying if you’re careful to do it right you can use a Bowie knife as a screwdriver.
Technically correct but practically will end in tears. If you’re not working with fixed-size strings (which few are these days) just copy over strscpy, or strlcpy, 5 lines won’t kill you.
https://elixir.bootlin.com/linux/v5.19.3/source/lib/string.c...
https://github.com/openbsd/src/blob/master/lib/libc/string/s...
It never got into glibc, which was I guess another. But of course it is all over the BSDs, and not going anywhere AFaIK, more's the pity.
So "all over the BSDs" rather glosses over the fact that whilst common in applications code, there's less need for these sorts of string functions in the kernel. (And it is the Linux kernel that is the headlined subject here.) strlcpy() is in libkern, but it's not anywhere near as frequent in the kernel as it is in applications code.