The Sad State of C Strings
symas.com
symas.com
This topic is a triviality of C programming. You run into performance (or usability) issues with standard strings - you roll out your own. It takes... what?... 20 minutes? OK, maybe an hour. Would it help to have an alternative standardized? Perhaps. Is it critical to have it? Hell, no.
The C standard says that "strings" are examinable buffers of bytes you can do arithmetic to and stream until NUL—and thus, every single library that uses C strings ends up baking in a bunch of pointer arithmetic and NUL-checks on buffers of bytes.
It doesn't matter how optimized your own code's string implementation is if 99% of your time is spent in, for example, OpenSSL's string implementation.
STPCPY(3) Linux Programmer's Manual
NAME
stpcpy - copy a string returning a pointer to its end
SYNOPSIS
#include <string.h>
char *stpcpy(char *dest, const char *src);
Feature Test Macro Requirements for glibc (see feature_test_macros(7)):
stpcpy():
Since glibc 2.10:
_XOPEN_SOURCE >= 700 || _POSIX_C_SOURCE >= 200809L
Before glibc 2.10:
_GNU_SOURCE
DESCRIPTION
The stpcpy() function copies the string pointed to
by src (including the terminating null byte ('\0')) to the
array pointed to by dest. The strings may not overlap,
and the destination string dest must be large enough to
receive the copy.
RETURN VALUE
stpcpy() returns a pointer to the end of the string
dest (that is, the address of the terminating null byte)
rather than the beginning.
CONFORMING TO
This function was added to POSIX.1-2008. Before that, it
was not part of the C or POSIX.1 standards, nor customary
on UNIX systems, but was not a GNU invention either. Perhaps
it came from MS-DOS. It is also present on the BSDs.
BUGS
This function may overrun the buffer dest.
EXAMPLE
For example, this program uses stpcpy() to concatenate foo
and bar to produce foobar, which it then prints.
#define _GNU_SOURCE
#include <string.h>
#include <stdio.h>
int
main(void)
{
char buffer[20];
char *to = buffer;
to = stpcpy(to, "foo");
to = stpcpy(to, "bar");
printf("%s\n", buffer);
}
SEE ALSO
memccpy(3), stpncpy(3), wcpcpy(3)stpncpy is equivalent to strncpy - it NUL-pads dest if the src is shorter than N. Frankly I have never found a use for this behavior. The desired behavior is to simply copy src without additional padding. That is what my proposed strecopy() does, and that's what would be required of a replacement for strlcpy().
// This code is licensed under CC0.
// A copy of the license can be obtained at https://creativecommons.org/publicdomain/zero/1.0/
// May you forever catenate in peace.
#define strcatm(...) strcat_multi(__VA_ARGS__, NULL);
char* strcat_multi(char *dest, ...) {
va_list srcs;
char *dest_end;
const char *src;
size_t src_sz;
va_start(srcs, dest);
for (dest_end = dest; *dest_end; dest_end++);
for (src = va_arg(srcs, const char *); src; src = va_arg(srcs, const char *)) {
src_sz = strlen(src);
memmove(dest_end, src, src_sz);
dest_end += src_sz;
*dest_end = '\0';
}
va_end(srcs);
return dest_end;
}
There. Now go and catenate, children. char buf[256]; buf[0] = 'x'; buf[1] = '\0';
strcatm(buf, buf, buf, buf);
producing a string of 16 'x's strcatm(..., NULL);
and it's 8 x/s, not 16. Also fix , to ; in the second for loop. // This code is licensed under CC0.
// A copy of the license can be obtained at https://creativecommons.org/publicdomain/zero/1.0/
// May you forever catenate in peace.
#define strncatm(dest, sz, ...) strncat_multi((dest), (sz), __VA_ARGS__, NULL);
char* strncat_multi(char *dest, size_t dest_sz, ...) {
va_list srcs;
char *dest_end;
const char *src;
size_t src_sz;
size_t dest_off;
size_t copy_sz;
va_start(srcs, dest_sz);
for (dest_end = dest, dest_off = 0; *dest_end && dest_off < dest_sz; dest_end++, dest_sz++);
for (src = va_arg(srcs, const char *); src && dest_off < dest_sz; src = va_arg(srcs, const char *)) {
src_sz = strlen(src);
copy_sz = src_sz < (dest_sz - dest_off) ? src_sz : (dest_sz - dest_off);
memmove(dest_end, src, copy_sz);
dest_end += copy_sz;
dest_off += copy_sz;
if (dest_off < dest_sz)
{
*dest_end = '\0';
}
}
va_end(srcs);
return dest_end;
}So,
copy_sz = src_sz < (dest_sz - dest_off) ? src_sz : (dest_sz - dest_off - 1);
...
*dest_end = '\0';To keep with the spirit of the article, calling strlen() just to allow convenient use of memmove() seems a bit counter-intuitive, I'd roll the two together into a copying loop instead.
http://c.learncodethehardway.org/book/ex36.html
I'm not an experienced enough programmer to be able to evaluate how much "safer" the bstrlib library actually is. I'm sure the various HN Engineers can chime in with their insights.
I think unicode support would be much more important, but not everyone agrees (for example, on embedded devices unicode support can be dead-weight). And adding to C without consensus leads to lousy decisions.
char *heap_ptr, *stack_ptr; ... fn_needs_heap_ptr(strcpy(heap_ptr, stack_ptr));
Another example (dma address space != buf a.s): char *dma_ptr, *buf; ... fn_needs_dma_ptr(strcpy(dma_ptr, buf)) ...
IOW, foo(strcpy(dst, src)) is VERY different than strcpy(dst, src); foo(src);You will indeed have a problem if you strcat() excessively, but then it will be a marginal case, so you will optimize for it in your own code rather than dragging it into the standard.
Cheap? 12 precious clock cycles is enough for up to 384 floating point operations (FMA).
Not sure how fast or slow rep sca/mov is these days. Regardless, strlen that is not scanning at least 4 bytes per clock is hopelessly slow. I think you can do at least about ~16 bytes per clock.
Anyways, the problem with strlen is unnecessary extra (potentially mispredicted) branches and memory accesses.
I have seen enough cases where strlen has dominated cost in the profile. Of course this depends entirely what you are doing, there are plenty of workloads where strlen just doesn't occur.
Though you still pay for the memory/branch prediction stalls.
There will be bottlenecks narrower (is that the right term?) than string operation in any project not dealing exclusively with strings. In the latter case, rolling your own string utilities may make sense
C string operations might have been elegant 40 years ago, but nowadays they're like rerouting a 747 to check if office lights are on.
Strlen is doing a lot of work for little benefit. It does slow and power hungry memory accesses. Because it's scanning for terminating 0x00 byte, it needs to contain a branch -- and loop terminating branch must be a costly mis-predicted one.
C printf and friends are even more insane. It scans "bytecode" instructions from a string and does dynamic formatting. You can do format options like this:
printf("% 0#*.*f\n", 15, 5, 1.234);
Or say print five chars of an unterminated string, left padded to total length of 10: printf("%10.5s\n", unterminated_str);
It won't (at least it shouldn't) crash even if 6th character is on an unreadable memory page.The last example is explicitly defined in the Standard.
Implementing all that subtle functionality correctly takes up a lot of CPU time.
According to my quick test, on Visual Studio 2012, even simplest sprintf with just one parameter seemed to take about 1 microsecond to execute.
Of course clang and gcc seem to sometimes compile whole format parser away. At least...
printf ("Hello World!\n");
... is optimized into a simple "puts("Hello World!");".Of course iostream << operator runtime performance is also pretty horrible. Each << invocation seem to call streambuf::sputn (or sputc) separately. On VS2012, simplest stringstream test...
ss << "value: " << intVal << endl;
... outputting one variable into it and turning the result into a std::string took about 3 microseconds. (Although it was a very quick test, a number of things might be suboptimal in the test code.)As an example of the alternate approach, the Rust language uses (ptr, size) to represent all arrays (called slices) and strings internally. This seems like a much better method, allowing you to reference substrings without allocating a buffer, and perform out-of-bounds checking at runtime.
Seriously though, a C programmer should know the downsides of C strings, and act accordingly, i.e. use a library if necessary.
strlcpy() also sucks. The solutions I proposed solve both the overflow protection aspect that strlcpy aims to solve and solves the inefficiency problems of strcpy/strcat/memcpy.
No, it's not.
"strcat(strcat(strcat(strcpy(buf, "This "),"is "),"a long "),"string.");"
Just checked 20 years of C/C++ repositories ... not a single strcat. Actually I can't remember ever using c-strings for more than args parsing or debug logging.
"strcat(strcat(strcat(strcpy(buf, "This "),"is "),"a long "),"string.");
len = strlen(buf);
> The above example executes in exponential time with the length of the strings.This statement is obviously not true. The example is linear with the length of the strings (assuming the number of strings is constant) and O(n^2) with the number of strings (assuming each string has non-zero length).
And of course you have the crowd that's been writing C for 20 years saying "it's not that bad". Folks, the god damn plane has crashed into the mountain.
Personally, my favorite for good performance is C++ with Qt. The QString class lets you do a lot of amazing stuff very easily.