Strcpycat
byuu.org
byuu.org
Firstly, strncpy() was not designed to be a "safer strcpy()" at all. It was designed for storing strings in the manner of original UNIX directory entries: in a fixed size buffer, null-padded only if the stored string is shorter than the buffer. This is why it fills the buffer with nulls, and why it doesn't necessarily null-terminate the destination.
Secondly, contrary to what the article implies, strncat() does not work this way: strncat() always null-terminates the destination, and it doesn't write any extra nulls.
The strncat() behavior is quite confusing. I can understand expecting the buffer to already be null-padded and skipping that step, but always null-terminating the result seems to go against strncpy()'s expected behavior. It also looks mildly dangerous, as the NUL-write can be one after the size given.
http://computing.fnal.gov/unixatfermilab/html/filesys.html
Having the size so fixed, storing the 15th character that would always be null on the disks (and also in memory) was a pure waste. Note that earliest versions of Unix run in less than 64 KB.
When copying the name to a fixed buffer of 14 characters which is to be stored on the disk, padding the buffer with nulls was also not a performance issue and had a benefit of not leaving "junk" in the buffer.
The problems only happen when people expect from such routines with very specific usage scenarios something for which they were not designed.
So the strncpy was never meant to be "a safer veriant for strcpy." Moreover, strcpy_s is also not a replacement for strncpy, but for a strcpy -- it was introduced to be replaced instead of strcpy in the old code bases as a part of reducing the number of security issues in the old code.
Of course there is (in Microsoft's CRT) strncpy_s, the rationale behind all _s functions is here:
http://msdn.microsoft.com/en-us/library/8ef0s5kh(v=vs.80).as...
"Sized Buffers. The secure functions require that the buffer size be passed to any function that writes to a buffer. The secure versions validate that the buffer is large enough before writing to it, helping to avoid dangerous buffer overrun errors that could allow malicious code to execute. These functions usually return an errno type of error code and invoke the invalid parameter handler if the size of the buffer is too small. Functions that read from input buffers, such as gets, have secure versions that require you to specify a maximum size.
(...)
Null termination. Some functions that left potentially non-terminated strings have secure versions which ensure that strings are properly null terminated."
The article is based on the wrong premises.
Advice: if you work low level with substrings and concatenations, where the performance matters, don't try to fix the str routines, use the pointers. Maintain the pointer to the last character of the area to which you concatenate. Traverse only the characters you need from the source. If you don't care about the performance or exact storage aspects, use much higher level approaches, some string libraries.
char* scatn(char* ss, …);
#define scat(args...) scatn(args, 0)
char* scatn(char* ss, ...)
{
char *s;
char *new;
size_t len;
size_t n;
va_list ap;
/* first find the lengths of the strings */
va_start(ap, ss);
s = ss;
len = 0;
while(s != 0){
len = len + strlen(s);
s=va_arg(ap, char *);
}
va_end(ap);
/* make a buffer big enough for all strings + '\0" */
new = (char *) malloc(sizeof(char) * (len + 1));
if(new == NULL){
perror("Out of memory");
exit(EXIT_FAILURE);
}
/* copy the strings into the buffer */
va_start(ap, ss);
s = ss;
n = 0;
while(s != 0){
len = strlen(s);
memcpy(new+n,s,len);
n = n + len;
s=va_arg(ap, char *);
}
va_end(ap);
/* terminate the string */
new[n]='\0';
/* return the string: remember to free it! */
return new;
}
use like this: s = scat("hello","hacker","news");
:
free(s);
I did have a version with only one loop and realloc but it turned out to be twice as quicker to go through each string twice. I have no idea why :)short answer: I did. It got slower!
They're not doing what you think they are. Each string only gets measured twice.
To factor them out I will need to start allocating extra arrays just to keep their values: the extra malloc and free is slower than measuring the strings twice. I also tried a Reallocing version that didn't need them but it was twice as slow.
I use this code for composing text in places where speed isn't my first concern; especially error messages. But its plenty fast enough for 90% of what anyone would need. You'd rarely be concating more than 6 things together.
note: it's really hard to talk about pointers on HN as you can't type an asterisk without a code block, but it's:
while(s != 0)
not while(*s != 0) std::map<std::string, int>
without paying for a extra copy of the string on insertion. How many years since C++ has existed before they finally allowed for a single copy? Everything is relative. You might never have had a need to worry about that double copy, but there are other people who have.Except, interestingly enough, Go code compiled with the standard compilers (5g, 6g, 8g, see http://golang.org/doc/go_faq.html#How_is_the_run_time_suppor...). The Plan9 expats are hell-bent on doing things their own way, I am 70% in agreement with them :).
The length value in that case would also be invalid for C usage, because std::string does allow storing nulls inside the string, but it would eliminate the need for allocation.
And much of libc is not well designed for modern code. Mostky I ptogram to the system call API in C.
char buf[65536]; /* 64k is enough for everything! */
and using strncpy w/ sizeof(buf) to populate it.I'll consider this idea and write some comparisons to see how it works, thanks.
I should note that I usually write articles to kick off discussions on my own forum, and then revise them after all input is gathered. They almost never appear on other sites, especially pre-revision.
I'm not saying the strmcpy/cat functions there now are the best solution, and I'm open to hearing what's wrong with them so that we can improve upon them.
//return = strlen(target)
unsigned strmcat(char *target, const char *source, unsigned length) {
const char *origin = target;
while(*target && length) target++, length--;
return (target - origin) + strmcpy(target, source, length);
}
//return = strlen(target)
unsigned strmcat(char *target, const char *source, unsigned length) {
const char *origin = target;
while(*target && length) target++; length--;
return (target - origin) + strmcpy(target, source, length);
}
How about this? //return = strlen(target)
unsigned strmcat(char *target, const char *source, unsigned length) {
const char *origin = target;
while(*target && length)
target++, length--;
return (target - origin) + strmcpy(target, source, length);
}
//return = strlen(target)
unsigned strmcat(char *target, const char *source, unsigned length) {
const char *origin = target;
while(*target && length)
target++;
length--;
return (target - origin) + strmcpy(target, source, length);
}
Granted, it's a tiny example, but you're going to be writing fewer bugs in the long run if you avoid being clever. So here are two rules to start with:1. Never put more than one statement on one line.
2. Never use the comma operator outside the "for" syntax.