strcpy: A niche function you don't need
nullprogram.com
nullprogram.com
• It takes a buffer size and truncates the output to the buffer size if it's too large.
• The buffer size includes the null terminator, so the simplest pattern of snprintf(buf, sizeof(buf), …) is correct.
• It always null-terminates the output for you, even if truncated.
• By providing NULL as the buffer argument, it will tell you the buffer size you need if you want to dynamically allocate.
And of course, it can safely copy strings:
snprintf(dst_buf, sizeof(dst_buf), "%s", src_str);
Including non-null-terminated ones: snprintf(dst_buf, sizeof(dst_buf), "%.*s", (int)src_str_len, src_str_data);
And it's standard and portable, unlike e.g. strlcpy. It's one of the best C99 additions.You probably know this, but sizeof is not a function. I prefer the easier to type
snprintf(buf, sizeof buf, ...);This is classic Linus--too emotionally invested in a preference. Except in this case it's particularly pointless and unjustified.
sizeof is an operator. Period. The point of not using parentheses is to continually drive that point home. It's praxis. Of course, it's not unreasonable to prefer using parentheses. And there's a middle ground: most C styles nestle function identifiers and opening parentheses in a function invocation, whereas they require a space between operators and binary operands. So if you prefer using parentheses, whether all the time or just in a particular circumstance, you can do:
sizeof (*p)->member
or sizeof ((*p)->member)
That's not entirely consistent. sizeof is a unary operator, and style guides tend to prefer nestling unary operators while spacing binary operators. But nobody is trying to be pedantic here. The issue is readability, minimizing typos, and dealing with the fact that the sizeof operator, while defined as and behaving exactly like a unary operator, doesn't look like one.Also, it's worth pointing out that not only does the C standard itself literally define sizeof as a unary operator, all the code examples in the standard put a space between sizeof and its operand. It's a stylistic convention, but hardly arbitrary.
By contrast, there are other constructs, like _Generic, where the code examples do NOT use spacing. _Generic is a specialized construct altogether, but syntactically it behaves somewhat like a macro, and it's customary to style macro invocations like function invocations.
From that lens, C has a design mistake. I'm inclined to agree.
In C++ most operators could end up invoking an actual, user-defined function at runtime. But, again, nobody is claiming that they therefore behave like functions.
I get the logic--if you squint really hard you can analogize sizeof to a function. It just doesn't work, as evidenced by his own example. C isn't LISP. It's C. It has a unique grammar with distinct classes of constructs. sizeof is a unary operator, parses exactly like every other unary operator (tokenization conflict with identifiers notwithstanding), which is quite different from how function calls are parsed, particularly wrt parentheses. sizeof has some unique characteristics, but every operator has unique characteristics; that's why they're operators, as opposed to some functional languages that try to subsume everything into function-like syntax.
Another point that I've been missing from Linus' post is that sizeof is special also in that its argument is never evaluated. sizeof launch_the_missiles() will never launch the missiles.
Yet another specialty is that sizeof applied to an array has no array decar. if char buf[1024];, then sizeof buf is 1024, not a pointer size like 4 or 8.
To me, it's not like a function at all. It's more like an assembler macro. I usually agree with Linus and learn a lot from his posts, but here I think he's unsuccessfully trying to find arguments for his (IMO, misleading) stylistic preferences.
It's not part of the C standard, but it's trivial to ship around an implementation with your code.
> If memory allocation wasn't possible, or some other error occurs, these functions will return -1, and the contents of strp are undefined.
I think most people would call its API safe but not trivially safe, where trivial means "when I see that function called in other people's code, I don't need to pay attention to that call because it can't do anything crazy like cause undefined behaviour."
After all, if you include "as long as you check" in your definition of trivial, it is also trivial to check the parameters to strcpy() if you are using it right. And yet here we are, in a discussion about how that's a risk because it isn't used right.
If asprintf() terminated the program when it failed instead of leading on to undefined behaviour when unchecked, I'd call that trivially safe in a more pragmatic way. If you're going to ship a function for portability anyway, that's what I'd recommend. And in fact, that's what is used in software like GCC, called xasprintf() there. It returns the pointer or exits the program on allocation failure.
Ugh. Maybe in some contexts that's ok, but some of us write code which handles memory allocation failures with a bit more finesse than abort().
But at the point where you're able to handle all memory allocation failures usefully, you're almost certainly doing something non-trivial to recover.
For example aborting some requested transaction, pruning items from your program's caches, delaying a subtask to release its temporary memory, or compressing some structures. At that point there's nothing "trivially" safe about a memory allocation.
Probably there are bugs in those recovery actions too. Even the obviously most simple recovery action of propagating an error return to abort a request: If that's a network request, just returning the error "sorry out of memory" is potentially going to fail. So you need recovery from the recovery path failing.
To take a rather convoluted example, if you dereference the pointer and then call a function that does a NULL check then writes to the pointer at some offset, it's possible that the compiler will in-line the function, then ellide the NULL check (since you've dereferenced it, the compiler assumes it's not NULL), then remove your dereference if it didn't have side-effects, so now the write goes through without any check. Granted, it would have to be a write to a massive offset to actually hit an allocated page, but I'm sure there are similar scenarios that are more realistic.
Oh and it's not compatible with UTF-8.
On the other hand, if people start to use snprintf in that particular form as a safe way of string copying, compilers could pattern-match this and substitute a direct implementation.
for (i = 0; i < ...; i++) {
offset += snprintf(
dest + offset,
sizeof(dest) - offset,
"%s...",
str[i]);
}
That will cause a buffer overflow: if iteration n gets truncated because the dest buffer fills up, iteration n+1 will write past the end of the buffer.- The project is new, in which case you can easily/safely ban functions, but then why are you starting a new C project in 2021?
- The project already exists, and now you need to refactor out all the compile-time errors in order to move forward (time-consuming).
Keep in mind the first is a real question that should be answered. If your goal is to avoid undefined behavior/potential security headaches, then C should be entered into after careful consideration of cost/benefits. There are better alternatives for some projects but not others YMMV.
If you need to deal with an older codebase, the all-or-nothing approach might not be appropriate - incremental improvements might be a better option. Yes, you will have more problems to deal with initially, but with time the situation will get better.
That does not follow.
If you do it, you ought to have a good reason, but very probably don't. Any reason you thought you had, if you had any at all, argues for C++ instead, because it works anywhere C does, and enables you to choose to use modern methods that are not inherently error-prone.
Whether you do choose to use modern methods, in any non-C language you end up using, is a whole other matter. But at least you can.
The BSDs have supported C++ user-space programs for long enough that "just working" should be well within range.
A new project involving a BSD or Linux kernel subsystem, or PostgreSQL or SQLite component, is a plausible interpretation, where C++ is actually forbidden for wholly non-technical reasons.
Interoperability? A tons of chips have a single C compiler forked from an old version of GCC without even full C99 support and that's it. Good luck getting anything else running on that hardware.
> If your goal is to avoid undefined behavior/potential security headaches, then C should be entered into after careful consideration of cost/benefits. There are better alternatives for some projects but not others YMMV.
I'd argue doing it globally but as warning would be a better alternative than doing it file by file.
extern void banned() __attribute__((deprecated("This function is problematic")));
#define strcpy(...) strcpy(__VA_ARGS__); banned;
Which would output something along the lines of: test.c:11:3: warning: 'banned' is deprecated: This function is problematic [-Wdeprecated-declarations]
strcpy(x, y);Of course if your code is just a thin layer of glue it's not worth it, you'd spend all your time getting into and out of the FFI. But if you write a non-trivial amount of code that actually does something because it isn't just a glue layer, C might actually not be the obvious choice after all.
In the particular case of Rust, somebody else apparently already did this work (not tested, and thus not vouched for): https://crates.io/crates/gstreamer
GStreamer is routinely exposed to hostile content in non-sandboxed contexts, such as your traditional Linux desktop; it even works as a browser plugin! Among projects you might want to write in C, GStreamer strikes me as one of the least reasonable. If GStreamer were more widely deployed on phones, it'd be prime attack surface for e.g. regimes attempting to surveil journalists.
Ehh..., we have written and deployed gstreamer plugins for a proprietary project at $dayjob which was written in Rust. In the case of gstreamer, they have a plethora of examples for their Rust bindings (I think it's also officially supported).
Why not?
Sure, moving to safer languages is good. But it is impossible to rule out the use of a language so established that every major operating system is written in it. It is practically impossible to not use C - safe languages are normally bootstrapped by it and eventually your “safe” code will run in an environment programmed in C.
Use a static analysis tool with proper, non-brittle tracking of new vs old findings (including proper move and copy file tracking) and forbid new findings, e.g. by preventing offending code from being merged to master with tools like Gitlab, Bitbucket, what have you. Works very well in practice. Bonus points if the tool can also tell you when you modified a function and didn't clean up the old findings in there. Then you can really enforce incremental improvement.
Disclaimer: I work on Teamscale (https://teamscale.com), which does all that.
Then when you open the box you discover that despite all the security talk, the only way to officially program the Azure Sphere is with a C SDK.
C is a sharp knife with no handle; this is it's purpose as a language and tool.
cc: theo@openbsd.org
Yes, they should. Or several other things depending on what exactly they need.
> C is a sharp knife with no handle; this is it's purpose as a language and tool.
Help me out here HN resident survivalists, carpenters, maybe circus knife throwers. What is the "purpose" of a "sharp knife with no handle" exactly? How often have you thought, "Man, it'd be so much easier to gather firewood, carve decorations or score a bullseye if only the blade would sink into my own flesh while I was using it because it doesn't have a handle" ?
Historically the argument was, "We're using C because alternatives like Java or Python or whatever aren't fast enough or capable enough". OK. But, somewhere in the last few years it moved to, "We're using C because alternatives aren't dangerous enough" and that's crazy.
The purpose is not to assume what the developer wants, but provide access to all the resources(and manipulation) they need.
If you need guardrails there are 10s of languages designed for specific purposes.
You should use a knife that comes with such a handle but gives you the option to take it off when you want to impress your friends with your knife-juggling skills and feel you have a few fingers too many.
C lacks a handle because its designers didn't care to provide one, it is after all based on a language whose main purpose was to Bootstrap CPL.
https://github.com/antirez/sds
Naturally turning on all security features of the C compiler being used, warnings, warnings as errors, static analysers, FORTIFY like libraries, the whole package.
And like almost everyone else, I end up having my occasional segfault regardless of the amount of precautions taken, because I am not elite enough.
Consider this before copying code if you are part of a large codebase. Note that sufficiently popular OSS typically is.
One of the last things that finally changed my mind was the observation that the length shouldn't live with the text, but with the structure describing the text. Some of you might be laughing now, because that was obvious to you, but I genuinely had gone years without considering that. I'd been imagining a hack like the length of the string lives in a few bytes "before" the text.
Once I was envisioning the mutable string as [length, pointer] itself, that seemed obviously better and I was onboard with abolishing NUL termination in software.
And a null-terminated linked list is different from a C-string.
That's normal, usually called a "Pascal string".
As I recall, the C standard makes no assumption of whether strings are null-terminated or not.
I'm not sure what you mean by assumptions made by the C standard, but it definitely says strings are null-terminated:
> A byte with all bits set to 0, called the null character, shall exist in the basic execution character set; it is used to terminate a character string.
and
> A string literal need not be a string [...], because a null character may be embedded in it by a \0 escape sequence.
(the second one is noting that if a string literal contains "\0", then it's not a string but contains a string with more stuff after it).
C's behavior of defining literal strings as null-terminated character arrays is already described in the 1978 K&R. The word "string" is used in the text of the section, but not in its title, "character arrays". Null termination is mentioned here, as the book works through an example of successively reading lines from standard input:
`getline` puts the character \0 (the 'null character', whose value is zero)
at the end of the array it is creating, to mark the end of the string of charac-
ters. This convention is also used by the C compiler: when a string constant
like
"hello\n"
is written in a C program, the compiler creates an array of characters con-
taining the characters of the string, and terminates it with a \0 to that func-
tions such as `printf` can detect the end:
| h | e | l | l | o | \n | \0 |
the `%s` format specification in `printf` expects a string represented in this form.
There is a hint, immediately prior to this, as to why null termination might have been chosen: The length of the array `s` is not specified in `getline` since it is determined in `main`.
(Where `getline` is a function defined in the example, and `s` is a parameter to that function.)https://en.wikipedia.org/wiki/Comparison_of_Pascal_and_C#Str... has an intriguing comment that seems likely to be related:
> In Pascal a string literal of length n is compatible with the type `packed array [1..n] of char`.
> Pascal has no support for variable-length arrays, and so any set of routines to perform string operations is dependent on a particular string size.
I suspect that I was remembering someone writing that how to represent strings was a live issue at the time of the creation of C, rather than, as I wrote above, being a live issue within C for some period after its creation.
----
On an unrelated note, it's interesting to see that the web convention of fixed-width type for code literals and variable-width type for natural text was already in force in K&R 1978.
So long as the format string is a literal you needn't care how it works.
Now, one of the places where C makes this nastier than it needed to be is that C built-in types are silly, and so any non-trivial program is using better fundamental types like uint32_t (or the more succinct u32), for which the built-in formatter offers no syntax. So you end up writing format strings like "There are "PRIu32" dogs\n" using macros to bring in the appropriate specifier for your literal. Blergh.
I first encountered that idea in this classic Joel on Software post, which rather put me off the idea of using them in production:
> Notice in this case you’ve got a string that is null terminated (the compiler did that) as well as a Pascal string. I used to call these fucked strings because it’s easier than calling them null terminated pascal strings but this is a rated-G channel so you will have use the longer name.
> Lazy programmers would do this, and have slow programs
char* str = "*Hello!";
str[0] = strlen(str) - 1;
Modern compilers understand strlen, and will replace the function call with a constant where possible. That code's not slow anymore: https://godbolt.org/z/Kjh8b44KfBut Joel was talking about C rather than C++, where your comments about std::string wouldn't apply, right?
Yes. Except that we aren't programming in a void. Particularly if you are writing C to begin with, you will have to interface with decades of existing code. Some of which has interfaces that crept into standards.
You eventually have to pass a string to some function that does not have a length argument and expects a null-terminated string, be it to a library function or the operating system itself (e.g. the `open` system call). You will still need to keep that null-terminator around.
Standardise support for pointer+length strings and the most active parts of the ecosystem will start using it. It will take a long time to get widespread but the sooner you start the sooner it will happen.
Sure, you will have to revert to traditional strings. Some times often. That's no big deal, there should be helper functions. In D you just add .toStringZ to any D string and you get a C string which makes interacting with C code easy.
Of course none of this will happen because C is dead from a evolutionary point of view. Hopefully new CS students will likely not have to deal with any of this bullshit in a few decades.
New CS students will always have to deal with this bullshit in all the decades to come, because the industry will keep relying on UNIX clones for its computing infrastructure until we switch to something else like quantum computers.
https://docs.microsoft.com/en-us/previous-versions/windows/d...
https://github.com/skullchap/chadstr
Maybe not though. Issue #6 is unresolved.
There are numerous safe-string libraries for C. I don't think anyone uses them much.
To me, null-terminated strings look like admirable restraint from Thompson, Kernighan and Ritchie.
Languages where a string type is possible to have usually make use of length and non-null-terminated, while I think C++ does length and C-string with longer texts but for short string can “hack” the text itself into the pointer.
But a more serious answer is to use C++ strings instead, lol. Writing "C-like C++" is probably more beneficial.
I do realize that a lot of people prefer to write in pure C (ex: Linux kernel team), but more and more people are realizing the benefits of C-like C++ code.
For passing references around, I use whatever works. Plain `const char *` argument is certainly a frequent choice for simple name or filepath arguments. That can even mean doing the occasional strlen() when making a copy of that string. It doesn't bother me at all; overall zero-terminated strings are very easy to use. Can't understand why people never stop bitching about it.
When the string is not just an opaque ID, but needs to be examined more closely, it's usually more of a "slicey" or a buffer-processing problem - then I'll add an `int len` to the list of arguments, or to the members in a struct.
Very rarely I'll create a String class, but usually I don't bother. It feels to me like going against the grain of the language. I don't want to create my own host of string processing functions that take this String as argument, when it's usually simpler to operate directly on the data.
Something that I close to never need is the "growable" string class with memory management like std::string. I have no idea right now why I would need such a thing. I tend to write my programs to work on fixed buffers. At most I'll create dynamically sized strings, but a generic string that can grow after creation isn't a frequent use case.
In our team c-string is prefixed with sz_, e.g. char sz_name[13], and we always use a safe subset(or a safe replacement) of strxxx functions with these sz_ prefixed variables. Using memxxx with sz_ variables is explicitly forbidden, since it may break the NULL-terminating contract.
The sz_ prefix convention is by no ways like the hungarian naming nonsense. Suppose that you have "char sz_name[13]" in a structure of configuration parameters, sz_ tells the guy changing the field to keep it NULL-terminated, if they don't, it's their fault. On the other side, users of this field can safely use printf("%s", sz_name) without the risk of crashing the program.
For safe replacements of strcpy, I recommend: https://news.ycombinator.com/item?id=27537900
OpenWatcom, which implements a late draft of the standard behaves as expected on their testcase:
Runtime-constraint violation: strcpy_s, s1max > RSIZE_MAX.
ABNORMAL TERMINATION
The reason not to use the _s versions isn't that they are bad, it's that basically nobody has implemented them (hence me having to use Open Watcom to demonstrate this example)[edit]
Just noticed that in the page they linked to at the top when they mention strcpy_s, it notes that the MSVC implementation predates even the original draft of the standard and lacks RSIZE_MAX.
(Rust crowd snickers as they unwrap<‘jk> &mut *foo_buf)
It all seems to depend on the compiler vendor.
The post seems to be arguing that in most cases where you would call strcpy, you should know the size already, because otherwise you wouldn't know whether the source string was short enough to fit in the destination buffer.
strlcpy.
Every use of it I have seen was subtly incorrect. To use it correctly takes more code than anybody wants to write, or (AFAICT) ever does.
strcat is worse than both, though.
I don't even know why the str* functions exist. They're just worse versions of mem* functions.
Yeah, some stuff like strcat, strtok, or maybe even strcpy is taking it a little bit too far and is too inviting of errors. 30-50 years ago there was often a practice of not minding security aspects of processing external data. But even today, you could use those safely if you know the data.
strtok especially is one arcane function that was probably used a whole lot more back then. Today, hardly anybody does fixed-character delimited fields. strtok was probably useful to "parse" /etc/fstab and formats like that with as little own code as possible :-)
Some functions I use from time to time: strcmp(), strncmp(), strcpy(), strncpy(), strstr(), strchr(), strrchr().
One can rewrite versions of them to work on different types of strings, but they're readily available (if you allow libc dependency). There might be a speed advantage to home-grown solutions, too - although that's not really the point, and if performance matters the bottleneck shouldn't be on such pedestrian string processing anyway.
Many projects just implement a safe_strncpy wrapper that always terminates the destination. Example: https://github.com/brgl/busybox/blob/master/libbb/safe_strnc...
linters and code reviewers commonly recommend alternatives such as strncpy (difficult to use correctly; mismatched semantics) […] Besides their individual shortcomings, these answers are incorrect. strcpy and friends are, at best, incredibly niche, and the correct replacement is memcpy.
But it isn't intended as a solution for buffer overflow bugs in your program, and so if you try to abuse it to solve that problem you likely introduce more problems.
Imagine your car's air conditioning doesn't work properly. On sunny days it's really much too hot in the car. So, you buy a sunroof. Says "Sun" right in the name, surely that will help right? No. That's not what a sunroof is for. The sunroof works fine as a sunroof but that is not what you needed.
I think it’s a travesty that these languages defined an API but didn’t provide an implementation. Hindsight is 20/20, but what a nightmare!
It is far more rational to provide an implementation using standard language features. It’s not like strcpy needs to make a syscall!
std::string does better facilitate copying strings. But string manipulation with std::string is really really bad.
I’m also saying that the concept of defining an API but not providing an implementation is insane. That is, imho, extremely inappropriate at the language level. Differences in behavior between C and C++ STL implementations on different platforms is infamous. And almost entirely unnecessary.