Fixing C Strings
thasso.xyz
thasso.xyz
struct str {
char *dat;
sz len;
};
It's the same solution D uses, except that it's a builtin type, and works for all arrays. I proposed this solution for C:https://www.digitalmars.com/articles/C-biggest-mistake.html
It's hard to overstate what a huge win this is. D has had 23 years of experience with it, and the virtual elimination of array overflow bugs is just win, win, win.
I will never understand why C keeps adding extensions consisting of marginal features, and ignores this foundational fix. I guess they still aren't tired of buffer overflow bugs always being the #1 security vulnerability of shipped C code (and C++, too!).
As well as most other languages and many C codebases right? Often with a separate length/capacity so the buffer can be larger than the string.
struct str {
sz len;
char dat[];
};You could say, well, forget binary compatibility and forget nasty code that bit-twiddles pointers. But then why are you even using C? Those are the things that set it apart.
Clang is trying to solve this with annotations that allow the programmer to construct fat pointers, either as structures or just implicitly by having the length in a variable somewhere, and enforcing those bounds in the compiler. Seems promising. https://clang.llvm.org/docs/BoundsSafetyImplPlans.html
Fixed with my proposal: https://www.digitalmars.com/articles/C-biggest-mistake.html
I'm also of the opinion that a backwards compatibility with null terminated string is actually terrible. Because you want people to eventually go, oh this code uses gross null terminated strings, lets fix that.
Well, I, for one, do like the idea of C (in contrast to D or C++) still being sort of the lowest-level high-level programming language - one that's just a notch above the assembler.
At this point, I think we might be in a better world if C simply did not offer any string API in the first place. If
That’s not true, this representation of strings (in the form of a literal) is built into the language itself.
That has never been true, unless you are writing programs against a PDP-11. C compilers can change even the O complexity of your algorithms, it's nowhere close to assembly.
I'd be curious to see an example...
GCC has the format attribute that lets you have printf type checking on your own variadic functions:
https://gcc.gnu.org/onlinedocs/gcc-14.2.0/gcc/Common-Functio...
https://dlang.org/spec/pragma.html#printf
It was a huge win, at least for me. I implemented it because I was really sick of mismatches. (Although I was careful to use the right formats, when refactoring I'd change a type and then the printf's would go awry. Having the compiler flag them made for quick fixing.)
print(f"{foo:8.2f}")
compared with C-style printf: printf("%8.2f", foo)Another one to consider is e.g. https://github.com/antirez/sds (used by Redis), which instead stores the string contents in-line with the metadata.
Choosing between these trade-offs just depends on what you're doing. I'd definitely choose this pattern if I were to write a parser for instance.
* One for the start of the source string, with an inline strong count * One for the end of the source string so you know how much to deallocate (only really applicable to Rust) * One for the start of the view * One for the end of the view
32 bytes for each string view is quite a lot. Depending on context you could use 32bit lengths instead of end-pointers if you're OK with <4GB strings, saving 8 bytes.
But views are also implemented as a plain pointer and a length, and that's where the memory safety issues from borrowing begin.
And reference counts are often atomic integer operations, so it might not be a regular memory increment, instead it would be an interlocked increment. And if there's multiple threads, the CPU cores will be competing over who gets to keep the reference counter in their L1 cache line. (There is a way around this where you can give threads an their own reference counter)
You could even do something crazy with packing a null byte with sz on 64-bit systems (since you will never have a string that long anyway...)
I don't include the null-terminator because I use this type in my own environment where I never use null-terminated strings so there is no need for it.
printf("%.*s", len, str);
lets you pass the length as an int argument.But I would say for 95% percent using a fixed length char array with strncpy will work just fine.
Where on OpenAI's site do I find a footer like that?
Once you go down the route proposed by many of the comments here - why not enhance it to deal with UTF8... Or rather implement a proper "array" type? What about the lack of multidimensional arrays instead of the pointer to pointer to ... approach? Idiosyncracies such as "int a[2][3];" being of type "int *" and not "int **"?
C was never intended to shield you from mistakes, but rather replace a macro assembler. ANSI C addressed some of the issues in the original K&R C, but that is about it.
If your use case would benefit from all of these protections, there are plenty of higher level language alternatives...
char *s = "Hi";
The compiler will not treat that as a simple mere pointer to char when allocating space for it in the binary. It will see that the rhs is surrounded by double quote characters, and allocate 3 bytes for it, instead of 2, and put a NUL byte after the bytes for 'H' and 'i'.Nul-tetminated strings are absolutely a part of the language. Certainly you can make and store strings in a different way if you'd like, but the language itself defines what a string and string literal is.
char s[3];
Which is then initialised with: 0x48,0x69,0x00
There simply is no such thing as a string type in C.
All the "string" functions work on a char pointer which is incremented until it points to a 0.
As per the OP’s example, a wrapper macro like their STR can work around this.
I see a problem with the separation between str and str_buf, though: you create new strings with the latter, but most functions take the former as arguments. Do you convert them every time? Isn't your code littered with str_from_buf()?
Put it in another way, it's like the mess with const that you mention in your article. If str is the type you use for a const read-only string, and str_buf for a non-const mutable string, you would like to pass a non-const even to those functions that "only" require a const. (I say "only" because being const is a weaker requirement than being mutable; the fact that it's more wordy is another thing that C's syntax makes confusing, but this is an entirely different topic!)
It would be nice if the compiler could be instructed to automatically cast str_buf into str and not vice versa, just like it does for non-const to const.
The only way out I can think of, would be to get rid of the two types and only use the one with the cap field, with the convention that if cap is zero, then the string is read-only. The drawback is that certain mistakes are only detected at run-time and not enforced by the compiler. For example, a function than takes a string s and replaces every substring s1 with s2 could have the following prototype in the two-type system:
replace(str_buf s, str s1, str s2);
And it would be immediate to recognize that you cannot pass a read-only string as the first argument. With a one-type system you loose this ability.Oh well, I guess if a perfect solution existed, it would have been adopted by the C committee, wouldn't it? /s
No, the article addresses this: since the memory layout of the first two struct members is the same in both structs, you can use a pointer to str_buf anywhere a function calls for a pointer to str, after casting it.
Yes, you could, but I see no function mentioned in TFA that wants a pointer to str, only functions that want a str: print_str(), print_fmt(), com_write(). At the same time, the functions that return strings return a struct, never a pointer: str_new(), str_from_range(), str_from_buf(), fmt_buf_new(), and the pseudo-function STR().
To use the memory layout trick you should go through reference + cast + dereference:
*((struct str *)&...)
My question still holds: is the code littered by such conversion artifacts?This article is about C strings FYI.
fgets will treat the length passed as the capacity of the buffer, and terminate the last byte with a 0.
scanf however treats the length as the number of characters to read, meaning that you need a capacity of n+1 to make sure the 0 terminator is stored properly as well.
Its quite easy to mess up placing the 0 terminator yourself too. It's an overall unnecessary burden that could've been fixed quite easily.
I don't actually believe you that you've been programming for 50 years and never misused a string or a string API in C or a language with similar string handling. But even if I did believe you, it wouldn't matter. Many people make mistakes, and those mistakes have cost people a lot of time, money, and stress. If you've not read about any of these instances, then I suggest you've been living under a rock and are incredibly out of touch.
Or you're just trolling.
The OP clearly stated that he did not mind the overhead (in terms of executable size, memory consumption, execution speed) in his particular use case.