Could we make C arrays memory safe? Probably not, but let's try
nibblestew.blogspot.com
nibblestew.blogspot.com
[1] http://animats.com/papers/languages/safearraysforc43.pdf
The main obstacle are people coming from MSVC or C++ not knowing variably modified types and people being convinced that VLAs are always bad. This then leads to many bad attempts at fixing the problem instead of simply using arrays which know their run-time length. While we still miss a bit of compiler support (I am working on it), this already helps today: https://godbolt.org/z/4a45xq5hr
(Update: Of course, the use of references in the proposal above and the motivation is a bit obscure. In any case, VM-types will not be optional in C23 anymore. And usage and interesting is going up.)
So will WG14 finally care about C's safety?
(Stack allocated) VLAs are (almost) always bad. How do you prove absence of stack overflow in their presence?
Your first reaction is "sure, even an old Intel 80286 chip can easily do memory-safe C arrays". Because if your code runs in '286 protected mode, using a (yes, scarce) separate segment register for each array, and you don't botch loading the array's base, limit, etc. into the segment descriptors - yep, your regular array-access code (in assembly, C, or whatever) can enjoy memory-safety array access for "free".
This got proposed in 2009 so i do not see things changing any time soon.
Edit: Typo
https://en.wikipedia.org/wiki/C23_(C_standard_revision)
> Variably-modified types (but not VLAs which are automatic variables allocated on the stack) become a mandatory feature
https://discourse.llvm.org/t/rfc-enforcing-bounds-safety-in-...
Supposedly it's already in use: "The -fbounds-safety extension has been adopted on millions of lines of production C code and proven to work in a consumer operating system setting."
It is only C if it lands on the standard, otherwise it is yet another compiler specific language extension, that portable code cannot rely on.
For that Clang extension above it looks like it's possible to annotate source code without breaking compilers that don't support the extension by defining a handful of dummy macros.
IMHO the actual strength of C is that compilers can (and do) explore beyond the standard on their own.
It is definitely a problem when said extensions don't have a counterpart in other compilers, specially a pervasive feature like bounds checking.
Also consider that this extension is designed in a way that existing code can be annotated without requiring drastic changes, and that it has been designed to remain ABI compatible.
A 100% watertight solution most likely requires new language features (or even a completely new language like a "Rust--") that would violate both of those requirements.
A solution is mostly possible with a segmented allocator, which is quite reasonable on 64-bit platforms (32-bit allocation ID, 32-bit index within the allocation).
But keep in mind that "buffer overflow within a struct" is often considered a feature.
I think the best approach is "design a new, 'safe' language that compiles to reasonable C code, and make it easy to port to that new language incrementally".
Yes this is used rampantly in any kind of code that's running in basically all of the world's network appliances that do packet processing.
I'd even say it's considered "best practice" by most C programmers.
Yes, one of those "features" designed to work around limitations of the language
C is "simple" in the same way a chainsaw without guards or brakes is "simpler" than one with those attachments
That's lying to the compiler, because there's no way to say what you meant.
That's why I proposed the syntax
typedef struct msgitem {
const size_t len;
char itemvalue[len];
};
struct msgitem firstmsg = { 100, 0 }; // empty msgitem, size 100
The last item of a struct could be a variable sized array. That's sound,
and enough for things such as network packets. It's important to avoid bikeshedding.It might very well be straightforward to obtain, just not located in that struct itself.
> I'd say just change the code.
And if the code isn't available to you to change?
That's not what I said at all. It's the exact opposite of what I said. Why this strawman? I already replied above and explained very clearly that I'm talking about when the size is known but not via the struct itself:
>>> the expression doesn't have to come from the same struct, though. It could be provided somewhere else.
>> It might very well be straightforward to obtain, just not located in that struct itself.
The size could be communicated in a different struct, no? Or passed back to the caller via a pointer argument? Or a million other ways beside the same struct itself?
Then you don't get the improvements in safety yet.
> Then you don't get the improvements in safety yet.
Huh? This isn't a limitation with current implementations like -fbounds-safety. It's just a limitation with the proposal I was pointing out this issue with [1]. The existing implementations decorate the function/usage sites rather than the struct, which gives you access to information outside the struct. And there's no need to change every single use of that struct, which you obviously don't generally have access to.
I'm saying to deal with it. Change the code to be compatible. It's not that important to keep it the way it is.
Now you're referring to better designs, which is great. Have the best of both worlds if that's possible.
But when you were just pointing out that difficulty, my response is that it's a very small difficulty so that's not a big mark against the idea. If it was that proposal or nothing, that proposal would be much better than nothing, despite the forced code changes to use it.
> But when you were just pointing out that difficulty, my response is that it's a very small difficulty so that's not a big mark against the idea.
In what alternate timeline do we exist where HNers believe you can just recompile the entire world for the sake of any random program? Say you're a random user calling bind() or getpeername() in your OS's socket library. Or you're Microsoft, trying to secure a function like WSAConnect(). All of which are susceptible to overflows in struct sockaddr. Your proposal is "just move the length from 3rd parameter into the sockaddr struct" because "it's not that important to keep these APIs the way they are"?! How exactly do you propose making this work?
Gradually, for one.
I can't believe you think changing the world isn't a big deal.
So say I'm on board and decide sockaddr Must Be Changed. Roughly how long do you think it will be from today before I can ship to my customers a program using the new, secure definition?
And how does the time and effort required compare against the more powerful implementation that's already out there?
And again, I wasn't comparing to any other implementations, because you hadn't brought them up yet!
char data[0];
The purpose is to give you a handle on the remainder of the data even though you don't know the size beforehand.Technically this is not valid under strict C, but it's also incredibly common.
char data[] is valid C in recent versions if it is the last entry of a struct.
int do_indexing(int *buf, const size_t bufsize, size_t index) {
if(index < bufsize)
return buf[index];
return -1;
}
buf[index] can hold the value -1, as buf holds signed ints, so that (corrected, by the author) function is totally wrong anyway.And I’m pretty sure I’ve seen similar stuff for clang.
(yes, a bit more compiler support is necessary to make this safe. I posted a patch to GCC, let's see)
I never knew that you got runtime bounds-checking with VLAs.
Do you have a link that explains this snippet in more detail? Why/how does it work?
It doesn't appear to work with non-VLAs though.
If you replace the sizes with constants it also works: https://godbolt.org/z/sKPW6zT87
The safety story is not complete though: If you pass the wrong size to 'foo' this is not detected (this is easy to add to compilers and I submitted a patch to GCC which would do this): https://godbolt.org/z/T8844e1z8
(ASAN still catches the problem in this case, but ASAN does not work consistently and has a high run-time overhead.)
Basic idea is to have a system call that allows library writers to get the bounds of a pointer. This way they can ensure they're not writing too much data to a location.
Another idea I've implemented in userspace is to create an allocator that allocates a page (via mmap) then set protections on the page before and after. The pointer returned aligns the end at the next page. If a write goes beyond the end of the pointer, it bumps into the protected page, and causes a fault. Then you can handle this fault, and detect an overflow.
A even more strict version of this is to add protection to the page the allocated pointer is assigned to. On _every_ write you get a fault, and can check that it's not out-of-bounds.
All of these methods are slow-as-hell, but detect any memory issues. While slow, they are faster than valgrind (not badmouthing it, it's an amazing tool!). So the recommendation is to use it in testing and CI/CD pipelines to detect issues, then switch to a real allocator for production.
Only if the stride is small enough to not skip over the guard page, surely? Unless you're setting the entire address space to protected, for any given base pointer BP there's a resulting address BP[offset] that lands on an unprotected page.
It might be interesting to expose an instruction that restricts the offset of a memory operation to e.g. 12 bits (masking off higher bits, or using a small immediate) to provide a guarantee that a BP accessed through such an instruction cannot skip a guard page; but that would of course only apply to small arrays, and the compiler would have to carry that metadata through the compilation.
We're almost 1/4 into the 21st century. As someone that has attempted this as an exercise in C, please, just learn a little bit of Rust and move onto modern problems. The analogy I like to use is adjusting all the springs on a mattress vs buying a memory foam mattress and letting better material do its job.
Since you seem to know how to do it, can you please tell me how to check the bounds of the arrays at compile time if the actual underlying value is not known until runtime? I guess I could make a type-level algebraic field and add, subtract, multiply, divide my generics but this seems like a huge pain in the butt.
If anyone knows how to perform comptime algebraic bounds checking in stable rust lmk
The annotations say "the bound of the pointer/array p is expression X", but if I subsequently do for (int i.....) { p[i]++; } you have to perform a bounds check unless you can prove that i never exceeds the expression given by the bounds.
e.g. (using the implemented and used in the real world bounds safety extension in clang[1])
void f1(int *ps __counted_by(N), int N) {
for (int j = 0; j < N; j++) ps[j]++;
}
void f2(int *ps __counted_by(N), int N) {
for (int j = 0; j < 10; j++) ps[j]++;
}
In the function f1 the compiler can in principle prove that the bounds checks aren't needed, but in f2 it cannot, and so failing to perform a bounds check would lead to unsafe code.Making the bounds of a pointer explicit does not mean that you no longer need bounds checks, it just means that now you know what the bounds are, so you can most enforce those bounds. Then as an implementation detail you optimise the unnecessary checks out. The enforcement aspect is required for correctness, if you cannot prove at compile time that a pointer operation is in bounds a bounds check is mandatory.
[1] https://discourse.llvm.org/t/rfc-enforcing-bounds-safety-in-...
It starts off with this function:
int do_indexing(int *buf, const size_t bufsize, size_t index) {
return buf[index];
}
Later it suggests changing just the declaration, adding a parameter annotation in this way: int do_indexing(int *buf,
const size_t bufsize [[arraysize_of:buf]],
size_t index);
It suggests that, with this function declaration and the former function body, the compiler should emit a warning:> Similarly the compiler can diagnose that the first example can lead to a buffer overflow, because the value of index can be anything.
This clashes with my understanding of why we would use the annotation in the first place. You seem to hold the same feelings, as this is analogous to your f1(), which you agree the compiler should consider safe.
Instead TFA seems to advocate that the correct function body should be:
int do_indexing(int *buf, const size_t bufsize, size_t index) {
if(index < bufsize)
return buf[index];
return -1;
}
even with the annotated declaration:> Now the compiler can in fact verify that the latter example is safe. When buf is dereferenced we know that the value of index is nonnegative and less than the size of the array.
On the other hand, many people have depended on interpretations of intent.
Changing the behavior of one is much easier for the community to swallow than the other.
An easy example is how early versions of IE failed to correctly implement the box model for sizing elements. For backwards compatibility reasons, IE6 and on would revert to the old, incorrect behavior if parsing the html document put it into 'quirks mode', and later on this behavior was added as an optional CSS box-sizing property.
https://wiki.edunitas.com/IT/en/114-10/Internet-Explorer-box...
In terms of validation: html5 has a specified parsing algorithm, you can validate html against it. That validation is not xml validation. Again, because html is not xml.
There is a very clear specification about how html tags are parsed, trying to reason about html as if it is xml is just as (if not more) incorrect than trying to reason about C as if it were C++. e.g.
void f();
Is a different type in C than it is in C++, but if you do this you can't turn around and complain that C and C++ see it as having a different type when the syntax has different meanings.https://hixie.ch/advocacy/xhtml https://www.w3.org/2003/01/xhtml-mimetype/
I am all for making C and C++ safer by adding mitigations for the most common security relevant bugs. Either on language level, or on tooling level or on OS level or on CPU level. Or a combination of them. But since making a language absolutely memory safe doesn't make it automatically impossible to have bugs at all (even security relevant bugs), we should consider everything a trade-off.
This is a singularly bad analogy. HIV used to be a death sentence, but now we have effective treatments and prophylactics.
What's the PrEP equivalent for buffer overflows? There isn't one.
I _love_ C and I've written one C program every year for the last decade or two. I generally write something really stupid, but there's going to be crashes that "prints the stack pants." As in: I'll fudge up returning the stack from a function, basically.
This proves nothing, of course. But Rust won't even let me fudge things up that way. This means that I have more time to spend fixing those other classes of bugs and that is the win!
For instance I find it incredible that you never had to debug and workaround a third party vendor DLL that you didn't have source for but that was leaking memory like crazy. This is the just one example of something that can be "fixed trivially" as you don't have source to modify, and is extremely common in some fields (i.e. embedded)
And no, we didn't have major issues with memory leaks either. About on par with what I have seen in garbage collected languages. RAII works quite good most of the time.
I get how that would be characterized as "trivially fixable", but I assure you that this is not the kind of projects people complain about when they discuss memory safety issues.
Consider that some people need to send emails to vendor companies begging them to stop segfaulting, writing the stack or leaking memory. You're lucky if it gets fixed in a few months, because that means you wouldn't have to seek alternatives which would be even more time consuming. In conclusion, there are many people out there dealing with memory safety issues which are anything but "trivially fixable".
Then people could decide which ones to use based on the label alone. This would be an incentive for people to fix their libraries.
Then the process of normal attrition would take care of all the sloppy libraries.
Evolution at its finest.
Had UNIX not been free beer, history would have been quite different for C and C++.