>
for my own code I think the biggest improvement to correctness would come from fat pointers with array size information, enabling bounds checking.Yes. In fact, I suspect that's the single biggest thing most C programmers could do to avoid CVEs. Rust does something similar with bounds-checked "slice types" like &[T], and it makes a huge difference.
For example, I wrote a VobSub subtitle decoder in Rust and I fuzzed it with 'cargo fuzz' and AFL. VobSub is a hairy binary format, and it's easy to make mistakes. AFL found 5 bugs:
- 3 were incorrect uses of Rust's &[T] type, all of which were detected by the bounds checks.
- 2 were integer overflows, both caught by the fact that Rust panics on overflow in debug builds. Neither of these looked exploitable—one was harmless, and the other one would have been caught the first time it tried to index a &[T].
So Rust's borrow checker is useful. But the single biggest win turned out to be having bounds-checked array slices, as you suggested.
> By the way, for a project the size of an XML parser I wouldn't feel shame for such a tiny list of CVEs :-).
I didn't write the XML parser. That was written by James Clark, who was something of a legend in the SGML and early XML world. So even if you pick your C dependencies carefully, it's still risky. (And my code has been fuzzed much less than Expat.)
But to paraphrase various recent medical claims, there is "no safe level" of CVEs.
> I've never done fuzzing (I might try it).
I recommend it to anybody who works with pointers or who parses untrusted data. It's fun to watch AFL generate a billion(!) test cases and slowly ferret out "impossible" conditions. But it may also make you completely paranoid about bugs.