Strict memcpy() bounds checking for the Linux kernel
lwn.net
lwn.net
C uses (pseudo)fat pointers all the time, any time an array(pointer) is passed with size (as a second argument). It's just not as hygienic. Just think of all the dev hours, all the painful debugging, all the exploits prevented, if in the 80's they put out a stdlib with a fat pointer buffer type. Can't argue they didn't know better, as plenty of languages had them by then.
Unfortunately Intel keeps borking their attempts to do the same.
For example if a routine wants to quickly zero fields pkt.mask pkt.sent pkt.copied pkt.recvd, pkt.ackd, and pkt.ready but "obviously" not the other members of the pkt structure, somebody will write code that zeroes out the whole stretch of the pkt structure from "mask" to "ready" inclusive, and then comment in the structure definition that it's important not to re-arrange these members. And that's faster, and it works.
Now, suppose I'm a bad guy & I'm able to fool a routine into calling the zero-out function with any range I like instead of just the range from "mask" to "ready". Memory tagging might prevent me clearing the adjacent structure, but all I want to do is smash pkt.uid to zero so that I can become root and that's just part of pkt...
I thought it was the programmer to make mistakes, not the language to be error prone.
Programmers making mistakes is a fact of life, that C is error-prone is what makes those mistakes into significant and recurring issues.
The problem with C is that it's a tool where you have no safety and I don't think there's a programmer alive that can actually use it in a safe manner. I guess some are close, but they will still end up with the occasional wound here and there on their bodies...
I would probably avoid using C for anything I put in production today, I still love the language though. It still feels special, sitting down with all that power at your fingertips knowing that you'll have a built in buffer overflow if you lose focus for a second, gets (no pun intended) my blood flowing! :-)
I strongly disagree. I think it is absolutely legitimate to blame a tool which is practically impossible to use correctly [1]. Some of C's design decisions were justified in its historical context, and some – like zero-terminated strings – were indefensible even back then.
[1] Alternatively, you could blame its creators, but that's not very useful.
Snippet from "https://dave.cheney.net/2017/12/04/what-have-we-learned-from...":
One can write a string copy routine using two instructions, assuming that the source and destination are already in registers.
loop: MOVB (src)+, (dst)+
BNE loop
The routine takes full advantage of the fact that MOV updates the processor flag. The loop will continue until the value at the source address is zero, at which point the branch will fall through to the next instruction. This is why C strings are terminated with zeros.The biggest issue with C isn't its footguns, rather the WG14 unwillingness to provide additional language or library features that would allow for a safer C outside the low level code where it pretends to be a portable macro assembler.
Besides, the x86 "repeat while" string instructions continued the PDP legacy.
Or like my spouse likes to joke when we get in the car to leave and forgot something in the house: "it's too late, the door's shut."
I get why null terminated strings once existed. It's baffling that they continue to exist 50 years later. Not to mention, they don't even work on data buffers, so you need fat pointers anyways!
It may be a good idea from a theoretical standpoint, but once you start calculating the cost it simply doesn't make sense.
Thankfully they are starting to pick up.
There's even plenty of means for backwards-compatible strings and arrays, such as sds.
But really the point was more to "this should have really been addressed decades ago."
Maybe liability and lawsuits are really needed to stop with such excuses.
Maybe it's better to just use a better tool for new code bases?
What I meant to say, in a rather roundabout way to be fair, is that the problem is both the language and the programmers. C is a tool that is too hard to use correctly, but the programmers who write crappy C code are to blame for their crappy code. It's another thing if an expert fails to use the tool safely, then one might blame the design of the tool used.
There seems to be 2 camps, one camp blames C and one camp blame it on poor programmers. Poor programmers have given C a reputation and the tool itself is too hard to use correctly, even for experts, so both camps are right and also wrong...
I think some of the decisions made for C back in the day where fine, C was designed to be lightning fast and close to the metal, but I think it's time to pivot. I don't think it's necessary to sacrifice security to squeeze out the last percent of "speed" today. C has a lot of legacy code still in use though, so I don't think it will ever happen, that's why I use other languages for production code today and only write C when needed for C code bases still in use.
But we'll live for a very long time with C code, it will most likely outlive us all, because important infrastructure code is never really replaced, it just becomes a new sediment layer.
However just like rotating blade covers, butcher metal gloves, seat belts, helmets,..., apparently external forces like government regulations are required to make WG14 act accordingly.
Unfortunately, the large majority to this day thinks they don't need such kind of tooling.
"Although the first edition of K&R described most of the rules that brought C's type structure to its present form, many programs written in the older, more relaxed style persisted, and so did compilers that tolerated it. To encourage people to pay more attention to the official language rules, to detect legal but suspicious constructions, and to help find interface mismatches undetectable with simple mechanisms for separate compilation, Steve Johnson adapted his pcc compiler to produce lint [Johnson 79b], which scanned a set of files and remarked on dubious constructions. "
I always find takes like this bizarre. We're constantly improving the safety of the tools and devices we use. Our history is filled with examples of tools designed without safety in mind leading to deaths or maiming or other injuries. Thankfully, the people who came before us saw that it was up to us to reduce the likelihood of accidents by improving the tools that we use.
Just look at the aerospace industry. They didn't say "well, don't blame the plane, the pilot should've gotten it right". Often times improvements in planes were to avoid common mistakes because of how fallible we human beings are, instead of holding us up to impossible standards.
The difference is that engineering over all is mature enough that you can make incremental changes over time and it still does not invalidate what you did previously. If you built a bridge 10 years ago its probably safe enough to leave standing even if you can build a safer bridge today.
The same cannot be said for programming yet. It's hard to improve the safety of the tool, the C compiler, without breaking your past bridges or having to redo the work again, i.e rebuilding the bridge.
So a modern railway engineer inspecting a Victorian bridge has a problem. The bridge was built in the usual fashion of the time, and it's impractical to fully inspect the load-bearing materials without dismantling the bridge. There is no detailed paperwork because the Victorians didn't keep any.
Still, it stands to reason that if the cast iron load structure exposed in one place has 15-20 years of life left in it, the unexposed structures you can't see are similar and this bridge can be scheduled for replacement in say 10-15 years. Right?
And then, one night, as a fully laden freight train crosses it, the bridge collapses. The driver feels something wrong on the bridge and then, a few seconds later, the locomotive automatically brakes to a full halt - unable to sense the rear of the train which is in fact now laying in the rubble of the broken bridge.
The Victorians saved a little money by using thinner metal for the unexposed girder, which had therefore failed earlier than the predictions based on the thicker metal.
Documentation is essential. If your 30 year old C project doesn't have adequate documentation chances are you don't know whether those unexposed elements are as strong as the parts you can see or if they're paper thin and likely to fail at any moment.
Additionally they use coding practices that would make the most hyped TDD advocates from Silicon Valley startups walk away from the projects without looking twice about what they were leaving behind.
When code kills, every line of code gets validated.
https://en.m.wikipedia.org/wiki/The_Power_of_10:_Rules_for_D...
This is not the kind of C you will find in FOSS or UNIX software.
The problem with C is the two generations of standard committee members have stubbornly shirked their duty to fix these problems.
Example #1 the array type looks useful, but oh, it's actually just coerced into a much less useful pointer
You can have a C function signature that says it takes arrays of sixteen integers. But it doesn't! The compiler cheerfully ignores this and uses the function for any pointer to integers. Sixteen integers, or Zero, or Sixty, the same function executes.
If you've used enough C you might just assume that's how it has to be. Nope. A Rust function declared for arrays of sixteen integers takes... arrays of exactly sixteen integers like you asked for. Not any other size.
Example #2 implicit narrowing conversions everywhere
This blew up in Linux recently, but C is OK with the idea of just assigning 64-bit values to 32-bit integers for example. The value doesn't fit in the narrower variable, but no problem just throw away the extra bits and it'll go in...
https://www.godbolt.org/z/7bcd9xPr3
For example #2, proper compilers have warnings for implicit conversions which would lose data, for instance:
https://www.godbolt.org/z/vEP8ra9MP
But for a proper solution to the array size problem you need array slice support built into the language, and this would be the first "opaque builtin struct type" in C (because slice types need either a pointer/pointer or pointer/size pair). At this point it's really better to switch to a different language.
And yes, it really is better to switch, I've been doing so, and I am hopeful that the kernel can follow in due time.
Aside: Matt Godbolt deserves some sort of award for Compiler Explorer. So much less awful to just link these examples in such conversations. It might well be as big a contribution to software engineering as git bisect or Tinderbox, or (really going back) Grace Hopper's "compiler". Yes Compiler Explorer seems "obvious" and people talked about things like this but did you build it? 'cos just talking about it doesn't make my life any better but building it does.
Here is one talk from Matt Godbolt about how Compiler Explorer was born,
CppCon 2019: Matt Godbolt “Compiler Explorer: Behind The Scenes”