Into the Depths of C: Elaborating the De Facto Standards [pdf]
cl.cam.ac.uk
cl.cam.ac.uk
...but...
The fact these curiosities are not an issue in day-to-day work and C is (one of) the most popular languages around today mean that they aren't too serious.
When you have a knowledge of the hardware and are working at that level day-in day-out then issues like this really don't bother you that much.
(I do realise that this is a slightly contrarian view these days, but there is an awful lot of unjustified C-bashing around currently).
8575 C or C++ packages in Wheezy. This tool found definite UB bugs in 40% of them.
How would you know that your code is sometimes misbehaving because of UB?
Some kinds of UB could be turned into something stricter, like reading a bad pointer. This either traps or returns some value. It won't format your hard drive unless you install a trap that formats your hard drive, but that's none of the spec's business. Traps can happen due to timers, so if arbitrary traps mean UB then every instruction is UB. Even if the spec punts on defining what a trap is, that's a progression over saying it's UB.
Same thing goes for division and modulo. In corner cases, it will either return some value or it will trap. It won't format your hard drive.
The most profitable "true" UB is stuff like TBAA, but smart people turn that off.
Do you know what the performance benefits are of other kinds of UB? Do you know how many of those perf benefits (like being maybe being able to take some shortcuts in SROA) can't be solved by changing the compiler (i.e. you'll get the same perf, but the compiler follows slightly different rules)? Maybe I'm not so well read, but I hardly ever hear of empirical evidence that proves the need for UB, only demos that show the existence of an optimisation in some compiler that would fail to kick in if the behaviour was defined.
Also, if there were perf benefits of the really gnarly kinds of UB, I would probably be happy to absorb the loss in most of the code I write. If I added up all of the time I've wasted fixing signed-unsigned comparison bugs and used that time to make WebKit faster, then I'd probably have made WebKit faster by a larger amount than the speed-up that WebKit gets from whatever corner-case optimisation the compiler can do by playing fast and loose with signed ints.
I suspect that UB is the way that it is because of politics - you can't get everyone to agree what will happen, nobody wants to lose some optimisation that they spent time writing, and so we punt on good semantics.
That's a good point, but arguably incomplete. It was there, to give the compiler writer leeway to implement the semantic in the way which is natural for the platform. It was not intended to play sophistic tricks on the programmer, in order to gain a few per cent of performance in some benchmark.
C became one of the most used languages, not necessary popular, thanks to the adoption of UNIX and the rise of FOSS/C culture (UNIX based) in the late 90's.
Back then it was just yet another systems programming language.
I only started to care about it when I moved from MS-DOS into Windows / UNIX, and even by then I was into C++ after a short (1 year) encounter with C.
As for not being an issue, the CVE list shows daily the cost of any programming language that "enjoys" copy-paste compatibility with C semantics.
Or the business opportunity for those that sell tools that help both developers (static analyzers) and users (anti-virus/firewalls) to overcome those shortcomings.
I think the last time a popular OS was built on something else than C was the original Mac OS, which had a Pascal API and used Pascal calling conventions.
On the desktop, C was chosen as the API language for Windows and OS/2 around 1986. That meant both Microsoft and IBM agreed that PC software is going to be in written in C.
On Mac OS (after the C transition), Windows and OS/2, they might had C as main implementation, but most of us that couldn't carry on using Turbo/Quick/HiSoft Pascal, Modula-2 or Basic compilers, moved to C++ instead.
We could still make us of improved safety and stronger type checking features, while being compatible with the C toolchains.
EDIT: Also IBM and Microsoft eventually had very good C++ support in the form of SOM, COM, C Set++ and MFC, with Borland providing the very good OWL and VCL.
Also any Windows 3.x old timer remembers the message and event handling macros alongside #define STRICT, that Microsoft used to bring some sanity to Windows programming with straight C.
Most significant DOS development seemed to be C.
On MS-DOS, on my part of the globe we only cared about Assembly, Turbo Pascal, Turbo Basic, Turbo C, Turbo C++ and Clipper.
C was hardly the only choice on MS-DOS.
All our stuff on demoscene related activities and game programming attempts were in Assembly and Turbo Pascal.
On MSDOS I never said it was the only choice - there were many choices, but most commercial development seemed focussed on C. DB work often ended up on Foxbase or Clipper. Turbo Pascal and C were hugely successful but didn't catch MS C. Somewhat surprising given how slow early MS compilers were.
Of course if you were doing DOS TSRs or games you'd be much more likely to use assembly in the mix.
Sadly the way Borland managed the company, let to us having to move to VC++ with MFC, instead of BC++ with OWL or C++Builder and VCL.
Only now VC++ is catching up with C++ Builder for UWP apps.
On MS-DOS besides the DB stuff, everyone I knew was either using Assembly, or a mix of Turbo Pascal with inline Assembly.
C and C++ only came into play on last high school year, just before getting into the university, but the majority of us already had almost a decade of coding experience by then.
I took a real dislike to Windows and MFC and moved back to the Unix side of things, so my Win programming was pleaingly brief. :)
Your experience is almost the inverse of mine - we had a few juniors comng on with Pascal as they'd learnt that in uni, but they were easy to convert to C. Just about everyone I knew in those days were C/nix or C/DOS, with just a few hanging on still trying to make a living on the Amiga - mainly games devs.
Yeah, I guess in the old days before the Internet and with expensive BBS connections, the technology had more silos than nowadays, because it was harder to move masses for any given technology.
My understanding of the narrative surrounding this is that Microsoft started working for IBM on OS/2 before switching to Windows, which - struggling here - was intended to have a certain amount of binary compatibility (?). Basically IBM decided, Microsoft went along with them, and the rest was history. Would be interesting to understand who made the decision to use C and why ... was OS/2 intended to be "unix-like"?
I literally can't remember the last time I had spent any significant time investigating one of these issues. In my experience when that a crash happen (usually in a unit test or the first time you start the app) because of these issues, the backtrace points you to the exact problem.
The pain start when the program and tests work correctly for all reasonable inputs and the underlying issue never manifests in during normal execution and can be potentially exploited by a malicious attacker with a carefully crafted input.
What I'm trying to say is that I don't want memory safety because it would improve my daily programming experience (in fact possibly the reverse would be true), but I want it because I want security.
I think you underestimate how much time is saved by not having to deal with these issues. Programming in C or C++ is frequently an exercise in writing the code, seeing a crash due to a memory safety problem, debugging it, and then repeating until you see anything resembling a working program. Writing in a memory-safe language lets you skip all that startup friction and go straight to "something resembling a working program".
Also, many of the data structures I deal with are highly intrusive (as in an object belonging at the same time in multiple containers) and short of full GC I doubt it would be easy to guarantee MS.
Then again, possibly I'm not representative of the typical C++ programmer.
This is about far more subtle issues than that - issues where there is some disagreement about whether it's OK to do or not. And the GP is right - these are often not such a problem in practice, if only because these are the kinds of issues where experienced C programmers know that they're sailing close to the wind, and there's almost always an alternative construct that's on more solid ground.
I had to deal with some of these before I was a compiler writer, and I would end up just kind of kicking my code repeatedly until stuff worked again.
Now that I'm a compiler writer, I know how to recognize what is happening, but I'm still not smart enough to avoid the bugs in general and I still spend time fixing bugs that result from these issues.
So, I'm with pcwalton: it is a problem. Maybe I don't see it every day, but I see it probably at least once a month.
https://www.cl.cam.ac.uk/~pes20/rems/
Cerberus main page and links are here:
https://www.cl.cam.ac.uk/~pes20/cerberus/
Work like this will eventually, if not already, be applied to other projects along the lines of CompCert, seL4, and static analysis. The models of real-world assembly and C come first. Then, other tools map C to assembly or specs to C. So, this is pretty fundamental stuff they're working on. That they try not to abstract away the dark corners is the real advance here as many try to cheat. :)
Note to ingve: One of those Jung-style coincidences that you submitted this around exact time I wrote up same project for Schneier's blog. I've only looked at it twice in its existence. Odds were slim we think & write around same time. Always find it interesting when that happens.
> If you zero all bytes of a struct and then write some of its members, do reads of the padding return zero? (e.g. for a bytewise CAS or hash of the struct, or to know that no security-relevant data has leaked into them.)
(and 14 other questions)
Webpage of the project: http://www.cl.cam.ac.uk/~pes20/cerberus/
And if you think you can answer the question without resolving the ambiguities, that answers some other questions.
Moreover, the study was explicitly not about ISO C: "We were not asking what the ISO C standard permits, which is often more restrictive, or about obsolete or obscure hardware or compilers. We focussed on the behaviour of memory and pointers. This is a step towards an unambiguous and mathematically precise definition of the de facto standards: the C dialects that are actually used by systems programmers and implemented by mainstream compilers."
Here is an actual example of a comment to this question:
I would expect this code to work:
struct foo
{
char a;
double b;
};
foo p;
foo q;
memset( &p, 0, sizeof( p ) );
memset( &q, 0, sizeof( q ) );
p.a = 1;
q.a = 1;
assert( memcmp( &p, &q, sizeof( foo ) ) == 0 );Q64. After an explicit write of zero to a padding
byte followed by a write to adjacent members of
the structure, does the padding byte hold a
well-defined zero value? (not an unspecified
value)
#include <stdio.h>
#include <stddef.h>
typedef struct { char c; float f; int i; } st;
int main() {
// check there is a padding byte between c and f
size_t offset_padding = offsetof(st,c)+sizeof(char);
if (offsetof(st,f)>offset_padding) {
st s;
unsigned char *p =
((unsigned char*)(&s)) + offset_padding;
*p = 0;
s.c = 'A';
s.f = 1.0;
s.i = 42;
unsigned char c3 = *p;
// does c3 hold 0, not an unspecified value?
printf("c3=0x%x\n",c3);
}
return 0;
}
Some of the questions have clear answers with respect to either the ISO or de facto standards, but many do not - that's the point.