“Fiercely resist any further broadening of the scope of the C UB problem”
lists.debian.org
lists.debian.org
Regardless of who is right and who is wrong in this matter, I think if everyone took a step back we could at least agree that it makes absolutely no sense to fix this on a (single Linux) distribution level. For Debian to configure/patch compilers on their platform to "narrow" undefined behavior is insane and ineffectual. Software isn't "validated" on/against a particular OS, it's validated on a compiler basis.
Breaking this assumption introduces a massive schism. While Debian is an amazing distro with plenty of clout (I'm a FreeBSD guy, but Debian comes second), it's terrifying to imagine a new generation of "cross-platform" C/C++ software that can only be verified working on Debian (or with Debian's fork/re-configured compiler). We've come so close to making truly cross-platfrom C++ code a reality (even bringing Windows, I repeat WINDOWS, into the fold) with C++11 (and the subsequent releases) and it's, in my humble opinion, utter folly to try and change the way code will fundamentally compile depending on the distribution you run.
If Debian cares, make a proposal to the C++ committee, bribe^H convince members to see their way (or threaten^H blackmail^H show them the dangers of continuing down the road they're on). Heck, fork C/C++ and call it E or C+++ or c-safe or something - or more reasonably - write a tool to convert C to rust or D-without-the-standard-library and announce only tools in that/those languages will be allowed in the standard distribution. But for Heaven's sake, please don't try to redefine C.
I am not a C++ expert but have tinkered with C++11. How will portable C++ work with the Windows UTF-16 (wchar_t) and other systems using UTF-8 (char_t)?
Is there a way to have standard C++ portable across Windows' UTF-16 and !Windows UTF-8 without #ifdef'ing a char wrapper?
edit: UTF-16 not UCS-16
For a more workable approach read this article, especially the section "How to do text on Windows".
So then the compiler developers give up and implement stuff 1) according to the standard document 2) such that performance on SPEC cpu is maximized.
That, or then everybody switches to Rust. :) (hey, I can daydream can I?).
If you want to get people using a Friendly C, you need to start convincing the biggest users of C and C++ that they should give up performance for a simpler dialect of the language. Like politicians responding to voters, compiler authors respond to what their customers demand. Up to now, their customers have demanded performance. It's not their fault for listening to them.
Why would would C++ programmers have to give up anything? You may notice that the frequent complaint (and title of the article) is "undefined behavior in C", and the hypothetical replacement language is "Friendly C", not "Friendly C++". From the point of view of most who are troubled, rightly or wrongly, C++ isn't considered relevant to the problem or the solution.
I think part of the "divide" is that compiler writers (and probably C++ programmers) are more likely to lump C and C++ together. This makes sense, as C++ is mostly a superset of C, and since many of the optimizations being questioned operate at the level of internal intermediate language that's the same for both. But many C programmers don't view them as being the same language at all, and have no opinion on how C++ compilers should operate.
Perhaps what needs to be questioned is assumption that the needs of C and C++ programmers can effectively be served by the same compiler?
(I realize I've responded strongly to two of your comments in a row. I care strongly about the issue, but don't intend this to be an attack. Rather trying to convince you that you're wrong, my goal is just to understand what produces the gap between our viewpoints.)
I presume there are some, but they aren't the sort of thing I have in mind. I'm talking more about the expectations that the users of each language have. For whatever reason, C programmers are much more likely to complain about "broken" optimizations than are C++ programmers.
Of course this isn't absolute, but I think it's undeniably a pattern. I'd guess this is because C has a heritage of being "portable assembly", and thus many programmers have an expectation of a 1:1 correspondence between the code they write and the finished product, and are startled when it doesn't.
In the case of explicit null checks being removed and loops being removed resulting in memory not being zeroed, I think they have a point. Perhaps there is some way to apply different levels of optimization to the code that the programmer writes versus code "generated" code?
I disagree - it is almost entirely their fault for listening to them. Customers usually don't know what they really want. "Faster horses". The whole point of having an oversight committee is to steer the language, and they failed at the task.
Rather than steering towards sanity - which C/C++ badly need - it was steered towards performance at pretty much all costs. The result is a broken language: it is extraordinarily difficult to write correct programs, and assure that a program will continue to compile as intended many years in the future. Sanity, correctness and fragility are exactly the kind of things which oversight should have tackled, but instead we got the opposite.
I say this as a decades-long user and fan of C/C++. But I've come to realize recently that it's a lost cause, because those in the position to fix the language are powerless against the wishes of their customers.
And this can work both ways. Friendly C would almost certainly forbid non-compatible type pointer castings, which would actually enable additional optimizations.
I doubt that. Autovectorization certainly does help performance of real code, for example.
> And this can work both ways. Friendly C would almost certainly forbid non-compatible type pointer castings, which would actually enable additional optimizations.
I'm not sure what that means; could you elaborate? That sounds like strict aliasing, which is one of the things "Friendly C" is opposed to…
I wasn't aware of that and my idea of it is different.
Aliasing rules help the compiler decide what values have to be reloaded. The stricter those rules less chance there is for pointers to alias the same memory.
Thanks!
What resources are you using to get into Rust? I've been curious for a while, but I'm not sure where to start.
But I think it depends a lot on your previous experience. If you have experience of other languages, particularly c/c++ and some functional style thing like Haskell there isn't that much of a mental barrier IMHO.
> There are two ways to evaluate the the C specification's rightness and properness. [...] The second is to ask what is most useful. And there again the C committee have clearly failed.
Just two days on HN we saw this article: https://news.ycombinator.com/item?id=11468603 In which it says:
> no matter the kind of software, performance is almost always worse than our customers would like it to be.
That is why all of this is happening. There is a market demand for performance. Compilers that increasingly exploit UB in C is just a manifestation of that market demand.
You can't just shrug this away with "Well, if you want performance...", because people in fact don't want abstract "performance"... they want the language they were truly writing in to perform well, not for what is de facto a different dialect of the language to suddenly appear and replace the language they were using.
Is wanting a for loop setting an array to zero to optimize into memset optimizing the language they were truly writing in? I think it is. But that optimization frequently depends on undefined behavior.
UB exploitation usually exists because people filed bugs on compilers complaining that they didn't optimize some case they expected to optimize.
Yes, this is a good optimization, since it efficiently does what the programmer intended. The bad optimization is removing the security-essential loop altogether when the compiler notices that result appears unused, and sensitive information is left susceptible to later attack.
UB exploitation usually exists because people filed bugs on compilers complaining that they didn't optimize some case they expected to optimize.
I doubt that anyone has ever filed a bug saying "I explicitly wrote a loop to zero memory, but the compiler failed to optimize it out." If you know of one, please point to it. I think you are throwing out the baby (intentional C) with the bathwater (autogenerated C++).
That optimization is a natural consequence of SROA and DCE. If you claim you don't want those optimizations, I don't know what to tell you. Those optimizations are some of the most basic, critical optimizations any modern C/C++ compiler does and throwing them away can easily result in at least a 2x performance loss.
I suspect it's less that they asked for that than they asked for some generic optimization on loops, and the implementation works great 99% of the time, and 100% of the test cases, but because it was approached from a performance perspective the idea that sometimes we want to change the state of the memory, even though it has no affect on a correctly functioning program, isn't considered. Zeroing memory in this way has nothing to do with correctly running programs, the only use is to defend against your program malfunctioning and exposing the data, or some external actor looking at memory. Neither of those have to do with the normal functionality of the code the compiler is running, but is is important. Unfortunately since it deals with things that have nothing to do with the actual instructions needed to make the program function, it's easy to overlook.
A lot of dead stores happen because of post-inlined, post-constant-propagated (etc.) code. Nobody writes that code by hand; it occurs due to inlining and other aggressive types of optimizations. You need those optimizations for performance; IPO is a huge win.
It's like the old joke: "Who writes that kind of code? Macros do."
If we disallow all those optimizations, we end up with hugely wasteful programs in the current C ecosystem. Enforcing this is useless, the market will route around your best intentions, because no company that's making money will be willing to cede a huge performance gap to their competitors.
If we require code be annotated to bypass optimizations, we run afoul of the fact that it's impossible to know what optimizations might be developed in the future that will affect portions of our code that might be security sensitive.
The problem is insidious and deep. Imagine an architecture that allows specialized instructions to speed up zeroing chunks of memory, but does so by hardware mapping different pages around behind the scenes (to swap in a pre-zeroed page for a filled page). Correctly detecting that the architecture has an efficient method for zeroing large chunks of space could be detected and used by the compiler, but the hardware is working against us because the original memory may still be accessible to an external actor. C is ensuring the memory addresses you requested are zeroed, it's just not ensuring that the data that existed within them is no longer anywhere on the system.
There's also the issue of the store changing the array base, as in Chris Lattner's example on "What Every C Programmer Should Know".
This is the example you're talking about in "What Every C Programmer Should Know":
float *P;
void zero_array() {
int i;
for (i = 0; i < 10000; ++i)
P[i] = 0.0f;
}
If you make the global variable static, then it gets converted to a memset with -fno-strict-aliasing, because the 'P' variable never has its address taken. If you used LTO with sufficient linkage information, the same thing would happen. Similarly, if you pass the array as an argument then the loop gets converted to a bzero call just fine without -fno-strict-aliasing: float *P;
void zero_array_impl(float* a) {
int i;
for (i = 0; i < 10000; ++i)
a[i] = 0.0f;
}
void zero_array() {
zero_array_impl(P);
}
I don't think this is a very good example for these reasons.What does this have to do with my comment, at all?
> You can't just shrug this away with
Who is shrugging?
You don't understand a complex problem until you understand both sides. The article only presents one side. I am presenting the other. It doesn't mean I think the problem is simple or easy.
That has not been my experience. Do you have an example of this?
Slower - the various methods of checking for integer overflow are slower than x = a + b; if (a < x) {/overflow*/}. It can also be easier to just use a slower, or more memory hungry way of doing things than a better but needlessly complicated method.
At my current job we've sacrificed memory (and in this case, performance as a consequence) because the complexity cost of keeping some critical code in the defined behavior realm was too risky.
As for integer overflow at least gcc and clang have built-in functions that generate optimal code by checking the CPU's overflow flag: http://clang.llvm.org/docs/LanguageExtensions.html#checked-a... and https://gcc.gnu.org/onlinedocs/gcc/Integer-Overflow-Builtins...
If the member used to access the contents of a union object is not the same as the member last used to store a value in the object, the appropriate part of the object representation of the value is reinterpreted as an object representation in the new type as described in 6.2.6 (a process sometimes called “type punning”). This might be a trap representation.
(copied from a comment on blog.regehr.org)
This is actually a really common area of misunderstanding. But ultimately, one of the main uses of unions is to provide multiple ways to access the same data.
where answers try to distill the standard into something legible by humans. Unsurprisingly, they come to different conclusions.
The overall interpretation in some situations, if you look at things just right, you can do type punning in C++. But it's pretty risky since you're a LOC away from the compiler deleting your function.
Or, you can't really do it at all safely. Which I would conclude from the above anyways.
[1] http://xenbits.xen.org/gitweb/?p=xen.git;a=blob;f=tools/libx...
P.S: In C++ std::less is guaranteed to do the right thing(tm).
That's true, if performance were the only consideration, we would all be writing assembly.
Edit: All the responses are right. Feel free to downvote this ill-considered comment into oblivion, now that I can't take it back.
This is patently a ludicrous argument. How can the speed of encoding an MP3 be used as a proxy for countless other applications of computers?
Programs that are slower than they need to be waste real energy and other industrial resources.
There are millions of computers in data centers around the world. They are working on problems that someone thought it was worthwhile enough to pay for. If C compilers emit code that is slower, that means buying and maintaining a lot more computers, using more energy, producing more greenhouse gas, etc.
Which was mostly GPU code, so it was either CUDA or OpenCL btw. Not actually C.
People do need to realize that the fastest, massively parallel supercomputers are in fact written in CUDA / OpenCL. And the most powerful supercomputers are mostly being used to display 3d-models of a sword-wielder fighting a guy with guns.
Daniel J. Bernstein made the same comment in his slides about "the slow death of the optimizing compiler"[ https://cr.yp.to/talks/2015.04.16/slides-djb-20150416-a4.pdf ]. Basically, a lot of programs are I/O bound. Those that aren't are often dominated by a tiny amount of super-hot spots that can be optimized by hand via assembly language. Most code is "freezing cold" -- it's hardly worth optimizing. C compilers are giving up safety and getting very little in return, except shiny new compiler benchmarks.
Just to stay within the domain of codecs, YouTube would not be a nice experience on a 366 MHz Pentium laptop with primitive compiler optimizations.
There is a market demand for performance.
You could fulfil that demand without breaking existing code by making the riskier optimisations configurable and turned off by default.-O3 is for optimizations that are mostly not risky.
-O2 is for optimizations that don't involve a speed/size tradeoff
-O1 is for optimizations that don't slow your compiler down too much
Obviously these are fuzzy categories, but -O4 is where "undefined behavior" optimizations belong. Alternately, each individual optimization has its own flag, and can be turned on individually.
Currently, there are no optimizations that fit in that category, but as recently as 2003 I remember seeing it in the man pages. I believe there was -O5 and -O6 too, but I don't trust my memory enough to be certain.
Maybe I remembered the history wrong. There are definitely cases of people using -O4 http://www.larcenists.org/Twobit/KVW/kvwbenchmarks.html .
There are definitely optimization options which are (or have been) available that are not turned on by -O3. For example:
-ffast-math
-funroll-all-loops
-fstrict-aliasing
Are not turned on (for good reasons: unrolling loops often slows a program down).
So there's no reason UB optimizations need to be turned on by default, or with -O3. You could throw them in -Ofast or use -Oub, or use optimization specific flags (they should have optimization specific flags in any case) and that was my main point.
That is a very different thing that turning on optimizations that might break non-standard-conforming programs!
I think the idea of making it controlled by optimization flags is interesting, but you would need to formalize what the more "safe" variant of the language is if people are actually going to depend on it.
You're talking as though the standard were well-defined. It's not.
I've had this same argument with cache management specialists. If you have to consider how the cache works for your product to work, you're doing it wrong.
At least in c++, this is absolutely 100% not the case, there's a huge difference between -O0 and -O2.
I asked a C++ and compiler expert about this once and he told me: "I believe both GCC's and LLVM's LTO will happily cross this barrier, so it doesn't offer you any real protection from their optimizers."
There is no single standard that defines what UB is for a mixture of C and C++. There are clearly some parts of C++ that are trying to improve interoperability with C, like "standard layout" classes. The best you can do is try to simultaneously follow the rules for both languages when you mix the two.
LTO is done on the compilers internal representation of the code (IE. GIMPLE or LLVM IR). This representation is generated based on the rules of each language, and optimizations are performed on this representation instead of the original source representation. Both C and C++ (and anything else) are converted into these representations. LTO simply keeps the GIMPLE or LLVM IR around until link time, and then when the program is linked optimizations are performed over the entire representation of the program. Crossing the language barrier shouldn't be a problem, because the GIMPLE has it's own rules to follow to make sure everything still functions the same. Once you reach this point both languages are already compiled in the practical sense, they're just not actual machine code yet. I would expect however, that because of the differences there are far less optimization opportunities to be taken advantage of.
// foo.c
#include <stdlib.h>
typedef struct { int x; } s;
s *make_s() {
s *ret = malloc(sizeof(*ret));
ret->x = 5;
return ret;
}
void free_s(s *val) {
free(val);
}
// bar.cc
// This class is standard-layout and matches "s" from C.
class C {
int getX() { return x_; }
private:
int x_;
}
extern C* make_s();
extern void free_s(C* c);
int main(void) {
C* c = make_s();
int ret = c->getX();
free_s(c);
return ret;
}
C++ says you can only call a method on an object whose lifetime has begun. Can C begin the lifetime of a C++ object?Yes. C++14 Standard §3.8p1:
...An object is said to have non-vacuous initialization if it is of a class or aggregate type and it or one of its members is initialized by a constructor other than a trivial default constructor... The lifetime of an object of type T begins when:
— storage with the proper alignment and size for type T is obtained, and
— if the object has non-vacuous initialization, its initialization is complete.
Since class "C" is standard-layout and has a trivial default constructor (in this case none at all), assuming that malloc allocates memory that is suitably aligned for "C", the function make_s returns a pointer to an instance of C++ class "C" whose lifetime has begun.
Edit: Thanks for asking btw. I use similar code in a project of mine but never actually checked whether it is actually standard-conforming.
1) Write your own damn spec with no "problematic" UB. Hookers and blackjack are optional
2) Write a compiler for your shiny new language.
3) Start writing/porting code to your new language
There you have it, UB problem solved once and for all.
I'm not sold over on Regehrs work, but at least he is doing something with his Friendly-C proposal. This has the added benefit that we get something more concrete to discuss and debate about instead of vague "optimizations breaking my code are bad".
It is completely unfeasible to try to turn back time on C and somehow magically make compilers deduce programmer intent from some random crap that you throw at them. Like it or not, computers are based on rules, and standards are the best way we have to establish those rules among large number of parties.
"Compiler authors will likely support this as long as possible, but the peer pressure of needing things to go faster and faster will likely push them to exploit more and more undefined behaviour to their advantage in the future. Their argument will be: 'After all, who has sympathy for those who don't follow the standard?'."
From: http://blog.robertelder.org/signed-or-unsigned-part-2/
Lots of people disagree with me, but its nice to have less UB issues to worry about far in the future.
We automatically distrust the compiler (synthesis tool) to do the right thing. You formally prove the 'compiled' output (logic gates) that will be manufactured matches with the source code of the design (verilog/vhdl) using tools written independently to the compiler.
This isn't easy, and I know the problem space is larger, but does anyone ever do this for software?
In the general case, no. The economics of software don't favor it. There are tools that inspect code and make helpful suggestions.
This is why you can't trust your synthesis tools, BTW.
But much worse in software is the cultural ... "zero" ( as in a zero in a filter) about the axis of provability in general. It's a point of despair.
I can tell you that in multiple cases, I was able to represent the "core" logic of systems much that I could build a test rig around it and do exhaustive testing ( with the caveat that it's only as good as the test framework ). Permutaitons are reasonably cheap these days. But to do this, you must nearly eschew the us of third party code and in cases, even large parts of standard libraries.
But the standard answer is to despair and moan "it's impossible." The "prove it correct" people and the "git 'er done" people are two different tribes.
In both cases the compiler is pretty much part of the trust base (which is a problem because they're annoyingly complex), but the issues discussed here are declared invalid by the subset (ie. you mustn't use statements that may lead to undefined behavior).
For seL4, mentioned in another comment, there was proof that interesting properties of the high-level code and the low-level binary were equivalent. That only works with a relatively static compiler version and for some optimization levels (anything that optimizes too globally will seriously mess up such attempts of showing equivalence), but it takes the compiler out of the trust base.
Typically the state space is simply too large to be practical. It is, admittedly, how a lot of verified software is done (because there are so few verified compilers), though often then the source-code to assembly correspondence is verified manually, which limits possible optimisations.
Perhaps you could offer a sense of how this compares to the hardware testing practices you use?
A secondary check is that the source that you functionally tested is logically equivalent to what you manufacture. This is where you are not checking your code, but the issue is trust of compiler/ compiler optimisations in synthesis. It needs to be redone if you recompile - that's the step I don't really ever see in software development - if I use a different compiler option or underlying instruction set architecture to the SQLite Dev team, do I still trust my binary?
Of course the level of paranoia is far higher in hardware where it costs multiple millions of dollars to crank out a new spin of a chip!
If that is critical, you can join the SQLite Consortium Membership for $75K/year and access to the test suite. There's also an option to pay SQLite developers to "run TH3 on specialized hardware and/or using specialized compile-time options, according to customer specification, either remotely or on customer premises." The TH3 test harness is an aviation-grade test suite.
The level of paranoia for aviation software is also rather high.
Ouch. I usually hear that about C++, not C. =[
I don't think I'm the only one.
Ada is as safe as Rust and lower-level but very fast and very mature. It gets a bit of a bad rap for being ugly, but after a couple hours the Pascal-ish syntax is not any uglier than C. The type system is much less expressive than Rust's, however, which makes a lot of things more awkward. What you lose in conciseness you make up for in performance and readability. Ada does require you to be extremely explicit about everything, which many people do not like. Still, if you're looking for a language around the abstraction level of C but safe and easily auditable, Ada is a quite well-designed language.
D is basically C++25. There's still some undefined behavior, but you hit it much less often. Also it has excellent metaprogramming support. It is not as safe as Rust, but a little safer than C++. Overall it's a very pleasant language if you're a C programmer and want something higher-level.
OCaml also has a syntax that many people find unpleasant at first. Personally I got over that pretty fast and now fine ML syntaxes beautiful. OCaml is essentially garbage-collected Rust with prettier syntax. It has predictable performance (in the sense of being fairly easy to predict what the compiler will generate—the GC will make actual runtime performance a little unpredictable, but isn't bad; probably because OCaml programs tend to involve much less GC pressure than Java or C#).
I don't think Go should be lumped in with those languages though. It is not very safe (it has a terrible error-handling mechanism and doesn't isolate unsafe code particularly well), it has a large runtime and slow generated code, it offers piss-poor abstractions, etc. If Go becomes the next popular systems language, language design will be set back another 30 years.
When programming Ada my thought process is relatively close to C (but with proper modules and generics). From what little Rust I've used, it was more like an ML—specifically more usage of higher-order functions and discriminated unions (which Ada supports, but is 100x less convenient than ML, but still 10x better than C's struct { enum { } what; union { } val; }; pattern).
In general, memory management is more manual in Ada than in Rust as well. Actually, Ada is technically less safe than Rust since you can use Ada.Unchecked_Deallocate to circumvent the accessibility checker and Ada doesn't have an equivalent of the unsafe { } block. But you almost always use RAII and memory pools and Unchecked_Deallocate for C types gets hidden in a RAII wrapper. Additionally, fat pointers are opt-in (though roughly as inconvenient as thin pointers), which can be useful for very low-level code. I believe Rust's pointers are always fat?
Also I went ahead and checked Rust performance again. I didn't realize you'd gotten so fast. Ada, like FORTRAN, does not allow pointer aliasing by default (there's an ‘aliased’ type qualifier though), so theoretically it can generate faster-than-C. Embarrassingly I don't actually know how alias analysis works in Rust, so maybe you have this advantage too.
One thing that you may not realize, talking about HOF and such: LLVM is very good at compiling them down. If you use a closure that doesn't actually close over anything, it will get turned into a regular old function, for example. Which means it can be inlined...
Rust pointers are not always fat. Slices are, but &T is a regular old pointer, as far as the assembly goes.
&mut T automatically gets `restrict` applied to it, basically, we do a lot of aliasing stuff as well.
Thanks for elaborating :) I should spend some time with Ada.
It also has support for hardware interrupts and real time programming.
A fellow stands up at the end and says, in detail, something that can be summarized as "it's all golden until you invoke the I/O monad."
Yes, most languages do, unless they are formally defined (ML is formally defined, but most other languages are not. In this book http://www.amazon.com/Masterminds-Programming-Conversations-... several of the language creators say that fully defining the language formally is not worth the effort).
Unlike C, Undefined Behavior is pretty limited in scope in Rust. All the core language cares about is preventing the following things:
- Dereferencing null or dangling pointers
- Reading uninitialized memory
- Breaking the pointer aliasing rules
- Producing invalid primitive values:
- dangling/null references
- a bool that isn't 0 or 1
- an undefined enum discriminant
- a char outside the ranges [0x0, 0xD7FF] and [0xE000, 0x10FFFF]
- A non-utf8 str
- Unwinding into another language
- Causing a data race
And all of those things require `unsafe`, so safe rust cannot do any of them (barring compiler or unsafe library or OS bugs).Edit: And I must admit I don't know much about language theory or formal definition, but there is also a self-described formal grammar [2].
[1]: https://doc.rust-lang.org/nomicon/races.html [2]: https://doc.rust-lang.org/grammar.html
let a = Int.max + 1
Produces a compile error as the overflow is within Swift's conception of a "literal expression".If we step out to a case where the overflow is not trivially a literal expression:
func foo(a: Int) -> Int {
return Int.max + a
}
foo(1)
This causes a runtime assertion failure in optimized or unoptimized builds.There is a (rarely-used) "Ounchecked" optimization level for which this program produces UB.
let x = std::u32::MAX + 1;
This line will generate a warning upon build, and in a debug build, will panic. In a release build, it will overflow.I think the C coding mindset should be this: if an approach requires code that isn't unambiguously defined, change the approach. If this means more boring coding, to circumvent cute tricks, so be it. If you need that extra 5% speed, use assembly instead of bending C.
A couple months ago, as a curiosity, I watched a few videos in a series on programming in C in Windows environment. The teacher was a serious programmer, but the first thing that went out the window was strict aliasing. After that assumptions of integer sizes and range started creeping in. It came apparent that the teacher knew C, but only superficially, signedness, integer promotions and usual arithmetic conversions were treated like a nuisance. If the code compiled and ran, it was good. Those videos were the first C programming experience for at least several hundred people.
That can be said almost any time people are trying to find fault. Better to just look for a solution, and not worry about whose fault it is.
I can see the C language being forked into the one we are getting and the one users actually want.
Wouldn't it be great if "memset(0,p,0)" was a harmless no-op? It was for time out of mind. But no more.
For bonus points: Can we have a #pragma that tells the compiler to abort with an error if the target machine uses any representation for signed integers other than twos-complement?
Here is a car analogy. It is like learning to drive and thinking that you are allowed to drive over a yellow-turning-to-red light. However the rules say, you should stop if you are physically able to. In reality almost everyone tries to get over than yellow. On a rare occasion they get spotted and pay the fine.
C strives for maximum performance, it will not check things for you if you don't ask it to. Having a library function that performs those checks for everyone, would go against that rule. Why would someone else have to pay the performance penalty for you? Write a wrapper or a macro that performs the check, couple of lines, it is that easy.
If you strive to write portable C code it should work regardless of signed representation. C defines all the range macros for types, and gives you types that guarantee certain ranges. Assuming you use those types and macros, I'm really curious what incantations, that couldn't been solved in a portable manner, require you to know the signed representation.
This should do it (off the cuff):
typedef int NOT_TWOS_COMPLEMENT[(unsigned int) -1 == UINT_MAX ? 1 : -1];
In C++ use static_assert.C has _Static_assert.
Recently: Use-after-free of a memory block. ASAN told me:
1. Where and in which thread the use-after-free occurred (including stacktrace).
2. Where and in which thread the memory was deallocated (including stacktrace).
3. Where and in which thread the memory had originally been allocated (including stacktrace).
Fixing the bug took about 30s with that info.