Null References: The Billion Dollar Mistake (2009) [video]
infoq.com
infoq.com
https://www.digitalmars.com/articles/C-biggest-mistake.html
It's simple to fix this in C, too.
That's the same kind of trade off as with zero as end of string marker.
On the other side: the platform might not have received any following without saving a few bytes here and there, computer memory was a scarce resource at the time.
Buffer overflows are the primary entry point for malware. Seg faults are not. Hence the former are far more costly.
I learned a lot from this and discovered that the real issue is that a process can fail spectacularly and do any damage at all. There are so many other concerns other than NPEs which need to be considered.
Yes. However, not all of those are trivially avoidable by just slightly changing the tech underneath.
I chose a battle with NPEs in a giant Java codebase and largely won. It wasn't trivial to do but the main take-away is that in order for Java to be backward compatible, javac can't enforce any mechanism to make a codebase consistent enough to be NPE free. One way to get this consistency is to change the semantics slightly to trace nulls accurately (some JVM languages attempt this and fail with 3rd party dependencies for eg) and enforce that you can't write code that will produce an NPE. There is even an errorprone plugin that does a pretty good (though not perfect) job of this called NullAway. For codebases that think they can do better and that "need" more flexibility in null handling, there are other frameworks. I'm hard-pressed to meet that need outside of legacy/untouchable code constraints.
Yes, you nailed it!
To quote Sean Connery in The Rock:
Losers always whine about their best. Winners go home and fuck the prom queen!
So we know there are other ways, we used systems with zero lines of C into them.
The prom queen came naked offering herself to everyone and the party was done for the other folks.
If you are old enough to remember those days, then you remember COBOL, Algol, Fortran, Pascal, BASIC, Ada, Oberon, Lisp/Scheme, Forth, O’Caml etc. They’re all great languages, some still have their uses. There’s a reason all of the major operating systems have cores written in C/C++. It’s entirely because they’re pragmatic and “work”, and not some conspiracy.
Edit: although now that I write it, what if C/C++ was planted on earth by an alien intelligence in order to slow down the development of the human race.
Thankfully governments have finally start paying attention regarding software liability.
And we'll see a great slowdown in the software industry.
The industry has miseducated them, and now it is finally happening, software products aren't a special snowflake.
Digital stores with returns, consulting contracts with warranty clauses with fixes at the expense of provider, and naturally cyber security bills.
Move fast and break things only works due to lack of liability.
Government regulations practically place an infinity price on selected characteristics. The result is less entrants in the market, less competition and naturally worse results across all parameters.
Occasionally the market isn't able to reach the optimal goal without some extra help.
No one cares about your code.
Right, because that's a one-bit truth value, not a scalar. The slope doesn't matter, it just has to be positive. Everything is so simple!
There's no such thing as high margins, or success in degrees.
Is there something adding friction to the process of making software/food/cars? Well, is it adding enough friction to make our profits go negative? No? Then literally no one cares.
No misogyny needed.
Guess where I'm taking that prom queen analogy
2) The proportion of prom queen fuckers among systems programming experts is probably much lower than among the general population.
3) The real Unix philosophy is a perverse form of "worse is better", i.e., it prefers making it easy to slap together something that kinda sorta works without worrying about "edge cases" that people are likely to run into, over crafting a reasonably complete, responsible solution. C is the result of this philosophy applied to language design. If it hadn't been for Unix becoming ubiquitous, we wouldn't have made that disastrous choice.
(Come to think of it, under this rubric, Electron is a Unix philosophy exemplar.)
Your calibration is off: he wasn’t trying to put-Chad him.. he was telling him to “git gud”
As for "is woke", I have no idea how to read that point.
No, C was really the only available systems options. C and C++.
>Tinder is irrelevant. The winners are the people who eventually she chooses.
In the dating market today, tinder has created essentially a shit show for males. With that much choice and selection female hypegamy becomes expressed to the max. The top 80% of the top females match with the top 20% of males; It has never been this way for most of humanity. Even the top female always had a limited pool to choose from and her feelings would naturally calibrate to that fact. With tinder everything changed. You get the top tier males having a lot of fun and never settling down, while the rest just give up. That's hook up culture for you.
C would NEVER have been chosen if there were more options.
>As for "is woke", I have no idea how to read that point.
It's the same thing. Wokeness is the result of more womens' rights. Historically, all women couldn't be as selective as they are now because they couldn't hunt for food or work corporate jobs. Survival depended on them finding a man. This is no longer the case and is responsible for much greater female selectivity.
Again C succeeded because of very few alternative options and necessity.
Anyway I can see why you or other people missed the analogy. You have to be familiar with the modern dating game AND familiar with what programming was like 20-30 years ago. Most people capable of even making the comparison have been out of the game for ages.
The state of programming languages that are chosen today is very very much like the hookup culture of today. Especially on the front end, the front end is a woke prom queen for sure.
That's throwing the baby with the bathwater :-(
void foo(char a[..]);
that causes an array argument to be passed as a "fat pointer", consisting of a pointer to the initial element of the array plus a `size_t` value for the array dimension.Tentatively, I like the idea.
As you acknowledge, this wouldn't fix existing code, but writing new code to use this new feature doesn't look difficult. (But converting all existing C code would take approximately forever.)
How does the function access the dimension? Is there a new syntax for extracting the length from a fat pointer, or do you just propose extending the semantics of `sizeof`?
Is there an existing C compiler that implements this?
A minor question: Given the above declaration, would
char c = '?';
foo(&c);
be valid, treating `c` as a single-element array?Finally, a very minor point: `...` is already a valid punctuator. I can't think of any ambiguities that would be introduced by adding `..` as a new symbol, but that might be just my lack of imagination. (Note that gcc uses `...` in its case range extension.)
So &c on a char should produce char, and &c on a char[20] should produce char[20].
char c;
&c; // type is char*
char arr[20];
&arr; // type is char (*p)[20], pointer to array of 20 char
(A small quibble: the term "type compatibility" in C doesn't mean implicit convertibility. Two types are compatible if they're literally the same type, and in just a few other cases. You can assign an int value to a long object, but int and long are not compatible, even if they happen to be the same size.)The problem with dropping implicit array-to-pointer conversion is that it would break most existing C code. It would be a great idea for a new language.
Suggested reading: Section 6 of the comp.lang.c FAQ, <https://www.c-faq.com/>. The relationship between arrays and pointers in C is admittedly confusing; this is the best resource I know of for explaining it.
https://go.dev/ref/spec#Slice_types
Slices can still be nil (null), but it isn't an unsafe memory access operation, just another type of potentially useful or potentially errant invocation to handle.
"var x []int; fmt.Printf(`%p %d`, x, len(x))" outputs "0x0 0"
Indexing "x[0]" results in: "panic: runtime error: index out of range [0] with length 0"
They can also be appended to and then produce a valid slice.
Null terminated strings were a horrible mistake though and really should have been fat pointers.
X86 restricting saved IP read/write on the stack to special instructions like call or ret would have been nice (mark stack memory as restricted when call saves eip for example$
how else you might structure errors without changing much of the rest of C left for an exercise
error_t oops;
int x = foo(10, &oops);
if(oops)
goto whoops;
int x = foo(10, 0); // yolo!sqlite C libraries apparently still do this. The person who wrote the code was both not very good and had no idea that an error was even occuring.
++countdown;Modern PC-style hardware is kind of a hybrid, acts like Harvard with regard to cache, and like Von Neumann with regard to RAM. Furthermore, it has a fancy MMU that allows for things like the NX bit. All C compilers/linkers I am aware of know the difference between code and data and are able to put each one in the appropriate section, what is done after that is the OS/hardware responsibility.
As for zero-terminated strings, I also think it is mostly a mistake, though it does have a few advantages. You can still work with size+pointer though, using mem- instead of the str- functions, and "%.*s" in printf(), not ideal though.
I disagree here, remember that at the time some strings with length implementations used only one or two bytes for the length, this can creates lots of issues that zero terminated strings don't have.
Of course nowadays zero terminated strings don't make sense anymore.
(WG14) is the ISO workgroup which maintains the C specification.
Every piece of code such as:
int a[SIZE];
foo(a, SIZE);
Would have to be rewritten to: int a[SIZE];
foo(&a[0], SIZE);
And this additional noise would just make C harder to write for no reason. Rather than making C harder to write, just pick a different programming language.If you don't add fat pointers (thereby requiring a major re-design of the language and causing an endless amount of pain with regards to ABIs etc) then the solution, which involves removing the implicit conversion of an array to a pointer to its first element in expressions where it is not used as an operand of the sizeof operator, unary & operator or as a string literal used to initialise an array, certainly DOES make it harder to write C.
No, it isn't. I implemented it in D, and know what is involved. I've written two C compilers, as well.
> the fact that the explanation fits within a short blog post doesn't make it an easy fix.
Correct, but since I've actually done it, I'm in a good position to say it is an easy fix.
The point is, it's an easy fix if you make a new language, it is a difficult fix if you want to call the end result C.
The key takeaway is that the implicit conversion from array to pointer to first element is not the problem with C. The problem with C is that it's a very limited language which would require a re-design (e.g. fat pointers and implicit range checks everywhere, which, while solving one problem, wouldn't really solve ALL or even most of the problems with pointers).
I think that, for its time, C was a decent improvement over assembly programming. I think that in the current day and age if you are writing C and use modern tooling as well as follow a set of good guidelines (as well as have the right mentality and don't ACTUALLY treat is as high level assembly) it is possible to write more or less safe C in small doses. For anything else I think the correct solution is not to ponder how to fix C, or pretend like there's just one problem with it, instead the solution is to use a better language.
You might think it's pedantic, but we're talking about C, it is important to be clear in language used to talk about C as it's an unforgiving language.
There is an untold amount of confusion surrounding how arrays work in C at least in part because of silly wording like "decay".
As a final note: When you write "f(foo.bar)" to call "f" while referring only to the "bar" field of the struct "foo" you are not losing "foo" and it is not "decaying" solely because the function which receives the result of the expression which formed the first argument of its invocation only sees the "bar" field of "foo". And now I'm not saying that it's a conversion either, but the point still stands. If "f(a)" where "a" is an array involves decay then so does array indexing or accessing a struct field.
If it was, we wouldn't be having this discussion and array bounds checking would have been added to C compilers 40 years ago.
struct s {
int bar;
int baz;
} foo = {
.bar = 5,
.baz = 6,
};
f(foo.bar);
Has "foo.baz" decayed? If the answer is no then you have an answer for why there is no decay involved in the implicit conversion.The fact the function does not receive the array size does not mean that decay has happened.
That being said, nothing I just wrote wasn't in the comment you replied to, so maybe explain what it is from that comment that you didn't understand about my argument for why "decay" is simply the wrong word for what is happening.
Also: This has nothing to do with array bounds checking.
Also: Given the nature of the C language, I would find it surprising if compilers 40 years ago would have implemented bounds checking irrespective of if the language started out with either fat pointers or some other way of carrying the size of a specific array around with a pointer to its first element.
It only impresses those that never used programming languages with modules during the 1990s.
Tony Hoare's Null References: The Billion Dollar Mistake - https://news.ycombinator.com/item?id=30719472 - March 2022 (13 comments)
Null References: The Billion Dollar Mistake - https://news.ycombinator.com/item?id=22019627 - Jan 2020 (150 comments)
Null References: The Billion Dollar Mistake – Tony Hoare (2009) [video] - https://news.ycombinator.com/item?id=11798518 - May 2016 (79 comments)
Tony Hoare / Historically Bad Ideas: "Null References: The Billion Dollar Mistake" - https://news.ycombinator.com/item?id=473158 - Feb 2009 (2 comments)
So it just took some time to get mainstream.
The basic issue exists - how to handle unknown values. It goes beyond programming language constructs.
But reflecting the underlying machine was not just a completely reasonable default it was the most viable performant option at the time.
And provided a clean enough means for writing cross platform compilers for more abstracted languages.
We probably took too long coming up with validated optionals as an abstraction, but so much is obvious in retrospect.
There is a lot that can be done, that must be done, but should only be done rarely - but the powers provided by a language that reflects the underlying machine are vital nonetheless.
You shouldn't be writing goto on a daily basis, but it's very useful for custom loop definitions.
The mistake was in not providing any language features to enforce or support that convention, or to communicate when the pointer is guaranteed to be valid vs when the null option needs to be handled.
Having an optional type as the language default is fine and sensible but could really have used aome language support to make it less of a footgun.
- null is a valid member of every type - there's no way to constrain a reference to only non-null members of the referenced type - the compiler allows dereferencing values of a nullable type
Weirdly, all three are true of C, C++ and Java.
If you forbid illegal states you end up with fake objects taking the place of non-existent real ones. (What's Find supposed to return if it's not there??) In most cases it's better to blow up when you try to work with a null than to mistakenly work with something that takes the place of a null. A null is easier to detect and thus the better solution.
Yeah, well implemented nullable and not-nullable types are good but the compilers didn't have that kind of smarts in the old days.
You don't use a fake object, you use a real one of a different type. In the cases where what you want is "maybe this thing, maybe nothing" you can use an option/maybe type. In the cases where you want something else (e.g. "either this thing or an error message") you can use a type that represents that.
> Yeah, well implemented nullable and not-nullable types are good but the compilers didn't have that kind of smarts in the old days.
ML had well typed optionals back in the '70s. There's no excuse for using a language that needs null today.
It's different because most things can't throw it. The problem isn't that your language has a way to implement a value that might be absent, the problem is that your language does't have a way to implement a (heap) value that won't be absent, except by convention.
Bad data should fail fast and null is good at doing that.
For example in C#, you would get something like "Use of unassigned local variable xyz".
If the performance cost of initilizing the variable is such an unberable cost, there is [SkipLocalsInit].
(Fundamentally, this was a case of a child operation that *might* be needed. It could have been rewritten polymorphically but that would have involved a fair amount of refactoring.)
Also such kind of situations are usually a sign that code blocks should be moved into a separate function.
What you really need is that the array keeps track of its own limits - something like std::vector or Java's ArrayList. But then you're back to the overhead of checking each access.
So... I'm still not seeing how you're making illegal states unrepresentable.
You can address that with some concept of lifetimes or ownership.
> What you really need is that the array keeps track of its own limits - something like std::vector or Java's ArrayList. But then you're back to the overhead of checking each access.
IME that's less good than having a dedicated type for "index into this array", even if the latter isn't perfect. A bounds-checked access using an offset from a different array might not be reading from uninitialized memory, but it's still overwhelmingly likely to be a programming bug.
A better solution: Implement a different language with a better type system instead. Or pick one of the hundreds that already exist and can represent the concept of a tagged union without having to implement it manually.
Null pointer checks are very, very cheap, no additional memory fetch since the pointer value is needed either way, easy to predict the slow path and if we are talking about languages that can deoptimize, then it is literally free (checked by hardware either way) -> a null value will cause a segfault, which will deoptimize the code on the slow path and continue from there. So it is not a problem from a performance point of view.
Yes, the NULL check (effectively, in most cases) happens via hardware either way, but it's a lot cheaper to have the MMU raise a page fault because you tried to read/write an unmapped page than it is to litter every pointer access (or at least every initial non-volatile pointer access) with a conditional jump.
I have no problem with someone adding compiler flags to enable slightly optimised NULL checking across every pointer access but I guarantee nobody will enable it for any code written in C for a reason as it will kill performance. Consider the instruction cache impact. You're also way over-estimating the branch predictor. Yes in hot code you will not see a major impact, but this is not something the kind of people who write high performance C really bet on anyway. And don't bring up __builtin_expect and friends, they don't do anything on AMD64, one of the most popular architectures in the world.
Now for software which is written in C for no good reason, there's a perfectly sensible target for this feature.