Falsehoods programmers believe about null pointers
purplesyringa.moe
purplesyringa.moe
All you need to know about null pointers in C or C++ is that dereferencing them gives undefined behaviour. That's it. The buck stops there. Anything else is you trying to be smart about it. These articles are annoying because they try to sound smart by going through generally useless technicalities the average programmer shouldn't even be considering in the first place.
#include <iostream>
int main() {
const char *p = 0;
std::cout << p;
}
You might answer "it's undefined behavior, so there is no point reasoning about what happens." Is it undefined behavior?The idea behind this question was to probe at the candidate's knowledge of the sorts of things discussed in the article: virtual memory, signals, undefined behavior, machine dependence, compiler optimizations. And edge cases in iostream.
I didn't like this question, but I see the point.
FWIW, on my machine, clang produces a program that segfaults, while gcc produces a program that doesn't. With "-O2", gcc produces a program that doesn't attempt any output.
std::cout << *p;
?I still think discussing it is largely pointless. It's UB and the compiler can do about anything, as your example shows. Unless you want to discuss compiler internals, there's no point. Maybe the compiler assumes the code can't execute and removes it all - ok that's valid. Maybe it segfaults because some optimisation doesn't get triggered - ok that's valid. It could change between compiler flags and compiler versions. From the POV of the programmer it's effectively arbitrary what the result is.
Where it gets harmful IMO is when programmers think they understand UB because they've seen a few articles, and start getting smart about it. "I checked the code gen and the compiler does X which means I can do Y, Z". No. Please stop. You will pay the price in bugs later.
Nope, I mean inserting the character pointer ("string") into the stream, not the character to which it maybe points.
Your second paragraph demonstrates, I think, why my former colleague asked the question. And I agree with your third paragraph.
I reckon we are generally in agreement. Perhaps I am not the best person to comment on the purpose of discussing UB, since I already know all the ins and outs of it... "Been there done that" kind of thing.
indeed. It is called UB because that's basically code for compilers devs to say "welp don't have to worry about changing this" while updating the compiler. What can work in, say, GCC 12 may not work in GCC 14. Or even GCC 12.0.2 if you're unlucky enough. Or you suddenly need to port the code to another platform for clang/MSVC and are probably screwed.
In another discipline you might ask what happens what happens when you stress a material near to or beyond its plastic limit? It's quite hard to find that limit precisely, without imposing lots of constraints. For example take a small metal thing eg a paper clip and bend it repeatedly. Eventually it will snap due to quite a few effects - work hardening, plastic limit and all that stuff. Your body heat will affect it, along with ambient temperature. That's before we worry about the material itself which a paper clip will be pretty straightforwards ... ish!
OK, let's take a deeper look at that crystalline metallic structure ... or let's see what happens with concrete or concrete with steel in it, ooh let's stress that stuff and bend it in strange ways.
Anyway, my point is: if you have something as simple as a standard that says: "this will go weird if you do it" then accept that fact and move on - don't try to be clever.
Some languages/libraries even make an explicit distinction between Undefined and Implementation-Defined, where only the latter is documented on a vendor-by-vendor basis. Undefined Behavior will typically vary across vendors and even within versions or whatnot within the same vendor.
The very engineers who implemented the code may be unaware of what may happen when different types of UB are triggered, because it is likely not even tested for.
I used to be one of the folks who defined the behavior of both languages and hardware at various companies. UB does not mean "documented elsewhere". Please stop spreading misinformation.
But not at all companies, orgs or even in Heaven and certainly (?) not at ISO/OSI/LOL. It appears that someone wants to redefine the word "undefined" - are they sure that is wise?
I don't think you know what undefined behavior is. That's a concept relevant to language specifications alone. It does not trickle up or down what language specifications cover. It just means that the authors of the specification intentionally left the behavior expected from a specific scenario as undefined.
For those who write software targeting language specifications this means they are introducing a bug because they are presuming their software will show a behavior which is not required by the standard. For those targeting specific combinations of compiler and hardware, they need to do their homework to determine if the behavior is guaranteed.
Often they use the word "unpredictable" instead. The behavior is perfectly predictable by an omniscient silicon demon, but you may not be able to predict it.
The effect that speculative execution has on cache state turned out to be unpredictable, so we have all the different Spectre vulnerabilities.
Hardware unpredictability doesn't overlap much with language UB, anyway. It's unlikely that something not defined by the language is also not defined by th hardware. It's much more likely that the compiler's source code fully defines the behaviour.
These would be fine interviewing questions if it's meant to start a conversation. Even if I do think it's a bit obtuse from a SWE's perspective ("it's undefined behavior, don't do this") vs. a Computer scientists' perspective you took.
It's just a shame that these days companies seem to want precise answers to such trivia. As if there's an objective answer. Which there is, but not without a deep understanding of your given compiler (and how many companies need that, on the spot, under pressure in a timed interview setting?)
I don't agree. They sound like puerile parlour tricks and useless trivia questions, more in line in the interviewer acting defensively and trying too hard to pass themselves as smart or competent instead of actually assessing a candidate's skillset. Ask yourself how frequent those topics pop up in a PR, and how many would be addressed with a 5min review or Slack message.
This isn't like timezones or zip codes where there are lots of unavoidable footguns - pretty much everyone at every layer of the stack thinks that a zero pointer should never point to valid data and should result in, at the very least, a segfault.
Trivially, `&E` is equivalent to `E`, even if `E` is a null pointer (C23 standard, footnote 114 from section 6.5.3.2 paragraph 4, page 80). So since `&` is a no-op that's not UB.
Also `*(a+b)` where `a` is NULL but `b` is a nonzero integer never dereferences the NULL pointer, but is still undefined behavior since conversions from null pointers to pointers of other types still do not compare equal to pointers to any actual objects or functions (6.3.2.3 paragraph 3) and addition or subtraction of pointers into array objects with integers that produce results that don't point into the same array object are UB (6.5.6).
Here's a classic article about the very weird and unintuitive consequences of null pointer dereferencing, such as "time travel":
https://devblogs.microsoft.com/oldnewthing/20140627-00/?p=63...
You’re being quite negative about a well-researched article full of info most have never seen. It’s not a crime to write up details that don’t generally affect most people.
A more generous take would be that this article is of primarily historical interest.
I don't think this is true. OP is right:
> These articles are annoying because they try to sound smart by going through generally useless technicalities the average programmer shouldn't even be considering in the first place.
Dereferencing a null pointer is undefined behavior. Any observation beyond this at best an empirical observarion from running a specific implementation which may or may not comply with the standard. Any article making any sort of claim about null pointer dereferencing beyond stating it's undefined behavior is clearly poorly researched and not thought all the way through.
This article is closer to category b. But the category a ones are most useful, because they dispel myths one is likely to encounter in real, practical settings. Good examples of category a articles are those about names, times, and addresses.
The distinction is between false knowledge and unknown unknowns, to put it somewhat crudely.
Is this actually a real optimization? I understand the principal, that you can bypass explicit checks by using exception handlers and then massage the stack/registers back to a running state, but does this actually optimize speed? A null pointer check is literally a single TEST on a register, followed by a conditional jump the branch predictor is 99.9% of the time going to know what to do with. How much processing time is using an exception actually going to save? Or is there a better example?
On every pointer deref in your entire program. Not for release mode.
No... And yes.
No: Because throwing and catching the null pointer exception is hideously slow compared to doing a null check. In Java / C#, the exception is an allocated object, and the stack is walked to generate a stack trace. This is in addition to any additional lower-level overhead (panic) that I don't understand the details well enough to explain.
Yes: If, in practice, the pointer is never null, (and thus a null pointer is truly an exceptional situation,) carefully-placed exception handlers are an optimization. Although, yes, the code will technically be faster because it's not doing null checks, the most important optimization is developer time and code cleanliness. The developer doesn't waste time adding redundant null checks, and the next developer finds code that is easier to read because it isn't littered with redundant null checks.
True for 99% of programming jobs, but if you are worried about the speed of null checks, you are in that 1%.
In high frequency trading, if you aren’t first your last and this is the exact type of code optimizations you need for the “happy path”
That's also rather... Redundant in modern null-safe (or similar) languages.
IE, Swift and C# have compiler null checking where the method can indicate if an argument accepts null. There's no point in assertions; thus the null reference exceptions do their job well here.
Rust's option type (which does almost the same thing but trolls love to argue semantics) also is a situation where the compiler makes it hard to pass the "null equivalent." (I'm not sure if a creative programmer can trick the runtime into passing a null pointer where it's unexpected, but I don't understand unsafe Rust well enough to judge that.) But if you need to defend against that, I think Panics can do it.
Thus, in the real world null can never occur--but there are a gazillion places where it would have to test for a null that should never happen. Or simply Assert in the routine itself.
I'm not. It's pretty clear that you don't know how null safe Swift, C#, and Rust work.
IE: Rust has the option type for when a value can be null, and the compiler forces you to do an "empty" check. Otherwise, null, does not exist in the language.
Swift and C# are more metadata based. You indicate if a value supports null or not. The compiler will give you a warning, or error depending on configuration, if you try to assign null to a value that does not expect it.
All three languages have a concept of values that can be null, (or empty in Rust if you want to argue semantics,) so your type system tells you when a value can be null, or when it is expected to not be null. In all three languages, the compiler can force you to make assignments before you use an object.
C# has https://learn.microsoft.com/en-us/dotnet/csharp/nullable-ref...
Of course, there's a big difference between doing it in a VM and doing it in a random piece of software.
But you're saving probably two bytes in the instruction stream for the test and conditional jump (more if you don't have variable length instructions), and maybe that adds up over your whole program so you can keep meaningfully more code in cache.
It's UB so you are forever gonna be worrying about what weird behaviour the compiler might create. Classic example: compiler infers some case where the pointer is always NULL and then just deletes all the code that would run in that case.
Plus, now you have to do extremely sketchy signal handling.
Unlike crummy barely more than a macro assembler compilers of yore. A modern compiler will optimize away a lot of null pointer checks.
Superscalar processors can chew gum and walk at the same time while juggling and hula hooping. That is if they also don't decide the null check isn't meaningful and toss it away. And yeah checking for null is a trivial operation they'll do while busy with other things.
And the performance of random glue code running on a single core isn't where any of the wins are when it comes to speed and hasn't been for 15 years.
Also, so long as the hardware ensures that you don't get away with using the null I would think that not doing the test is the proper approach. Hitting a null is a bug, period (although it could be a third party at fault.) And I would say that adding a millisecond to the bugged path (but, library makers, please always make safe routines available! More than once I've seen stuff that provides no test for whether something exists. I really didn't like it the day I found the routine that returned the index of the supplied value or -1 if it didn't exist--but it wasn't exposed, only the return the index or throw if it didn't exist was accessible.) vs saving a nanosecond on the correct path is the right decision.
At least from a C/C++ perspective, I can't help but feel like this isn't great advice. There isn't a "null dereference" signal that gets sent--it's just a standard SIGSEGV that cannot be distinguished easily from other memory access violations (memprotect, buffer overflows, etc). In principle I suppose you could write a fairly sophisticated signal handler that accounts for this--but at the end of the day it must replace the pointer with a not null one, as the memory read will be immediately retried when the handler returns. You'll get stuck in an infinite loop (READ, throw SIGSEGV, handler doesn't resolve the issue, READ, throw SIGSEGV, &c.) unless you do something to the value of that pointer.
All this to avoid the cost of an if-statement that almost always has the same result (not null), which is perfect conditions for the CPU branch predictor.
I'm not saying that it is definitely better to just do the check. But without any data to suggest that it is actually more performant, I don't really buy this.
EDIT: Actually, this is made a bit worse by the fact that dereferencing nullptr is undefined behavior. Most implementations set the nullptr to 0 and mark that page as unreadable, but that isn't a sure thing. The author says as much later in this article, which makes the above point even weirder.
C does actually have arrays (don't let people tell you it doesn't) but they decay to pointers at ABI fringes and the index operation is, as we just saw, merely a pointer addition, it's not anything more sophisticated - so the arrays count for very little in practice.
Too general, too much trivia without explaining the underlying concepts. Questionable recommendations (without covering potential pitfalls).
I have to say that the discourse here is refreshing. I got a headache reading the 190+ comments on the /r/prog post of this article. They are a lively bunch though.
Alternately, stop writing code in C.
no serious alternative
For the bigger systems where that's not appropriate, you'll value a more expressive language. I recommend Rust particularly, even though Rust isn't available everywhere there's an excellent chance it covers every platform you actually care about.
I think you could argue that Zig is still very new so you might not want to use it for that reason, but otherwise there is no reason to use C for new projects in 2025.
one compiler, two platforms, no spec
> Zig
dollar store rust
Oh you would like a byte? Is that going to be a 7 bit, 8 bit, 12 bit, or 64 bit byte? It's not specified, yay! Have fun trying to write robust code.
The problem is that C++ is a huge language which is complex and surely not easy to implement. If I want a small, easy language for my next microprocessor project, it probably won't be C++20. It seems like C is a good fit, but really it's not because it's a high level language with a myriad of weird semantics. AFAIK we don't have a simple "portable assembler + a few niceties" language. We either use assembly (too low level), or C (slightly too high level and full of junk).
In reality, if you're writing C in 2025, you have a finite set of specific target platforms and a finite set of compilers you care about. Those are what matter. Whether my code is robust with respect to some 80s hardware that did weird things with integers, I have no idea and really couldn't care less.
Because I want the next version of the compiler to agree with me about what my code means.
The standard is an agreement: If you write code which conforms to it, the compiler will agree with you about what it means and not, say, optimize your important conditionals away because some "Can't Happen" optimization was triggered and the "dead" code got removed. This gets rather important as compilers get better about optimization.
Still, while I acknowledge that this is a real issue, in practice I find my C code from 30 years ago still working.
It is also a bit the fault of users. Why favor so many user the most aggressive optimizing compilers? Every user filing bugs or complaining about aggressive optimizing breaking code in the bug tracker, very user asking for better warnings, would help us a lot pushing back on this. But if users prefer compiler A over compiler B when you a 1% improvement in some irrelevant benchmark, it is difficult to argue that this is not exactly what they want.
The big weak region seems to be in-order machines with smaller numbers of general purpose registers.
GCC at least seems to do its basic block planning entirely before register allocation with no feedback between phases.
In my experience, if you don't try to be excessively clever and just write straightforward C code, these issues almost never arise. Instead of wasting my time on the standard, I'd rather spend it validating the compilers I support and making sure my code works in the real world, not the one inhabited by the abstract machine of ISO C.
> In my experience, if you don't try to be excessively clever and just write straightforward C code, these issues almost never arise.
I think these two sentiments are what gets missed by many programmers who didn't actually spend the last 25+ years writing software in plain C.
I lose count of the number of times I see in comments (both here and elsewhere) how it should be almost criminal to write anything life-critical in C because it is guaranteed to fail.
The reality is that, for decades now, life-critical software has been written in C - millions and millions of lines of code controlling millions and millions of devices that are sitting in millions and millions of machines that kill people in many failure modes.
The software defect rate resulting in deaths is so low that when it happens it makes the news (See Toyota's unintended acceleration lawsuit).
That's because, regardless of what the programmers think their code does, or what a compiler upgrade does to it, such code undergoes rigorous testing and, IME, is often written to be as straightforward as possible in the large majority of cases (mostly because the direct access to the hardware makes reasoning about the software a little easier).
Like it or not, C can run on more systems than anything else, and it's by far the easiest language for doing a lot of low-level things. The ease of, for example, accessing pointers, does make it easier to shoot yourself in the foot, but when you need to do that all the time it's pretty hard to justify the tradeoffs of another language.
Before you say "Rust": I've used it extensively, it's a great language, and probably an ideal replacement for C in a lot of cases (such as writing a browser). But it is absolutely unacceptable for the garbage collector work I'm using C for, because I'm doing complex stuff with memory which cannot reasonably be done under the tyranny of the borrow checker. I did spend about six weeks of my life trying to translate my work into Rust and I can see a path to doing it, but you spend so much time bypassing the borrow checker that you're clearly not getting much value from it, and you're getting a massive amount of faffing that makes it very difficult to see what the code is actually doing.
I know HN loves to correct people on things they know nothing about, so if you are about to Google "garbage collector in Rust" to show me that it can be done, just stop. I know it can be done, because I did it; I'm saying it's not worth it.
The articles that started this genre are about oversimplifications that make your program worse because real people will not fit into your neat little boxes and their user experience with degrade if you assume they do. It's about developers assuming "Oh, everyone has a X" and then someone who doesn't have a X tries to use their program and get stuck for no reason.
Writing a bunch of trivia about how null pointers work in theory which will almost never matter in practice (just assume that dereferencing them is always UB and you'll be fine) isn't in the spirit of the "falsehoods" genre, especially if every bit of trivia needs a full paragraph to explain it.
> For all intents and purposes, UB as we understand it today with spooky action at a distance didn’t exist.
The first official C standard was from 1989, the second real change was in 1995, and the infamous “nasal daemons” quote was from 1992. So evidently the first C standard was already interpreted that way, that compilers were really allowed to do anything in the face of undefined behavior. As far as I know
> An integer constant expression with the value `0` , such an expression cast to type `void *` , or the predefined constant `nullptr` is called a null pointer constant ^69) . If a null pointer constant or a value of the type `nullptr_t` (which is necessarily the value `nullptr` ) is converted to a pointer type, the resulting pointer, called a null pointer, is guaranteed to compare unequal to a pointer to any object or function.
C 23 standard 6.3.2.3.3
Also this is point 6 in the article.
Calling this a "falsehood" is utter bullshit.
Let me unpack that for you. Old compilers didn't recognise undefined behaviour, and so compiled the code that triggered undefined behaviour in exactly the same way they compiled all other code. The result was implementation defined, as the article says.
Modern compilers can recognise undefined behaviour. When they recognise it they don't warn the programmer "hey, you are doing something non-portable here". Instead they may take advantage of it in any way they damned well please. Most of those ways will be contrary to what the programmer is expecting, consequently yielding a buggy program.
But not in all circumstances. The icing on the cake is some undefined behaviour (like dereferencing null pointers) is tolerated (ie treated in the old way), and some not. In fact most large C programs will rely on undefined behaviour of some sort, such as what happens when integers overflow or signed is converted to unsigned.
Despite that, what is acceptable undefined behaviour and what is not isn't defined by the standard, or anywhere else really. So the behaviour of most large C programs is it legally allowed to to change if you use a different compiler, a different version of the same compiler, or just different optimisation flags. Consequently most C programs depend on the compiler writers do the same thing with some undefined behaviour, despite there being no guarantees that will happen.
This state of affairs, which is to say having a language standard that doesn't standardise major features of the language, is apparently considered perfectly acceptable by the C standards committee.
Then later 3. "But we implemented it that way for the benchmarks, can't regress there!"
I also like to point out that not all modern compilers behave the same way. GCC will (in the case it is clear there will be a null pointer dereference) compile it into a trap, while clang will cause chaos: https://godbolt.org/z/M158Gvnc4
Both have sanitizers to detect this scenario.
The point is the old K&R era compilers did that 100% of the time, whereas the current crop uses UB as an excuse to generate some random behaviour some percentage of the time. Programmers can't tell you when that happens because it varies wildly with compilers, versions of compilers and flags passed. Perhaps you are right in saying it's "most scenarios" - but unless you are a compiler writer I'm not sure who you would know, and even then it only applies to the compiler you are familiar with.