NULL is just a tool. It's just a way for you to express some property of your API.
That said, some of the strategies outlined in the article could be very useful for avoiding use of pointers. Use of a class type implementing an interface as a "null type" is a very good strategy. Such a type could guarantee successful instantiation, and so avoid the need to throw exceptions or otherwise crash entirely. You could architect an application with those strategies and completely avoid the use of NULL and nullptr entirely.
But where your program is dependent on a finite resource, you must have a way to represent the absence of that resource. NULL and nullptr are convenient ways to do that without a ton of extra code. So no, don't avoid NULL.
It's been more than a decade since I did any C/C++, so I wonder if you were referring so something else that I'm not aware of.
But modern hardware is frequently designed so that dereferencing a zero pointer causes a crash. Even embedded systems will often treat the first couple hundred bytes or so after 0 specially, since it catches so many bugs.
I think it's generally safe to assume you'll get a crash if you actually dereference a null pointer.
I do agree that there's no fundamental difference between using a "might-be-valid pointer" vs. a sum-type (Option/Maybe/etc.) to indicate nullability. However, it's usually easier to get static analysis of the sum type—most type-inferencing compilers can ensure that when you destructure a sum type, you provide match-clauses for every possible instantiation of the type (i.e., for an Option type, they ensure you do the equivalent of a NULL check, and handle both the NULL and non-NULL cases), and so forth.
What I'd really like to see, though, is a language with data-parameterized "types" that unfold (as a last resort, when the concrete type cannot be inferred at compile-time) into runtime data-comparison branches. You should be able to have a raw may-be-valid address with no metadata tag, and then treat that data as having a separate type (a Just or a Nothing) depending on whether the memory address itself is nonzero. You write match clauses; the compiler turns those into NULL checks.
(As an aside, such a language would also allow for things like a type for "Float that only has an integer component", or "Even number", or "Positive number", or "Number in the range 5..23", or "String of length 5", or "String that encodes a valid Unicode codepoint sequence", etc. The type system would become your swiss-army knife for expressing invariants.)
In Rust, the compiler optimizes pointer types wrapped in Options to nullable pointers. Is this the kind of thing you meant?
That's sort of what I mean, here: the nullable ptr as the "canonical, machine-level" representation, with the type serving as a language-level abstraction, effectively a "lie" because it has no real effect on the type system, and only serves as syntax sugar for generating runtime checks.
Depending on the type of semantics you are attaching to the pointer, you can use something like this[1] function, which does exactly what you're asking. It is a runtime no-op, but it coerces the unsafe pointer type so that memory safe code can use it (based on your word that it can be interpreted in this way).
[1]: https://doc.rust-lang.org/std/primitive.pointer.html#method....
Such an FFI function declaration really can't require that the thing returned inside the Option be a dereference of the pointer: in the happy case, it's still an opaque FFI handle, so dereferencing it just exposes an opaque struct you shouldn't touch; and in the unhappy case, it's still an invalid pointer, just one that's invalid because somebody else (like the library that gave it to you) freed it. The raw pointer itself, even if non-NULL, is still "tainted"; you still want to apply manual runtime checks to it before casting it to a Rust reference. (And, past that, in both cases you still probably want to deal with the pointer mostly just by passing it back to other FFI functions from the same library, treating the pointer itself as an identifier, a key the API associates with something.)
I'm emphasizing this distinction because the NULLness of the pointer is, from what I can see, "outside" the lifetime of the reference, and even outside the "identifier-ness" of the handle. Knowing the validity of the pointer requires knowing about ownership/lifetime; the pointer might point to memory that has been free()ed. But you can know that the pointer is intentionally NULL (and therefore not a valid handle/identifier from the FFI's perspective) without attempting to construct a reference.
Effectively, a [nullable] raw pointer is the moral equivalent of a tagged data structure. There's a piece of metadata you can get from the tag: when it's zero, the contained raw pointer is intentionally invalid; when it's nonzero, the contained raw pointer is only maybe-invalid. You can destructure and act on the "tag" without doing anything else to the [non-nullable] raw pointer "stored inside".
I wouldn't expect most languages to care about a distinction this fine, because if a language has references, it really only wants to think about "pointers: maybe invalid" and "references: always valid". But Rust is in the unique position of trying to work "directly with C" without the impedance mismatch of glue like SWIG. And C provides many libraries that return nominally opaque handles, but expect their consumers to notice if those handles are NULL (which happens often, when e.g. a factory-function from such a library is called with NULL arguments, or unexpectedly receives a NULL from a library it itself consumes.)
The function I linked takes a possibly null raw pointer, and vouches that:
1. It points to null or something. 2. The object it points to (if it is not null) will live at least as long as the current scope. 3. The object it points to will not be mutated as long as the current scope is active.
Practically speaking, as long as you don't do anything with the resulting Option<&T> other than check if it is empty, you're safe. In addition, Rust will preserve the internal pointer address, so you could bring it back to the FFI. So you could treat it consistently as such, you'll be fine. I'm not sure if the language spec guarantees that, because that's not how Rust references are designed to be used, but it won't matter in practice.
For what you want, you could just create a generic structure that holds a pointer to the given type, and lets you ask if the contained pointer is null. That way you won't add any false guarantees to the pointer (e.g. immutability, lifetime, etc.).
> Use of a class type implementing an interface as a "null type" is a very good strategy.
Have you ever programmed Objective-C? It uses the null object pattern (called "nil") for all Objective-C objects. This makes it really easy to put big logic errors into your code, since calling upon a nil will just give you more nil. So now you have to wait even longer until a fatal error pops up, and then find where on earth the original nil came from.
> NULL and nullptr are convenient ways to do that without a ton of extra code.
How do you figure? If malloc can return NULL and you want a sane error message, you need to write a NULL check whenever you do a malloc. Otherwise, your user will just get a segfault, instead of a notification that it was memory they ran out of. So why not do that by default, and for the few places where you have prepared yourself to handle out of memory conditions use a different malloc which returns a value that must be explicitly checked?
NULL breaks the type system. Now every pointer type is actually two types: the normal type, and the NULL type. You can dereference or call methods on one, but not the other, and only a runtime check will tell you what you're dealing with. Any time you see:
Object *
You cannot read that as "pointer to Object." You have to read it as "pointer to Object, or NULL."This causes three big problems. One is that it's way too easy to forget to check for NULL. The second, related problem is that without a way to express a non-nullable pointer type, you have to rely on documentation and convention to express the difference. The third is that since only pointer types are nullable, you have to reinvent the wheel any time you you want to express the concept of "null" for a non-pointer type.
To illustrate, you point out malloc which returns NULL when it runs out of address space. What's the appropriate return type?
void *
Now let's say you had some other memory allocation call which could never return failure to the caller. Maybe it aborts execution on failure, or maybe it retries continually until memory becomes available. What's the appropriate return type? void *
Oops. Same problem with parameters. Here's a function: void f(void *ptr)
Are you allowed to pass NULL? Who knows, go RTFM. If the answer is "no," what happens if you pass NULL anyway? The compiler certainly won't catch it, and who knows what will happen.NULL is "just a tool" but it's not a very good tool. It breaks the type system in a fundamental and painful way.
Edit: whew, HN's markup makes it kind of hard to talk about C pointers!
Part of programming is about memory management. You can hide it with a sophisticated compiler but it's always going to be there, and there are always going to be two types of memory: that which is available, and that which is not.
In the context of C and C++, dereferencing a null pointer does not guarantee a crash, or anything else for that matter. In fact, not only does it not give you any guarantees, it nullifies any guarantees you might otherwise have had, because the behaviour of a program that dereferences a null pointer is undefined. Now, an implementation is of course free to make guarantees about programs that have undefined behaviour, but none I know of does.
>Part of programming is about memory management. You can hide it with a sophisticated compiler but it's always going to be there, and there are always going to be two types of memory: that which is available, and that which is not.
This is not true. It is perfectly possible to write useful programs that never have to deal with unavailable memory or even memory management at all past compile time, even in C. This is quite common (as are implementations that don't crash programs which dereference null pointers) when working with embedded systems, for example.
Here's another potential problem. Let's say you write a little bit of code to copy an array into some dynamic storage, basically strdup for non-strings:
void *memdup(const void *source, size_t length) {
void *destination = malloc(length);
memcpy(destination, source, length);
return destination;
}
But then you're a little paranoid about a NULL dereference not crashing, and you want to call some special failure handler that will log more information about what happened, so you add this bit of code in the middle: if(destination == NULL) {
log_failure();
abort();
}
Oops, you've written a bug! It's possible for malloc to return NULL when you haven't run out of memory, specifically when you pass 0. The FreeBSD man page calls this "a silly response to a silly question." Zero-length arrays are a perfectly valid concept, and there's no reason this memdup function should have trouble with them, but now it does. Maybe. Depending on your malloc implementation.Using an optional type would do away with all of these problems. It would guarantee that dereferencing a null pointer would actually halt execution, rather than just "probably," and it would highly encourage distinguishing between "an error occurred" and "no error occurred, but you didn't actually ask for any memory."
Yes, the underlying NULL value will always be there. Any reasonable system will represent the "null" value of an optional pointer type using the underlying NULL value. But in terms of the type system, NULL is terrible and optionals are much better. There's no real downside to it, and you can rest easy knowing that you're still working with 0x0 in the end.
Null-checks only make sense when you expect a null value. If you don't, and you get one, then either your code should be expecting one and has to do some processing differently, or something else is broken.
If the answer is "no," what happens if you pass NULL anyway? The compiler certainly won't catch it, and who knows what will happen.
In all likelihood, a crash when it tries to do something with whatever the pointer should be pointing at, in which case you should be asking why your code is giving it a null.
Beware programming errors which reveal themselves as "in all likelihood," because that hides an entire universe of difficult-to-debug behavior, and sometimes behavior that can be subverted.
Optionals solve these problems, and solve them far better than the "be careful not to do that" approach you describe here.
But option types do the same thing. The option type is actually two types. You can use one, but no the other, and only a runtime check will tell you what you're dealing with.
It may be more convenient than dealing with NULL, but it's doing the same thing...
If you have an interface which returns a pointer that's never NULL, you still have to return a nullable pointer. The non-nullability of the value can only be encoded in documentation or naming conventions. Because of this, the language has to make it easy to use pointers without first checking for NULL, otherwise all these interfaces would become a giant pain in the ass to use. And that means that when you're working with an interface which can return NULL, you have nothing that forces you to check for that. You just have to remember. It's similar for function parameters.
Optional types also work for non-pointer types, which C's NULL simply cannot do. That's a major plus, beyond the problem of all pointers being nullable.
You're right that an optional type is actually two types. But it's designed such that you can't get at the underlying runtime type without first checking to see what it is. That's a huge difference.
It's also error numbers. Yeah, they're just an int as far as types go, but they also have nullable semantics, with 0 being "no error". So it's not just pointers (quite).
> You're right that an optional type is actually two types. But it's designed such that you can't get at the underlying runtime type without first checking to see what it is. That's a huge difference.
Agreed.
If you want an example of a non-pointer nullable type, float and double probably qualify. Both can be NAN, which is effectively NULL for floats. Like NULL, NAN behaves completely differently from any other value and only convention and documentation tell you whether any particular place might produce or accept it.