Null References: The Billion Dollar Mistake – Tony Hoare (2009) [video]
infoq.com
infoq.com
Perhaps the examples I've seen in Rust just don't do that, because part of the appeal of Rust is that you sort of don't have to.
That makes the cognitive load acceptable and documents all the UB prevention in source code. It's a habit arrived at from using formal instantiation protocols that were driven from external to the device being made - starting with RFC1695.
Null is a damn useful construct and while there may be better solutions to using a null reference they're often not as clear as using null or worth the time investment to develop properly.
So yea, I generally try to avoid null's and completely understand how they've caused a lot of problems, but I'm not willing to say we shouldn't have them or use them when appropriate.
Allowing nulls explicitly, like with a "String?" type, is the way to go in languages where Option is not an option.
1. Option is a regular data type/object that represents "a collection of at most one value". It's similar to a list, set, heap, tree or whatever other collection types you have, except it can not contain more than one value.
2. With explicit nulls I guess any data type (e.g. Integer) will automatically get a clone of itself (a different type named Integer?) where also null is a valid member.
From a purely theoretical standpoint, I like #1 better.
Here's a Java example where I want the type system to enforce that a method in `ClassB` can only be called from `ClassA`. However, the fact that `null` circumvents the type system makes this pattern just wishful thinking.
class ClassA {
private static final Witness witness = new Witness();
final static class Witness {
private Witness() {}
}
void callClassBMethod(ClassB classB) {
classB.onlyClassACanCallThisMethod(witness);
}
}
class ClassB {
void onlyClassACanCallThisMethod(ClassA.Witness witness) {
// ...
}
} Option[Option[String]] != String??It might not be very interesting in the Option[Option[String]] case but imagine Try[Either[String, Int]] or List[Future[Double]].
It's a very important distinction.
Collapsing cases is one of the primary thing why exceptions sometimes get a bad rap, and Kotlin (and Ceylon) do the same with ? (and |, &) at the value level.
By encoding `null` in its type system, Kotlin lets you manipulate these values directly which leads to code that is much less noisy and just as safe.
The main strength of the first approach is that Option is only one type out of many error-handling structures.
Not every error is handled appropriately by Option/?.
If you have a language like Kotlin where they hard-coded one way of handling errors, it feels very unidiomatic to pick a better fitting error handling type, while in languages where errors are handled by library code, it's a very natural approach.
Which is expected since these two constructs are not aimed at handling errors: they manage missing values.
> If you have a language like Kotlin where they hard-coded one way of handling errors
No, no. `?` is not for handling errors.
Kotlin is as agnostic as Scala for managing errors: you are free to use exceptions, dumb return values or smarter ones (`Either`, `\/`, `Try`, ...).
> Which is expected since these two constructs are not aimed at handling errors: they manage missing values.
Which is a very small part of handling errors in general. As Kotlin offers special syntax for only this case, developers tend to shoehorn many errors into the "missing-value" design to get the "nice" syntax even if a different approach would have been more appropriate.
> Kotlin is as agnostic as Scala for managing errors: you are free to use exceptions, dumb return values or smarter ones (`Either`, `\/`, `Try`, ...).
That's not true in practice:
Just have a look at funktionale: Despite providing almost the same as Scala's error handling types (partially due to the blatant copyright violations) almost nobody uses it. This is a direct result from having a "first-class" construct in the language: It turns library-based designs into second-class citizens.
That's the thing Scala got right, and many of the copy-cat languages got wrong.
Missing values are not errors.
If you look up a key on a map and that key is not present, it's not an error.
> partially due to the blatant copyright violations
Uh copyright what? On an API?!?
Call it whatever you want. ? only covers a small subset of interesting "conditions" while tremendously hurting "conditions" which could be handled in a better way.
> Uh copyright what? On an API?!?
Implementation. The copying of slightly buggy exceptions strings makes it even more obvious that files were copied verbatim with just enough syntax changes to turn Scala code into Kotlin code while replacing the original license and authors with different ones.
PS: Feel free to comment on actual the points I made.
Sure.
I think the idea that API's (or implementations as you said) can be copyrighted is completely insane and I can't believe any software engineer would be okay with it. Which makes me think you're not a software engineer, and that's okay, but please read up on the issues, this is super important for our profession.
I can't belive the US made that a law and it makes me sure that I will never want to move there.
I think you are super confused here. This is not about APIs. Copyright is what allows software developers to enforce a license of their choice. Without copyright, the license is just a text file without meaning. I suggest you read up on the FSF's position on this if you want to have an example.
> Sure.
(Still waiting for you to comment on the points I have made.)
https://www.eiffel.org/doc/eiffelstudio/Differences%20betwee...
Storing None for an Option<&T> as a null and Some(ref) just as the reference is an special optimisation that the rust compiler does. Usually you have a selector which tells you which variant of the enum it is.
Tree *a = new Leaf();
Tree *b = new Leaf();
Tree *c = new Node(a, b);
a.parent = b.parent = c;
null has other uses, e.g. as above for cyclically dependent initialisation. data Tree = Leaf Tree | Node Tree Tree
a = Leaf c
b = Leaf c
c = Node a b
But more importantly, there is no reason you can't do this with Optionals. Tree *a = new Leaf();
Tree *b = new Leaf();
Tree *c = new Node(a, b);
a.parent = b.parent = new Some(c);Java has the checker framework that supports `@Nullable` annotations and will allow only those vars/fields to contain `null`.
If you write the `send_mail` procedure you might only want it to throw `MailException`s of various kinds. But if you use a TCP procedure inside, you'll have to declare `send_mail` to also throw `ConnectionError`s, thus revealing its implementation. To correctly hide/abstract over the implementation, you have to internally catch any `ConnectionError`s thrown by the TCP procedure, and re-throw them as `MailException`s. That's a lot of manual work that shouldn't be needed.
Another similar problem is what someone else mentioned: if you don't actually know which specific exceptions your method can throw until runtime, what do you declare?
----
Note that this is also the case for some "null alternatives". Sure, the `Option`/`Maybe` type is easy to compose, but as soon as you start inserting error information in there (with an `Either` type) you run into some of the same problems. This is acknowledged by language communities where that is practised, and some of them prefer unchecked exceptions to `Either`-style types for that reason.
And how can you not know the specific exceptions you need to handle? Don't you need to test all of those?
I am sure you didn't mean it to, but this reminds me of meetings where people say "oh, we don't have to worry about that. TCP is reliable."
Twitch :)
You're going to have to handle them as TCP errors anyway, even if they're lifted to a different type, so why not just throw them as they are?
TL;DR: Unlike a checked exception, an Option type doesn't break the normal control flow (easier to reason about, not as prone to not cleaning things up) and allows each layer of code in a call chain to do as much or as little error-handling as it deems appropriate.
You imply that this behavior is unreasonable, but that's the approach Scala takes (all exceptions are unchecked), and a lot of users seem to like it. So I think the parent's question still stands.
I didn't actually want to imply that lacking checked exceptions unreasonable. I was merely trying to paraphrase isthe null problem in terms of checked exceptions, which the person I was replying to seems to favour. It was just an attempt to create an "aha" moment, showing what kind of pain is created by allowing null references.
A proper answer would probably have a lot more substance to it, but in short I get the feeling that null-free programming is a lot easier to achieve in general than exception-free programming. In fact, in most cases, not allowing nulls seems to be the implied default, yet it's hard (or impossible) to declare this explicitly (in most languages, like Java, etc). It just happens to be an extremely common problem that can be solved in a nice way, unlike exception checking in the general.
TL;DR: The added safety that is given by allowing the type system to reason about nulls - when looked at in the light of the amount of boilerplate code that it creates - compares favourably to checked exceptions.
With option types, we check whether it is empty once, and if it is not, we unwrap it and use it as if it were never nullable. Any function that uses the unwrapped value downstream is oblivious to the fact.
But they provide a useful social convention: if it is a pointer then it is your job to check for NULL. If it is a reference, then it is the job of the other guy to ensure it is not null. And the rules of the language mean that null references (as opposed to pointers) are rare.
Well then! That just goes to show how much safer C++ references are than I had previously supposed.
"A reference shall be initialized to refer to a valid object or function. Note: in particular, a null reference cannot exist in a well-defined program, because the only way to create such a reference would be to bind it to the “object” obtained by dereferencing a null pointer, which causes undefined behavior"
Optional<WidgetFriend> GetWidgetFriend(Widget widget)
and WidgetFriend had an Optional<FriendName> property, we could simply write [1] from widget in _widgetRepository.FindById(widgetId)
from friend in _widgetRepository.GetWidgetFriend(widget)
from name in friend.Name
select name;
and get an Optional<FriendName>, without having to repeat all the tedious pattern matching of Some or None for every call. If any of those calls returns a None, the entire evaluation short-circuits and we get a None result at the end.Of course, we can still write
Optional<Widget> widget = null;
which is the problem with not fixing this issue on a language level.[1] You'd have to implement Select and SelectMany, instead of Map and Bind, to get query syntax like this.
Who was the guy at the end from the Erlang community? (The one who drew a chart with the axes useful/useless versus unsafe/safe that he claimed he got from Simon Peyton-Jones?)
There are some trivial cases where the compiler could figure it out for you but most of those cases are not very useful in practice.
Is this a trend in newer languages in general or an FP specific thing?
It's not necessarily new since the ML family of languages has had it for a while now. It's also not limited to FP languages but it best known from ML the ML family. Rust has it though as well as Swift i think.
non-shitty
> a trend in newer languages
yes, only discovered a few decades ago
:-D
A bad trade-off – i.e., one where the downsides outweigh the upsides – is a mistake.
Nulls are not a necessity: Several languages demonstrate that there are better alternatives.
Drawbacks: numerous
How is this a trade-off?
Consider this piece of pseudo-Java/Spring code and think how you would do it if the platform you were using forced you to either declare a and b as optional (which they are not) or assign them some value between the IoC container instantiates the class X and wires the values for a and b.
public class X {
@Autowired
ServiceA a;
@Autowired
ServiceB b;
@PostConstruct
void init() {
// a and b are instantiated by the IoC container and
// are fully valid for the rest of the execution
}
}2. Do you not consider cyclic dependence a code smell? At least I have always avoided cyclic dependencies.
But in general yes; cyclic dependencies are a bad idea.
[1] https://en.wikipedia.org/wiki/Option_type [2] http://slick.lightbend.com/
If NULL value is in a field that is used for JOIN-ing, I don't think you would "represent" the NULL value as much as you would simply have a lack of data. For example, if you had some set of results from a query that contained a JOIN on a field and some values were NULL, those records with the NULL value would not be in the result set.
If we do receive results with possibly NULL values I believe they could be either be represented with an appropriate zero value -- e.g. "" or 0 -- or with an optional type like another commenter suggested.
Also 6th normal form?? What is that?
Not something I'd like to use.