Why Maybe Is Better Than Null
nickknowlson.com
nickknowlson.com
In doing so, Java makes it just that much easier to create bottom values when they are not desired (e.g., dereferencing a null pointer).
E foo; ... foo=foo.bar().bat()
without having to do a null check after each function.
I think that Java's approach has the worst of both worlds, because there is no way to make foo.bar().bat() safe when a function could return null.
In C, for example, calling a method on a null object does not cause an error. Rather it passes in null as the 'this' value, allowing you to do you null checks within the method.
case a of
Nothing -> Nothing
Just a -> case b of
Nothing -> Nothing
Just b -> a + b
This is quite a bit of boilerplate hiding the expression that actually matters--a + b! Moreover, whenever you have code that creeps steadily to the right, it means you either messed up or missed an abstraction.It turns out that this pattern--do a computation if all the values are present, but return Nothing if any of them are Nothing--happens very often. Happily, we can get some nice syntax for Maybe computations like this using do-notation in Haskell or for-comprehensions in Scala:
do a <- a
b <- b
return (a + b)
This is much better! It makes even more sense for more complicated expressions, especially when later results depend on values of earlier ones. However, for a simple example like a + b, it's still quite a bit of boiler plate; we can certainly do better! Here are two alternatives using functions from the Control.Applicative module: (+) <$> a <*> b
liftA2 (+) a b
The important idea here is that both versions somehow "lift" the (+) function to work over Maybe values. This just means they create a new function with checks for Maybe built in. This is great because it saves all the boilerplate above and nicely abstracts away most of the null checks while preserving safety. But the syntax is still a bit awkward. Happily, if you don't mind using a preprocessor[1], you can get some very nice syntax called "idiom brackets": (| a + b |)
[1]: https://personal.cis.strath.ac.uk/conor.mcbride/pub/she/The computation inside the (| and |) is lifted over Maybe, just like the two previous examples. I think this is the clearest option here: it has the least syntactic overhead, and the base expression--a + b--is very easy to read. They also have the advantage of nesting, so you can express a + b + c, where you want a null check for all three variables, as:
(|(|a + b|) + c|)
This isn't perfect, but I think it's still very easy to follow. It might be better if the (| and |) were a single character, something like this: ⦇⦇a + b⦈ + c⦈
However, some people really don't like Unicode symbols in their code :(. Happily, you can have the source look like (|foo|) and have Emacs replace it with ⦇foo⦈ without actually changing the code. It's basically Unicode syntax highlighting. I think this leads to the most readable code so far.So my main point is that you can abstract out the common case where you check for Nothing and make the whole expression Nothing if any sub-expression is. This saves quite a bit of typing and much more importantly makes the resulting code far easier to read.
Another really cool part is that all these syntax forms and functions are not specific to Maybe--they actually work for a whole bunch of different types. So you would not be bloating your language by including special features just for safely checking nulls; these features are much more general.
(fromMaybe 0 a) + (fromMaybe 0 b)
This will always produce a number and it's pretty clear why you'd get the one you'd get. do x <- a
y <- b
return $ x + y
can result in Nothing if either a or b is Nothing, but (fromMaybe 0 a) + (fromMaybe 0 b)
will always result in a number, with 0 being used in place of Nothing. If only one of them is Nothing, you'll get the value of the other, the Nothing having been treated as 0.Nobody said they were. That's idiomatic Haskell for something like this in Java when unboxing Integers:
return (a != null ? a : 0) + (b != null ? b : 0);
As you can see, there is no adding Nothing to integers. instance Num a => Num (Maybe a) where
(+) = liftM2 (+)
(-) = liftM2 (-)
(*) = liftM2 (*)
abs = liftM abs
signum = liftM signum
negate = liftM negate
fromInteger = Just . fromInteger
> Just 4 + 2 * Just 6
Just 16
> Nothing * 42
Nothing
Notice how the fromInteger method allows you to freely mix Maybe and non-Maybe numbers.However, sometimes a Maybe represents an inherently "optional" part of your data model, such as "a Foo may have zero or one Bar". In that case, you'll probably hold onto the Maybe until the point where you'd actually read and use that field.
In practice, this means that you write a decent part of your program using these techniques, which creates a block of code that produces a Maybe value after taking a bunch of Maybe inputs. Then you only use a case statement at the very end, when you need to plug the Maybe value back into normal code.
All these functions are useful for one particular case: you don't know what to do with a Nothing value, so if you see one anywhere, you just pass it on: your final result is Nothing. That pattern just turns out to be very useful.
Like NaN, it can be hard to track down where things went wrong.
This is a great reason to avoid huge chunks of code stitched together staying inside Maybe, while still being convenient on the small scale.
data Employee = Employee { name : Text, spouse : Maybe Text }
You may also want to use Either to store two possible outcomes of an operation, though I would recommend using your own sum type for clarity: data MyOwnEither a b = MyLeft a | MyRight bUnlike NaN, you have the freedom to not use it when you don't want it.
Haskell's type classes, unlike ML functors (I think), are "coherent", which means that you can't have scoped or multiple instances for the same type, lest you risk breaking the type system. With that in mind, extending Maybe to Num would mean a reduction the number of type errors caught in other code that uses Maybe and Num near each other.
http://stackoverflow.com/questions/3079537/orphaned-instance...
There are occasions where it can be useful to have a module export an orphan instance for compatibility before it makes it into the more appropriate spot in the standard libs.
And you're right you can't have multiple instances for the same type; a newtype wrapper is required.
But here we start with a nice ring like Integer and end up with a type that has this weird, extra element that has no inverse with respect to addition, etc.
0
do a <- a
b <- b do a <- getA "foo" "bar"
b <- getB "foo" a
...
I used a deliberately overly simple example so I could go on from do-notation--which many people are already familiar with--to applicatives and idiom brackets.Besides, it looks like any normal program, except you're using <- to define variables rather than =. Can't see how it could be any clearer than that.
I like the Maybe concept and non-nullable types; I just think being able to overload operators like "=" and ";" in C++ and Haskell is optimizing writability over readability and in most cases, readability is by far the more important attribute.
In this case, think of Maybe as a box containing one or zero instances of a type. For Maybe
do x <- maybeAnInt
y <- maybeAnotherInt
return (x+y)
Lists are 'boxed' values containing any number of elements do x <- [1,2]
y <- [10,20]
return (x+y)
The monadic semantics for lists means this returns the sum of each combination of values, namely [11,21,12,22]. This is essentially like a database join, which is why a monadic structure was used for LINQ.For Promises
do x <- intPromise
y <- anotherIntPromise
return (x+y)
This returns a new promise containing the sum of the result of two promises instead of using callbacks.Essentially all these types live in a Maybe/List/Promise box. There are many more examples. The nice thing about monads is that the semantics of how these things work is abstract enough to allow a variety of interpretations, but constrained enough (by the Monad laws) that you get a nice intuition of how things work after using a few different instances.
do
a = getFoo x y z
b = getBar p q
(which won't compile), you're saying do
a <- getFoo x y z
b <- getBar p q
Which is not an overloaded operator. In fact, it's not even meant to suggest equals (which in Haskell rather strictly means mathematical equality) but rather assignment ("=" in the expression "x = x + 1;"). addM :: Maybe Int -> Maybe Int -> Maybe Int
addM ma mb = do
a <- ma
b <- mb
return (a + b)
As an aside, the inferred type would be much more general, because (+) works on any number and do/return work on any monad: addM :: (Monad m, Num a) => m a -> m a -> m a
Which means you could use it like this: addM (Just 5) (Just 10) == Just 15
But also like this: addM [1, 2] [4, 8]
== [1 + 4, 1 + 8, 2 + 4, 2 + 8]
== [5, 9, 6, 10]
Or like this: addM (Right 42) (Left "NaN") == Left "NaN"
And in many more interesting ways. :)1. By default all types are non-nullable. 2. You explicitly mark if you expect that a value could be null. 3. The compiler helps you by either making sure you check for null on those values you mark, or by doing some magic (like Scala does) that will always return null for expressions where one of the values is actually null.
Is the reason we call it Maybe/Option/whatever just to disambiguate with the traditional use and lack of safety in what pretty much all languages use "null" for? Or is there a distinction I'm missing?
Also, most languages require something to be a reference or pointer to be null. You can't have a plain int or structure be null.
So to use Maybe, all you need from the compiler is to not have nulls everywhere. Since you don't need language support for it, your language is simpler and the Maybe behavior is part of a library.
This also ensures that you Maybe values behave as first-class citizens. You can do anything with the Maybe type that you could with any other type, because that's all it is. For example, this means that you can nest them: have a Maybe<Maybe<A>> value, for example. It also means Maybe can play well with other libraries; for example, inn Haskell, it works immediately with the alternation operator:
result = tryA "foo" <|> tryA "bar" <|> tryB
Part of the beauty is that <|> is an operator that represents alternation for a whole bunch of other types as well. There are a whole bunch of other functions like this.So: yes, you can have language support for it. But just having it as a normal type makes the language simpler and ensures you have full generality. The only thing that you need from your language is to get rid of null.
Following your argument, we need special language support for any sort of abstraction, because everything has to be built in terms of some built-in language features at some point.
Also, on a largely unrelated note, I think that there is no reason for modern languages not to have sum types in this day and age. (cough Golang cough)
We have a convention that all recoverable errors be captured in Either or Maybe. This policy is paying off in a huge way because the type system forces us to think about what error cases mean. This is driving us to write substantially higher-quality code than I've written with other tools.
Just x -> something x
Nothing -> halt_and_catch_fire # impossible
In Haskell there is even a standard library function fromJust which does exactly that.One introduces this hack knowing that this particular variable will always be Just and having absolutely no way to deal with Nothing and later somebody else sees the type Maybe Foo and figures that it must be OK to put a Nothing in there.
Yep, problem solved, 'billion-dollar mistake' corrected (at least for future times).
Typesafe null and flow-dependent typing
There's no NullPointerException in Ceylon, nor anything similar. Ceylon
requires us to be explicit when we declare a value that might be null,
or a function that might return null. For example, if name might be null,
we must declare it like this:
String? name = ...
Which is actually just an abbreviation for:
String|Null name = ...
An attribute of type String? might refer to an actual instance of String,
or it might refer to the value null (the only instance of the class Null).
So Ceylon won't let us do anything useful with a value of type String?
without first checking that it isn't null using the special if (exists ...)
construct.
void hello(String? name) {
if (exists name) {
print("Hello, ``name``!");
}
else {
print("Hello, world!");
}
}
From http://ceylon-lang.org/documentation/1.0/introduction/That is, if you are going to use Maybe, you either want to do so from the beginning, or you want a good layer of abstraction between where the value is optional and where it is not. Moving something from guaranteed to optional is a bit more combersome with Maybe.
Also, I have gotten really used to "truthy" values.
They added a statically-compiled mode to Grōōvy last year. Most users like Grails don't use it yet, but it's there to try out and you can report any problems to the Grōōvy issue tracker.
While pointers which point into the wrong place for various reasons (off end of array, previously freed memory) cause horrible issues to this day, I can't personally remember ever having a serious issue with a null pointer (they tend to crash quickly and loudly, because in all modern OSes dereferencing NULL segfaults)
In some scenarios, this is a serious issue all by itself. My day-to-day work is mostly on Android, and eliminating nullable references altogether would eliminate some crashes, which are highly visible to the user.
This is especially a problem in dynamic languages that sling nils around....like any major modern scripting language. Checking if a value is nil before proceeding is aping what a language like Haskell does when it pattern matches against Maybe (Just a, Nothing) albeit in a post-facto bad way. Granted, you can't really make any assertions about reflecting a maybe value in the type of a language that doesn't care about types before runtime.
Making types non-Nullable by default is nice, though. Even in a dynamic language you can have a syntactic distinction to make interfaces more explicit.
I'm sort of a fan of Haskell already, to tell you the truth.
Not sure that's so far from Haskell, really... :-P
> and a type system that can handle external hardware poking around in its memory
Can it be in isolated places? That's doable...
Maybe it's my lack of experience talking. Haskell feels heavier in a lot of places. I think the focus on compilation in the backend is almost a downside here, too.
Can it be in isolated places? That's doable...
Depends on the application. There's a lot of hardware out there that does really weird stuff. There's some interesting work in Haskell-space (eg Atom) but I don't know how comprehensively mature it is.
if (argumentX == null)
return null;
at the top of function signatures. It was just defensive programming. Maybe argumentX couldn't be null, but it would take time to figure that out (sometimes I did that, though). More code means harder to read and maintain, thus costing dollars.This would also be contagious: If a piece of code checks if X is null, you'll assume that X can be null, whether or not that's true.
I'd certainly prefer to be able to reason about the code with the safe assumption that certain things cannot be null.
You can do this, but if you want it to be maintainable, you'll also want to detail in the function comments this technical debt. If you don't, someone else will come along and see your sweet method (looking only at the comments) and use it where the input can be null.
Ex:
/**
* This does some stuff.
* @param entry Does something with this
* DEBT: Assumes the input entry is not null.
*/
void doSomething(SomeObject entry) { }For anyone who needs to look it up like me: http://en.wikipedia.org/wiki/Reference_(C%2B%2B)
If people want to pass lazy NullPointerExceptions, it's their fault.
Depending on your user-base (i.e. if you distribute headers to other developers with precompiled code), the documentation may be necessary on its own... but it's much weaker than an assert.
You should verify arguments at a top level, then let the code underneath blow up with a NullPointerException if something unexpected happened. Your stack trace will point at where the problem lies.
I find for some reason that a lot of people want to use null instead of empty lists where you'll end up with this ugliness:
if (list != null) {
for (Thing t: list) {
processThing(t);
}
}
Instead of just passing in a Collections.emptyList() instead.They had to "fix" it by blocking userspace memory mappings at the 0th page.
OP - I noticed you were thorough enough to mention both Fantom and Kotlin in one section, so for the preceding section you might want to note that CoffeeScript also has a safe-invoke operator like Groovy's.
http://en.wikipedia.org/wiki/Null_Object_pattern
It needs static analysis tools such as Code Contracts (C#) or SpringContracts (Java) to make it really robust.
Is this instead arguing Maybe vs pointers? No, that can't be right; pointers are just type-safe as Maybe, albeit less general.
I guess it's a dynamic vs static typing argument in disguise?
The reason we can compare the type and a value is because they serve exactly the same purpose in different ways. The core argument of the article is that we should get rid of null because Maybe does the same thing in a safer way.
1. Opt-in
2. Not tied to reference types
3. Safe extraction/"dereferencing". ie case expressions rather than Java's implicit "Maybe t -> t"
All of which aren't restricted to a static type system.
The problem for dynamic type systems is how do you enforce 3? Do you care to?
Null means unknown.
Take a look at the difference between open and closed world systems.
Given an object a of type Optional<MyObject>, we write:
if (a.isPresent()){ MyObject o = a.get() }
Of course, I could still do a.get() without evaluating isPresent() and end up with an java.lang.IllegalStateException. Here, we are "reminded" to do the null check.
I would agree that the real world can bait one in the ass, out of the blue, just, what the, things are trying to eat my ass!
As for herding spherical cows and the vacuum, I'll do my best:
val maybeCow = Some(1) // SomeOne is the speherical mu (i.e. you)
maybeCow getOrElse 0 // Cow here becomes 1 with the universe
Were maybeCow initialized with a None value, then into the vacuum it goes...and out it comes with a safe 0 to keep order in our [application] universe.