Why we've banned null in our codebase
blog.stackmob.com
blog.stackmob.com
public int getIntFromHeader(Request req, String headerName) {
if(req != null && req.getHeaders() != null) {
String header = req.getHeaders().get(headerName);
if(header != null) {
try {
return Integer.parseInt(header);
} catch(Exception e) {
return -1;
}
}
}
return -1;
}
as public int getIntFromHeader(Request req, String headerName) {
return Integer.parseInt(req.getHeaders().get(headerName));
}
That way it either returns with a valid int or throws an exception of some sort (null or otherwise).Also, -1 is a perfectly valid integer and so this function can't distinguish between a header with -1 in it and an error. The fact that there's a req != null check seems silly. Surely, the caller should be responsible for that and if they aren't they should get an exception.
If the caller would handle the not-present case by assuming a value then use the default-value version; if it's an error then use the throwing version.
Also, Java is statically typed. If you need such dynamic capabilities, look elsewhere?
I'm not sure what you mean by your static typing comment; I don't see where dynamic typing would offer a better alternative to anything we're talking about here?
I suppose when you talk about a large codebase used by people who are unaware of your conventions then you can run into issues, but I'd imagine that either null serves a purpose in your system, or not. If it does, handle it with the appropriate logic, if not, you shouldn't have to worry.
Dynamic typing would offer you the convention of inherently being uncertain about what you're receiving and being overly defensive about everything.
But that has nothing to do with "null". There are other paradigms (linked lists or weak references, say) where you want the ability to distinguish between a "valid" object and an invalid/empty/not-present one. And whether you call this thing "null" or not is just a semantic question. It's no less an error to operate past the end of the list or try to access an expired cache empty just because you decided to allocate memory to store it.
1. Decide on some property you want the compiler to enforce for you. E.g. I care about only operating on values that exist.
2. Encode this in the type system. E.g. use the Option type to distinguish those values that may not exist (Options) from those that definitely do (everything else).
It's not about being able to distinguish valid and invalid values, it about getting the compiler to enforce it for you. This is the real power of modern type systems -- not bookkeeping to keep the optimizer happy but getting the computer to check the things you care about.
a) Every function checks all its parameters for null, and you do null checks before calling a method on any object, wasting time.
b) It's down to the programmer to figure out where a null check is necessary and where it isn't, and they sometimes get it wrong.
(Of course, better code will avoid calling isDefined or get, but even if you just use it as a like-for-like replacement, using optionals in place of null can improve your program)
If your code doesn't check for a null, JIT adds in an implicit check in case it needs to throw a NullPointerException. When you put a check in your code, it knows to eliminate the implicit check.
Putting null checks everywhere does lead to ugly, complicated-looking code though; it'd be nice if there were a simple syntax for non-null parameters.
The problem with null is that it's a semantic gremlin. Depending on context it could mean all sorts of things including but not limited to, "the value has not yet been set", "the value is undefined", "the value is nothing" or "the software is broken". Which one of those is the correct interpretation at any given time is a question that can only really be answered by a programmer with a good knowledge of the codebase.
Worse yet, even in cases where null isn't a desired behavior it's still inescapable. If your procedure accepts reference types as parameters, it doesn't matter if it has no interest in taking null as an argument, or that it doesn't even make any sense (semantically) for null to be passed as an argument. It still has to be prepared to be handed null as an argument. At a language level, there's simply no support for the concept of, "Guys, I really do need something to go here." Which is simply insane, considering how simple a concept it is.
In languages that don't have null references, it's not that the various semantic meanings of null have been thrown out. All that's been removed is the ambiguity: For each possible (desirable) meaning, a construct is provided to represent specifically that meaning. Meaning that you've gotten rid of the problem of programmers having to magically know what null means in that case, and the worse problem of programmers introducing bugs resulting from them failing to understand null's meaning in a particular context.
This doesn't seem like a particularly good idea if any of your code needs to run reasonably quickly. Even relatively tame uses of Option[T] in the collections that ship with scala can have a terrible impact if you want to run the code rather than just typecheck it. For instance, if you have some loop in which you want to perform I/O and call getOrElseUpdate on your scala.mutable.HashMap[T, Long], your code will probably spend a majority of its time creating millions of java.lang.Longs, stuffing each one into its own newly-minted Some[Long], and GCing these two things.
Of course if this is an issue you can always fall back to Java's collections or a collection class specialised to Longs. I think the Scala implementors made the right tradeoff for a general purpose library.
scala> def doFoo(x: Some[Int]) = x.get*2
doFoo: (x: Some[Int])Int
scala> def doFoo(x: None.type) = None
doFoo: (x: None.type)None.type
scala> doFoo(Option(3))
<console>:9: error: type mismatch;
found : Option[Int]
required: None.type
doFoo(Option(3))
^
But it will work as either of these: scala> def doFoo(x: Option[Int]): Option[Int] = x match {
| case Some(n) => Some(n*2)
| case _ => None
| }
scala> def doFoo(x: Option[Int]) = x.map(_*2)
doFoo: (x: Option[Int])Option[Int]I'm not saying it's a bad trick. I'm saying that thinking of it as "banning null" is an incomplete understanding of the issue.
The Haskell equivalent to this Scala Option fiasco is Maybe, which is defined as:
data Maybe a = Just a | Nothing
"Nothing" is not a null reference and "Just a" is not an object with a pointer to another object. "Just a" is a constructor with a single value parameter. "Nothing" is a constructor that takes no parameters. These are values, not references or values with references or references to references. "Nothing" always has a specific type (Maybe Integer, Maybe String, etc.) even though it has no auxiliary data of its own. This means "Nothing" can't show up anywhere you have values--it can only show up where you expect a value of type Maybe. Indeed, the whole point here is that you are forced to explicitly reckon with the possibility of null where it can happen, and you are freed from responsibility to worry about it everywhere else.
Haskell does just fine without the kind of null we're talking about, in a way that I suspect is fundamentally different from what you're intuiting.
The problem is that for the most part needing to check whether an object is null/"valid" should be an exceptional case, not a common one. Most object variables do not need to have the optional/null state. For languages that allow every object to be null you create a situation where you need to have null checks for every object you encounter unless you come up with some error prone convention outside the language itself. Every method/function has to handle the nulls, it's not clear whether an object will be null when passed into a method or not, and if you are writing APIs for people to use, you must check for nulls because you cannot expect much from the user of the apis.
TLDR; Having non null objects by default reduces the amount of code, which in turn increases the correctness of the code.
Handily enough, C#, Java and similar languages actually throw an exception when you try to work with a null reference as if it wasn't null. These exceptions are generally only cryptic when the null references are stored somewhere other than the stack before being manipulated. That is, so long as foo(null) throws an exception before foo() returns, it's no big deal to diagnose, and mostly unnecessary to check the value of the parameter as it passes through the graph of method calls.
(To a degree I'm arguing a devil's advocate position, as I have a lot of sympathy for encoding not-null-by-default into the type system.)
This isn't something easily fixed by a compiler or language feature. There's no magic bullet.
Bugs "lying in wait" are a minor problem; if they lie in wait long enough, they are almost by definition not a problem at all.
It's possible to code around it if you're presenting a packaged executable or product, but if you're creating an API of any sort, or dealing with user input, it should be a standard case.
Blank inputs or null cases should always be checked in that case.
Dealing with user input is a separate issue, yes you need to validate it. The problem with every type allowing null is that it infects EVERY object for a situation that isn't always needed or wanted.
I'd be careful - that's all. I've seen systems which have done exactly this for core pieces of logic and they have suffered miserably during long running processes.
Mind you I'm referring to Java/.NET here - I'm not familiar with Scala's subsystem so can't comment on that...
[1] http://www.codinghorror.com/blog/2006/01/flattening-arrow-co...
[2] http://martinfowler.com/refactoring/catalog/replaceNestedCon...
[3] http://c2.com/cgi/wiki?GuardClause
The most important point: handle negative conditions first, so checking for null first and putting a return statement right away when detected.
[1]: https://blog.stackmob.com/2013/01/free-yourself-from-the-tyr...
What you're actually doing is using the Option monad in Scala, but the article doesn't tell you that to avoid scaring you. And it succeeds: the pattern described in the second part of the blog post is easy to follow and clearly has no magic.
[1]: https://blog.stackmob.com/2013/01/free-yourself-from-the-tyr...
Essentially, part 2 shows how you can very easily overcome the issues with the approach in part I using some nice Scala features. You can get elegant code without sacrificing the more expressive types.
Here, you’ve just traded one kind of checking (!= null) for another (isDefined). This is more type-safe, to be sure, but has a high legibility cost. There is also no gain in exception safety: NoSuchElementException is not appreciably more meaningful than NullReferenceException.
There's a history of elevating useful programming patterns (constructors, generic references, and dozens of other C++ features come to mind) to the level of being enforced by the language or runtime, and in most of these cases I've found the cost of doing so (in language complexity and inflexibility) far outweighs the benefit of preventing programming mistakes.
The car doesn't need to blow up if the air conditioning isn't working.
The @NotNull annotation seems like it should be the answer, but I don't know of anyone that uses it. Would love to hear if anyone has figured out how to get value out of @NotNull!
> We’ve now, strictly speaking, solved our problem
No, now you have two problems. What if request is null?
It still baffles me why C# and Java does not have this.
At least in C# you can use code contracts to constrain functions to not accept "possibly" null values. I haven't used this myself but the demos I've seen look impressive. Anyone got hand on experience of this or similar?
TLDR: Careful, there could be some one with last name or first name = NULL"
There doesn't seem to be a google cache available.
Made me laugh.