The Logical Disaster of Null
rob.conery.io
rob.conery.io
The essay is conflating 2 different concepts of "null."
The Tony Hoare quote is about null _references_ which is about aliases to computer memory. (E.g. the dreaded NPE or null pointer exception.) He's not talking about _logical_ nulls such as "tri-state booleans" or "missing data".
The logic-variant of Null usage for "unknown" "missing" "invalid" values is unavoidable in programming. This is why that null concept shows up repeatedly in different forms such as NaN in floating point, NULL in SQL language, etc. If you invented a theoretical language without logical Nulls, the users of your language would reinvent nulls using worse techniques such as homemade structs with an extra boolean "HasValue" field. E.g.:
struct NullableInteger {
int x;
bool hasvalue;
}
The "hasvalue" field becomes a "null by convention". Other programmers might not even code a verbose extra boolean variable and instead use (dangerous) sentinel values such as INT_MAX "32767" or negative value such as "-1" as the "pseudo null" to represent missing data. A lot of old COBOL programs had 99999 as some out-of-range value to represent missing data. We need nulls in programming languages because they are useful to model real-world (lack of) information. All those other clunky techniques will reinvent the same kinds of "null" programming errors!The orthogonal aspect is making the type system more powerfully aware of nulls so that the compiler sees that possible null conditions were not checked. A compiler error forces the programmer to put in the explicit defensive code to handle the possibility of null values.
It is now common when creating a library to rely on manual checking and coders explicitly writing code to check and throw NullArgument exceptions for each object parameter in every public method, and then writing dozens or hundreds more pointless unit tests to ensure this null exception behaviour works as expected; this is totally unnecessary in languages that demand explicitly stated Option<T> when an Option is meaningful.
Instead of requiring all the above for safety -an approach which will notably fail to report any errors if the null argument checks are partially or completely omitted- we have -and we should prefer- languages where the compiler or interpreter will not let us pass null into a function where null makes no sense, because that fact is encoded in the function definition.
At the library boundaries, which are comparatively rare to internal code, we can simply use the explicit Option type.
I should note I have one objeciton with the F# implementation - in that I think that ideally the properties { .IsSome, .IsNone, .Value } should not exist, as they promote bad programming practices.
[1] https://fsharpforfunandprofit.com/posts/the-option-type/
For instance a `std::string` or a `double` in C++ can never be null ; a pointer to either can be but that's far from idiomatic.
My understanding so far is that option types have two advantages over the way current mainstream languages handle the null case: a) you can find the potential uses of null values statically, and also ensure that the programmer writes some code that at least nominally handles the case; b) you can avoid this burden in cases where null is not an option. Does null-as-its-own-type improve on this?
In Crystal we go a step further and allow type unions between arbitrary types (of which nil is just an ordinary empty struct in the stdlib which you can call methods on). I love how it works in practice.
(I don't like inclusive union types because they're noncompositional - code you write with them isn't parametric any more - which seems particularly awkward in a compiled language - how do you compile code that forms unions involving type parameters?)
I'm not 100% sure how you can integrate sum types with a flow-typing system, perhaps with pattern matching?
So inside the body of a generic method (or class) you can't form unions involving the method's type parameters? That works but makes them much less useful as a language feature - most language features work as normal within a generic method.
> I'm not 100% sure how you can integrate sum types with a flow-typing system, perhaps with pattern matching?
Whatever you do for union types should work, surely. Indeed it ought to be simpler since you have more information to work with - with a sum type if the thing is B then you know it's not A, whereas with a union type it's possible for the thing to be both A and B.
And regarding sum types I was just thinking syntactically instead of in terms of the type system.
Hmm. Does that mean you can't have a generic method in a separate compilation unit from where it gets used?
> I really am interested in knowing more about what you said about union types being non compositional because I'm definitely not a type theorist.
I'm not really a type theorist, I'm just thinking about e.g. if you have a method like:
def doSomething[A](userSuppliedValue: A) = {
val x = if(someCondition) Some(userSuppliedValue) else None
x match {
case Some(a) => y ...
case None => z ...
}
x match {
case None => v ...
case Some(a) => w ...
}
}
then the compiler knows that if we take branch y we will also take branch w and if we take branch z we will also take branch v. Whereas if you do the same thing with an inclusive union type, that's true most of the time but not when userSuppliedValue is null.I'm confused by your example too. If you mean replacing matching None with testing for nil, then you'd have to replace Some(T) with an `else`, then the compiler absolutely can prove that if you take one branch you take the other - given that x is not assigned to (not sure why the compiler would want to prove that though). And after all that I'm not sure how your example relates to, or what you even mean by "noncompositional".
If you use an inclusive union, the type of x becomes A | nil (or however you write it). The trouble is that you don't know whether A might also be a type that includes nil, and userSuppliedValue might be nil already. So if you write in match/case style your branching behaves inconsistently in that case (because x is both an A and a nil) - and if you write in "if(x == nil)" style then you can't infer that branch y will be taken when someCondition is true.
> And after all that I'm not sure how your example relates to, or what you even mean by "noncompositional".
What I mean is that viewing these types as type-level functions, they don't compose. For option/maybe-style types, composition works as expected: you can know the behaviour of X[Y[A]] by reasoning about X and Y independently. That doesn't work for inclusive unions, because whether A | nil is a different type from A depends on what type A is.
Being able to move all your concerns into the type system is really nice, and the difference between a type system that does what you expect 95% of the time and one that does 100% of the time is huge - but I don't know any way to get there except gradually. I've been doing Scala for 8 years and it probably took 4 before I realised how important this stuff was even in the simple case of Option - and that's coming from a background of Java before then and already being dubious of things that step outside the type system.
Language inconsistencies matter, but only in large codebases, and by the time you have those large codebases it's probably too late to change the language. I think Scala gets a lot right, I think something like Idris gets more right. I can make the case for why those things are the right things in terms of theoretical purity/consistency. But when it comes to practical differences I don't know how to convince anyone other than by saying "go maintain a 300kloc system in an ML-family language for 5 years and see how much easier it is" :/.
Best of luck. I hope Crystal finds some things that work well, and that we all add something to the language design state of the art. I do think we're getting gradual progress and consensus - if nothing else, almost every new language has at least some form of sum typing and pattern matching these days - but I guess progress on the aspects that matter for big systems is inherently going to be slow, because it's only after we've built those big systems that we can see what's wrong (and what's right) with them.
And yeah, nobody has any experience maintaining 300kloc+ systems in crystal. And i'm sure we get loads of things wrong and there are lots of pain points. I hope we find some type theorists to pick through the problems when that happens too, because crystal's type system is really quite different to anything I've ever seen before. And it would be nice to see what comes out of that theory wise.
Well, the use cases will be there, and any general-purpose language will need answers for them. One can of course achieve the same effect with other language features (e.g. Java's famous use of the "visitor pattern" for those cases), but there's a cost in verbosity and maintainability.
And pattern matching isn't really needed if you have flow typing + OO (at least I haven't found a usecase for pattern matching which crystal doesn't already have an answer to)
Sum types usually only have two or three entries too.
> And pattern matching isn't really needed if you have flow typing + OO (at least I haven't found a usecase for pattern matching which crystal doesn't already have an answer to)
How does "flow typing + OO" solve the problem? You can use OO to emulate pattern matching via the visitor pattern, yes, but that's notoriously cumbersome; I don't see how flow typing makes any difference to that?
Crystal has method overloading (dynamic dispatch), which is a form of pattern matching (but not destructuring) but that's it. So I guess we have pattern matching but definitely not to the extent say haskell or elixir has it.
Visitor pattern in Crystal is just a class with a bunch of `def visit(node : Type)` and then just call `visit(foo)` and let dynamic dispatch handle the rest. The visited data structure has no knowledge about the visitor.
I'd take it a step further, and posit that The Fine Article's rage comes from its author not grokking the distinction, or even recognizing there is one.
Of course null is going to seem frustratingly and dangerously wrong if that's your concept of it.
PS: To disembiguate I was speaking of the null a value in a nullable type. Maybe I got you wrong.
Also I guess Null is way less a problem in strictly typed language than in JavaScript for instance.
Null values and null pointers aren't fundamentally the same thing, though some programming languages may be implemented in a way such that use of one exposed the other. A language can have a Null value that never results in NPEs.
The problem with nulls isn't strictly null references, in the sense of "a pointer that points to nothing" or "a pointer that throws a null pointer exception", it's that languages with nulls implicitly make null an element of every type. They tend to do this because variables are allowed to have null references, but if you changed the semantics of variables to technically forbid null references, but still had a null value as an implicit element of every type, you'd still have the problem of needing to do runtime checks for null everywhere.
sum = 0
for entry in entries:
case (val):
null: pass
number n: sum += n
endcase
in pseudocode. Several languages have offered this since forever (eg. Haskell and Maybe).The fun part is that the typesystem enforces that you check all possible cases -- or only lets you explicitly un wrap potentially null values.
So it's not like a C programmer checking if (x == NULL) manually.
Alternatively, you can only use languages which offer actual tagged unions, and avoid the ridiculousness of these 'hasvalue' fields.
I tried to address explicit-null-handling programming concepts such as Option<T> or tagged unions in my last paragraph.
Nevertheless, it didn't seem like the essay was about non-nullable concepts but instead, he was writing about the idea of "null" as a data representation for unknown values to be fundamentally flawed because of various boolean operations strangeness.
My point is that "null" as a data representation mapping to "missing" is an irreducible complexity which is orthogonal to Option<T>. E.g. F# Option<T> still has ".None" which is a "null by another name". Yes, the programming language support is superior because it forces the programmer to handle possible nulls instead of crashing. Safety around nulls enforced by the compiler is progress. (The null handling as a verb assisted by compiler checks.) However, it didn't remove the need to represent a null as a state -- or null as a noun. Why do I separate those 2 verb-vs-noun concepts? Because he used examples with constants such as "10* null" returning a "null" instead throwing an error as violating math logic. And "10< nil" violating 2-value boolean logic. Therefore he plays with the idea that nulls in the language shouldn't exist. Let's rewrite one of his examples using variables instead of constants:
x = Option<int>
y = Option<int>
x = null
y = 10 * x
What should "y" be? The essay says that's not the question to ask. He's wondering why multiplication doesn't throw an error. The idea of tagged unions and Option<int> really doesn't address "data" / "mapping" / "noun" concept of null and how it fits into certain intuitions about math and booleans.This irreducible complexity is why "null as noun" will get reinvented as "99999", "-1", etc if you don't have null (whether spelled as ".None", "Nil", or some other standard way). Even if you have an array of Option<int>s and want to serialize to disk, you still need to write "null as data concept" in some fashion (empty string, sentinel value, extra "hasdata" field etc).
As side note, I find it interesting that some commenters believe the author is talking about Optional<T> or similar concept and that I missed his point. I ask those readers to look at the essay again and confirm that it makes no mention of "Option", "Optional", "Maybe", "sum types", "tagged unions", etc. His essay seems to be based on the flawed concept of data representation of unknown values. A key sentence near the top:
>Logically-speaking, there is no such thing as Null, yet we’ve decided to represent it in our programs.
But he's right. It shouldn't be an inherent part of any data. Sometimes, something not being present is a reasonable logical state to be in. For these times, there are tagged unions. In other cases, it is absolutely insane to simply add on the possibility of the NULL value to every data type in a language.
Again, tagged unions do not solve the author's examples of Aristotle logic puzzles and inconsistencies around null/None/Nil/Nothing/etc.
He's not talking about "type safety" enforced by static compiler checks, or opt-in nulls as the default vs opt-out, etc. Therefore, compiler enforced explicit pattern matching code on tagged unions misses the point.
I suggest you click on the Stackoverflow q&a link in the author's essay to get a better idea of what the his philosophical conversation is about.
>In other cases, it is absolutely insane to simply add on the possibility of the NULL value to every data type in a language.
That's a totally valid position but the author isn't talking about that. He's questioning the _meaning_ of "null" in light of observing (in his view) non-intuitive behavior with math and boolean operators.
Your "tagged unions" doesn't address his philosophical conversation at all.
My point is that if the author removed "null" from a hypothetical language because it's "weird", the users of such a hypothetical language would recreate the "null" again. This has nothing to do with discriminated unions or Option<T>.
Your example would be nonsensical in any language I’m aware of.
It makes some sense for interoperation when living inside .NET or the JVM, but this was a known problem when those systems were designed.
2. Restricting a language to non-circular data structures will not be popular.
3. Having a nullable and non-nullable version of every pointer or reference type is not usually super popular either.
Phrased differently: nullable types are actually a fairly reasonable compromise from a set of not fully appealing choices.
This was discussed in rust some time ago: https://mail.mozilla.org/pipermail/rust-dev/2013-March/00330...
And you may start needing fancier versions of polymorphism to express things like "this accepts a T=Foo or Optional<Foo>, and returns Bar<T>".
In passing, I remain to be convinced that that x?.length/etc syntax (and equivalents in other variations of Optional) is really a good idea - to me, it feels like an error-prone way of papering over the fact that non-explicit null(aka absent-optional) checks are just too painful to program with. Phrased differently, I see no reason to believe that the sufficiently-common-to-warrant-special-support appropriate handling of null is just to feed it up through as the computation of the computation if it occurs in any part.
I'd assume that while Kotlin's String? is equivalent to Optional<String>, Kotlin has no equivalent to Optional<Optional<String>>.
Which, personally, I think I'd prefer Kotlin's approach, because that means you can do `maybeString = "exists"`, which is more readable than `maybeString = Optional::Exists("exists")`
Also, you can still nest them in Kotlin, since you can have a nullable type inside generics.
Also, regarding your third point, IMO the question mark syntax in for example TypeScript and C# is fine.
ftp://ftp.cs.princeton.edu/reports/1989/220.pdf
There are various possible solutions for circular structures that do not require destroying all static safety with nulls everywhere.
Doing so *nullifies*
pun intended?[0] https://channel9.msdn.com/Blogs/Seth-Juarez/A-Preview-of-C-8...
But essentially they're not solving it: you will get a warning on explicit default-ing of a non-nullable reference, you will get a warning on implicitly initialising/defaulting class fields, but you will not get warnings when doing so for structs or arrays.
As an aside here, there is also a school of thought that NULL in SQL was a mistake, see e.g. https://www.dcs.warwick.ac.uk/~hugh/TTM/Missing-info-without... and http://thethirdmanifesto.com/ (Hugh Darwen and Chris Date)
Worse, instead of blowing up and yelling "this is wrong, fix it!", they'll pretend to work and give you the wrong results.
I imagine if null is the "billion dollar mistake", then null never having existed - and relying on sentinels - would've been the trillion dollar mistake.
Doing things that way means every time you check the value of a date, the very first check needs to be "is it the sentinel value?". But that's kind of the concept of a sentinel value anyway.
That doesn't work as most of the program logic simple checks of a date is between two dates. It basically just treats 2025-12-31 as the end of all time. I suspect they will patch it before the end and change that date to something else.
The whole idea is that you don't have to rely on sentinels -- and with optionals and pattern matching we didn't have, ever since the 80s or so.
On the contrary it’s perfectly avoidable. You just need to use an option type with a language that supports comprehensive pattern matching and won’t even compile if you don’t explicitly consider the case of a missing value. If your model is flawed and you don’t want to explicitly specify that some value is optional then the problem is in your way of thinking. And there is no need to invent a language that doesn’t support nulls, they already exist.
As the op says - is a "logic-variant of Null"
No, it's not.
> Or a "typed null" to be more precise.
To the extent this is true, it is a fundamental difference.
> It's got some nice properties, but really it's the good old tri-state.
No, it's not. Because the Boolean type contains only True and False, not Null and the Option<Boolean> type (likewise, Option<Option<Boolean>>, etc.) does not contain True or False. There's no weirdness in doing logic operations on an Option<Boolean>, or a pair of them, because it is a type error to attempt it because Option<Boolean> is not Boolean.
I agree that for other reasons it's better, but conceptually they are used to represent the same thing. (Logic variant of null)
Yes it is, the point they make is that Option<Boolean> is explicitly tri-state in contrast to regular Boolean which should be two-state.
void f(char *x) {g(&x); printf(x);}
It's possible that calling g(&x) might cause x to no longer be null, even if x originally was null. It would then be inappropriate for the compiler to complain that the next instruction attempts to use x without checking whether x is null.
So... yeah I dig it man :) I just left some things out to keep on a pace.
Languages which lack a way to signify an error other than null can be an issue. Most modern languages have option types and/or exceptions, both of which provide good ways to deal with error conditions.
NULL in SQL is also problematic -- e.g. as it behaves wrt aggregations and such. Explicitly user defined unknown values are better.
I'm struggling to understand what you mean here. My understanding is that there is no distinction between "logical null" and a "null reference" in terms of the problem we're discussing--as soon as you introduce nulls as a placeholder for values that are not yet initialized you have to deal with the logical implications of null being a member of those types, no? It's been a while since I watched the talk, but scanning the transcript from the talk we're discussing (https://www.infoq.com/presentations/Null-References-The-Bill...) the way Prof. Hoare talks about them seems to be the same in terms of their logical impact on the type system. I quote:
25:55 One of the things you want is to be able to know in a high level language is that when it is created, all of its data structure is initialised. In this case, a null reference can be used to indicate that the data is missing or not known at this time. In fact, it's the only thing that can be assigned if you have a pointer to a particular type.
...
27:40 This led me to suggest that the null value is a member of every type, and a null check is required on every use of that reference variable, and it may be perhaps a billion dollar mistake.
And another thing you wrote which I don't understand:
We need nulls in programming languages because they are useful to model real-world (lack of) information. All those other clunky techniques will reinvent the same kinds of "null" programming errors!
Isn't that exactly backwards? The entire point Prof. Hoare was trying to make is that using nulls to model the "real-world (lack of) information" causes us to have to contend with the logical problems inherent in making null a member of every type. Additionally, sum types seem to model this in a logically consistent way very well, as other commenters have noted.
These are not trivial distinctions; knowing the semantic intent of a NULL can either hamper or aid in reasoning about a code-base and avoiding subtle errors
That said, the memory layout of Maybe/Optional (assuming it can be stack allocated) will look much like what you outlined, though of course the API can be more user friendly.
But it's really not. First of all, most Haskell compilers will optimize away usages of Maybe that are known to be Nothing or Just. Secondly, GHC at least uses tagged pointers for small sum types, including Maybe.
It catches so many stupid cases where I'd have call bar on foo where foo could be undefined.
TypeScript handles it pretty well given the constraints it has to work in.
Any other number of states automatically ensure that the logic is non-Boolean or non-Extended Boolean.
The interesting thing is that positive powers of 2 will also allow consistent extended Boolean logic and semantics.
But anyway, that's history, now we have nice languages that have safe nulls, so the interesting questions moving forward are: why are people still creating new languages that have unsafe nulls (looking at you Go), and why are people still choosing to use languages with unsafe nulls?
Worse is better. Ease of deployment over ease of development.
As for the Q: why design languages with null (that is - incomplete or flawed type systems) now? I have o idea. I think the reasoning is that “worse is better” succeeded for JS, php, C, so it’s a viable path.
“Maybe an X” and “Definitely an X” are more different than string and number. If a language pretends string and number are distinct but at the same time has no distiction for “maybe X” - then it doesn’t have a very good type system.
Note that it doesn’t necessarily need to avoid null values for this. Non-nullables is mostly equivalent although less elegant. That is, “String s” means a string or null, while “String! s” means a non-null string (example from a C# vNext syntax).
Different explanation of history: Something "wins" because it is precisely what is needed.
"My goal was to ensure that all use of references should be absolutely safe, with checking performed automatically by the compiler. But I couldn't resist the temptation to put in a null reference, simply because it was so easy to implement."
References in C++ are non-nullable so the precedent already existed.
I hope C# gets around to solving this problem -- Microsoft is working on it -- but it's much harder to retroactively solve it.
Wrong analogy. References in C++ are not first-class entities. Conceptually, they're _alias names_ for existing objects. Unlike C# and Java "references", you cannot "reset" a C++ reference to "point" to another object. Because it's not a pointer. It does not "point". It's an alias. `void f(int& ref)` in C++ is the same as `void f(ref int x)` in C# (the equivalent doesn't exist in java). Inside `f`, you cannot change `x` to "point to some other int" because it's not a pointer!
C# and Java "references" are semantically the same as C++ pointers minus arithmetic.
This could've been done since 2003 but wasn't. There are attempts to bring it into language as part of C++ Core Guidelines project.
This is possible because unlike Java and C# not everything can be null.
But you can entirely write a shitton of C++ programs that won't use pointers at all - not even shared / unique. In contrast, you can't avoid C# and Java references.
template<typename T> class List<T> {
protected:
List();
template<typename U> class Visitor<U> {
public:
U visitNil();
U visitCons(T current, U accumulator);
}
public:
virtual <U> U visit(Vistor<U> visitor);
}
template<typename T> class Nil<T> extends List<T> {
U visit(Visitor<U> visitor) = visitor.visitNil();
}
template<typename T> class Cons<T> extends List<T> {
private:
T head;
List<T> &tail;
public:
Cons(T head, List<T> &tail) {
this.head = head;
this.tail = tail;
}
U visit(Visitor<U> visitor) = visitor.visitCons(head, tail.visit(visitor));
}Then you've thrown in the towel when it comes to optimality. Not saying this is necessarily bad (maybe you're inventing the next Python), but acknowledge this isn't really an acceptable option for the next C.
type X struct{}
func (x *X) y() {}
func main() {
var x *X = nil
x.y() // no error
}
Also its not unsafe. Go doesn't allow unsafe memory access unless you use the unsafe package.It is incorrect though. An Go does allow you to write bugs.
res, err := http.Get("wcpgw")
defer res.Body.Close()
...
panic: runtime error: invalid memory address or nil pointer dereference
[signal SIGSEGV: segmentation violation code=0x1 addr=0x40 pc=0x5ed9af]Yes we did.
It's also wrong to suggest that Null has no place in "logic". Boolean logic is one type of logic, but it's not the only type.
Finally, the examples of how Null works in various languages are really poor. "Why doesn't Ruby coerce nil into zero?" Because it doesn't coerce any types implicitly. Why should nil be the exception? How is that expectation "logical"?
What they don't have is the global concept of Null that sits at the root of the type hierarchy. This means every type in a language like C# or Javascript has to have Null as one of its members. If you want to define some operators/functions on the members of the type, have have to always consider Null as a member. Like the author listed: How do you compare Null, how do you negate Null? Those questions shouldn't even be asked because Null is not negatable or comparable. Maybe the problem is that Null shouldn't have been admitted as member of a type that you consider comparable or negatable.
So in Swift/Haskell, when you define a type you don't consider "nothing" as one of its members.
The mental model in Null-using-languages is that a Type is a sort of blueprint that can be stamped out to make instances of that type. From that point of view its natural to consider you may be missing an instance.
In language like Swift or Haskell the concept of a Type is closer to a set in math; in this case the set of all possible values of that type. This is why I said above that Null has to be considered as a member of a type. I’m using math language to re-interpret the type-as-blueprint world view, to show why it leads to illogical results.
Seriously, Swift demoting of the concept of “nothing” to just another enum is alone worth the price of admission. This has cascading consequences that lead to safer code. You also get simpler code as you strive to resolve the uncertainty of a missing value as early as possible in your code. Swift is the most impressive language I've seen in a while, but think its origins at Apple have overshadowed the its incredible technical value.
e.g.
public Foo GetFooWithId(int id){...}
When I call that, I might get a Foo but what if I don't... Will it be null, is Foo a struct or a class? I have to know what Foo is and check if I need to check the result for null.As opposed to
public Maybe<Foo> GetFooWithId(int id) {...}
I know just from reading that I might not get a Foo back so I had better match over the result.IMO optimals are just a crutch to make up for a too weak type system.
I think this is becoming a trend in modern language design. However as this comments section demonstrates it’s hard to understand its benefits or why it’s an important improvement, if your only experience is from C/C++/C#/Java etc
Consider:
x : Int32 | String
if x.is_a? Int32
# typeof(x) == Int32
else
# typeof(x) == String
endI should have stuck to terminology I know.
If it stayed true to its word, and fully embraced 3VL, there wouldn't be an issue.
Type systems can be viewed as a lattice, usually as a bounded lattice. If you look at Java's reference type system (ignore primitive types), there is a top type, aka Object. There is also a bottom type. It doesn't have a name you can spell, but it does have a single value--null. This means that null can satisfy any type, even if the type is impossible to satisfy--and as a result, there is an unsoundness in Java's type system.
Obviously "nothingness" exists and needs representation. The problem with Null is that it's not part of the type system but is part of the program, so it undermines the value of having a type system.
I'm not optimistic enough to think it would always be a good thing, though. I've actually seen some where the lack of a value was not noticed because people just mapped over the optional and did not code for the missing case. Effectively coercing the value to whatever the zero was.
Now, I can't claim empirically that these would outnumber null pointers. I just also can't claim they don't exist.
That was exactly what I would've guessed with those examples, even without knowing Ruby. Likewise, C# being the one where things "get strange", especially after Javascript? From the example, the C# version acts really close to NaN - a "this is unknown" value that contaminates whatever it touches. I'd call C# and Ruby equally logical, with different intentions, given the examples.
Sounds good, waiting for an example to support this...
>It's also wrong to suggest that Null has no place in "logic". Boolean logic is one type of logic, but it's not the only type.
What other types did you have in mind? Given that we work as programmers in a world defined by true/false 1/0 logic, I think you might want to reconsider this blanket dismissal.
And there's an infinite other number of topoi to choose from! Topoi can be custom-made to categorical specifications. We can insist that there are three truth values, and then we can use a topos construction to determine what the resulting logical connectives look like. [3]
Finally, there are logical systems which are too weak to have topoi. These are the fragments, things like regular logic [4] or Presburger arithmetic.
To address your second argument, why do we work in a world with Boolean logic? Well, classical computers are Boolean. Why? Because we invented classical computing in a time where Boolean logic was the dominant logic, and it fits together well with information and signal theory, and most importantly because we discovered a not-quite-magical method for cooking rocks in a specific way which creates a highly-compact semiconductor-powered transistor-laden computer.
Computers could be non-Boolean. If you think that the brain is a creative computer, then the brain's model of computation is undeniably physical and non-classical. It's possible, just different.
Oh, and even if Boolean logic is the way of the world, does that really mean that all propositions are true or false? Gödel, Turing, Quine, etc. would have a word with you!
[0] https://en.wikipedia.org/wiki/Topos#Elementary_topoi_(topoi_...
[1] https://ncatlab.org/nlab/show/topos
[2] https://ncatlab.org/nlab/show/two-valued+logic
true || 3 ==3 || panic()
will, through short-circuiting, evaluate to “true” even though the final element doesn’t evaluate to either true or false.(As you may observe, they also follow a weird logic where AND and OR are not commutative...)
Operators are lifted[0] over nullability in C#. Every value type T can be converted implicitly to Nullable<T>. When `10 * null` is typechecked, both sides of * are typed as Nullable<T>. The * operator then acts like the pseudo-Haskell
(*) <$> 10 <*> null
or, I guess liftA2 (*) (Just 10) Nothing
The semantics of Nullable<T> are similar to those of an optional type, with some implicit mapping and lifting. In that context, null acts less like null. (Thanks to convenient conversions.) Note that the following doesn't throw: int? value = null;
value.HasValue // == false
[0]: https://blogs.msdn.microsoft.com/ericlippert/2007/06/27/what...[0] https://basarat.gitbooks.io/typescript/docs/options/strictNu...
With Java we used @NotNull annotations and IDE support to give warnings. Luckily we are converting our legacy code base to Kotlin.
With conditional type conditions Typescript will have better support than earlier:
type NonNullable<T> = Diff<T, null | undefined>; // Remove null and undefined from T
Source and the examples in: https://github.com/Microsoft/TypeScript/pull/21316Computer systems are purely logical, but the applications that we write, if they embrace Nulls, are apparently not. Null is neither true nor false, though it can be coerced through a truthy operation, so it violates Identity and Contradiction. It also violates Excluded Middle for the same reason. So why is it even there?
And suddenly a bunch of type theorists just winced. It is possible to have trivalent logic that's coherent. SQL does this -- and it makes sense to do it in that application.
However, this person begins their article poisoning the well, saying,
> I find that the people I talk to about this instantly flip on the condescension switch and try to mansplain this shit to me as if.
Proceeds to speak in that very same condescending tone, presenting a smarter-than-thou trundle down from the mountain. All the while including some nonesense about logic.
This person clearly has a chip on their shoulder: tacking on things about community-regulation in an article about Null, like all technical public engagement authors do. Let's signal "what good behaviour we expect" in the midst of a book pitch and a confused discussion of a technical matter.
One of these goals can be failed in its attempt with good nature. When you smush all this together it seems an exercise in performative intellectualism, or more accurately, pseudo-intellectualism.
Saying they don't work because Aristotle didn't use them is like throwing out Newtons laws because Aristotle didn't use them.
There have been whole schools of mathematics that reject the law of the excluded middle: https://en.wikipedia.org/wiki/Constructivism_(mathematics)
Newer logics also don't use that law but instead deal with undecidability, which is less than a century old but is the most import result in logic since it's invention.
That said, I think a large portion of the problems caused by null could have been avoided by making one small change to its behavior: accessing a member (or element) of a null should evaluate to null instead of throwing. This is deeply intuitive and is the biggest cause of null pointer exceptions. If you try to access a.b, and a is null, it makes perfect sense that b is also null. Many languages have recently started adding an "optional chaining" operator that works this way, but a lot of pain could've been avoided if things were this way from the start.
The billion dollar mistake is about not being able to say that a type (specifically a reference/pointer) does not include NULL/nil/whatever as an element.
Solution to the problem is already there: Kotlin, TypeScript, Rust, Crystal, Swift, Haskell just to name few.
We just need to push the industry toward better and safer languages. We keep building an abstractions on top of unsafe languages and then we are suprised that it suck.
Oh but there is.
The author is probably considering only propositional calculi with two truth values, such as Boolean algebra and, er, well, they're probably only considering Boolean algebra because that's what we use in computers, because it maps nicely to 0s and 1s.
However, two-valued logics are by no means the only possible logics, neither are they the only ones that have actually been described in the framework of mathematical logic. Probably the most well-known many-valued logics are Łukasiewicz's and Kleene's that have three truth values (i.e. values assigned to literals): true, false ...and unknown.
... which is to say, "null".
And just to blow your mind, there are also infinite-valued logics, like fuzzy logic and, I'd argue, the good old probability calculus of the reverend Bayes, which is nothing if not a many-valued logic.
The mistake is, I think, that the author is taking "logic" to mean Aristotelian logic, however that is not at all the logic we use in computers. Like I say above, computers use Boolean algebra which is an entirely different formal system with its own axioms, separate to grandpa Aristotle's own. For instance- Aristotle never said anything about functions, mappings from the set of literals to {0,1}, neither did he formally define the algebra of the Boolean operators AND, OR and NOT (a.k.a. conjunction, disjunction and negation). Although you can project Aristotelian logic onto Boolean logic, they are far from the same and I would really struggle to see how one would implement a programmable computer using Aristotelian logic.
P.S. Am I mansplaining now? Wouldn't that be a little ...weird?
Regardless of what logics exist the statement you make - that "unknown" is "null" - is actually wrong and the heart of the problem. "Unknown" is only one semantic interpretation of null; there are many others, and problems arise in software development because of those (sometimes slight) differences in the interpretation of what "null" should reflect in the real world.
Some sources suggest as many as 129 possible semantic meanings for "null", which is why Date rails against "null" values in SQL. With Option/Maybe we're constrained to the universe of values of that type plus exactly one more value (None); with null who knows how many of the 129 possible meanings of the null value we are dealing with in addition to the base type? Null is computationally and mentally expensive to deal with.
That null is implemented oddly doesn't negate the fact it's a meaningful and well defined mathematical construct.
I don't see people saying false doesn't exist because forth implements it as -1.
Edit: removed unnecessary blablah.
The only catch, is now you have to unwrap your function results somehow (functional languages provide ways to do that, like the scary monad).
I like this website, which demonstrates the concept: https://fsharpforfunandprofit.com/rop/
Result<,> has the same semantics as checked exceptions, BTW. The conversion between the two representations of code is mechanical.
Checked exceptions are a poor idea the more dynamically bound your language is, and are generally anti-abstraction in any case (failure modes are implementation specific, i.e. non-conformant with an information hiding interface). An error result is only useful if you can make a decision based on the specific value; that's not the case for almost all sources of error in most user (i.e. non-system) programs, where complete coverage of error cases with specific handlers is outside their design parameters, and termination (of program or request or whatever) is preferable.
Equivalent C# code is to add checks for nulls all over the place, aka “defensive coding”. C# has an Option type sort of with nullable types, but they are for value types only, and developers can still reach in and just grab the value, eliminating the safety that option types provide.
> Null is neither true nor false, though it can be coerced through a truthy operation, so it violates Identity and Contradiction. It also violates Excluded Middle for the same reason.
Intuitionistic logic doesn't have the law of excluded middle and is the form of logic underlying the simply-typed lambda calculus.
There is a more nuanced relationship between logic and programming languages than is being discussed here.
OP should check out trinary logic and first-order predicate calculus. Just because OP doesn't appreciate "null" doesn't mean that it should be removed from programming languages or from existing programs.... false means not true, zero means no things, null means absence.
That there is a logical truism, isn't this fun?
Also: false means not true, zero is a number and null doesn't exist, by definition. We can model true/false/zero easily as they exist. Null is made up, so every language gets to think about what it means in an abstract made up way. Thus the pain, thus the post.
It is very much still an open question if a statement is absolutely undecidable and if we need to add something like null in all logic [0].
As for your arguments on why we don't need null, they sound exactly like the arguments against zero from the middle ages [1].
>Just as the rag doll wanted to be an eagle, the donkey a lion and the monkey a queen, the zero put on airs and pretended to be a digit.
[0] http://logic.harvard.edu/koellner/QAU_reprint.pdf
[1] Menninger, Karl (1969), Number Words and Number Symbols. Cambridge, Mass.: The M.I.T. Press.
It's that in assembly Null's just a zero.
People claiming the industry is wrong and that we should use some logically pure language are usually impractical and deny the compromise that must be made for systems to be usable by the masses.
"this" should have been a reference, not a pointer. Strostrup admits that was a mistake.
/pendantic
Value types can work without null because 0 is valid.
Could the language designers implement object types without null? It would be very difficult for fields, especially in structs, because it's nearly impossible to force memory initialization. You couldn't do "default (StructType)" if it had a not null object as a field.
That said, people are _also_ free to experiment with null-free languages, or to avoid null. If it works for their use case, fantastic!
Just the attitude of the post is terrible. For example the user creates a sock puppet and posts a trolling/edgy question to Stack Overflow... just to see if it's toxic.
Sorry, these posts are toxic, in my opinion.
On the actual topic, I think a good compromise is what typescript and I think rust do, which is, a value is only nullable if you explicitly declare it as such. It forces you to be judicious, while still allowing you the option when you need it.
This is to say: it could be good, or it could be another "exceptions in Java" moment
If programming is about managing complexity and status is ideally specified explicitly, then null is simply the option in a mathematical set defining that status which corresponds to the option commonly seen on surveys: Other (please specify): ...
While it can be handy to have this 'extra value' in a set, and it is most commonly used to denote special meanings within a carefully controlled context (eg. SQL column in a result set is empty, a variable is of no type - ie. not defined at all, etc.), issues arise when people accidentally carry context or presumptions about the meaning of null across contexts, creating a leaky abstraction.
Most of the author's article appears to deal with differences in these context-specific assumptions.
Perhaps the Java approach: throw an exception.
Unix approach? Nonzero return values and arbitrary stream or file data to clarify. In edge-case leakiness, very similar to the Java approach.
The functional and dedicated non-OO procedural programming approach: define exit parameters to your function, specifying complete precision and ending any ambiguity.
Since a type is a formal context (set), then using a typed language is another solution, although that adds overhead it brings benefits in rigor.
Sometimes, the elegant implementation is just a function. Not a method. Not a class. Not a framework. Just a function. - John Carmack
Setting aside how different languages (mis)treat null as a concept, I think the whole discussion is about convoluting boolean logic with memory addresses and some languages do a better/worse job at being 'intuitive' than others.
I had great success with this[1] whenever I felt things were going in the wrong direction or when the language constructs of dealing with nulls were in the way of the domain design.
[0] https://www.b4x.com/android/forum/attachments/unbenannt-jpg....
Null values propagating through expressions SQL style is an approach to handling unknown values. Whether it's desirable is somewhat besides the point; databases have null values, and LINQ enables capturing expression trees that get converted into SQL, so keeping similar semantics makes a kind of logical sense.
Nulls are undesirable in a language until you want to initialise large structures in a simple language with a simple compiler. Staying away from the temptation then takes ingenuity and discipline.
0 is not an option, as that will return incorrect mean/median calculations. So what would be there? Or am I misunderstanding?
If you do anything else, such as using nulls, then you are just creating a problem that WILL come back and bite you in the nether regions of your psyche.
Of course, there will other opinions about how to handle this. As far as I am concerned, Nulls are a curse foisted on us by those vendors and standards bodies who took the the easy way out.
That will create real problems and bite you immediately.
When nulls are stored there are no guarantees that anything you ask of the database will ever turn out right. Especially when there are millions of records stored. Two queries that should give you the same answer give different results when nulls exist. Seen it too often.
I have also worked for companies that didn't use relational database theory for their products and they had far more issues. In a couple, I was able to hive off the database designs from the main systems and got the applications to actually work and work properly.
Ideally, your database system would support tagged unions, and have a strong type checker. Then, you may have a column of type 'Maybe Int'. If you want to take the median, the median function would be of type 'Agg Int Int' (or something like that), so you wouldn't be able to use this column directly as an input. Instead, you'd have to only supply rows whose values have been checked to be 'Just'.
So
CREATE TABLE MyTable ( group TEXT, myColumn MAYBE INT )
SELECT group, MEDIAN(myColumnValue) FROM MyTable WHERE Just MyColumnValue = myColumn GROUP BY groupI wonder, if one were to query the DB, what would show up in the "birthdate table" column? I'm kind of hoping it would be zero, which would in turn be interpreted as the epoch, but sadly it would probably be null...
Not deliberately, perhaps reflexively. It just seemed the most straightforward (and interesting) interpretation of the comment in question. I hadn't seen the other comments.
> I wonder, if one were to query the DB, what would show up in the "birthdate table" column?
As I understand it, what would show up in the "birthdate table" column is a table. You would have to do a select on that table to get actual data out, and in that select you'd get either zero or one row returned.
instance Monad Maybe where
return = Just
Nothing >>= _ = Nothing
Just x >>= f = f x
(And is not really related to the discussion)Because Ruby doesn't automatically coerce arbitrary types to numbers when you try to perform arithmetic on them. This is a Good Thing.
2 * "4"
> TypeError 2 * "4".to_i
> 8Interestingly, it is possible to multiply a string by a number, but no coercion takes place.
"2" * 4
> "2222"I think implicit coercion is bad, but that's a separate issue from null.
In data, on the other hand, null is a very useful concept. It explicitly means that there is no value in a field. And that prevents such mischief as interpreting missing data as zeros.
1. What should `malloc()` return on failure?
2. How should I indicate that I don't care about certain out parameters, e.g. the parameter to `time()`?
3. How should I initialize variables that are going to be used as out parameters, e.g. in `strtol()`?
Other languages address these at the cost of significantly complicating their type system and ABI: generics, multiple return values, etc. In a language intended to be small like C, I'm not sure how you can do better.
Initialization for out parameters does not matter. In C++ you'd return a tuple or structure or class instead of having multiple return values.
It does complicate ABI some.
1. Why not just return meaningful values? Wouldn't that help in debugging?
2) ???
3) That makes sense to me.
But isn't ambiguity the tradeoff?
Computer SCIENCE is that ugly thing that tells us that three-valued logic gives rise to 19683 distinct binary logical operators, while two-valued has 16.
Computer SCIENCE is that ugly thing that tells us that if you want a computer language over a three-valued logic to be expressively complete, then you need to implement all of those 19683 logical binary operators one way or another. In the worst case, that's 19683 operator names for the programmer to remember. And you come here claiming that it's "trivial" because you have a "sense" of what the results ought to be ? That proves just one thing but site policy probably won't allow me to spell that out.
(In case you were wondering what the 16 names are in two-valued logic : they aren't needed because the system being two-valued gives rise to certain symmetries that gracefully allow us to reduce the set we need to remember to just {AND OR NOT} (or some such) which beautifully parallels the way we communicate in everyday life.)
First, the same argument above is also an argument that, say, integers, are not "Computer SCIENCE".
More to the point, you might enjoy reading the work of Charles Pierce and other logicians of that era who began to explore many variations on formal logic. Note that just as many operations arise from trinary relations in bivalent logic. Are binary relations "Computer SCIENCE", but not trinary or higher relations? Before you answer, you might want to look into whether all possible relations can be expressed using only binary relations (hint: nope).
Look deeper into the concept of functional completeness (with respect to a subset of operators), which you reference above without naming. You might be able to understand how many of those many trivalent operators are actually necessary to reason with (hint: not very many, hardly more than for bivalent logic, where, as you note, we only tend to use a few, and need not worry about it).
Consider also the relationship between operators folks have identified as useful in bivalent vs trivalent logic (hint: they not picking at random).
Could it be that just as with the 16 binary operators, many of which have relations to one another (e.g. inverses and complements, among others) that the trinary operators could fall into similar groups, which, making the 3^9 number you mentioned seem a whole lot less complex? Could that be why it's neither necessary nor customary to work with all the operators in either sort of logic?
Once you've caught up to state of the art in formal logic as of the 1930's you might have a new perspective -- perhaps you might even begin to let us know when "Computer SCIENCE" will catch up!
Oh, and if you want to know why people don't want to find more "useful" operators than what they're used to from good old two-valued logic then I have a hint for you too : it's because they all immediately sense that their brains are not up to it as soon as they actually try (and my actually doing the maths has very clearly shown me why - so as you suggested to me "perhaps give it a try").
Don’t get me wrong - theory is important. The models that you develop though are only an approximation of the real world.
The way I see it there are practitioners who care about theory (and take the bother to try and understand some of it, and even more so the consequences on their practical tasks) and there are those who don't.
(Aside : the world of fact-oriented modeling and especially FCO-IM - that's communication oriented modeling - has this stance that our databases are actually not modeled according to how the world is, they are modeled after how our comunications about the world are. An interesting distinction you might want to ponder a bit more deeply.)
> Logically speaking, there is no such thing as Null
Yes, there is: "Did you pass your test?"
"I haven't taken it yet."
"Okay, will you pass it?"
"Null."
Suppose there was a database table of students, with a column called "passed," which is boolean. If it's in the middle of the semester, that column must be null. True means they passed. False is put there when they fail.---
Another example, in a table of help tickets, suppose there is a timestamp field called "closed," for when the ticket was closed. If the ticket is open, then that column must be null.
---
"Which show is currently playing on channel 3?"
"It's just static. It's midnight, and channel 3 has stopped broadcasting."
(This of course must have been back in the '80s.)If the TV guide were a database table, with time slots for each channel, then the columns for which show is scheduled for midnight on channel 3 must be null.
In fact, that's how I picture null: ever-shifting static inside the little cell in my database table. That's why I'm fine with how if you ask if two null values equal, the answer is null (at least in Postgres).
Someone may say that there is a show currently on channel 3, and the name of the show is "Static." But that's not quite true. If channel 3 recorded static and broadcast it, then yes. But channel 3 has literally turned off its tower. It is broadcasting nothing. Your TV is showing static because you tuned to a station that is not there.
---
The writer complains about inconsistent implementations of null, like in JavaScript. I agree that there are misimplementations. I would not have had both null and undefined.
In fact, that's another way of understanding what null is. Suppose you say my earlier examples with SQL are flawed, because SQL is flawed, because it insists on every row having columns it doesn't need, that the field for "passed" should even exist in that row until the class is over. Say instead you were using objects. Within the semester it may look like this:
{
name: "Edgar Smith"
}
Then at the end you add: {
name: "Edgar Smith",
passed: true
}
Well, what if you asked for the value of student.passed in the middle of the semester? You might say it is a mistake, but it is a question that can be grammatically formed: "What is the value of the 'passed' key in the object?" The only answer is null. To me null often means "not applicable."---
Again, another example. In school we played a game of guess-the-object, where you could ask only yes-or-no questions. There were three possible answers: yes, no, or "does not compute." Clearly the answer of "does not compute" was needed for when the answer was neither yes or no. "Does not compute" was a fancy synonym for "null."
---
Null is the answer to the infamous loaded question, "When did you stop beating your wife?"
---
Null is quantum foam. It is what is happening at the base of the universe, if you look close enough. Is the particle here or there? Null.
You may say that doesn't make sense, that it's just that we don't have the right instruments. But that is Quantum Mechanics. If you wish to disprove Quantum Mechanics, Einstein and I would both be interested, as he was uneasy with it too. But going after quantum mechanics would be a better use of your time than going after null. Right now, null and the scientific consensus about physical reality agree.
We traded goto for callback hell and ten-layer inheritance hierarchies. A null-purge would probably end similarly.
Don't like null? Don't use it.. and you find null propagation ruining your abstraction then treat it as an error and fix your program.
You know Rust has null too, right? https://doc.rust-lang.org/std/ptr/fn.null.html
Rust has both references (which can't be null†) and pointers (which can be null, but can only be dereferenced within "unsafe" blocks or functions).
† Actually, they sort of can: an Option<&T> takes the same space as a &T, with the None variant of the Option being represented as a null behind the covers.
There is no “sort of can” here. It just so happens that None uses the same representation as null, but by that logic you’d say that u64 can be null because 0 has the same representation as null.
To paraphrase Carmack, "Any syntactically valid code, that the compiler will accept, will eventually make it into your code base." [1]
Take a database of people. There is literally no sensible default for name, age, gender, height, weight, social security...
If you’re amazon, your products have no sensible default for manufacturer, shipping weight/size, delivery address...
In fact, for just about any real-world data, there simply is no sensible default for anything at all. Most “sensible” defaults will eventually bite you in the arse. The only sane way to keep nulls from your DB is to refuse inserting incomplete data in the first place, and propagate the error to the user. Heavens save your team if you’re dealing with batch data and insist on not allowing nulls in the DB, though.
You can sweep this mess under a rug and pretend you have no nulls by turning things into relations that are allowed to be empty — “there are no delivery_address rows for this user” — but that’s a null in sheep’s clothing. Either your application knows how to deal with the query coming up empty, or it doesn’t.
Where have I said any such thing ?
(BTW I doubt very much that "Codd designed null into the RM". Even his 12 rules mention only "a systemic way to deal with missing information", not "null".)
You don't.
If a value may not be present for an entity, it's not an attribute of the entity in question, it's an attribute of another entity that has a (0..1):1 relationship to the entity in question.
Normalization eliminates NULL.
Or maybe I don't do a join. Maybe I do a separate query. If the query comes back with zero rows, then I... what?
Then I ... what ? Then you do what needs to be done as specified by the business in the case the queried piece of information is unknown.
In the join case, don't I get a NULL in the row that comes back if there isn't an entry in the other table? Or do I just not get a row?
> Then you do what needs to be done as specified by the business in the case the queried piece of information is unknown.
Sure, but how do I represent that condition in my software? With a different class/structure? With a flag that indicates that the other field isn't valid? Or with a null?
From where I sit, normalization doesn't make the problem go away at all.