Having used Scala extensively, None as its own type is very much a mistake; it's not a type that you ever want and it only serves to get in the way.
Having used Scala extensively, None as its own type is very much a mistake; it's not a type that you ever want and it only serves to get in the way.
E.g. it's very common to define a helper type like
type Maybe<T> = null | undefined | T
and then with TypeScript's flow analysis I can easily determine when it's safe to call properties on a value, e.g. const foo: Maybe<Foo> = ...
if (foo) {
// Typescript now guarantees foo is not null or undefined
// thus it's a Foo, so I can call any Foo methods on it
foo.bar();
}Have you ever used a language with a good implementation of sum types?
> and then with TypeScript's flow analysis I can easily determine when it's safe to call properties on a value
The problem isn't calling methods on null or undefined. The problem is when the value is null or undefined, but not for the reason you think (e.g. you think your code set it to null, but actually the generic code you called came back with null: T), and so you mix up the semantics.
The argument against union types would be, that they are an advanced feature compared to sum types, and sum types are simpler to use right. The argument against sum types is that they are moving the role of types from descriptive to prescriptive, and that is what leads to the problem Hickey is pointing out.
Why? If you write a function that takes T | null and returns T, then you're going to need to filter out your null.
Now, Typescript isn't safe, so you can have f: (T | null) -> T and give it a (null | null) and its return will indeed be null (at least at the moment, see [0]).
But your function's type guards will, at one point, filter out the null case. You might try doing this through some assertion, but if you do the runtime check, it will in fact blow up at runtime.
So Typescript will allow invalid code, but if you are actually doing type assertions correctly you'll get a runtime error at the "type narrowing" step. Worse than static verification of course, but better than carrying around a null that you think is not null.
Anyways, yeah, the conclusion is that in TS at least, T | null narrowing to T doesn't mean that T is not null. TS is at least smart enough to handle that.
[0]: https://www.typescriptlang.org/play?#code/GYVwdgxgLglg9mABGO...
// returns null if post doesn't exist
async function loadPost(id: string): Promise<Post | null>;
interface ResourceState<T> {
// null if no data has been loaded yet
data: T | null;
}
function createResource<T>(load: Promise<T>): Observable<ResourceState<T>>;
Both are largely designed to fit together (`createResource(loadPost(id))`), but if you do then you now have no idea whether `data` is null because the post couldn't be found, or because it's still being loaded. You would either have to use different nulls (null vs undefined) or have `ResourceState` box T somehow.§um types are safe from this because they can distinguish between `None` and `Some(None)`.
Higher level, I think that it's fairly rare to see Optionals outside of the _very_ basic cases, and downstream of that you're actually not flinging around nulls or undefined as data. Instead everyone reaches for a kind attribute. Especially in cases like you're talking about (where there's this notion of not finding something, but also this notion of something still loading).
Not to say the distinction doesn't matter, but TS feels well designed in the sense that all of its unsafety is in places that end up not coming up in many "normal" codebases
The problem is the other side: if you filter out null then you accidentally end up filtering out part of the T case as well, when you only meant to filter out the case that was not T.
More generally, I submit that if you're doing your types properly you never want to collapse T|S into something different from a sum type. You have code that handles T which you expect to handle the cases that come from the code that yields T. You have code that handles S which you expect to handle the cases that come from the code that yields S. When a T turns out to be an S, or vice versa, that can only ever be a nasty surprise, or at best some code that works by accident.
I'm not sure what this would imply.
const foo: Maybe<boolean> = false;
It seems to me that your code wouldn't behave they way you'd expect.You could say the same with 0, "", NaN.
That being said, I wonder what was the logic behind the decision to implement that.
Typescript supports this because it is a strict superset of Javascript.
They certainly could have made TS strict mode complain about it though.
That being said, I agree it's an unfortunate footgun and think tsc should yell at you for using non-boolean types in conditionals if strict is on.
If it's None, don't put it in the cache. If T is None, don't put it in the cache.
if (typeof input != None) putInCache(input)
> if (typeof input != None) putInCache(input)
Exactly, now you've just written exactly the bug I was talking about.
I'm not an expert on programming languages, but asaik this is how typescript operates
Now have a function whose argument is T|None. Substitute in the definition of T and you get (int|None)|None. If the function's argument is None, what does that mean? Was it given a T that happened to be None, or was it not given a T? Nobody knows.
Hmm. It _feels_ like a function shouldn't ever need to differentiate between "what the caller thinks it has" (that is - being passed a "T-which-is-None" should have the same significance as being passed a "None"), but I'm not confident claiming that would ever be the case.
Also consider a cache lookup function that takes some key and returns a generic T|None for T the type of cache entries, with None signifying the key was not found.
Both are, on their own, pretty reasonable things to have.
If you are using union types and you put the results of your expensive computation into your cache, you will not be able to tell when you've done an expensive computation for some key and got a None result, or when you haven't got a result in the cache for that key.
This is certainly an issue, but that's an issue for the _overall system_ of a cache which stores T. My claim wasn't that "Union types that Union with None can never cause issues", but rather that, for a function whose argument is `(int|None)|None`, there shouldn't ever be any different behaviour _of that function_ between "passed a T (int|None), which was None", and "passed a None (not a T)". The function itself should still behave the same way.
You're right, of course, that _returning_ an (int|None), where "None" might mean "the answer is definitively known to be None" or might mean "the answer is unknown, and that is represented as None" can lead to unnecessary recomputation - but that's an issue of return types, not of parameter types.
Right, neither do I - like I said, "that's an issue for the _overall system_ of a cache which stores T.". The bug is that a given type (None) has different meanings to different components of the system, and this bug only arises _because_ "One function's return value is [being passed, directly, without any interpretation, as ] another function's parameter". If the type signatures were changed so that interpretation was required - so that "the answer is None" could be distinguished from "I don't have an answer" - then the bug in the overall system goes away.
Which is the normal way of programming. If you can't safely compose functions without adding an extra layer of interpretation between them, programming becomes much harder.
> If the type signatures were changed so that interpretation was required - so that "the answer is None" could be distinguished from "I don't have an answer" - then the bug in the overall system goes away.
Which is something that using sum types rather than union types achieves by default. Wherever you want to localise the problem, union types add a big, easy class of ways to shoot yourself in the foot that just aren't there if you use sum types.
As the other posters have said, it (automatically flattening the Maybe monad) is non-composable and should be considered a bad language design, like it was a bad idea to automatically flatten lists in Perl.
sentinel = object()
value = dictionary.get(key, sentinel)
if(value == sentinel) ...
Which, sure, it works, but it's working around a problem that didn't need to exist in the first place. Why not just have Option work the way you expect, and be able to contain None the same way as it can contain any other value? >>> T = int|None
>>> T|None
int | None
Yep, quite simply this is solved flattening the unions and removing duplicates :)Which is confusing and introduces subtle bugs. It makes it impossible to reason about any part of the code in isolation, because you can't understand the behaviour of None|T unless you know what T is.
Same with None. None means “absent value”. If you use it like that, no confusion at all. Absent value or int is absent value or int. However, if in some places None means “absent value”, in others “error value”, in others “infinity”, in others “empty collection”—due to programmer’s laziness instead of using proper types, then _surely_ it will be impossible to reason. But not due to union types, I think.
Which is essentially every use of None? Like, the whole point of an option type, an X | None, is that None is not an X, and means something different.
> I think it is known that magic int values always come back to bite you later.
But the reason a magic int is a problem is because it's also a valid value. If you use -1 as a magic value, you will get confused because you can't tell whether it was whatever magic meaning you meant or the actual value -1. The whole point of using int | None is to avoid that problem, because None is never a valid int. Unfortunately if you have inclusive unions then that breaks down - you use T | None because None is never a valid T, but then if you try to use that with T = int | None, whoops.
> Same with None. None means “absent value”. If you use it like that, no confusion at all. Absent value or int is absent value or int.
What does that mean? I don't think there's any universal notion of "absent" that applies to every function in a single program, much less every program.
What None means is context dependent, sure, but that's fine as long as that context is local; after all, what e.g. 3 means in your program is also context dependent (maybe it means "3 users" or "position 3 in the array" or "file not found"). If you have proper nesting options, then maybe you'll compose together three layers and in the end you have some value where None means "not cached" and Some(None) means "error computing" and Some(Some(None)) means "empty collection" - but that's absolutely fine, each layer knows how to handle its own option and knows what None means in that context. The problem only comes when you have union types, because then you can't compose your layers without them interfering with each other.
Absolutely not! X|None does not mean X is never None. It's just a logical union (∪) from school. In this case it means "all possible values of X and also None if it was not in X".
For example, a function may take something like "Indexable<T> | Iterable<T>", and it is fine if it is passed Vector<T> which is _both_ Indexable<T> and Iterable<T>, the sets are not disjoint.
I think this is the root of our mutual misunderstanding.
> I don't think there's any universal notion of "absent" that applies to every function in a single program, much less every program.
I agree it's sometimes hard to maintain the same semantics over a codebase, but that's because software architecture work is hard.
Number 5 should mean approximately the same over the entire codebase, and for sure programmers will find a way for it to mean different things in different parts of the code, but that's what makes it a sloppy code that is difficult to maintain!
Similarly, if None means different things in different parts of a program, that's not a fault of mathematical logic or union types, that's just sloppy programming!
Types are not sets, and thinking of them as sets will lead you astray.
> For example, a function may take something like "Indexable<T> | Iterable<T>", and it is fine if it is passed Vector<T> which is _both_ Indexable<T> and Iterable<T>, the sets are not disjoint.
Only because you don't care whether Vector<T> is processed as an Indexable<T> or an Iterable<T> - which is because you know that it implements both in a way that's consistent with each other, which is because you know there's a relationship between those two interfaces. But that kind of relationship ought to be expressed in the type system (in this case Indexable<T> should probably be a subtype of Iterable<T>), at which point you don't need to use a union at all.
The key use case for a union U | T is when the two types U and T are unrelated. And in that case, if you passed a type that happened to implement both U and T, you would very much care about whether it was processed as a U or as a T.
> I agree it's sometimes hard to maintain the same semantics over a codebase, but that's because software architecture work is hard.
> Number 5 should mean approximately the same over the entire codebase, and for sure programmers will find a way for it to mean different things in different parts of the code, but that's what makes it a sloppy code that is difficult to maintain!
> Similarly, if None means different things in different parts of a program, that's not a fault of mathematical logic or union types, that's just sloppy programming!
A codebase has to work up from the generic to the specific. Ultimately programming is the art of translating a business problem into a bunch of 1s and 0s, it would be absurd to demand that every 1 or 0 has the same semantics everywhere in your program. Just as at the very low levels you have code that interprets a bitpattern as a number or a character or an enumeration, at a slightly higher level you'll have code that interprets a collection as meaning exclude/exclude/transform or a value as meaning target/default/.... Particularly in library code, you don't necessarily know what the objective semantics of the values you're working on are. And all that's fine and normal - for most code, the internals of the value you're working on are and should be a black box - e.g. a sort function doesn't and shouldn't know or care whether the values it's sorting are numbers or strings, or whether one string is alphabetically before or after another - all it knows is that it has a collection and a way to compare elements of that collection. You should be able to use the same sort function to sort a collection forward that value withor in reverse, even though those are the exact opposite of each other.
It only becomes a problem if you mix up your layers - e.g. if you somehow pass a bitpattern that was meant to represent a number to a function that thinks it was meant to represent a string, or if your sort function confuses the magic value that was returned from the comparator with one of the values it was meant to be sorting. That's not (just) sloppy programming, it's poor language design, because you shouldn't even be able to make that kind of mistake.
You can, Rust has sum types rather than union types.
> That’s really bad!
I agree, but some people seem to like union types for some reason.