I claim Rich Hickey is wrong about non-null arguments to functions (2020)
blog.jonstodle.com
blog.jonstodle.com
This means that relaxing the precondition is, by definition, not a breaking change because the old inputs are a subset of the new ones. Similarly, strengthening the postcondition isn’t a breaking change because the new outputs are a subset of the old ones. (If and he use-site assumes that it gets a specific member of the function’s codomain for specific input values, you have an abstraction leak and the use-site has a bug [or the function should be inlined because it’s a false abstraction])
> You are changing the behavior of the function. That is a breaking change.
So this guy is declaring someone else "wrong" because he unilaterally decides to redefine what a "breaking change" means? This is one of the dumbest "all Internet arguments are over semantics" examples I can think of.
As an aside, I've seen a number of recent instances (here's another, https://news.ycombinator.com/item?id=36826111) where a flat out ridiculous blog post makes it to the front page, only for the top comments to all point out how dumb and ridiculous it is. To stay within the HN guidelines, I'm going to refrain from speculating why, and I certainly know that HN audience is not some monolithic group, and that people have many different opinions. Would just like to know if others have noticed this or have other specific examples.
1) Agree with the article (virtually no one has) 2) Point out that the breaking changes Rich Hickey pointed out might be worth it, or not a big deal
A few people have responded with sentiments of #2. Not because it's not annoying to change these things, but that it's a smaller price to pay for relatively good type systems (I'm thinking more F# than Haskell)
Unfortunately, the talk Rich Hickey gives explicitly says Clojure spec doesn't have a solution for a flexible type system either, and the solution they proposed is still in alpha 4 years later. So while I appreciate the insights he offered, I'm torn between the problem of refactoring with the extra work of a finicky type system, or the work of hoping my tests are good enough to refactor my Clojure code.
Types in Clojure is the number one requested feature in their surveys, year over year, and despite the reputation he has, Rich recognizes that it can really help refactoring a code base. Good variable names aren't going to help you with this. Pushing Spec2 along would be a big win without having to implement a full blown type system.
Is there a language that you feel solves this problem better than Spec could ever hope to?
(defun strlen (s)
(assert (stringp s))
(len s))
(defun map-len-test ()
(mapcar (lambda (arg)
(ignerr (strlen arg)))
'("abc" nil "defg")))
(pprinl (map-len-test))
(defun strlen (s)
(if (null s)
0
(len s)))
(pprinl (map-len-test))
Output: (3 nil 4)
(3 0 4)
The caller had one idea of substituting a value in the null argument case. The function changed behavior, implementing a different idea.Only in a statically typed language in which the call with nil is impossible, and which provides no dynamic access to the compiler, can we be 100% confident in saying that extending the function to support nil is a non-breaking change.
If the is impossible, then the function has no behavior for that case; a program invoking that case doesn't exist in a translated, executable form and so there is nothing to break as far as the program is concerned.
It's possible for another program to break: a compilation test case which validates that code which calls that function with nil cannot be compiled. If the language supports dynamic access to the compiler, then an application can be written which can behave differently due to the change: an application which dynamically compiles a call to the function and now gets a valid result (compiled code) rather than an error.
In the dynamic world, it's a non-breaking change for applications that wisely don't rely on exceptions being thrown on bad inputs, and instead handle those themselves.
If you're changing the function, you cannot know that is not the case; at beast you can decide not to care.
Not caring is not the same as knowing.
This just means that the preconditions are “the value is not null” even if there are no types to capture that precondition. I was pretty careful not to use the word “type” in the comment you’re replying to.
Programmers will make code dependent on anything that is visible in an API, even if the specification tells them not to, an ISO standard tells them not to and even if their mother tells them not to.
Here, the spec may even be saying: if you call this function with a non-string, it throws an exception.
There are no _checked_ preconditions. It’s only meaningful to talk about the correctness of a program relative to a specification of its behavior and a specification needs to specify preconditions and postconditions. So, if there are no preconditions, any behavior of the function being called is as correct as any other behavior.
The fact that some programming languages allows the callers of a function to pass values that don’t satisfy the spec is irrelevant to the point I’m making.
No matter what consequence you give to the failed check, something can break when the check is taken away.
We could have a True(r) Scotsman's checked precondition like this
(unless (stringp arg)
(abort)) ;; bail the entire process image with abnormal status
If we check it like that, nobody will try to depend on the behavior within an application, in the way that my test program did.It can still be a breaking change if the function is written that way, and then changed so the abort is taken away. Because the application can do this:
(defun my-stupid-abort ()
(strlen nil)) ;; this calls abort!
So now if the program relies on (my-stupid-abort) to abort, that will break; suddenly that function call returns zero and the process keeps going.You really need:
;; ... ffi definitions here ...
(unless (stringp arg)
(acpi-power-off))
But again, dumb program can depend on (strlen nil) to power of the machine, which could be part of some critical system that breaks as a result of it not powering off.Only a compile time check can do it, by preventing the program from existing. A program that never compiled and therefore was never deployed to its installation will never break.
Umm, no. If the program is supposed to unconditionally produce the output "Hello", but it instead produces "Goodbye", then it is incorrect. Unconditionally means not that there are no preconditions but that the precondition is T (true). If the precondition is constant truth, it means that there are no circumstances under which the program can be excused for not satisfying the postcondition (e.g. producing "Goodbye" rather than "Hello").
Correctness is almost irrelevant here, because the question is about breakage. It is in fact the change in specification which is behind the breakage. The function is correctly implementing its specification at all times; the correctness of my example strlen is never being called into question. What it means for it to be correct has been changed, and the function's definition followed suit.
Most accurate and practical definition of a breaking change is - change dependents can't detect and complain.
You replace O(n) algo by a O(n^2)? Breaking if someone complains. Vice versa as well.
Is fixing a bug a breaking change? Yes. Ask Microsoft about porting SimCity that relied on a bug in MS-DOS.
There would be very little point in doing changes that are completely undetectable. Surely, there is Hyrum’s law and fixing even a bug can be a breaking change, but I believe it is not as black-and-white as in the case of compiler optimizations. For example, the reverse of your aforementioned algorithm from O(n^2) to O(n) is still detectable, but should not cause any harm in any reasonable program. Surely it can break a program full of race conditions that just so happened to work with the former implementation, but it should be similar to how compilers handle UB — if you are in the wrong, then I really can’t save you (at least that is my opinion).
Also, relevant xkcd: https://xkcd.com/1172/
That's why Microsoft was in a bind. You fixed an issue. Your change obeyed all relevant theoretical stuff, but to a consumer it's a breaking change.
But customer sees that going from Win X to Win X+1 broke their stuff and they will harass your support.
From the other side, people picked your function in particular not because it fills that type hole but because it does something in particular. If you meaningfully change the thing it does then that probably breaks the callers.
The author's argument is something along the lines of (1) if you're changing the type signature, that's probably because a new implementation requires it, and (2) if the new implementation is sufficiently different to warrant a change in the type signature then it's going to break somebody. I don't particularly like the example since in my experience somebody adding a nullable input will leave the old behavior alone and just wants to call the same function somewhere else, so it probably actually is a non-breaking change for all existing callers, but it's not hard to imagine stronger examples where the author would be right.
The purpose of interface contracts is exactly to precisely specify what can change and what can’t change without breaking clients, and conversely which assumptions clients can or can’t make. This is a precise give and take. The clients are expected to play by those rules just like the implementations of the interface are expected to maintain their promises — nothing more, nothing less.
In other words, if you break prod, "but we never promised to behave this way" is a poor excuse. When changing something, it's your responsibility to check for breakage, not just throw up your hands and say "not our fault" if it occurs.
Nobody does that, right? It's undefined behavior.
Downstream distros broke; it had to be backpedaled out.
Edit: another nice consequence of this design would be None being its own type, which can be transparently converted to an Option<T> for any T, allowing for type inference to go in only one direction, while still avoiding the boilerplate of Option<T>::None. Unidirectional type inference would make for more intelligible compiler errors. Full Hindley-Milner can get confusing.
You don't need to qualify the None, the following works:
fn foo<T>(x: T) -> Option<T> {
None
}I'll disagree on "spooky". If the type is ambiguous, the code will fail to compile. Contrapositively, that means that if the code compiles, the types aren't ambiguous. I find bidirectional type inference to be absolutely lovely (or at least Rust's implementation of it), and I wouldn't give it up.
I would love to see examples of this. We've gotten much better on this front over the years, but I'm sure there are plenty of cases yet to be addressed.
The "problem" with improving error messages is that the common and easy cases are addressed early, leaving only the uncommon and difficult to address left after a while, and people get used to the understandable errors which leaves them baffled when they encounter one that isn't.
Having used Scala extensively, None as its own type is very much a mistake; it's not a type that you ever want and it only serves to get in the way.
If it's None, don't put it in the cache. If T is None, don't put it in the cache.
if (typeof input != None) putInCache(input)
> if (typeof input != None) putInCache(input)
Exactly, now you've just written exactly the bug I was talking about.
I'm not an expert on programming languages, but asaik this is how typescript operates
Now have a function whose argument is T|None. Substitute in the definition of T and you get (int|None)|None. If the function's argument is None, what does that mean? Was it given a T that happened to be None, or was it not given a T? Nobody knows.
Hmm. It _feels_ like a function shouldn't ever need to differentiate between "what the caller thinks it has" (that is - being passed a "T-which-is-None" should have the same significance as being passed a "None"), but I'm not confident claiming that would ever be the case.
Also consider a cache lookup function that takes some key and returns a generic T|None for T the type of cache entries, with None signifying the key was not found.
Both are, on their own, pretty reasonable things to have.
If you are using union types and you put the results of your expensive computation into your cache, you will not be able to tell when you've done an expensive computation for some key and got a None result, or when you haven't got a result in the cache for that key.
This is certainly an issue, but that's an issue for the _overall system_ of a cache which stores T. My claim wasn't that "Union types that Union with None can never cause issues", but rather that, for a function whose argument is `(int|None)|None`, there shouldn't ever be any different behaviour _of that function_ between "passed a T (int|None), which was None", and "passed a None (not a T)". The function itself should still behave the same way.
You're right, of course, that _returning_ an (int|None), where "None" might mean "the answer is definitively known to be None" or might mean "the answer is unknown, and that is represented as None" can lead to unnecessary recomputation - but that's an issue of return types, not of parameter types.
Right, neither do I - like I said, "that's an issue for the _overall system_ of a cache which stores T.". The bug is that a given type (None) has different meanings to different components of the system, and this bug only arises _because_ "One function's return value is [being passed, directly, without any interpretation, as ] another function's parameter". If the type signatures were changed so that interpretation was required - so that "the answer is None" could be distinguished from "I don't have an answer" - then the bug in the overall system goes away.
Which is the normal way of programming. If you can't safely compose functions without adding an extra layer of interpretation between them, programming becomes much harder.
> If the type signatures were changed so that interpretation was required - so that "the answer is None" could be distinguished from "I don't have an answer" - then the bug in the overall system goes away.
Which is something that using sum types rather than union types achieves by default. Wherever you want to localise the problem, union types add a big, easy class of ways to shoot yourself in the foot that just aren't there if you use sum types.
As the other posters have said, it (automatically flattening the Maybe monad) is non-composable and should be considered a bad language design, like it was a bad idea to automatically flatten lists in Perl.
sentinel = object()
value = dictionary.get(key, sentinel)
if(value == sentinel) ...
Which, sure, it works, but it's working around a problem that didn't need to exist in the first place. Why not just have Option work the way you expect, and be able to contain None the same way as it can contain any other value? >>> T = int|None
>>> T|None
int | None
Yep, quite simply this is solved flattening the unions and removing duplicates :)Which is confusing and introduces subtle bugs. It makes it impossible to reason about any part of the code in isolation, because you can't understand the behaviour of None|T unless you know what T is.
Same with None. None means “absent value”. If you use it like that, no confusion at all. Absent value or int is absent value or int. However, if in some places None means “absent value”, in others “error value”, in others “infinity”, in others “empty collection”—due to programmer’s laziness instead of using proper types, then _surely_ it will be impossible to reason. But not due to union types, I think.
Which is essentially every use of None? Like, the whole point of an option type, an X | None, is that None is not an X, and means something different.
> I think it is known that magic int values always come back to bite you later.
But the reason a magic int is a problem is because it's also a valid value. If you use -1 as a magic value, you will get confused because you can't tell whether it was whatever magic meaning you meant or the actual value -1. The whole point of using int | None is to avoid that problem, because None is never a valid int. Unfortunately if you have inclusive unions then that breaks down - you use T | None because None is never a valid T, but then if you try to use that with T = int | None, whoops.
> Same with None. None means “absent value”. If you use it like that, no confusion at all. Absent value or int is absent value or int.
What does that mean? I don't think there's any universal notion of "absent" that applies to every function in a single program, much less every program.
What None means is context dependent, sure, but that's fine as long as that context is local; after all, what e.g. 3 means in your program is also context dependent (maybe it means "3 users" or "position 3 in the array" or "file not found"). If you have proper nesting options, then maybe you'll compose together three layers and in the end you have some value where None means "not cached" and Some(None) means "error computing" and Some(Some(None)) means "empty collection" - but that's absolutely fine, each layer knows how to handle its own option and knows what None means in that context. The problem only comes when you have union types, because then you can't compose your layers without them interfering with each other.
Absolutely not! X|None does not mean X is never None. It's just a logical union (∪) from school. In this case it means "all possible values of X and also None if it was not in X".
For example, a function may take something like "Indexable<T> | Iterable<T>", and it is fine if it is passed Vector<T> which is _both_ Indexable<T> and Iterable<T>, the sets are not disjoint.
I think this is the root of our mutual misunderstanding.
> I don't think there's any universal notion of "absent" that applies to every function in a single program, much less every program.
I agree it's sometimes hard to maintain the same semantics over a codebase, but that's because software architecture work is hard.
Number 5 should mean approximately the same over the entire codebase, and for sure programmers will find a way for it to mean different things in different parts of the code, but that's what makes it a sloppy code that is difficult to maintain!
Similarly, if None means different things in different parts of a program, that's not a fault of mathematical logic or union types, that's just sloppy programming!
Types are not sets, and thinking of them as sets will lead you astray.
> For example, a function may take something like "Indexable<T> | Iterable<T>", and it is fine if it is passed Vector<T> which is _both_ Indexable<T> and Iterable<T>, the sets are not disjoint.
Only because you don't care whether Vector<T> is processed as an Indexable<T> or an Iterable<T> - which is because you know that it implements both in a way that's consistent with each other, which is because you know there's a relationship between those two interfaces. But that kind of relationship ought to be expressed in the type system (in this case Indexable<T> should probably be a subtype of Iterable<T>), at which point you don't need to use a union at all.
The key use case for a union U | T is when the two types U and T are unrelated. And in that case, if you passed a type that happened to implement both U and T, you would very much care about whether it was processed as a U or as a T.
> I agree it's sometimes hard to maintain the same semantics over a codebase, but that's because software architecture work is hard.
> Number 5 should mean approximately the same over the entire codebase, and for sure programmers will find a way for it to mean different things in different parts of the code, but that's what makes it a sloppy code that is difficult to maintain!
> Similarly, if None means different things in different parts of a program, that's not a fault of mathematical logic or union types, that's just sloppy programming!
A codebase has to work up from the generic to the specific. Ultimately programming is the art of translating a business problem into a bunch of 1s and 0s, it would be absurd to demand that every 1 or 0 has the same semantics everywhere in your program. Just as at the very low levels you have code that interprets a bitpattern as a number or a character or an enumeration, at a slightly higher level you'll have code that interprets a collection as meaning exclude/exclude/transform or a value as meaning target/default/.... Particularly in library code, you don't necessarily know what the objective semantics of the values you're working on are. And all that's fine and normal - for most code, the internals of the value you're working on are and should be a black box - e.g. a sort function doesn't and shouldn't know or care whether the values it's sorting are numbers or strings, or whether one string is alphabetically before or after another - all it knows is that it has a collection and a way to compare elements of that collection. You should be able to use the same sort function to sort a collection forward that value withor in reverse, even though those are the exact opposite of each other.
It only becomes a problem if you mix up your layers - e.g. if you somehow pass a bitpattern that was meant to represent a number to a function that thinks it was meant to represent a string, or if your sort function confuses the magic value that was returned from the comparator with one of the values it was meant to be sorting. That's not (just) sloppy programming, it's poor language design, because you shouldn't even be able to make that kind of mistake.
You can, Rust has sum types rather than union types.
> That’s really bad!
I agree, but some people seem to like union types for some reason.
E.g. it's very common to define a helper type like
type Maybe<T> = null | undefined | T
and then with TypeScript's flow analysis I can easily determine when it's safe to call properties on a value, e.g. const foo: Maybe<Foo> = ...
if (foo) {
// Typescript now guarantees foo is not null or undefined
// thus it's a Foo, so I can call any Foo methods on it
foo.bar();
}Have you ever used a language with a good implementation of sum types?
> and then with TypeScript's flow analysis I can easily determine when it's safe to call properties on a value
The problem isn't calling methods on null or undefined. The problem is when the value is null or undefined, but not for the reason you think (e.g. you think your code set it to null, but actually the generic code you called came back with null: T), and so you mix up the semantics.
The argument against union types would be, that they are an advanced feature compared to sum types, and sum types are simpler to use right. The argument against sum types is that they are moving the role of types from descriptive to prescriptive, and that is what leads to the problem Hickey is pointing out.
Why? If you write a function that takes T | null and returns T, then you're going to need to filter out your null.
Now, Typescript isn't safe, so you can have f: (T | null) -> T and give it a (null | null) and its return will indeed be null (at least at the moment, see [0]).
But your function's type guards will, at one point, filter out the null case. You might try doing this through some assertion, but if you do the runtime check, it will in fact blow up at runtime.
So Typescript will allow invalid code, but if you are actually doing type assertions correctly you'll get a runtime error at the "type narrowing" step. Worse than static verification of course, but better than carrying around a null that you think is not null.
Anyways, yeah, the conclusion is that in TS at least, T | null narrowing to T doesn't mean that T is not null. TS is at least smart enough to handle that.
[0]: https://www.typescriptlang.org/play?#code/GYVwdgxgLglg9mABGO...
// returns null if post doesn't exist
async function loadPost(id: string): Promise<Post | null>;
interface ResourceState<T> {
// null if no data has been loaded yet
data: T | null;
}
function createResource<T>(load: Promise<T>): Observable<ResourceState<T>>;
Both are largely designed to fit together (`createResource(loadPost(id))`), but if you do then you now have no idea whether `data` is null because the post couldn't be found, or because it's still being loaded. You would either have to use different nulls (null vs undefined) or have `ResourceState` box T somehow.§um types are safe from this because they can distinguish between `None` and `Some(None)`.
Higher level, I think that it's fairly rare to see Optionals outside of the _very_ basic cases, and downstream of that you're actually not flinging around nulls or undefined as data. Instead everyone reaches for a kind attribute. Especially in cases like you're talking about (where there's this notion of not finding something, but also this notion of something still loading).
Not to say the distinction doesn't matter, but TS feels well designed in the sense that all of its unsafety is in places that end up not coming up in many "normal" codebases
The problem is the other side: if you filter out null then you accidentally end up filtering out part of the T case as well, when you only meant to filter out the case that was not T.
More generally, I submit that if you're doing your types properly you never want to collapse T|S into something different from a sum type. You have code that handles T which you expect to handle the cases that come from the code that yields T. You have code that handles S which you expect to handle the cases that come from the code that yields S. When a T turns out to be an S, or vice versa, that can only ever be a nasty surprise, or at best some code that works by accident.
I'm not sure what this would imply.
const foo: Maybe<boolean> = false;
It seems to me that your code wouldn't behave they way you'd expect.You could say the same with 0, "", NaN.
That being said, I wonder what was the logic behind the decision to implement that.
Typescript supports this because it is a strict superset of Javascript.
They certainly could have made TS strict mode complain about it though.
That being said, I agree it's an unfortunate footgun and think tsc should yell at you for using non-boolean types in conditionals if strict is on.
Bidirectional type checking has good error messages and a very intuitive implementation.
My personal experience has been very much in line with OP's: nonlocal type inference (in Rust specifically) frequently makes code hard to modify because a change can cause inference to fail in unexpected ways and places. Local inference doesn't have the same tendency.
I'm using it for my personal lang project and I'm quite satisfied as well, it translates easily into code and interacts well with subtyping.
Really? This surprises me to the point that I'd like to ask what you are coding.
I'm having a very difficult time picturing "You have A|B|C|D and you send in an A and extract a D" (union type without tag). The only way to do that is to guarantee memory layouts, which is something Rust explicitly does not do.
What am I missing?
Making assignment easier and compile time at the cost of making access costlier and runtime seems like an unusual tradeoff. I'm very interested in the case that makes that worthwhile.
There is no shifting of cost from compile time to runtime, or slowing down access. Just like with sum types you're writing match clauses from time to time, with union types you do the same. If you need to inspect the discriminant, then you need to inspect the discriminant.
With both enums and sum types, if you want to differentiate the types at runtime, you need to include a discriminant in order to be able to do that. Enums include that discriminant explicitly (that's what the `Token::Identifier(...)`/`Token::Number(...)` part represents). The memory layout for, say, a base identifier can consist only of the necessary fields for that identifies, but when it gets wrapped it will get an extra field that identifies which variant of the en it is.
With sum types you have two options:
* You always include a type tag directly in the representation (so a base identifier always has an extra byte or so of information with it). This is typically used in dynamic languages where the types are known at runtime anyway. I believe it's also the case in Java and I suspect other VM-based languages are similar. * You dynamically generate the enums at compile time (so a base identifier doesn't have a tag, but if it's used in a `Identifier | Number` union, then a tag for that specific union gets added in).
The problem with the first is that it adds unnecessary runtime cost, because now every instance of any type has extra information in it that probably isn't needed. The problem with the second is that the creation and unwrapping of enum values becomes implicit, and you'll probably need to juggle your discriminants around more. Say I gave a function of type Identifier | Number -> Boolean | Identifier | Number, then should the discriminants for the Identifier and Number cases stay the same? Or do we need to add extra code to munge the input discriminants into output discriminants? In the first case, this adds a lot of complexity to the compiler (and may not in the general case be possible), in the second case we have runtime performance issues.
This is, I think, the point that the previous poster was getting at. If you have implicit discriminators, this will have an implicit runtime cost that may not always be easy (or possible) to avoid. That is the value of explicitly wrapping types in enums - the runtime cost can be minimised, and the developer is always aware of where that runtime cost appears.
Suppose I have an Identifier or a Number or a StringLiteral. To work intelligently, this needs Identifier and StringLiteral to be disjoint. This isn’t just a type theory issue —- it’s a semantic issue. What is “foo”? Is it an Identifier or a StringLiteral? What happens when I add a Filename (for including/embedding files in the future)?
Logically, the fact that an Identifier is encoded by a particular type is a bit of an implementation detail, and a good type system will catch errors. It’s the fact that it’s an identifier — that is, it’s purpose — that matters for semantics and for understanding the code, and union types (and C++ variants, for the most part) are very bad at exposing this.
> What is “foo”? Is it an Identifier or a StringLiteral?
Normally, it's an identifier, since there are no quotes. If there is ambiguity about it, the lexical categories you're working with are ambiguous, which is a real issue. When designing a grammar, that should be avoided. Sadly, sometimes we're given a poor-quality grammar like in C, where some tokens can be interpreted both as headers and as string literals (but neither is a subset of the other) and can't do anything about it, but then you just need to decide which category you're going to assign to ambiguous cases (either one of two, or make a third one) and later you can disambiguate based on context. In any case, I don't consider it an issue of the tool, but an essential domain problem to be solved explicitly by the programmer.
> Normally, it's an identifier, since there are no quotes.
Sorry, I was vague. I don't mean the sequence f o o in the input or the sequence " f o o ". I mean the string (String, str, whatever your language calls it) containing the three letters f o o, which is stored in a union type.
Sure, one can lexically-scoped-newtype it if the language supports it, but at that point, unless I'm missing something, it might as well be a sum type.
fn do_something(value: impl Into<Option<T>>) {
let value = value.into();
// remaining function body using `value` as an `Option<T>`
}
Existing calls to `do_something` will continue working because `Option<T>` implements `From<T>` (and `Into` implementations are generated in the opposite direction whenever `From` is implemented). Passing in `None` will also work because every type `T` implements `From<T>`, and therefore `Into<T>` as well.I'd even argue that in a lot of cases, this is worth doing even when creating new library functions (where there are no existing callers to worry about). In my experience, it's generally better to keep the API clean from a caller's perspective by taking on extra boilerplate internally in the implementation rather than shift the burden to the callers of the API in order to keep the library's implementation clean.
If client code passed a null value before, an error would have occurred. Now it is handled with a default that the original code was not accounting for, and this might be bad.
Thus the change in interface _is_ a change in behavior.
No client could, even. He's talking about static types so it wouldn't have been possible even to express a caller not providing a value.
If I were a hardware designer and some changes to an IC I'm working on are going to make the next batch, say, work properly in hotter temperatures than it could before, and one of my coworkers comes and says, "Don't make that change, what if someone using that IC is deliberately operating it out of spec expecting it to fail and now the behaviour is going to change!" I'm going to start polishing up my resume, because I'm working for an organization that employs lunatics. Happily, that wouldn't happen in the hardware world; unhappily, I'm a software developer rather than a hardware designer.
Old situation: called function says: "i would crash if you gave me a null value, so my interface says you cannot give me a null value"
New situation: called function says: "i no longer crash if given a null value, which you couldn't do before anyway, so you won't notice any difference"
Also, thrown exceptions are values. Consider for example a function that searches for a file a return null if it's not found, which you are composing with a function that opens a file but throws on null. You might very well write something like open(search(...)) and catching exceptions above that level. Now if I make open(..) able to accept null (maybe it opens a temp file?) then you now need to add a null check on the return value of search to get the old behaviour. That is 100% a breaking change!
As a C++ programmer (oh, the horror!) I'd say that's an issue with Clojure rather than the general principle. And also one of the reasons I like the pointer/reference distinction in C++ (the former allows null values, the latter does not).
But sure, that changes my interpretation somewhat. It doesn't really change program safety, however.
If at the same time as increasing the domain of the function, some of the return values changed for some possible values, then it is a potentially breaking change (although one not necessarily reflected in the API), and needs to be communicated. Whether it really is considered a 'breaking' change will depend on a fuzzy evaluation of whether it was legitimate for the client to make such assumptions about the return value, and will most likely depend on how the function was documented.
That case, where the function returns different values to what it used to, and which may break code that had expectations about those values, is a change that can happen any time and is not really related to the question of what to do when a function domain increases.
That makes almost everything breaking.
Going back to the original article, it seems reasonable to me that the following qualifies as such a user-observable change:
foo(null) # throws InvalidArgumentException
updated_foo(null) # the same as foo(0)This isn't a breaking change because up until that point any code that was written to use that method already wasn't passing null (because it couldn't, because it wouldn't compile). The method's behavior hasn't changed, just its type signature, and so for any of the arguments that the existing code might possibly pass to it, it will still handle all of those exactly the same as it would have before (because, again, the implementation did not change).
Therefore, it is not a breaking change.
Yes, in most programming languages/environments it is always possible to write your application in such a way that any change of a dependency is I’ll break your app. Heck, you could throw an exception if someLib.version != “1.0.2” and then complain that a patch to 1.0.3 is actually a breaking change. Or your app could test that some API method does not exist, then complain when that method gets added later. Or you could complain when performance improvements reveal race conditions in your app (or break your usage of your computer as a space heater).
We can divide those breaking changes into two:
1. The call obeys the specification.
2. The call, though apparently successful, circumvents the specification, relying on undocumented behavior.
The question is how much we care about 2, and there is no 100% answer.
There can be situation in which some kinds of 2 breakages are such that we care about them more than 1 breakages.
Suppose there is a certain correct, documented way of using the API, but almost nobody out there uses it that way. And suppose there is some undocumented way of using the API, which millions of installations use, thousands of times per second.
Suppose we need to implement something new, or even ore importantly, fix a critical bug, and suppose that the work boils down to either breaking one, or the other. It may be better to break the former to keep the latter working (and possibly elevate the latter to documented status).
In other words, in a perfect world we'd like to say that breakages of type 1 are non-negotiable, whereas 2 can be debated. But we can't even do that.
I both agree with the author and disagree.
I disagree that Rich Hickey is wrong when it comes to whether those changes should be breaking or not. Those curves can be non-breaking, unlike what the author of the article claims.
But I agree with the author that the fact that they are breaking doesn't really matter.
Rich Hickey mentions this himself in his talk. He says that no one talks about the costs of using Option or Maybe. So then he lays out the costs.
And I was...not that impressed.
Sure, this change that should be non-breaking is breaking. And your downstream clients will have to change their code. That seems bad.
Until you realize that the only change they would be forced to make is to delete what is now dead code.
Sure, it's annoying; I don't deny that. But Rich Hickey claims that it increases code maintenance to have to make the change. I whole-heartedly disagree because any time you can safely delete code, you are making maintenance easier.
Plus, despite all of his complaints that people don't considers costs, he never really considers the costs of his proposal, the biggest of which is the complexity of the language.
I think the industry has rightly come to the conclusion that, all else being equal, and even sometimes when they are not, a more complex language is a worse language.
I think proper union types, the kind needed here, would add enormous complexity to the language for little benefit in rare cases. (Because how often do programmers relax constraints? I don't think they do that often.)
In my language, I'll keep Option as it is, and make such changes breaking, in order to avoid the complexity of union types.
That was a theory behind golang, and the observed result is that you're better off having a little bit more language constructs that may be misused, than having essential complexity in your problem space that you can't easily describe in your solution space, for lack of tools.
In “Simple made easy” he defines complexity as entanglement and simplicity as the opposite of that.
A proper union Option = T|None, is simpler than a sum type Option<T> = Option{T, None}, because the latter complects the types into a container. It’s even worse with nominal types because it complects both yhe structure and the name.
I think you think of complexity of language implementation (compiler) or runtime complexity?
That’s a fair argument! It highlights a different preference.
Yes, it "complects" the types into a container, but then you write code to handle the container, and only the container, until you are ready to open it up.
With union types, you must write code to handle all types in the union always.
This complects the code written in the language. It is also an obvious entanglement and meets his own definition of complex.
When every one of 100 libraries you depend on, does that once a month, you are forced to perform hours of daily unproductive and annoying busywork unrelated to your goals, just because every lib author wants to make your maintenance easier!
Of course, at some point you get fed up and just freeze all your library versions and stop ever updating your dependencies—which is arguably the opposite of what everybody should strive for.
> I think proper union types, the kind needed here, would add enormous complexity to the language for little benefit in rare cases.
Yes, and there's another, simpler way which proved to be working well: dynamic typing.
That's not the rebuttal you think it is. I mention this specifically in my post: how often do programmers actually relax constraints? I don't think they do it very often.
> Yes, and there's another, simpler way which proved to be working well: dynamic typing.
Funny thing: I implemented something close to Clojure's map stuff in C. It really is dynamic, with checks at runtime.
Having only dynamic typing in a language is as much a mistake as making it complex
I think every time a new parameter is added to a function, or some new "mode" of calculation gets supported, or a new field is added to a struct, etc. I personally do that all the time, but it's very hard for me to estimate how often other programmers usually do that.
> Having only dynamic typing in a language is as much a mistake as making it complex
I am interested in your thoughts why.
The reason: he specifically mentions that tightening the contract should be a breaking change, and adding a parameter is tightening the contract. Same with adding a new field.
I'm only concerned with changes that should not be breaking changes, but are because of implementation issues.
That isn't to say that static type systems are automatically good. If a static type system cannot express dynamic types, it is no good because sometimes, dynamic types are needed.
Speaking of dynamic typing...
Having only dynamic typing means that you have to wait until runtime to catch every mismatch. That's a bad deal when that mismatch may happen in production.
The very existence of TypeScript, the existence of Spec in Clojure, and the fact that type annotations have been added to Python should be enough evidence of this. These were all languages that prided themselves on their dynamicism; if they reneged, there's probably a good reason.
I mentioned above that having static types with no way to have dynamic types is a mistake. I believe it's just as much of a mistake as having only dynamic typing. They are both needed.
But personally, I would make static typing the default.
If a function took 3 parameters, and I've added another optional 4th one, how is that a tightening? All old callers are still using 3 parameters, their contract has not been broken.
Of course, adding a required 4th parameter is an incompatible change, if you meant that. I wouldn't call it strictly tightening though, since it's incompatible both ways. But yeah, excuse me for being vague, I meant adding optional parameters.
Similarly, adding a field to a struct which the caller does not have to initialize or mention at all, is, in my opinion, also not a breaking change and does not tighten. Since all callers don't need to change their code or semantics.
Adding an optional one should not be a breaking change, and I have an idea or two of how to do it without breaking callers, even in a static language.
Adding a field to a struct should still be a breaking change because that field could be used. If the caller does not initialize it, you might have a bug.
If you're taking about Clojure maps, and not structs, that's completely different. But Clojure maps are open. Structs are closed.
Adding something to a closed thing is a breaking change. Why? Because of what I said above: the new thing needs to exist. Rich Hickey calls this "place-oriented programming," and he hates it for that reason. He's not entirely wrong either because yes, it can easily cause breaking changes.
But when you add something to an open thing like Clojure maps, that should not be a breaking change, and having implemented something like Clojure maps in C, I can tell you that it is not breaking, even in C.
Whilst existing callers who obey the contract will not be broken by this callers who depend upon this code breaking for null values will no longer do so.
so it is a breaking change for them. You have just extended their behaviour.
That said if we defined the functionality of code as being valid within a scope of input and not defined anything outside of this we are safer in a sense. That said you do require feedback to show you are out of range.
But the point I want to make is that the line of argument that attempts to ascribe changes to behaviour on values that were out of scope as breaking changes is not helpful to anyone who is trying to co-exist via contracts in order to understand what work is expected as a result of change.
See for yourself the "Fig. 3. Clojure codebase—Introduction and retention of code" chart at page 26 of https://download.clojure.org/papers/clojure-hopl-iv-final.pd...
Rich Hickey is right and has a track record to show.
If there are some documented circumstances under which the function returns null, then you can write a program which produces those circumstances and expects the null value. That program will break if there is no more null value.
The only way it can be a non-breaking change is if there are no such circumstances; the function never actually returned null, but only threatened to do that in its documentation or the way it was declared, or both.
In fact, changing a function that returns a string to one that returns a string or nil, can be a realistically useful non-breaking change.
It can be a non-breaking change if the null value is not produced for any client which adheres to the existing API documentation. So that is to say, the circumstances by which the function returns null are entirely new.
For instance suppose we have a function like this:
number identity(number)
It's an identity function, that only works for numbers. We cannot call identity(nil); that is an error. Since it returns number, it will never return nil.This is a totally non-breaking change:
any identity(any)
the function now just returns its argument, no matter what that is. It will now return nil---but, only if invoked as identity(nil), which was previously erroneous.The conditions are paramount. When we analyze the call, the conditions are easy to think about: is the situation that the program is passing nil, which was previously not allowed, or not? When we analyze the return, it's not so easy: we have to think about: is this returning nil under existing conditions under which it previously could not have done that? Or is it only returning nil under new conditions?
The conditions which pertain to the call are always those of the call. The conditions which pertain to the return are also those of the call, plus the semantics of the function!
Because you're restricting or relaxing the types, and you want to type system to know this.
This doesn't have to translate to a change in behavior.
When it comes to types, a subtype is one that can be transparently substituted for its super type. If S <: T, a term of type S can be used where a T was expected just fine. For union types, it's a given that T <: T|a for any type a, and a T will behave exactly like a T|a whose value is a T.
The Liskov principle states that a subclass should be a subtype.
Hickey seems to be talking about breaking the compiler. That is, he is saying that changing the return signature from null|int to int should not break the compilation. In effect, this would create a warning a la 'You are checking for null, but var x cannot be null'.
The author seems to be talking about breaking changes in the behaviour/operation sense.
Those are different types of breaking.
Some new ways-of-calling will now be legal, but that is not a breakage, because no code that already existed could have been using those ways before, and so no code exists that can be broken.
Clients _may_ choose to write new code to take advantage of the new behaviour - but there is nothing that they _must_ do (unlike with a breaking change, where they _must_ take action). It's for the purposes of this categorization, for communicating about client's responsibilities, that the concept of breaking changes is valuable.
> But that doesn't seem to be what you were, or are, saying. If you issue is specifically with "an epicurean theme park of enlightenment and joy to span the stars", and not with the dramatized reaction to (against) it, then...I must say I still don't understand what the problem is. An experience which is _definitionally_ pleasant, fulfilling, and non-harmful (to self or others) is....well, it's capital-g Good, no?
Technological advance is fine on its own. The risk with "an epicurean theme park of enlightenment and joy to span the stars" is our own mental faculties, presuming no one is left behind. A variety of examples of this turning bad exist in fiction, a particularly pertinent one to me is Asimov's "The End of Eternity", and another one is the post-immortality centuries of Niven's "Known Space", but there are many others. Psychologically theme parks are meant to be visited for a break. If we spend our entire lives in them it does stuff to our capacities.
In everyday life animal species spend most of their problem solving effort on dealing with other members of their species (and some similar species). We are our own greatest competitors. In part, I think animals do this because members of each species are approximately on equal footing, so our individual problem solving is sufficient to deal with other individuals, and our group problem solving sufficient to deal with other groups of us. But nature itself, and the universe itself, is the biggest player. The biggest risk.
I fear that themepark life, as opposed to themepark visit, will either spark ennui and a bunch of interpersonal shenanigans, or will invite people to forget that we're just living in a fragile bubble. If the opportunities are not equal for every person, the situation is even worse. And I like to think that we have a responsibility to our co-inhabitants in the universe, the other species. A themepark life ultimately risks navel-gazing (at the individual or group level) and ignoring these responsibilities as well.
No thank you. I want a nice life, and a nice place to live it. But I don't want ennui. I don't want self-involved navel-gazing. I want people paying attention to the real issues in the universe. We spend enough effort already on dealing with other humans.
> What is it about the real world that is more important, inherently, than "the experiences that arise from living in it"? If a "better" (I'm hand-waving the complexity of comparison, because it's assumed as part of the discussion) set of experiences can arise - with perfect certainty, no trade-offs, no utilitarian cheats, just "everything is better for everyone" - then, as the other commenter said, it would be abhorrent to deny it 'because of some philosophical quibble about a difference between “simulation” vs “reality”'.
Sure. I agree. As long as we take pains to address the inevitable externalities before they impact others (including other species) this is fine. But make sure you know what's "better" for everyone before it is implemented. It's not obvious that even a god could manage this, much less humans (or even the humans/angels/demons of The Good Place).
"What is it about the real world that is more important, inherently" - It's everything, and we're just a part of it. "importance" is a value judgement, so is inextricably caught up in the individual so judging, but subtract out the value judgement and it's obvious that a human and a squirrel are more "important" than a single human. Now expand this to the "real world" entire.
Is there something special about Clojure making that change safe?
I prefer sum types.
But this is a legitimate point for union types.
Assuming implicit conversion for union types, you could either widen argument or narrow return without changing that implementation, but not both. A different implementation may not be able to handle either modification, though.
If the function started to provide different results for the same arguments, that’s different from what he’s talking about and would be a breaking change.
Worrying about the internals of the function is a violation of encapsulation. We care what we provide and what we receive back.
Observable to new callers who might want to pass None, but not to the existing callers who don't pass None as it is because it was not possible prior to the change.