The way we're thinking about breaking changes
welltypedwitch.bearblog.dev
welltypedwitch.bearblog.dev
The elephant in the room is changing business requirements. If the change requires the code to consume new parameter without reasonable defaults, no automatic migration would solve it reliably. Let’s say you deal with the money and you introduce currency field. The compiler of the client code cannot just assume the currency, neither it can take it from anywhere by just applying the migration. It is a manual operation to fix the broken code.
I think part of our hang-up is that we tend to think of code and data in very separate categories (except for literals, of course)--and only allow treating code as data after a lexer starts processing it (or via reflection or first-class functions at run time).
Correcting this category error and adding metadata (not just types) could yield all sorts of interesting ideas.
Like - as TA suggests, why not encode the notion of how a function signature has changed over time to enable automatic migrations?
Let’s say you have this change:
BigDecimal doSomething();
to Money doSomething();
In theory metadata can capture the migration from a plain number to Money#amount() and compiler will generate a facade for the client code. But simply discarding the currency will have disastrous effects at some later point. Even relatively simple migration String->Number may go wrong without semantics. Language feature allowing such metadata and migrations will be a minefield.That doesn't mean migrations in general aren't a good idea.
And it doesn't invalidate the notion of metadata for code in general and all the possibilities that opens up.
The node approach is very fine grained libraries, but that makes its own headaches
Really every function or interface invocation has an implicit version number. It should be explicit in the code, and that assumes major minor breaking change versioning.
Second, database migrations are notoriously tricky for developers to manage correctly. They often require significant domain knowledge to avoid breaking assumptions further down the line. Applying that paradigm directly to compiler-managed code changes feels like it would amplify the same problems—especially in languages that rely on strong type inference. The slightest mismatch in inferred types could ripple through a large codebase in ways that are far from obvious.
While it’s an interesting idea in theory, I think the “fix your old code with a macro-like script” approach just shifts maintenance costs elsewhere. We’d still be chasing edge cases, except now they’re tucked away behind code-generation layers and elaborate type transformations. It may reduce immediate breakage, but at the expense of clarity and predictable behavior in the long run.
I was assuming that the migration would actually alter the call site in a way that is reviewable and committable, not implicitly do it every time it's needed at compile time. If that's the idea, it doesn't alter anything behind the scenes, it does so explicitly and visibly one time to change everything to meet the new requirements.
The larger problem I see is that the applications for this would be extremely limited. There's a reason why they used a simple s/a/Some a/ as an illustration—once you get beyond that the migrations will become a pain to write, a pain to validate, and likely to break tons of code in subtle ways. And since most library code changes aren't this simple kind, it's likely that it's not worth the effort of building and maintaining this migration syntax for something with so narrow an application.
In code, if you want to handle a breaking change transparently, you don't need the moral equivalent of a linker fix-up table. You need to keep the old definition around and redirect it to the new function. This can be as simple as having, say, the old version of your code wrap something in a list or closure and then call the new version. In languages with overloading, this can be done transparently; in others you'd have to give the new function a new name. Maybe API versioned symbols and a syntax for them is what you would want?
[0] The original idea with SQL is that multiple applications would store data in the same place in a common format you could meaningfully interchange.
The example the author uses is modifying a function to return an int or null instead of an int. Let's say you implemented the function a bit naively and in the null case your function would crash the program. Now you are going to refactor your codebase so the caller gets to decide what happens in the null case. Some callers may be unable to handle the situation and will need to implement the crash/exception some may use some sort of fallback behavior.
I think haskell's type system probably would probably prevent someone from having the issue I described above but the problem still stands. Let's say you had a function that returns a union type and you have areas in the code that are supposed to handle every case of that type. If you had a case you need to handle it everywhere and the compiler can't know how you want to handle it. A lot of type systems will catch the missing case which is awesome but you still need to handle each one.
This assumes that you're approaching the breaking change problem from the perspective of immutable names and deprecations: don't make breaking changes to existing names, create a new name for the new behavior and communicate that the old behavior is no longer going to be supported.
[0] https://kotlinlang.org/api/core/kotlin-stdlib/kotlin/-replac...
The proposal essentially expects the dependency to provide a migration path. However, to the extent described this is only necessary in languages with no overloading. And it doesn't help for cases when the change is not as simple as providing a few rewrite rules or things like that.
Also, the jab about static type systems is misplaced. Their purpose is specifically to reject certain invalid programs, at the cost of rejecting a few valid. You can't have one without the other. With dynamic typing, backwards compatibility breakages fly under the radar until they gleefully blow up in production.
If writing such a function is not possible (because there is no automated way to convert from the old function call style to the new) then a database-style automated migration isn't going to be possible either.
I agree that attitudes towards security are generally very poor, but breaking working infrastructure sounds like a crazy practice. Like any sensible system, a good/robust design should allow staged upgrades / hot reloading for anything but a very tiny core of critical functionality. Erlang/BEAM is a great example; it just requires software engineering to adopt a different mindset.
Yes, breaking infrastructure is bad. But letting already broken infrastructure continue can be worse.
The point is that we want a better way to detect when breaking changes happen so that security fixes can be applied without breaking anything, while permitting optional upgrades on our own schedule for other features. There doesn't seem to be a great solution yet, it's either "it never breaks but you're possibly vulnerable to security issues that can't be easily patched", or "things can break at any time due to updates so we have to manually verify this doesn't happen".
I don't think it's fundamentally incompatible with a secure design though, you just need to reify the authority to do those things so you can explicitly grant them to specific programs as appropriate.
The most used desktop OS on the planet has allowed this since forever and the world hasn't ended
It's not a good model, and that's why this only forced in commercial software or in particularly obnoxious projects, like earlier versions of Ubuntu Snap. Every other case is user's choice - package managers have lock files; automated updates can be disabled; docker images can be referenced by SHA; etc...
That's not to say that infrastructure does not break - there plenty of horrible setups out there... but if you discover you "must drop every other priority and upgrade", then maybe spend some time making your infra more stable? Commit that lockfile (or start saving dev docker containers, if you can't), stop auto-deploying latest changes and make sure you keep previous builds around, instead of blaming software ecosystem and upstream authors.
- Imagine that in my project, I have a "FOO" function, which is called by many others
- I've decide to change type of one parameter of FOO function. This would be a breaking change in regular language, but in Unison, nothing breaks - I push the new definition, but every caller is still using old version.
- New callers come up, and they use new version. So far so good.
- Some times later, I've discovered a critical business-logic bug in FOO function! So I fix it, and I have to update all the callers to use the latest version... except for half of them I can not, because the parameter types do not match. Seems like I cannot ship the fix to the customer until I spend a bunch of time rewriting the existing code to accommodate argument type change.
As long as there are functions, there are always some kinds changes to them that require one to fix up the callers. How this is enforced can be different - in strictly typed languages, code may fail to compile; in dynamic languages, you may see runtime failures; and in Unison, things will work until you try to edit the caller, at which case it'll fail to compile (unison docs call that operation, converting from text to internal language's representation, "typecheck").
I am not convinced that the "postpone failure until you edit the caller" is the best approach here. When I refactor, I normally want to see any problems surface right away, while I still have the context for the change.
none of the above as always true.
I think most of all one must have discipline to thoroughly think of old code and how it should behave now. I found that one can already do this by versioning code and types, similar to how an API endpoint might go from /v2/… to /v3/…. But again, it requires discipline and it’s not something I’d do for anything but very critical code.
There are some quirks you have to work around for big projects since the free tooling has some limitations.
—and—
From the hip: I think this is encompassed by dependent types.
The stuff that just falls naturally out of dependent types includes tools that grasp… this. And other concerns.
If I put it like so?:
Software is automation.
Dependent types is the automation of the automation.