Unison: A Content-Addressable Programming Language
unisonweb.org
unisonweb.org
But I'm already completely convinced that every language should have these features! Especially languages used for the web.
I can think of: avoid binary code duplication - cause everytime you see a hash you've already come across, the compiler can jump to the already defined code. But that sounds like a lot of jumping around.
The website says "it eliminates builds and most dependency conflicts, allows for easy dynamic deployment of code, typed durable storage, and lots more." but I don't understand this.
If your code says "I depend on that hash", then the runtime needs to locate where the binary code that corresponds to that hash is located. And that's a dependency problem to resolve.
If someone fixes a bug in a dependency, your program may not be able to locate the hash anymore. You have to "re-build" your hashes everytime a dependency changes.
Can someone write the benefits more clearly?
It's not a dependency conflict though.
> If someone fixes a bug in a dependency, your program may not be able to locate the hash anymore. You have to "re-build" your hashes everytime a dependency changes.
Again, that's not a conflict. A conflict goes like this: Dependency A has a breaking change, but Dependency B transitively depends Dependency A as well, so you cannot update your own code until Dependency B also updates. Even if A and B are updated, you are prevented from adding any dependency that hasn't updated yet. You can't mix and match to use the old code in one place when you need it.
This wouldn't be such a problem if programmers didn't break interfaces for dumb reasons all the time, but they do, so lots of people just run older versions of the software.
My impression is, that all terms that used a wrongly implemented term also ned to have a new definition appended.
The hash of the main function describes an unambiguous ast, and also needs to be updated.
All this initially seems tedious, but I actually think it makes sense given that it is based in a strongly typed functional paradigm.
I definitely need to play around with this.
Dependency conflicts usually occur when the set of dependencies the original author wrote his code against are not the exact set of dependencies which are on your machine, which can result in build errors, or even runtime errors when code changes behavior. There are already projects which aim to tackle this - Nix and Guix - if you alter a codebase in any way, it results in a different hash for the entire package, and all dependants of the package need to update the hash they refer to in order to reference the alterated codebase. But importantly, they can still reference the old one explicitly by hash! This is important because it allows unlimited number of versions of the same codebase to exist on the same machine without any chance of the system assuming the wrong dependency based on its name and version number.
Unison takes the concept of Nix further and instead of just giving each package a unique identifier, it gives one to every semantic unit in a codebase. You will be able to directly refer to a function, rather than have to import an entire codebase.
Deployment can be done in a similar way to Nix too. Have a look at `nix-copy-closure` for example. It bundles together a package with its entire dependency tree ready to send over the wire. You no longer encounter the case where you download some code, and before you build it, it errors "you must have abclib-1.2.3". Also, it is possible to optimize the amount of data transferred with set reconciliation. You work out which dependencies a remote host already has and then just package up the missing dependencies.
Unison can compact this process even further, because instead of downloading entire packages, you only need to download the individual functions you need. The hash of a unison expression becomes not only its identifier, but an argument to some function which will be able to reproduce an exact, minimal copy of the program with its complete dependencies.
> If your code says "I depend on that hash", then the runtime needs to locate where the binary code that corresponds to that hash is located. And that's a dependency problem to resolve.
This is trivial. The important part is that hash collisions are essentially impossible, so every dependency has a unique entry. The expression corresponding to a hash can be looked up in a database. These lookups can be optimised with b-trees and such, and can also make use of bloom or cuckoo filters for quick set membership testing. There could be scalability issues if the databases start containing many billions of hashes.
> If someone fixes a bug in a dependency, your program may not be able to locate the hash any more. You have to "re-build" your hashes everytime a dependency changes.
All existing hashes are kept, so your program will continue to work as it always has, but it will not automatically inherit the bug-fix. This is the trade-off it makes versus existing systems, which have shared dependencies, but also dependency-hell.
A rebuild of the codebase can take the latest versions of each identifier and build with complete bug-fixes. This should also mean you don't have to wait around for individual package maintainers to update the dependencies of their projects (which can take several years), but you should be able to incorporate any change by rebuilding all of the dependants of that change which are contained in your local database.
Also, the system can incorporate garbage collection, where any hashes which are no longer referenced by any code can be removed from the local storage.
Unison is a functional language that treats a codebase as an content addressable database[2] where every ‘content’ is an definition. In Unison, the ‘codebase’ is a somewhat abstract concept (unlike other languages where a codebase is a set of files) where you can inject definitions, somewhat similar to a Lisp image.
One can think of a program as a graph where every node is a definition and a definition’s content can refer to other definitions. Unison content-addresses each node and aliases the address to a human-readable name.
This means you can replace a name with another definition, and since Unison knows the node a human-readable name is aliased to, you can exactly find every name’s use and replace them to another node. In practice I think this means very easy refactoring unlike today’s programming languages where it’s hard to find every use of an identifier.
I’m not sure how this can benefit in practical ways, but the concept itself is pretty interesting to see. I would like to see a better way to share a Unison codebase though, as it currently is only shareable in a format that resembles a .git folder (as git also is another CAS).
[0]: https://news.ycombinator.com/item?id=22010510
[1]: https://www.unisonweb.org/docs/tour
[2]: https://en.wikipedia.org/wiki/Content-addressable_storage
https://www.unisonweb.org/docs/tour/#names-are-stored-separa...
lost me right here. fetishizing software "craftsmanship" isn't going to make the software run better. It might make it more maintainable. But even then it's better to have a well-designed, efficient system with poorly crafted components than artisanal for-loops