The Unison language – a new approach to Distributed programming
unison-lang.org
unison-lang.org
They already have a very innovative way of managing source code, with a database of definitions that keeps the hash of the syntax tree instead of actual source. That's a very neat idea that solves many problems (read their docs to understand why).
But instead of developing that well enough so that it works with source control tools, IDEs, can be deployed easily and painlessly on existing infrastructure... no, they decided to ALSO solve distributed computing, a really, really complex space with a pretty crowded space of solutions... and seem to be focusing on that now instead of the "original" ideas. Looks like a huge issue with scope creep to me... unless they are kind of pivoting to distributed computing now only because the original ideas were not attractive enough for people to embrace it, but I have not heard of anything like that, everyone seems to be pretty vibed by those things.
I think the distributed computing problem is pretty related once you have "content-addressable" source code. Agreed that it's a lot of work but I hope it pans out!
Also I don't think that you need to create a new language to have 'content addressable' source code distribution..
Creating yet another language ensure that this will get nowhere, too bad.
There are a few things about unison that make this easier:
* content addressed code * unison can serialize any closure
I can ask for a serialized version of any closure, and get back something that is portable to another runtime. So in a function, I can create a lambda that closes over some local variables in my function, ask the runtime for the code for this closure and send it to a remote node for execution. It won't be a serialized version of my entire program and all of its dependencies, it will be a hash of a tree, which is a tree of other hashes.
The remote node can inspect the tree and ask for or gossip to find out the definitions for whatever hashes in that tree it doesn't already know about. It can inspect the node for any forbidden hashes (for example, you can'd do arbitrary IO), then the remote node can evaluate the closure, and return a result.
In other langauges, perhaps you can dynmically ship code around to be dynamically exeecuted, but you aren't going to also get "and of course ship all the transitive dependencies of this closure as needed" as easily as we are able to.
Since everything is made up from car and cdr, it's easy going from there. There is no difference between running locally, or anywhere. Just look up the data by request or gossip, as you said.
(For performance reasons, one might want to let a cons which represents "source" smaller than the size of a hash, be a literal representation instead of hashed. No need to do a lookup from a hash when you can have the code/data in the cons itself. Analoguous, don't zip a file when the resulting zip would be larger.)
Lets say you wrote a imaginary program to sum a column in a csv:
(defun my-program () (let* ((raw-data) (s3-load-file "htpps://...")) ((parsed) (csv-parse raw-data)) ((column1) (csv-column parsed 1)) (mean column1))
In this pretend program we are using some 3rd party s3 library to fetch some data, some other 3rd party CSV library to extract the data. Now I want to run this on some other remote node. I know I can just sent that sexp to the remote node and have it eval it, but that is only going to work if the right versions of the S3 and CSV library are already in that runtime.
I want to be able to write a function like:
(defun remote-run (prog) (....))
That can take ANY program and ship it off to some remote node and have that remote node calculate the result and ship the answer back. I don't know of some way in lisp you could ask the runtime to give you a program which includes your calculation and the s3 functions and the csv functions.
In unison, I can just say `(seialize my-program)` and get back a byte array which represents exactly the functions we need to evaluate this closure. When the remote site tries to load those bytes into their runtime, it will either succeed or fail with "more information needed" with a list of additional hashes we need code for, and the two runtimes can recurse through this until the remote side has everything needed to load that value into the runtime so that it can be evaluated to a result.
Then, of course, this opens us up to the ability for us to be smart in places and say "before you try running this, check the cache to see if anyone ran this closure recently and already knows the answer"
The car/cdr nature of Lisp makes it almost uniquely suited to distributed runtime IMHO.
As for your last sentence, this touches on compilation. Solve this, and you solve also dependency detection and compilation. I feel there are so many things which could/should converge at some point in the future. IDE / version control, distributed compute, storage.
An AST (or source) program could have a hash, which is corresponding in a cache to a compiled or JIT-ed version of that compilation unit, and various eval results of that unit etc etc. So many things collapse into one once your started to treat everything as key/value.
What has happened before, consistently, is that research or proof of concept- style languages pave the way for bigger players to take the ideas and incorporate them into existing or future mainstream languages.
I love that metaphor btw. Gonna use that.
They have VC funding and employees. I presume he's told investors that it will see at least fairly wide adoption!
There are a bunch of cool implications for distributed computing, namely that you can easily distribute fine grained parts of your application across servers and that you can cache the result of expensive calculations.
The first time the language or compiler changes such that the same code generates a different syntax tree they'd have to do something pretty fancy to avoid rebuilding the world. (That, plus all the usual caveats about what happens when old hash algorithms meet malicious actors from the future.)
I don't know what the social and legal issues might possibly be, though I might be missing somehting, what do you have in mind there?
hard pass. but thanks for the suggestion. I'll start with the home page and the examples and most likely end there.
It makes it simple to transparently ship the code (not just data) around to any worker node. The point of Unison Cloud is to disappear the difference between AWS EC2 and Lambda.
That said, it came out around the same time CoffeeScript did, and nobody uses CoffeeScript anymore either. So probably the fate was inevitable regardless of whether the framework was included in the language or not.
The one thing identified as “the big idea” of Unison seems to me to be conceived as a solution to a distributed computing problem that incidentally also solves a number of problems that are issues outside of distributed computing (but also within distributed computing, such that fleshing out how it can solve them enhances unison as distributed computing solutions as well as providing side benefits.)
The Hello World example introduces one of their concepts that you might not see every day, then they show a little algorithm, then a practical "stuff you need to get work done" example.
I got an immediate sense that the language has some familiar "ML family" type features (like F# or Scala) but also some distinctive aspects.
I plan to use it for a slightly different purpose. I want to implement an interpreter that is multithreaded similar to Java or Erlang that can send objects between shared memory without marshalling or copying.
I had a talk with someone on HN https://news.ycombinator.com/item?id=32907523 about python's Global Interpreter Lock and we talked about how objects are marshalled between subinterpreters due to object identity. The identity of an object is defined at creation time as the hash of that object.
If the hash of the object was the sourcecode, two interpreters could load the same Object hierarchy and send data by hash reference.
I wrote a multithreaded interpreter that uses message passing to send integers and program counters to jump to code in other threads
This is at https://GitHub.com/samsquire/multiversion-concurrency-contro...
- a year ago https://news.ycombinator.com/item?id=27652677
- 8 years ago https://news.ycombinator.com/item?id=9512955
Unison Programming Language - https://news.ycombinator.com/item?id=27652677 - June 2021 (131 comments)
Unison: A Content-Addressable Programming Language - https://news.ycombinator.com/item?id=22156370 - Jan 2020 (12 comments)
The Unison language - https://news.ycombinator.com/item?id=22009912 - Jan 2020 (141 comments)
Unison – A statically-typed purely functional language - https://news.ycombinator.com/item?id=20807997 - Aug 2019 (25 comments)
Unison Language March Update - https://news.ycombinator.com/item?id=19528189 - March 2019 (1 comment)
Unison: a next-generation programming platform - https://news.ycombinator.com/item?id=9512955 - May 2015 (128 comments)
Unison: a next-generation programming platform - https://news.ycombinator.com/item?id=9512955 - May 2015 (128 comments)
- Dependency management handled the same way the Nix handles it.
- Some kind of object storage system that uses content-addressable structures as the schema.
- Hyperlinked codebase.
- Human-readable function names as (essentially) git tags.
These ideas are all pretty nice. A dedicated IDE for this language would be a lot of fun to work with. The debugging story likewise seems like it will be pretty solid. I'm not sold on the zero config storage layer: basic object retrieval is different than schema prepared for query performance. I'd like to learn more about the concurrency and synchronization story.
Naming and resolvers, in order to be human friendly. This isn't easy to get right, but we have a lot of prior art in dependency management systems.
Persistence layers with GC, like cache and db. You're gonna want to fetch and prefetch in ways that are quite advanced. You don't want to be be blocked in a critical section by network fetching the leftpad function.
What does having a separate language give us as opposed to taking say, Kotlin, and having a compiler that stores the AST/whatever in a database and does all the interesting goodies?
Or is that the end goal, but we're using a basic language to test it out and work out all the quirks before writing a compiler for existing languages?
I remember many years ago, i had the same idea on how to correctly version a dependency.
And git commits "are" (versions of) programs.
So what does Unison have that git does not?
I can't think of any cases where it would make a difference in which order they were handled, however. Can you?
I think perhaps it might in the case where abilities themselves were able to make requests of other abilities, but that's not something allowed by our type system currently
Presumably
someAction : '{Choose, Abort} a
if it's handled by `Choose.toList` and then `Abort.toOptional` in that order you end up with `Optional [a]` whereas if you do in the other order you have `[Optional a]` right?N.B. the reason this is theoretically important is that `someAction` may be written with the assumption of e.g. certain short-circuiting behavior in mind and the "wrong" order of handlers might cause different short-circuiting behavior. In other words there's no consistent semantic interpretation you can assign to `someAction` even if you establish certain invariants that your abilities and ability handlers individually satisfy, since the global configuration of your ability handler changes what `someAction` means.
I think the jury is still out on whether this is a practical issue for any language that doesn't try to focus too hard on code having formal semantics (which is most real-world languages). I can definitely craft "real-looking" code that would be buggy depending on the order of handlers, but I'm not personally sure how much of a problem that actually would be for people familiar with the issue.
> if it's handled by `Choose.toList` and then `Abort.toOptional` in that order you end up with `Optional [a]` whereas if you do in the other order you have `[Optional a]` right?
I mean, not necessarily. Something that handles a `'{Abort} a` doesn't necessarily produce an `Optional a`. It could produce a Boolean, it could produce an Int, whatever. This is really up to the handler. But I still don't see how you produce a "bug" because you can dispatch the requests to abilities in either order. You couldn't, for example, ever produce a (well-typed) situation where a call to abort doesn't abort, or a call to Choose.toList would fail to produce a list. (Perhaps if toList were allowed to also use {Abort} but it is not)
I think if I pass something that has e.g. `{Abort, Exception, IO}` to `bracket` and then handle Exception before Abort, my Abort handler can break out of `bracket` before the finalizer action can run.
More generally I must always process any ability with a "bracket"-like function last to prevent this from happening right? And if I have multiple abilities that all have "bracket"-like functions they can step on each other's toes?
Even more generally I think any sort of "scoped" function in an ability has this problem.
This theoretically seems scary (imagine you have some big complicated action that does some bracket deep under the covers to e.g. release file handles; if I handle abilities in the wrong order the file handles might not ever be released, even if I have individually reasonable ability-handler pairs that locally don't do anything silly), but I'm not sure practically how often this comes up, and how much "just always handle Exception last" fixes that (i.e. how unlikely it is for any other ability to have a bracket function).
So if you want to be really sure of resource cleanup, you should use something like the `Resource` ability to acquire resources:
https://share.unison-lang.org/@runarorama/code/latest/namesp...
Also I would suggest removing Exception.bracket from base then, or at least changing the file handle example for Exception.bracket, since it seems a tad dangerous.
Just like VisualAge in the last century. Sign me up, that was great stuff.
https://www.unison-lang.org/blog/jit-announce/
We are expecting to be a monumental speedup for us. We have some promising results so far, but we haven't yet ported all of the runtime.
That said, it wouldn't be as bad as it sounds. Unison is a purely functional language, so if you don't explicitly provide the ability to e.g. do arbitrary I/O, then other nodes will not be able to send you code that does I/O. It will not type-check.
So in our cloud runtime, we blacklist EVERY IO function, then we can in our cloud runtime give you back the ability to do, for example, http requests, but not any other network IO. We won't let you open arbitrary files, but we'll provide ephemeral / persistent block storage through some other runtime ability.
This could also be used to do something like blacklist functions with known security vulnerabilities to catch people that aren't applying their patches!
There is probably an opportunity here to do interesting work around authenticating and validating computations from remote clients.
A word of caution -- javascript engines routinely get hacked, and once attackers can execute native code in the engine process they can call syscalls and have all of the rights of the underlying process. It may be useful to have some form of sandboxing of the language runtime for clients exposed to the internet or intranet. Additionally, Java & Ruby web servers routinely suffer from code injection when deserializing objects, which it seems like your language may be prone to as well.
I downloaded and ran it, it created folders NOT where I told it to, and started.... but no command I type seems to result in anything other than an error message.
How can I get it to add 5+5, without using an external editor?
https://www.unison-lang.org/learn/quickstart/
If this quick start tutorial doesn't work for you then there's a bug on unison's end.
Plus you're significantly limited in how much you can memoize if you can have different versions of the "same" function.
Something very much like Spark map-reduce can be implemented in ~100 lines of Unison code:
https://www.unison-lang.org/articles/distributed-datasets/
Some videos on Unison's capabilities over and above Spark:
Distributed programming overview: https://www.youtube.com/watch?v=ZhoxQGzFhV8
Collaborative data structures (CRDTs): https://www.youtube.com/watch?v=xc4V2WhGMy4
Distributed data types: https://www.youtube.com/watch?v=rOO2gtkoZ3M
Distributed global optimization with genetic algorithms: https://www.youtube.com/watch?v=qNShVqSbQJM
- pandoc: convert (almost) any document format to (almost) any other document format
- The Elm Compiler: Possibly the most widely used statically typed pure frontend language
- XMonad: A tiling window manager for linux that I enjoyed using even without knowing much Haskell
- Purescript: the other widely-used (for an FP language) statically typed pure compile-to-javascript language
Anyway, you'll very rarely see completely novel ideas in this space. Good combinations, compromises and applications matters a lot. Look at your favorite applied Merkle tree tool.
An understatement — I’d say when it comes to language design, this is pretty much the whole game right here.
So I guess its different from apache beam in all of the ways.