Why Do ML on the Erlang VM?
underjord.io
underjord.io
Not that this article makes a good argument for the benefit of any of those things for ML. But at least generally while I am a pretty strong believer of "eh language doesn't really matter" erlang-based languages I do entertain a possible exception for.
Agree wholeheartedly. Elixir/BEAM/Erlang are a league of their own.
That's flat out false. Only the syntax is taken from Ruby, that's it. Even the added metaprogramming isn't anything like Ruby's. It otherwise just like Erlang.
Other than that if you look at how the language works under the surface it is completely different from Ruby which uses Runtime Reflection to add methods to objects in memory.
(Elixir is a little more special with its hygienic macros, whi allow for for libs like NX to exist.)
The BEAM/erlangVM is probably closer to a OS than a traditional-ish VM like the JVM.
Another way to say this is: "no other language will ever be better than X", which is clearly not true.
Eventually something better will arrive. Maybe that's Elixir, maybe not, but is sure won't be Python forever.
It will be Python for all of our lifetimes here.
Is there a concrete reason for that? A lot of the approach in Elixir is to not go the route of other languages and merely slap an adapter interface on Python libraries. The approach is to approach things from a very Elixir and BEAM specific point of view. This is how Livebook was developed and how the machine learning libraries are being developed.
In fact, one could view it as Python being the one that's going to struggle to overcome the limitations set forth by its own language design. Elixir comes with very tightly integrated tooling (projects, testing, type analysis, static analysis, notebooks, scripting, etc.) that basically has 100% adoption rate without a plethora of third-party competitors and lack of integration with the language such as Python has. With immutability and concurrency built into the language, Elixir has a leg up, and there is active research into bringing static typing to Elixir and also binding Elixir to faster runtime environments, such as Rust.
Python simply cannot make the jump to Elixir, Erlang, and BEAM's language and VM capabilities like they can to a machine learning ecosystem akin to Python's.
These two are so terribly, laughably bad in Elixir that you listing them as advantages makes me doubt the integrity of your entire post.
I hope you can tell me.
Dialyzer and Credo for Phoenix does require some workarounds, as that project isn't interested in providing typespecs or base compliance with Credo and is very heavy on macros.
https://news.ycombinator.com/item?id=36282419
https://elixir-lang.org/blog/2022/10/05/my-future-with-elixi...
This strongly suggests that Dialyxir is not working as effectively for most users as e.g. Typescript or Mypy.
† I can imagine people quibbling with 'entirely different', but the proposed type system does have a new(ish) theoretical basis that's still the subject of active research.
Static analysis in Elixir is quite more accessible because you rely way less frequently on dynamic dispatch than Python (or Ruby or JS). You typically know which module you are calling to, structs and pattern matching tell you about fields (which we verify at compile-time) and primitive type information. Things like deprecated and undefined functions are part of the compiler, while most other dynamic languages typically require type systems or linters (or do it exclusively at runtime).
I agree on type analysis though. It is better than nothing but that's not much to talk about. We are working on it. :)
I’m excited to hear that Elixir may be getting a bit closer to that, and I agree with your observation that because typical Elixir code uses dynamic dispatch very little, there’s a lot of room for improvement in ways that eg Ruby will never be able to offer.
I’m sorry about the “terribly, laughably” bit of my comment.
For what is worth, I was not considering Dialyzer as part of my reply. You should be getting many warnings related to typos from the compiler!
Anyway I agree with the sentiment: we have a high ceiling but we are currently far away from it. If we had a Dialyzer that runs all the time, is fast (Erlang/OTP 26 already improved here), and has good error messages, it would already be great.
We are currently researching a proper type system into the language. Fully integrating and improving the Dialyzer experience would be our plan B.
I mean when 2 NodeJS processes communicate via JSON over HTTP then those aren’t typechecked either, even if their codebases are all TypeScript. Except, of course, when using the same TS rpc lib on both sides of an interface, with some code sharing between them so they use the same types.
I feel like even when all typechecking happens only on regular static synchronous function calls, and message passing is left out of scope, it’ll be so extremely useful. And like you said, the low amounts of dynamic dispatch makes it tractable. Still a huge amount of work I bet though :-)
Yes, Dialyzer could be better, but it's pretty good if you write typespecs. But Credo is quite good and one of the best such tools I have ever used. And Elixir g enerally has excellent error messages. Why do you say it's laughably bad?
Also, the point was primarily that all these tools are effectively built-in to the language without several different competing libraries.
And I neglected to mention the official formatter.
People will build huge C and C++ libraries for realizing distributed systems that they can then use from Python, instead of getting to know a language on the Beam VM, because they only want to stay in the mainstream ecosystems or are unwilling to deal with new concepts like in Elixir or Erlang (for starters lets say lightweight processes, distributed system, pattern matching, proper recursion (TCO), the functional programming view of things. Some of those are very scary for many Python developers.).
Fortunately there are those on the other side as well, stubborn enough to try to implement things in those more niche languages that they prefer, based on technological merit of basic principles and concepts of a language, rather than just doing what everyone else does. So sometimes something really great and interesting comes out of that.
Often however, the shoehorning masses are simply too many, the people knowing the other ecosystems too few, to keep up with them or the newest developments, further enabling the "but this is what everyone else is doing!" kind of mindset. Something something "beating the averages" here.
Or want to have a large easily accessible pool of developers to pull from. There are real, practical considerations here more than just feelz.
Not everyone is always looking for the next thing, better or not.
I would absokutely write the post "Why I prefer Elixir" and I think I have. But this was not intended to be that.
And while Elixir may be as well suited as Python to being the interface/scripting language for ML tasks, Python has a pretty unassailable lead here with an overwhelmingly dominant ecosystem.
Elixir could be used to build and manage pipelines (like a roll-your-own Jenkins) in this area, and in this regard (concurrency, reliability, async) it is probably a significantly better language than Python, but the actual tasks on the pipeline will need to be defined in Python, just because it's use is so entrenched.
Also, the "reinvent something another language can already do" pitch is going to become a much harder sell in today's funding market as the silly money dries up.
This was in early 2017, and my search essentially went from a variety of JS frameworks to RoR to Scala + Play and then eventually to Elixir + Phoenix. The ecosystem was much smaller then but of the options that were viable, it was by far the most productive.
* http://neuralnetworksanddeeplearning.com/
I decided to try to implement everything from scratch in Elixir (after initially doing all the math with pen and paper on a trivial example to get the feel of it). Obviously pure elixir was extremely slow, so I started creating NIFs to pass over matrix multiplication to OpenBLAS. Then I was thinking more and more of what things I can pass to C code and just have Elixir as a "frontend" for it. My enthusiasm died down when I realised I was simply implementing things in C with the pretext of "doing elixir", a nice learning experience but I could see I was not doing the things that initially got me pumped up.
Don't get me wrong, I loved the discovery part of it, reading research and trying to understand so I can implement the different new (at the time) deep learning techniques, like convolutions, LSTM, and the different nuances of it. I think it gave me a better understanding of how things work and why it works. But it deviated from the initial scope and I lost interest once the learning phase was over and I knew I could simply use tensorflow or pytoorch as I did not actually need the advantages BEAM offers for this type of workload.
Code is still available here:
* https://gitlab.com/sdwolfz/experimental/-/tree/master/exlear...
That Elixir project is nice, but I already have enough on my plate with JVM, CLR, V8.
Noted for follow-up though.
That is because "ML" is a proxy for an entire domain (data science) that has both very distinct requirements and patterns developed over decades and is now evolving very fast as it broke out of the whitecoat lab and (for good or bad) into our lives.
The ascent of python is as the article noted "accidental" but that doesnt mean it is suboptimal given all the requirements of the domain and the available alternatives (here and now).
Its natural to focus on an exceptional attribute (like concurrency) to motivate even having the discussion but this addresses only one of three critical requirements if you decide to do everything within one langauge.
1) Ability to experiment interactively (REPL and friends), expressivity, human oriented language design etc
2) Ability to build production scale models (training or estimation phase) utilizing heterogeneous and fiddly compute
3) Ability to deploy to diverse devices and footprint / latency requirements both natively and as part of a web application. NB: Its only here that the intrinsic qualities of erlang might put it above the rest
Historically there was never an ecosystem than could do all that optimally within just one language. Some of the requirements are arguably even incompatible.
The current preferred solution is not "python" but python/c++ which does some of that well (1, 2) and 3 only passably.
The only declared single language contender is julia but its very far from being a complete ecosystem in the above sense.
Another emerging possibility is mojo, as a superset of python. But that is not even available yet.
Back to the erlang/elixir combo. Its not even clear to me if it can address all of the above in a passable way.
But on the other hand ML is really shaking the game and what is mayne more imprortant is the vitality and resourcefulness of different communities to develop in the direction of the new rewuirement.
We have not seen the end of ecosystem evolution.
I like your breakdown of the requirements in three abilities. I think Elixir provides all three of them (with a caveat) and I will refer to two videos for those who want to learn more:
* This video shows how you can use Livebook (open-source computational notebook environment) to explore pre-trained neural networks (from Hugging Face), learn how they work, and then deploy them as part of a Phoenix web application: https://news.livebook.dev/announcing-bumblebee-gpt2-stable-d...
* This video shows how to scale a neural network serving, first to run it concurrently, and then have it run distributed over multiple CPUs and GPUs: https://news.livebook.dev/distributed2-machine-learning-note...
The caveat is about your second ability. The production scale models are, at the end of the day running native/cuda code, using the same backends as Python (Google XLA and LibTorch). What we do is that we compile your Elixir code into a graph which is then compiled to the CPU/GPU just-in time (in an approach similar to JAX). You can find more information here: https://github.com/elixir-nx/#why-elixir
What will it take? I am just an armchair philosopher in this respect. Like they say languages are only in part about communicating with the computer, they are about communicating with other people.
So these three major tasks are really about winning over three different types of people and their peculiar brain wirings, motivations and aspirations. 1) and 2) are more introvert: mathematical types that want to express their inner abstract worlds in computer instructions in the least painful way. The famous data point is that pytorch became the tool of choice of academic ML publishing over the first mover tensorflow. Then you have the true computer nerds that love making the most of hardware and see in ML a major new stage of computing: HPC becoming mass market. 3) is more extrovert (keeping in mind the norms of the tech industry :-) developing and scaling delightful products and apps to be used by non-technical people.
Erlang/elixir clearly offers something special in terms of 3) and that is important because in the end its all down to serving useful software to people.
My experience is that Python has so many high quality packages to deal with cpu intensive tasks that unless you absolutely cannot use them, for some reason, performance is rarely a problem. But I'd like to be informed about the exceptions.
One involved custom data augmentations on spatial graphs. I couldn't find any python libraries that offered performant implementations of the functions needed and ended up writing custom C++ extensions to call from python.
The other involved running video encoders and decoders inside the dataloader. The problem was more Pytorch specific than python per se, but the way Pytorch handles multiprocessing and CUDA streams made efficient parallelism very difficult. It would have been much easier in C++.
I've also needed to ditch python for ML inference in production when extremely low latency was required.
One thing we can blame Python for is that the data scientists can't ever loop over their data with a regular loop. Simply impossible; it's too slow and would immediately wreck performance. They're more Pandas programmers than they are Python programmers; everything has to be done with Pandas operations for performance to be acceptable.
This limits our (the computer programmers) ability to help out, because the languages we use are not slow like that, so we may give suggestions from our world that simply can't be used in Python because it's too slow. "Can't you just loop over the data and do XYZ?" Nope, they can't.
We still use Python because it's all the data scientists know, and it's easier/cheaper to just write a bigger check to AWS than to hire people with both data science and computer programming backgrounds.
I wrote the same thing in Julia in about the same number of lines of code -- avoiding the sins above -- and I could run it on my laptop in a few seconds.
Julia is great, but sometimes a bit less user friendly than Python, it's quite easy to shoot oneself in the foot performance wise using Julia.
Generally though… While Python might have many distinct popular/well-developed tools, they tend to compose very badly. IME, it feels relatively painful if you’re trying to do something new which hasn’t been solved and wrapped for you.
As for Julia footguns, it has been relatively trivial IME to get substantial improvements compared to Python. Whether that is the best one can do, I dunno.
The user experience of trying to set up one of the Python libraries for LLama vs llama.cpp is night and day.
I get how array languages such as APL suit ML, and distributed languages such as those which run on the BEAM (I am now looking at Gleam, an ML-like language that runs on the BEAM. I prefer the syntax to Elixir's the same way I tried LFE first) and seem to fit ML. Python's angle is that it utilizes C/Fortran libs like LAPACK and BLAS and hooks to do what it does in ML, and Elixir's Nx is similar in that sense to get the performance in number crunching that the BEAM languages are horrible at. You can use Zig now to write NIFs for Erlang/Elixir so maybe this is how Elixir moves in the direction of Python's approach to speeding up the number crunching for ML along with Nx.
I prefer to try and use a more suitable language rather than try to make it work like Python/Numba/Numpy/Pandas and Elixir/Nx do. Funny thing is that Pandas by Wes McKinney was inspired by the J programming language. So I think you may not be with the crowd in my approach, but there is some harder-to-reach fruit in this approach that is more rewarding to me, and sometimes for others.
(I really enjoyed this article though!)
As far as I know theres a few implementations of ML like languages on the Erlang VM
https://github.com/llaisdy/beam_languages
caramel and alpaca are worth checking out.
Gleam doesn't look like a ML lang but has a lot of the same semantics of a ML lang
Yes, downvotes. ML used to mean the language which would be a worthwhile thing to render on the Erlang VM.
I'd rather optimize the current julia web app stack a bit and use with existing julia ML libraries.
Can’t be the only one to dislike these languages. I don’t even care about languages that much, but I avoid AI stuff because I just can’t stand the python or elixir ecosystems. So much crap is required to get something basic working, although it’s less annoying with elixir.
Oh well
I just stick to Go for all my backend needs. It does everything elixir is capable of and provides more than just raw IO handling without blowing up memory or struggling with computations. Never got into ruby either the syntax just makes me lose interest.
So you are not the only one. Python is just good in tying CUDA libraries together while suffering through Python's syntax.