It's more or less as close to a middle-ground as I could imagine at the time.
2,658 karma · joined March 17, 2009
It's more or less as close to a middle-ground as I could imagine at the time.
That it _also_ ships with other ways of doing things in no way constrains or limits your decisions, and most modern Erlang (or Elixir) applications I have maintained ran the same way.
You still get message passing (to internal processes), supervision (with shared-nothing and/or immutability mechanisms that are essential to useful supervision and fault isolation), the ability to restart within the host, but also from systemd or whatever else.
None of these mechanisms are mutually exclusive so long as you build your application from the modern world rather than grabbing a book from 10-15 years ago explaining how to do things 10-15 years ago.
And you don't _need_ any of what Erlang provides, the same way you don't _need_ containers (or k8s), the same way you don't _need_ OpenTelemetry, the same way you don't _need_ an absolutely powerful type system (as Go will demonstrate). But they are nice, and they are useful, and they can be a bad fit to some problems as well.
Live deploys are one example of this. Most people never actually used the feature. Those who need it found ways (and I wrote one that fits in somewhat nicely with modern kubernetes deployments in https://ferd.ca/my-favorite-erlang-container.html) but in no way has anyone been forced to do it. In fact, the most common pattern is people wanting to eventually use that mechanism and finding out they had not structured their app properly to do it and needing to give it a facelift. Because it was never necessary nor totalizing.
Erlang isn't the only solution anymore, that's true, and it's one of the things that makes its adoption less of an obvious thing in many corners of the industry. But none of the new solutions in the 2023 reality are also mutually exclusive to Erlang. They're all available to Erlang as well, and to Elixir.
And while the type system is underpowered (and there are ongoing area of research there -- I think at least 3-4 competing type systems are being developed and experimented with right now), that the syntax remains what it is, I still strongly believe that what people copied from Erlang were the easy bits that provide the less benefit.
There is still nothing to this day, whether in Rust or Go or Java or Python or whatever, that lets you decompose and structure a system for its components to have the type of isolation they have, a clarity of dependency in blast radius and faults, nor the ability to introspect things at runtime interactively in production that Erlang (and by extension, languages like Elixir or Gleam) provide.
I've used them, I worked in them, and it doesn't compare on that front. Regardless of if Erlang is worth deploying your software in production for, the approach it has becomes as illuminating as the stacks that try and push concepts such as lack of side-effects and purity and what they let you transform in how you think about problems and their solutions.
That part hasn't been copied, and it's still relevant to this day in structuring robust systems.
Well there you go, I guess the pattern is equivalent but incidental.
It is therefore a bit more general than Elixir's 'with', and it would be interesting to see if the improvement could feed back into Elixir as well!
The initial inspiration for the 'maybe' expression was the monadic approach (Ok(T) | Error(T)) return types seen in Haskell and Rust, and the first EEP was closer to these by trying to mandate the usage of 'ok | {ok, T}' matches with implicit unwrapping.
For pragmatic reasons, we then changed the design to be closer to a general pattern matching, which forced the usage of 'else' clauses for safety reasons (which the EEP describes), and led us closer to Elixir's design, which I felt was inherently more risky in the first drafts (and therefore I now feel the Erlang design is riskier as well, albeit while being more idiomatic).
So while I did get inspiration from Elixir, and particularly its usage of the 'else' clause for safety reasons, it would possibly be reductionist to say that "the good ideas were stolen from Elixir." The good ideas were stolen from Elixir, but also from Rust, Haskell, OCaml, and various custom libraries, which have done a lot of interesting work in value-based error handling that shouldn't be papered over.
I still think these type-based approaches represent a significantly positive inspiration that we could ideally move closer to, if it were possible to magically transform existing code to match the stricter, cleaner, more composable patterns that they offer.
In the end I'm hoping the 'maybe' expression still provides significantly nicer experiences in setting up business logic conditions in everyday code for Erlang user, and it is of course impossible to deny that I got some of the form and design caveats from the work done in the Elixir language already :)
Also as a last caveat: I am not a member of the Erlang/OTP team. The design however was completed and refined with their participation (and they drove the final implementation whereas I did the proof of concept with Peer Stritzinger and wrote the initial EEP), but the stance expressed in my post here is mine and not the one of folks at Ericsson.
There is an objectively quantifiable disagreement. But its nature (and even whether it is desirable or not) are possibly camped in subjective terms. Of course you could argue that I am objectively wrong — though trying to prove that with my own writings is risky since we’ve established I’m not a trustworthy source — but that in itself does not resolve the overall disagreement from existing.
This sort of situation can also happen in software where an ambiguous specification yields two distinct compliant implementations that nevertheless do not work together.
Also blame and accountability and responsibility are all subtly different.
Here’s another one: should the incident actually have an impact on anyone of the companies employees yearly reviews? Either positive or negative? Why?
An open-source database is being used and operated as a service by a vendor, which a SaaS company relies on to provide a feature that your organization uses to manage data on behalf of users.
We now have a chain that includes: users <- customer organisation <- SaaS vendor <- DB as a service vendor <- OSS DB maintainers <- Linux maintainers <- Driver writers <- Hardware vendors.
There is suddenly a power outage at the DB as a service vendor (because of an unmaintained powerline falling over) and their UPCs appear not to be functional for yet unknown reason (cost cutting or supply chain issues during covid time may receive some blame). Your users lose their data regardless.
What is the bug? Who is at fault? Is it the engineer? The team who wrote the code? The QA folks? The organization that hired them? Who should fix the issue? Who should be on charge with repairing data corruption? Whose backups should be trusted most? The least? Have you been lenient in your usage of a SaaS vendor? Has the SaaS vendor been lenient in the services they use? Which actors can be considered liable from a legal standpoint? Which actors can be considered liable from an ethical or moral standpoint? Are your customers the one who made a bad decision contracting you? Can there be more than one party responsible? Do any of these answers changed based on whether the power loss is caused by an act of god or bad maintenance? Based on which jurisdiction you're in? How do you define honest mistakes? Negligence? A bug is a bug because the software did not meet the expectations that were set. Were the expectations reasonable? Who should have managed them?
Events happen. The meaning we attach to them is of course based on expectations and standards and the environment and context, but the way we build our explanation, the ways we attach blame and accountability varies. You can sometimes decide to assign accountability to individuals, sometimes to systems, sometimes both. Sometimes only some or sometimes neither.
So sure, you can point at the actual technical lines of code and say "these aren't doing what they should", but if you do this in a vacuum without also wondering who decided what these lines should be doing and what pressures were at play when they were written, are you necessarily learning a lot about how events unfold and how they might unfold in the future?
A systemic perspective will yield different reactions than one based on personal engineer responsibility, which will be different from one that looks at it from an insurer's point of view, which will be different from one which looks at it from an education point of view, etc. So the lens you take to look at the events surrounding the error and the interpretation you make of it are absolutely crucial to the corrections and learnings that follow.
The biggest gains to be had are obviously systemic, and what I do as a consumer is far more limited in its scope and impact. I still limit my flights, no longer attend in-person conferences, try to travel more local, because what else am I going to do? I'm aware this is like putting out a cigarette when the whole town's already on fire, but I can't deal with the dissonance otherwise. It's still an individual luxury that can have an oversized impact compared to everything else I do.
Advocating for it is not going to be sufficient at all, but it's still the most impact I can have when all the big stakeholders who have to fix their powergrid are not even in countries I live in.
So even though aviation is a small part of it all, it is one of the individual actions able to have outsized impact, usually for leisure or at least often for non-essential reasons.
However, we didn't spit on having a package manager (that wasn't a bad lazy index hosted on a github repo), and it became a very interesting bridge across communities that we don't regret working with. Our hope now is to try and make it possible to use more Elixir libraries from the Erlang side, but the two languages' build models make that difficult at times.
But you'll also get issues with staffing and getting people interested in working in mainstream languages that are less cool than they used to be, frameworks that are on older versions and hard to migrate, deployed through systems that aren't as nice on your resume than newer ones, or on platforms that aren't seen as positively.
I don't have a very clear answer to give about why Erlang specifically wasn't seen positively. The VP of Eng at the time (now the github CTO) saw Erlang very positively (https://twitter.com/jasoncwarner/status/1287383578435780608) but I know that some specific product people didn't like it, and so on. To some extent a lot of the work pushing us aside was just done by very eager Go developers who just started doing work on replacing our components with new ones on the other side of the org, and then propagating that elsewhere.
Whether the roadmap or other policies ended up kneecapping our team on purpose or accidentally is not something I can actually know. I kept pushing for years to improve things for our team, but at some point I got tired and left for a different sort of role.
Changing the instance means having to re-transfer all of the data and re-establish all of that state, on all of the nodes. You could easily see draining of connections take 15-20 minutes, and booting back and scaling up to be taking 15-20 minutes as well, if you can do it for _all_ the instances at once (which may not be a guarantee, and you could need to stagger things to be more cost-effective).
You start with each deploy taking easily over an hour. If you deploy 2-3 times a day and that your peak times line up with these, you can more than double your operating cost just to deploy, and that can take more than 4 figures to count.
Some of the systems we maintained (not those we necessarily live deployed to, but still required rolling restarts) required over 5,000 instances and could not just be doubled in size without running into limits in specific regions for given instance types.
If a blue/green deploy takes a couple minutes, you're probably not having a workload where this is worth thinking about that much.
For the first system, it was deprecated without replacement, and just let to run by managers and people who had moved on to other teams (but used to work on it) who did the minimal maintenance required, and former employees were given emergency contracts in weird circumstances to deal with things. Roughly 3-4 years after, they finally replaced it after 2-3 attempts at rewrites that had failed before. The old design with minimal maintenance for years finally approached limits to how it could be scale without bigger redesigns; I consider this to be extremely successful.
It wasn't exactly bulldozed nearly as much as declared "done" and abandoned without adequate replacements while major parts of the business was just being rewritten to use Go and a more "standard" stack. Obviously these migrations always start with something easier and by replacing components that are huge pain points to its contemporaries, and you're left with more legacy stuff in the end that is much harder to replace done in the final pushes. I felt that the blog post would have veered off point if the whole thing became about that, though.
The people on these teams left in part because the hiring budget was redirected towards hiring on the new projects. The idea was that everything could be done in Go for these stacks (by normalizing on tools and libraries developed in another project and wanting to have one implementation for both the private and shared platforms), and the rewrites were to start with Ruby components.
You knew working on the Erlang side of things that no feature development would ever take place again, that no new hands would be hired to help, and that you would be stuck on call 24/7 with no relief for years. All efforts were redirected to Go and getting rid of Ruby, and your stuff fell in between the cracks. I was one of the people who left on the long tail there. After my departure, I was brought back on a lucrative part-time contract as a sort of retainer for years to help them in case of major outages (got 1 or 2 in 4 years) since that was the only way they could get expert knowledge once they drove us all away.
I'm still on good terms with the people there, it's just that "maintaining a self-declared unmaintained legacy stack without budget or headcount until we get to rewrite it in many years" is not where any of us wanted to drive our careers.
Interestingly, we tried very hard to add new developers. We wrote manuals, tooling, a book on operating these systems (see https://www.erlang-in-anger.com), wanted to set up internal internships so developers from other teams could come and work with us for a while, etc. Whereas our team was very willing, internal politics (which I can't easily get in a public forum) made it unworkable and most attempts were turned down. These things were not always purely business decisions, and organizational dynamics can be very funny things.
I'd advance the theory that picking an off-the-shelf popular solution is going to be beneficial in that case because you externalize the costs of maintaining your expertise and knowledge to the rest of the ecosystem or industry, to other companies, and just never develop that fiber within your organization. It is, however, worthwhile to develop it for more things than just your tech stack. Everything having to do with on-boarding and dealing with legacy code is improved, along with broader dynamics if you try to be a learning organization that develops expertise in its people.
2. You always design your system in accordance to its deploy system, regardless of whether you realize it or not. If you ship binary artifacts and signed packages to customers who install them on their own devices, you will have a different development practice than if you do CI/CD with a single pipeline that always goes forwards. This will also be somewhat different if you work with open-source components that require paying attention to version schemes rather than just pushing a hash in a container image. If you use feature flags to help merging but also to control deployment, adopt A/B testing, and all these practices, you're intimately adapting your development approaches to deployment mechanisms that are available to you. It's not an extra cost, it's a cost you already pay today.
Tell me that for every person getting yelled at on Twitter you couldn't find countless more groups of minorities who have been denied justice over the years, whether because they are aborigines, black, lgbtqia+, or any other group of the kind. That open criticism and denial of cultures and ways of life wasn't just the default mode of operation. That one's life being valued less than someone else's property, beatings by police, harsher criminal sentences, and lack of equal rights wasn't just the mode of operation.
Getting yelled at on Twitter by people fed up with someone's bullshit is not even close to actual mob justice. It's just angry people shouting. Sometimes people shout enough that it turns to direct action (like letter-writing, which was used at least as far back as the 1800s), boycotts, and stuff like that. Today's cancel culture isn't mob justice any more than it was before, and it's not new.
Again, it's just a bunch of people who usually were never on the short end of the stick seeing its shadow pointed their way and freaking out.
For example, Kelly Loeffler, claims to have been "cancelled" while being a sitting US senator.
Cancel culture by popular action is not new (see letter-writing to TV stations, Frank Zappa having to testify in front of congress about censorship for his music albums). The only thing freaking people out is that people who have traditionally been structurally shielded from criticism and direct action (and often behind the cancelling itself) are just now on the receiving end of it.
I left because stackoverflow is not a community, I don't give a crap about the gamification, the incorrect answers most upvoted because they provide a quickfix rather than a proper fix, and I'd rather go in places where you can have an actual discussion when you need to re-frame problems since newcomers aren't generally used to the paradigms of the VM and OTP.
Stackoverflow felt like working for free just spending time answering things for imaginary points whereas other fora have more reciprocity available through richer dynamics that can help build communities.
Stackoverflow, to me, is a last resort more than anything else.
Far more Elixir developers I know have just given up on using Dialyzer altogether than the Erlang devs I know.
This means that from the time someone is infected until they die, a month and a half can have gone through.
By the time you take new isolation measures, you have to wait roughly two weeks to see if daily infection rates start to vary. But the lethality at this point, assuming no change in healthcare system overload, can be more or less statically predicted for weeks after the fact.
Those are called "lagging indicators"; the main thing you try to prevent are deaths, which you can lower by reducing the overload on the healthcare systems, which you can lower by controlling infection rates and keeping them low.
Due to the duration of the disease's evolution, you get to see if your policies really worked more than a month and a half after enacting them, which is extremely slow for a disease that propagates at exponential rates. So people look for proxies like infection rates that are still lagging, but far less so.
Do note that the one good metric you want in the end is going to be "excess deaths", which counts the impact of not just the diseases, but of all other side issues that the disease may have caused. This can usually take months up to years to properly account and analyze.
An E-program is written to perform some real-world activity; how it should behave is strongly linked to the environment in which it runs, and such a program needs to adapt to varying requirements and circumstances in that environment
Furthermore, as a part of its own environment, the software's own existence has an impact on what people expect and want to do with it.
Their mental models and their understanding of everything is not fungible, but is still real and often what lets us shift the complexity outside of the software.
The teachings of disciplines like resilience engineering and models like naturalistic decision making is that this tacit knowledge and expertise can be surfaced, trained, and given the right environment to grow and gain effectiveness. It expresses itself in the active adaptation of organizations.
But as long as you look at the software as a system of its own that is independent from the people who use, write, and maintain it, it looks like the complexity just vanishes if it's not in the code.
That being said, Mix does have its advantages in other ways (easier to extend for one-off scripts, for example, since it won't actually need a whole plugin; the ability for mixed-language projects with Elixir and Erlang), and they're clearly the reason (with Hex) why we have good package management today.
Again, I'm not ready to say one language has better tooling than the other; it's just that I feel it's a far cry from the situation from a few years ago where Elixir was just miles ahead -- the gap has closed since then.
Less “magic” (fewer macros that may change semantics, straightforward config handling), focus on declarativeness, no variable re-binding, more powerful test framework (common_test), seamless Dialyzer support, better logging framework for structured logging, simpler/minimal syntax, personal preferences for tooling (until recently, direct support for releases was lacking in Elixir), and force of habit.
In a nutshell, docker and k8s are to OTP what region failover is to k8s tech. They operate on different layers of abstraction and impact distinct components.
OTP allows handling partial failures WITHIN an instance, something K8s can’t help with.