We used Elixir's Observer to hunt down bottlenecks
blog.sequin.io
blog.sequin.io
So I'm always intrigued by some of the more BEAM-specific things that folks do, like using `observer` on a remote (production??) node here, or distributed Elixir where the nodes communicate with each other, or "hot" code updates.
How do companies deploy Elixir in such a way to take advantage of all those things? Does Sequin talk anywhere about their deploy process and how their infrastructure looks?
Hot code updates for most applications aren't really worth it in my opinion, assuming you do something like blue/green rollover deployments. It's cool that it's possible though. But it requires appup files and afaik Distillery is one of the release tools that has support for it built-in.
In Erlang/Elixir you can actually override how instances of the BEAM find each other (instead of the standard EPMD daemon), so we have a module that does some DNS queries, finds the IPs of the other containers and says “hi, here’s your cluster, discovery done.” (Your setup may preclude all that, I know this all depends on how a system’s architected.)
After doing that we were free to use all of Erlang’s cool cluster stuff! In our case we have in-memory caches for a few things, and if a given instance does a lookup because of a cache miss it broadcasts a message to all the other nodes saying “I just looked up $expensive_thing, here’s its value” so they don’t have to do the lookup themselves, they just cache that value, so you end up with a little distributed cache with a few lines of code. In our case, btw, these cache entries are short lived and a little inconsistency does us no harm if one of our instances misses the message, networks are networks, but it’s been great!
Anyway, I think it’s super cool and I’d encourage you to play around if you get the chance.
Also the observer is just amazing. We’ve debugged some pretty weird memory and cpu usage issues with it, I have some internal blog posts, maybe I should see if I could make them public.
You can skip to "Let’s use something else" if you've already got a good grasp of the knobs to tune with epmd.
https://www.erlang-solutions.com/blog/erlang-and-elixir-dist...
It took some work to piece together but wasn't too complicated in the end. It's internal code that I can't quickly sanitize or I'd just dump it in a gist :/ Someone on the Elixir Forum might have a template or library handy though.
it's funny how 20 years ago "N containers" would be normal, but now everyone can read both
I'm not criticizing the methodology as much as the useless performative nature of compliance work.
1. We had a breach. A factor in this was insufficient oversight on a process that granted privileged access to customer data. We fixed the problem, promise that your data is safe, and don't believe this will happen again.
2. We had a breach. A factor in this was due to a gap in an existing control around customer data that had a problem we had not anticipated. These were the people involved. This is exactly how this problem occurred. This is the data that was exposed. This is documentation of our response to this incident. This is our existing policy around how we handle data and how we respond to breaches.
Customers, partners, regulators, and law enforcement respond a lot better when you can demonstrate good intent and at least imply that you have some kind of process. Of the two scenarios I outlined, the latter provides those assurances.
Compliance isn't the only way to do this, but it's often the easiest.
Hard to say without knowing much about the data in question, but my recollection is that large Erlang/Elixir/BEAM "binaries" are actually not copied around. That might be a strategy for sharing larger things in some cases.
Marshalling data is pretty easy in Erlang:
2> Bin = erlang:term_to_binary([1, 2, 3]).
<<131,107,0,3,1,2,3>>
3> erlang:binary_to_term(Bin).
[1,2,3]iex>term = for _ <- 0..1000000, into: [], do: :rand.uniform(0x1111111)
iex>bin = :erlang.term_to_binary(term)
iex>for _ <- 0..1000, do: spawn(fn -> x = bin; :timer.sleep(1000000) end)
# Memory usage exploded in line below
iex>for _ <- 0..1000, do: spawn(fn -> x = term; :timer.sleep(1000000) end)
I never understood what in that lib was causing the leak but I fixed it (or more accurately mitigated it) by wrapping the call in a Task.async/1
Maybe that will help someone else one day.
Running the function (which probably parses large binaries) in a separate process ensures that it's properly garbage collected in time.
Yes that could be it.
A related, but different refc binary hazard is a Process that obtains large refc binaries somehow, and makes a subbinary that it sends to another process (or ets!). The large binary is still referenced from the subbinary so there's a significant amount of excess memory. You can also run into this when binary creation is optimized to allow for appending [1], because that makes a binary of much larger than the required size (either double or 256 bytes, whichever is more). Either way, if you have a use case that naturally results in long term storage of binaries or subbinaries that allocate much more space than is really required, binary:copy/1 can be used to make a clean copy that's the exact size and isn't (yet) shared.
I've seen mnesia (ets) nodes where due to the code structure, the memory use ended up at 4x what was needed, and binary:copy added before storing with ets fixed things up with no other code change.
[1] https://www.erlang.org/doc/efficiency_guide/binaryhandling.h...
Interesting to see this approach to article graphics after I first read about it on HN recently.
> Painting of a detective from the 1800s, portrait, looking at a magnifying glass at a computer monitor, digital art
Our solution now uses trigger functions. These trigger functions fire whenever a create/update/delete happens on a Sequin table. They insert a row into a log table. That log table is processed by our workers to send changes to the upstream API.
The advantage of using trigger functions + a log table are all about ease of use and compatibility: our customers don't have to do anything fancy to setup Sequin, we just need a role with `create` privileges in the database. The log table also makes it easy for both them and us to debug issues, as the stream of changes that we captured is right there in the database.
I'm using Elixir to listen to change events via https://github.com/cpursley/walex (which I basically ripped off from Supabase).