Rustler – Safe Elixir and Erlang NIFs in Rust
hansihe.com
hansihe.com
As an Erlanger, and maker of a NIF -> Rust library I wasn't allowed to open source (sadly), I can tell you I think this is excellent. My current science project is trying to rebuild some small components of BEAM and the Erlang Runtime System in Rust that are currently in C.
Mostly the above is an exercise in understanding the internals of BEAM better, because my ideal end state is to figure out what it would take to make an Erlang/OTP-like library for Rust (with task/thread/process/whatever preemption)... which is actually what originally drew me to Rust 3-4 years ago back when it was shaped more like a systems-Erlang (task supervision hierarchies, etc) and less like a renovated-safe-C/C++.
Presuming that this happened on company time, I'm curious to know if there's a company using Erlang+Rust in production or if this was just a fun side project. :)
I've always been a fan of Twitter's mutual opt-in requirement for DM chatting. You're being followed. Which feels really weird and funny to write. But it'll be reasonably obvious.
Wow, any chance Erlang-like supervision and actor model in general might make a comeback in Rust? I'm sure the Erlang guys has thought about this, but no static type checking makes me a bit nervous. Erlang can be surprisingly strict (a good thing) in some ways, but as far as I can tell, you only discover any failures at runtime.
For example, if I have the following Elixir function definitions:
def foo({:bar, bell}), do: IO.puts("got a barbell: #{bell}")
def foo({:baz, bell}), do: IO.puts("got a bazbell: #{bell}")
Subsequently calling `foo({:bat, "some value for bell"})` should be detectable by the compiler and thus generate at least a warning. We can luckily catch it relatively easily and painlessly at runtime with Erlang's normal insistence on process supervision trees, but it's still a crash that could be easily avoided without having to jump into the realm of full-blown "defensive" programming.But finding race conditions is another problem entirely. Have a look at Concuerror and PropEr.
As a library, sure, but like green threads, actors aren't ideal for a systems language.
-Wunderspecs
-Wunknown
-Wunmatched_returns
-Woverspecs
-Wspecdiffs
And depending on a whole slew of things that may or may not make this useful, or more likely tractable for your codebase, you can enable this to track down some shapes of race conditions: -Wrace_conditions
The part where there's still a hole (now that most of the hole around Maps has been plugged) is in the message passing semantics. You can be conventional about how you write an API around the message passing to alleviate this, but there's still nothing stopping any random process from sending any shape of data to any other random process and thus essentially breaks Dialyzer's ability to enable static type checking through all paths. But as long as you have a catch-all matching clause implemented that dumps anything that doesn't explicitly match one of your types you're mostly fine.(Ports, on the other hand, are inherently safer due to being isolated from the Erlang VM, such that even if they crash they don't bring down the whole VM. They also communicate over STDIO, meaning that I'm free to use them to interface with components written in languages with less-than-pleasant methods of achieving C ABI compatibility, like Perl 5 and Ruby when not augmented with extra packages.)
On that note, Rust also has some wonderful implications for the Perl 5 world in terms of providing some semblance of safety around native Perl extensions. Same deal for Perl 6 (though its FFI mechanisms are infinitely nicer than Perl 5's).
https://github.com/hansihe/html5ever_elixir
It shows how create a threadpool and avoid the 1ms maximum nif execution time too.
My understanding is that if I still want to ensure that the Beam VM can continue to schedule all processes efficiently, these NIFs shouldn't calculate forever, but return quickly to avoid blocking all other processes. So the straightforward way to ensure this for longer calculations would be to split the calculation in smaller, parallelizable jobs, if that is possible. But then the overhead of calling NIFs might actually matter, if you split them into chunks that are too small.
I've no idea how big the overhead actually is, I'd be interested in a rough estimate. Is it small enough that I can just ignore it entirely?
There's some overhead in type-casting/data-type-interpretation, but not much. The internal representations of the core data-types in Erlang are mostly already represented in C in the runtime. There's some overhead in memory copies across that boundary too unless you make your data opaque to Erlang and always manipulate it via the NIF in which case you can get away with using NIF "resources" and more or less just pass a pointer and environment back and forth.
I wouldn't be surprised that at least one of this library or the Erlang VM itself develop an official way to easily use a NIF to run a longer-running process safely. There hasn't been a need up to this point because there hasn't been a such thing as a long-running NIF.
At the moment, if you're looking at something that may take several seconds you're probably still better off coordinating something over a port to an external process or something. Or working with this project to make it feasible to run a long-running native process.
Trying to write a NIF that could somehow yield back and then be "continued" later would be a royal pain. Cooperative scheduling was enough of a pain when we were in the hundreds of processes mostly doing nothing on OSes; trying to use it on an Erlang server is probably just infeasible.
nifs are ideal for operations which mutate binaries, eg hashing or unmasking a websocket frame
There's some some parameters in there that are going to be tricky to set up correctly (number of workers in the pool for a given NIF has a lot of implications at scale), but in the end it's not significantly different than communicating to a separate process on the some machine; it's all the same resources being used.
I know... nothing... about the NIF interface, but the way that I would want this to work wouldn't involve NIF resumption at all, but rather for NIFs to be able to put responses back on a channel.
Unfortunately neither quick internet searches nor searching the rustler docs revealed anything about a NIF/channel interface. I'm not sure how feasible it is to pass a channel into a nif and allow it to pass that into some thread and return, nor how good of an idea that would be.
If you want the behavior you're talking about you would communicate via ports, or you'd write a NIF that maintained its own off scheduler thread(s), returned immediately, but kept a reference to the calling pid, and then sent a message to that pid at some later date. Several IO related NIFs do that today.
That's actually an excellent use for Rust in this case because it can help you make that multi-threaded implementation safer and more reliable.
It's worse than that, overlong NIFs can lead to scheduler collapse with your schedulers being put to sleep and stuck there.
> these NIFs shouldn't calculate forever, but return quickly to avoid blocking all other processes.
Recent BEAM (R17 and later) has the experimental feature of dirty schedulers which allows running long NIFs without risking scheduler collapse. There is one set of dirty schedulers for IO-bound NIFs and one for CPU-bound NIFs, they should be marked accordingly.
> I've no idea how big the overhead actually is, I'd be interested in a rough estimate. Is it small enough that I can just ignore it entirely?
https://news.ycombinator.com/item?id=10141945 and the linked article https://medium.com/@jlouis666/erlang-dirty-scheduler-overhea... measure regular nifs at ~1.2µs and dirty nifs at ~10µs by default though they can be as low as ~3µs for core-bound schedulers and tuned wakeup options.
If it's something that really can be split into small parallel jobs, then NIFs make sense here; the overhead is pretty low, since the NIFs are linked into the VM itself.
If it's something that can't be easily split into small parallel jobs, then you might want to go with a port, such that the long-running job doesn't hang your whole Erlang VM.
As for not blocking, the OTP team is working on getting dirty schedulers to be a first class system. The overhead there is also pretty low, but now you don't have to worry about blocking the schedulers anymore.
What is pricey in a NIF is conversion of data back and forth more than the overhead of the call itself. it is usually quite fast.
This is not quite accurate; you can still segfault rust writing 100% safe code, for example if you have a large stack overflow ever since __morestack was killed.
Though these cases are fortunately rare
Do you think I should have written it differently? I might want to add a clause in the caveats section at the very least.
I wouldn't say "allowed" to segfault. :P The behavior you're referring to is currently due to a deficiency in LLVM for non-Windows platforms. Here's the bug on the Rust repo tracking it: https://github.com/rust-lang/rust/issues/16012 Its resolution is long-awaited, to say the least, but it requires someone proficient with LLVM to do the legwork...
> an in-rust panic might be recoverable?
Anytime you're writing an interface where code is going to be calling into Rust via C FFI, you ought to be using https://doc.rust-lang.org/std/panic/fn.catch_unwind.html , which is specifically intended to prevent panics from crossing FFI boundaries.
How do you expect a stack overflow to unwind? I have difficulty imagining the implementation
As kibwen says, the ways this can happen currently are bugs. They'll be fixed.