An Introduction to Ractors in Ruby
blog.appsignal.com
blog.appsignal.com
* Run a bunch of background workers (Sidekiq/Resque), and queue up a job for each item you want processed in parallel.
* Provide relatively granular HTTP APIs, and have your JS frontend call them in parallel with AJAX, instead of having the server handle concurrency.
I think this is just the nature of Ruby being widely used for web apps where performance isn’t a big concern. That said, I’d love to see Ractor catch on, since it’s a pattern built into the language everyone could standardize on.
My info may be dated, I haven’t kept up with developments in Ruby 3. Is the GIL not going to be a thing any more?
Ruby 1.8 had green threads, 1.9 onwards has native threads, but with a GVL. And 3.2 may have N:M threads.
> Is the GIL not going to be a thing any more?
On paper it's already done in 3.0. What used to be the GVL is now one lock per Ractor. Every object belongs to a Ractor, and the Ractor lock need to be acquired to access the objects, except for mutable objects that can be shared across ractors.
So in practice if you are not spawning any more ractor than the main one the VM execute your program in, then it's technically still a global lock, but you now have (limited) ways to go around it.
Edit: as the sibling comment mentions, that’s part of what Ractor is trying to solve.
It's entirely orthogonal concerns. Job queues are for "fire and forget" tasks. Using Sidekiq and co offer tons of advantage over queuing this in a local thread pool / coroutine or similar.
It provides efficient backoff retries, work distribution, durability, backpressure, and a tons of other goodies.
If you were to just queue these things in-process, you may queue more work than your process can handle which may hurt latency.
It also make deploy of web application awkward. If you know jobs are externally queue, you can stop sending traffic to that process and once all in-flight requests are completed you know it's safe to stop the process. If that process may contains queued work, well who knows when it's safe to restart it.
In process concurrency (or parallelism) is useful for other things (e.g. parallel queries to a service), and that's were you use threads / fibers / ractors / async.
While learning, I worked through Practical Object Oriented Design in Ruby but doing the examples in Pony. It really highlights some of the short-comings of Ruby, that are implemented better in Pony. Off the top of my head: Pony has named arguments, default arguments, real interfaces (no need for "duck-types"), it's strongly typed... There are probably more but those are a few I noticed specifically in a book that could otherwise stand as promotional material for Ruby.
Definitely recommend anyone interested in OO, or the Actor Model to check it out.
Pony is a great language though. And yes it has an amazing concurrency and parallelism model and showcases the actor model very nicely.
However, it has been a beast to get it working and it is definitely not anywhere near "done". Sharing even simple config objects code between ractors is extremely difficult, because most of the code was never written with "sharing between ractors" as a design concern. It is like the red/green functions thing but with ractors instead of async. :/
Another thing to note is that famous actor-based platforms, like Erlang's OTP and Scala & Java's Akka, offer a lot for building network distributed systems with features for dealing with faults. That's a big part of their appeal. Not sure whether Ractors offer that.
I wrote a model checker in Python and it's slow (it's not finished yet!) as it tests every possible interleaving of the threads. (I test some interleavings I was worried about explicitly with manual specifications) I've run it a lot but there's the possibility there is a rare bug. But my model checker needs parallelising to run in reasonable time to prove the algorithm correct. So I need to parallelise my parallel algorithm checker.
It's kind of based on https://www.cs.technion.ac.il/~erez/Papers/wfquque-ppopp.pdf in terms of work stealing (or helping)
Every thread can communicate with every other thread and I load balance the pool with a simple atomic counter and modulo.
https://github.com/samsquire/multiversion-concurrency-contro...
For 100 parallel threads on a hexacore I get 1.75-2.3 million communication events a second (depending on CPU load, when I'm running the model checker, I get around ~1.5 million requests a second) on an Intel Processor Intel(R) Core(TM) i7-10710U CPU @ 1.10GHz, 1608 Mhz, 6 Core(s), 12 Logical Processor(s). Yes it's an Intel NUC. Communication costs 500 nanoseconds. I plan to implement bidirectional communication so threads swap messages with each other, for double the throughput. The plan is that threads/actors have multiple event streams or sources of work, so they always have work to do even if there is contention.
I don't know if it is fast or slow, as I haven't benchmarked golang channels, erlang actors or akka or any other actor framework but I do wonder what is the raw performance of a multithreaded actor model.
If you found this comment useful or interesting, check out my profile. I write of this stuff everyday.
When exactly can you send to or receive from a ractor? Are you expected to wrap every send and receive in some kind of error handler, or can you at least check with some "is_closed" method? When exactly do ports open and close, and how long do they stay open for?
They explain it with a lot of prose but give no definitions, properties, or proofs.
Maybe they should leave out the “mathematical” part if that isn’t what they meant.