Ruby’s GIL and transactional memory
mikeperham.com
mikeperham.com
http://pypy.org/tmdonate.html (has the best executive summary of what the branch intends to accomplish)
Writing single multithreaded apps with low-level locking hardcoded everywhere is now quite clearly NOT the right way to build software. If you don't use locks, i.e. only use lock-free data structures and immutable state, then you won't care about the GIL. And you can use multiple processes and interproccess communications in place of threading. On Linux, the difference in performance between threads and processes is very small. Most people who complain about the GIL have not even profiled multithreading versus multiprocessing. They are just bound and determined to reinvent the wheel in their own code base.
There is no reason why you can't leverage C++ (ZeroMQ) or Erlang (RabbitMQ) to do the hard bits and write the rest of your app in nice simple Ruby (or Python) scripts that are designed according to the Actor Model.
This is a fair point, but the issue might be about memory usage, not speed. A unicorn setup might have two to eight worker processes to service HTTP requests. Even with copy-on-write-friendly garbage collection, the memory usage of each additional process is significant. On the other hand, a thread-based solution (using JRuby, for example) can maintain a threadpool with hundreds of worker threads because the cost of an additional thread is nearly negligible.
There is definitely annoying memory overhead with multiprocess (vs. multithreaded) architectures, but it's on the order of 2x-8x, not 100x. And that's 2x-8x the code size of the application, not data size - you only need duplicate interpreter objects, anything at the app or framework level (like templates or data files) can be stored in read-only shared memory or just COW'd with no writes. (It's technically not even every interpreter object - a number of function objects are completely static data that will never have additional references made, and so COW means they'll be shared perpetually between processes.)
Shrug why not? You get a different (easier?) programming model where you can use blocking IO rather than an event loop.
> on the order of 2x-8x, not 100x
Yeah, but that 2x might be the difference between one virtual machine flavor and the next price up.
Unicorn is a pre-forking multiprocess server so I don't know why it would be using an event loop.
Why threads over processes? Because memory isn't cheap when you don't own it yourself.
It all depends on how much you want to get out of your hardware.
So it's a trade-off where the downsides often get pushed off into another group. As a developer, I really miss the CPython solution, which was a lot simpler and seemed more robust. But then, I wasn't the one responsible for pushing out new code or monitoring. I do think there were various optimizations we could've made to our other tools that might've compensated for the need to run more server processes, and wish we'd tried that before jumping to "Let's use a GIL-free language."
Unless you're trying to do something that requires parallel processing. Crypto cracking was one example recently.
>> And you can use multiple processes and interproccess communications in place of threading.
This adds complexity
>> There is no reason why you can't leverage C++ (ZeroMQ) or Erlang (RabbitMQ) to do the hard bits and write the rest of your app in nice simple Ruby (or Python) scripts that are designed according to the Actor Model.
This also adds complexity.
Basically it would be nice if threads in python acted like they do in other languages, rather than just pretending to.
It is a very opinionated view that writing multi threaded code with locking is not "the right way". Of course, using locks requires disciplined engineering but there are applications where you might want to use threads in a single process rather than multiple processes. Although it might be less common for domains where Ruby is popular.
Multiple processes and IPC is not a whole lot easier than writing multi threaded code, especially if you have nice a nice concurrency framework with channels, etc available.
> If you don't use locks, i.e. only use lock-free data structures and immutable state, then you won't care about the GIL.
Unfortunately, lock free data structures or immutability will not avoid the GIL. The GIL is used to guard internal data structures in the interpreter and it's practically held always when code is being interpreted and released only to when blocking I/O is happening (at least that's the way it works in Python).
Not being able to write multi threaded code is a weakness in CPython and CRuby and finding ways to avoid locking the GIL would make them better.
Using scripts running in separate OS level processes -- no matter what infrastructure you use to connect them -- isn't necessarily the ideal solution.
Look at the first slide from this talk by Professor Michael Stonebraker to see why more locking is not such a good idea. http://blog.jooq.org/2013/08/24/mit-prof-michael-stonebraker...
Just because we can do it doesn't mean that we should do it.
I think you've got it slightly wrong. What do you think happens when one request is being worked on in effectively one green-thread-equivalent in node.js? Everything else is blocked until the current operation yields. That's an equivalent of a GIL, yet it doesn't "reduce performance". If you switched to multithreading + locking (making it run in a M+N scheduling model), you'd gain performance, not lost it.
You (or your standard library) need ruby-level locks in ruby code when you do things like increment a counter. E.g.: "obj.count = obj.count + 1".
The interpreter needs interpreter-level locks when doing things like looking up attributes from an object's dictionary or hash table or whatever ruby calls it. That's why you don't always need a lock around a simple assignment to avoid interpreter crashes: "obj.count = 0" is safe. Without a ruby-level lock around that, there is a race condition with the previous example, and resetting the count to zero may not have an effect if you're unlucky.
Without a GIL, however, you could have problems like the interpreter segfaulting (instead of throwing a nice execption that you can handle). You may also have problems like that assignment turning into an infinite loop, depending on ruby's implementation of hash tables.
And if accessing a hash table isn't safe, you can't even use ruby-level locks. How would you create a lock? You'd need to access classes or functions in a module. The module stores those things in a hash table, which you can't access without a lock. That's what the GIL is for.