Proposal of a new concurrency model for Ruby 3 [pdf]
atdot.net
atdot.net
* Guild has at least one thread (and a thread has at least one fiber)
* Threads in different guilds can run in parallel
* Threads in a same guild can not run in parallel because of GVL (or GGL: Giant Guild Lock)
A guild can't access the objects of other guilds.
About channels:
* We have Guild::Channel to communicate each other
* 2 communication methods
1. Copy
2. Transfer membership or Move in short
Copy is a deep copy and the object is duplicated into the destination guild. A transfer removes an object from a guild and makes it available to another.
There are also immutable objects that are available to all guilds. An obvious example are numbers, which are objects in Ruby, booleans and symbols. I think that other objects are frozen with https://ruby-doc.org/core-2.3.1/Object.html#method-i-freeze
They already did some encouraging benchmarks.
> 2. Transfer membership or Move in short
How is this enforced? What exactly happens at runtime if a guild tries to manipulate an object that belongs to another? (Absent a compile-time check, this is always a possibility.)
It would seem that:
(1) Guild ownership would have to be tracked in the runtime, obviously.
(2) Any access from Ruby code in the runtime, the runtime would also know what Guild the access request came from as well as the Guild the object belonged to against which access was sought.
(3) The runtime would be required to fail in some well-defined way (presuming, raising an exception in the requester) when the rules were violated.
It should be reasonably straightforward to assure this for all accesses within the runtime, since you can just make sure that there is no method to request access which isn't always attached to the Guild that the request comes from. It may be possible to break the runtime with poorly-behaved extension code that subverts the normal mechanisms, and it may be impossible to fully protect against that, but that's pretty much always a potential with extension code.
I'm not particularly worried about C extensions. I already know those are a lost case.
Very carefully?
More seriously, I think with guilds, what you absolutely don't want to do is build yourself into a position where you ever want to move a big linked data structure (that is, unless you know you are only going to use it in one guild, you never want to build a big linked data structure of mutable objects.)
Big structures of mutable objects should be guild local (or external, in a store that has its own controls for concurrent access.)
Check the code at http://rustbyexample.com/scope/move.html and run it (inside the page). There is a commented out println towards the end. Comment it in, run the code again and see the compiler error.
More about transfer of ownership at https://doc.rust-lang.org/book/ownership.html
For example, how would you transfer ownership of a big linked data structure from one guild to another?
(0) In Rust, this is as easy as handing ownership of the root node (an O(1) operation) to another thread. Ownership is transitive: whoever owns the root node also owns whatever the root node owns.
(1) In Ruby, I don't see how transitive ownership could possibly work. If I understand the proposal correctly, every object is owned directly by a guild, never by a parent object. Thus, you would have to traverse the entire linked data structure to transfer ownership of every node. This is O(n) work. Making things worse, you would have to make sure that the ownership transfer is atomic - no other part of the program should see the data structure in a “partially transferred” state.
Another example: Say you initially have a guild with three objects, Foo, Bar and Qux, where Foo and Bar point to Qux. If I transfer ownership of Foo, should Qux be transferred as well?
(0) In Rust, the type system forces me to explicitly distinguish between the following possibilities:
(0.a) Foo exclusively owns Qux, and Bar merely borrows it. In this case, Foo and Qux are frozen, and thus can't be transferred until Bar's borrow ends.
(0.b) Bar exclusively owns Qux, and Foo merely borrows it. In this case, transferring Foo doesn't change the fact Qux is owned by Bar.
(0.c) Foo and Bar jointly own Qux (using an Arc). In this case, transferring Foo doesn't change the fact Qux is jointly owned.
(1) In Ruby, what exactly should happen here? Have these guys really thought about the possibilities?
Correct.
> Making things worse, you would have to make sure that the ownership transfer is atomic - no other part of the program should see the data structure in a “partially transferred” state.
Correct. Guild::Channel#transfer_ownership() does it.
Basically, share big linked data with multiple threads is difficult (simply we need to lock every access).
> Another example: Say you initially have a guild with three objects, Foo, Bar and Qux, where Foo and Bar point to Qux. If I transfer ownership of Foo, should Qux be transferred as well?
Foo -> Qux; Bar -> Qux
Yes, Qux also moved. Programmer can know by "exception" when accessing Qux via Bar after transfer.
This "realization" is the key of Guild. On threads, we can't realize that Qux is shared mutable.
Slowly.
Objects well-suited to transfer would be shallow mutable structures with immutable data "underneath". Immutable data can be shared, so they'd be ignored by the transfer logic.
Another approach likely to be common would be not transferring objects at all, but sending proxy objects to other guilds which transparently marshal method invocations between those guilds and the object's home guild. For large mutable complexes which are intertwined with everything a guild does, that's probably a more manageable approach.
> Making things worse, you would have to make sure that the ownership transfer is atomic - no other part of the program should see the data structure in a “partially transferred” state
This is easy. You just run the transfer operation with the guild's mutex locked. No other guild can have a reference to a mutable object, and an immutable object doesn't need to be transferred.
> In Ruby, what exactly should happen here? Have these guys really thought about the possibilities?
It seems obvious that every linked mutable object is invalidated, so Qux will be transferred.
How often are data structures designed like this in Ruby?
> Another approach likely to be common would be not transferring objects at all, but sending proxy objects to other guilds which transparently marshal method invocations between those guilds and the object's home guild.
What if said “home guild” ends up overburdened with requests coming from all over the place?
> This is easy. You just run the transfer operation with the guild's mutex locked.
You mean both the sender and the receiver's mutexes? I'm worried about the receiver being able to observe partial transfers.
• Accessing from the source guild is invalidated
• Cause exceptions and so on
• ex) obj = “foo”
ch.transfer_membership(obj)
obj.upcase #=> Error!!
p(obj) #=> Error!Moving data between guilds is cheap because data does not have to be copied. Referencing frozen (immutable) data is cheap to.
It seems it will track ownership of objects to make sure guilds don't access other guilds data. But it doesn't seem that it uses OS-level data protection.
1. This Ruby 3 proposal says that Ruby 2 compatibility is mission critical, therefore this proposal rejects concurrency solutions from other languages (e.g. Erlang) and concepts (e.g. functions) and data structures (e.g. immutable collections).
2 Instead the proposal is to create a fast copy-on-write with rules to "deep freeze" some kinds of objects and primitives into an immutable sharable state.
Matz has been very public about his fear of a "Python 3" situation occurring in the Ruby community.
I can easily see how people might (rightly or wrongly) say "Ruby 3 broke my code, I'm rewriting in Go".
Transfering ownership would probably also mean that Ruby not only needs to move one object but probably all subobjects recursively as well. I assume here that "moving" just means updating the guild field for each object.
Is this really feasible or wouldn't just copying the object be faster... I don't know of any system with gc that uses moving to transfer mutable objects between threads. Do such systems exist? Are there better ways of implementing this?
Ruby is already checking the class of the object on every access. You could combine the guild and the class into a tuple and compare against that instead, so it adds no extra overhead.
There is a paper at OOPSLA this year on doing just that http://2016.splashcon.org/event/splash-2016-oopsla-efficient...
Adding a guild word to each object header is certainly a way to check ownership, and should be a cheap check to perform in the interpreter, but will obviously add some extra overhead to standard program execution.
The thing that concerns me is that explicit ownership passing can introduce as many bugs as it solves. If I have two objects A and B, with A holding a reference to B, then I can freeze A and freely pass it between guilds, but if I try and touch B I'll get an error until that too has been frozen or its ownership transferred. The same problems occurs with explicit ownership transfer of a non-frozen A, which leaves you with the slower option of a deep-copy or a recursive ownership transfer which can have equally unexpected consequences.
The "Ruby global data" slide also gives me the scream heebie-jeebies, as did finding stack overflow answers on how to unfreeze objects in MRI. I'm sure nothing will go wrong. :-)
Having said all that, it probably can work nicely for the common use cases of balancing requests between a group of worker guilds where the request is a simple data structure whose ownership can be safely transferred, but it would be hard to do a general work stealing solution that was always safe.
It needn't be done this way. When an object is invalidated it could have its class pointer sneakily changed into a special "invalid object" class. Any attempt to do anything concrete with the object would be rebuffed, but normal object accesses wouldn't be changed.
> The thing that concerns me is that explicit ownership passing can introduce as many bugs as it solves. If I have two objects A and B, with A holding a reference to B, then I can freeze A and freely pass it between guilds, but if I try and touch B I'll get an error until that too has been frozen or its ownership transferred.
At least you get an error. Ultimately, the only alternative with comparable performance is sharing mutable references. That avoids this specific problem but is open to the full assortment of problems that can occur with concurrent mutable state, many of which aren't automatically detectable in principle.
> The "Ruby global data" slide also gives me the scream heebie-jeebies, as did finding stack overflow answers on how to unfreeze objects in MRI. I'm sure nothing will go wrong. :-)
If this proposal is adopted it's a simple matter to prohibit unfreezing objects that have been shared. :)
Not that it would really be necessarily. If you're reaching into MRI's internals to unfreeze an object, then it's up to you to make sure that things don't break.
I really hope they make this work in Ruby 3. If you program in a "functional" style anyway, this approach ("relaxed Erlang" style messaging) should fit nicely, as you would not be mutating things willy-nilly anyway. Of course, FP practices are a hard sell to the "OOP is the one true abstraction" (COBOL with encapsulated DATA DIVISIONSs) crowd.
Since the requests would not be in the "main" guild, it might be painful to call into gems.
It all really depends on how much overhead there is to create and destroy guilds. If it's easy then ideally you could start 100s of guilds or 1000s should your hardware allow it.
I see guilds as a subprocess with its own isolated resources.
On par with That's "crates". Gives the impression some people just want to be remembered as inventing names.
Pretty much. I mean, if there is a standard name for a thing between a process and a thread that is not a thread group, I haven't heard it.
Here's my reasoning. Since the GVL is insufficient to guard against data races on Ruby 2, under the guild system, locks would be needed to guard against concurrency issues if multiple threads are present.
It would seem like the intention would be to replace usages of Thread with Guild to avoid the concurrency issues inherent with threaded code. Will there be API support to create a Guild that only allows a single thread?
It seems to me that is the intent; that is, any Ruby code that exists now is single-guild Ruby 3 code -- if its multithreaded, it needs locks, for the same reason it does now.
> It would seem like the intention would be to replace usages of Thread with Guild to avoid the concurrency issues inherent with threaded code
I think that'll be a common use case, though running what amount to multiple "legacy" Ruby 2 multithreaded systems in separate Guilds in the same Ruby 3 process seems also to be an intended supported use case.
> Will there be API support to create a Guild that only allows a single thread?
It certainly sounds like a good idea.
That is correct. You'll still need to use locks if doing multi-threading inside of Guilds.
It looks like Guilds can have 1 to X threads.
Python's attempts to remove the GIL are not going anywhere really.
Seems like a guild is just a subprocess with its own resources. And you copy objects over as needed. And when the guild is done it will get garbage collected. Like other objects.
In an ideal world the guild is the interpreter state which would be very far from processes. How far down you can go there largely depends on what promises the API made to C extensions and other things in the past.
That likely would be faster than having OS threads in each guild that use PS locks to prevent running >1 of them concurrently.
Actually, I think Go is some kind of multi-thraeding (goroutine is only a useful mechanism on the "threads" and can't help to avoid multi-threads difficulties (but this design helps to reduce difficulties)).