For my python multi core code I like using fork with sending objects over a socket with pickle. It gives more control than multiprocessing.
It works pretty well. One downside of fork with python is that the reference counters are scattered about, leading to big COW memory churn, however. Big numpy arrays should be shared, however.
With sub-interpreters, I wonder how different it is from normal multithreading like in JVM. Need to read up on this idea. I can't see how it wouldn't require the same locking and mutexes or message passing if multiple interpreters are to work on logically related data.
So my immediate thoughts were about leveraging it as a replacement for multithreading and event loop clustering by treating the interpreters as lightweight processes and having them communicate over some kind of protocol. Like how the BEAM does built-in process supervision.
The "supervisor" process itself is async in the way it coordinates the tree.