fork() isn't awesome with CPython, as its reference counting kind of destroys the utility of copy-on-write (reading an object creates a reference, which increments the object's reference count, which is a write). And there's really a lot of shared data that you can have between processes, since every module you import is data (and expensive data!) -- I'd like to see it feasible to create extremely short-lived processes, for instance (why have a worker pool at all?)
Once you've created a new microprocess I suppose the actual concurrency would be up to the interpreter to best decide (or at least there would be a lot of flexibility because that concurrency would not be so overtly exposed to the Python program) -- it would be reasonable for a Python microprocess to be an OS-level thread, for instance, or for it to simply be cooperative multitasking at the VM level, or even be a real OS fork (though without an OS fork it seems reasonable to share immutable objects, but once forked the objects must be immutable and serializable).