A long time ago IronPython was released which showed that you could build a high performance Python interpreter that was GIL-free. It had thread-safe containers (so multiple threads could work against the same lists, dictionaries, etc) and in some cases was faster than CPython (it was implemented using CLR and .net)
When I saw IronPython I was immediately convinced that CPython should be the same way- already low-cost dual and quad core Intel machines were becoming available and it seemed like core counts were going to increase faster than clock rates. I figured that a small hit to serial performance would be more than acceptable if people could write multithreaded systems in Python, in much the same way as I wrote multithreaded C++.
Over time after watching nogil not going anywhere in CPython (the python leadership didn't want to do nogil), with concomitant speedups in the single-processor implementation, along with the increasing use of C++ code that releases the GIL, and seeing that many people just weren't good at multithreaded programming, I have started to conclude that the multithreading/multiprocess in Python today is about the best we can get. That is, instead of having threaded containers and multiple interpreter threads all banging on the same underlying data, it's a lot easier to just use threads as work queues that have minimal interaction with other threads.
So that's where I've ended up: some of my code using multiprocessing, typically the Pool or ThreadPool, with the concurrent future API to handle result-gathering, barriers, etc. Other code has an external system that starts many python processes from the command line and waits for those processes to complete. Other code is single-threaded in python and launches C++ cores that launch multiple threads just to return a computational result to python faster.
And I think trying to do both gil and nogil interpreters, rather than committing to one or the other, the python leadership will sign us up for untold inconveniences around packaging. We already see this in the move to async and it will only be worse with threading.
So sad to say I think sticking to GIL and using the approaches I mentioned above (along with others that work well for concurrent, rather than parallel, computing, like coroutines) is the best thing to do right now and I'm a bit bummed that the leadership signalled their intent to accomodate both.