I do not have time to check the Python code to see how it is actually implemented (hence my question) but based on my knowledge it would imply at least a CAS operation to check and take the lock, writing the register values into memory (cache) and applying a memory barrier.
You can not keep the values in the registries (elimination of optimization possibilities) and you add considerable overhead by needing memory barriers and CAS operations.
I am not claiming that Python does it like this, I am just assuming that it should do it like this to obtain the guaranties of GIL.
But I had an impression that it is acquired before every atomic Python instruction and it looks that it is actually acquired for group of predefined number of instructions (100) that then are executed inside one GIL time frame [0].
Therefore it actually should not be a big obstacle to make Python code run fast by a JIT compiler.