So I'd expect that even on MRI the original code could in fact fail to yield 0 once in a while. That's what'd happen on CPython
The global lock is a feature of MRI that basically wraps a big mutex around all of your code. That's right, even if you're using multiple threads on a multi-core CPU, if your code is running on MRI it will not run in parallel.
Technically the GIL is there to protect the VM's internal state, so its operational semantics are that it's taken and released for each bytecode instruction.
However the efficiency would stink (short of awesome automatic lock elision or merging), so GIL-based VMs usually have a higher threshold, either based on some sort of instructions count (which may or may not map 1:1 to bytecode instructions) or an internal "wall clock" timer. I think MRI's the latter but don't quote me on it.
There are some very specific caveats that MRI makes for concurrent IO. If you have one thread that's waiting on IO (ie. waiting for a response from the DB or a web request), MRI will allow another thread to run in parallel. thread #1:
local a = x[i]
a = a + 1
context_switch()
thread #2:
local b = x[i]
b = b + 1
x[i] = b
context_switch()
thread #1:
x[i] = a
resulting in a wrong (or merely unexpected) result.