I just tested this on a Linux box I happened to have running, using the "perf" tool. By my measurements, each iteration takes about 104 instructions, of which 23 are conditional branches, and completes in about 31 cycles.
(That's after subtracting about 30 million cycles of startup overhead. Tested with Python 2.7.9 on an Intel i3-4160 processor.)
Remember, Python is a bytecode-interpreted language. Each iteration of the loop involves multiple bytecode operations:
2 0 SETUP_LOOP 20 (to 23)
3 LOAD_GLOBAL 0 (xrange)
6 LOAD_FAST 0 (NUMBER)
9 CALL_FUNCTION 1
12 GET_ITER
>> 13 FOR_ITER 6 (to 22)
16 STORE_FAST 1 (_)
3 19 JUMP_ABSOLUTE 13
>> 22 POP_BLOCK
>> 23 LOAD_CONST 0 (None)
26 RETURN_VALUE
Executing each of those instructions means fetching it, jumping to the implementation of the appropriate opcode, and then updating the VM's state -- or, in the case of FOR_ITER, calling the native C function that advances the iterator.
Frankly, it's impressive that it's as fast as it is.