f3
------
92 LOAD_FAST 2 (a2)
94 LOAD_FAST 3 (c)
96 BINARY_ADD
98 LOAD_FAST 3 (c)
100 LOAD_CONST 1 (2)
102 BINARY_ADD
104 ROT_TWO
106 STORE_FAST 2 (a2)
108 STORE_FAST 3 (c)
f4
------
92 LOAD_FAST 2 (a2)
94 LOAD_FAST 3 (c)
96 INPLACE_ADD
98 STORE_FAST 2 (a2)
100 LOAD_FAST 3 (c)
102 LOAD_CONST 1 (2)
104 INPLACE_ADD
106 STORE_FAST 3 (c)
And perhaps the culprit is that INPLACE_ADD requires more work than BINARY ADD https://stackoverflow.com/questions/15376509/when-is-i-x-dif...I strongly believe that the idea that it's running on multiple cores is not the case. Python doesn't implicitly schedule any user code on separate threads, and if it did, the coherence cost of doing so would swamp any benefit of doing two simple operations in parallel.
Disassembly is as simple as:
with open('f3', 'w') as f:
dis.dis(factor_fermat3, file = f)
with open('f4', 'w') as f:
dis.dis(factor_fermat4, file = f)
Edit: Whoops, I see that I used Python3, while the post was about 2.7.4, and there's some wrangling over the differences below. I reran the disassmbly under 2.7.4. The bytecode was more verbose, but differed in the same way.