Test 1: Boring, small array of integers
In [28]: arr = range(0, 300)
In [29]: %timeit [(arr[3*x], arr[3*x+1], arr[3*x+2]) for x in range(len(arr)/3)]
10000 loops, best of 3: 27.2 us per loop
In [30]: %timeit numpy.reshape(arr, (-1, 3))
10000 loops, best of 3: 45.2 us per loop
In [31]: %timeit zip(*([iter(arr)]*3))
100000 loops, best of 3: 6.25 us per loop
This roughly matches the article's timing ratios, so far so good.Test 2: Use numpy's random number generation to get a small array of floats
In [32]: arr = numpy.random.ranf(300)
In [33]: %timeit [(arr[3*x], arr[3*x+1], arr[3*x+2]) for x in range(len(arr)/3)]
10000 loops, best of 3: 54 us per loop
In [34]: %timeit numpy.reshape(arr, (-1, 3))
1000000 loops, best of 3: 1.06 us per loop
In [35]: %timeit zip(*([iter(arr)]*3))
10000 loops, best of 3: 39.7 us per loop
numpy is two orders of magnitude faster here; it's evidently using a highly optimized internal codepath for random sequence generation, which I'd guess is a common thing to do in numeric analysis. I assume it's using a generator, so there's no actual array being created, blowing up the CPU cache lines etc.Test 3: Verify that analysis by interfering with numpy
In [36]: arr = [x for x in numpy.random.ranf(300)]
In [37]: %timeit [(arr[3*x], arr[3*x+1], arr[3*x+2]) for x in range(len(arr)/3)]
10000 loops, best of 3: 26.2 us per loop
In [38]: %timeit numpy.reshape(arr, (-1, 3))
10000 loops, best of 3: 48.5 us per loop
In [39]: %timeit zip(*([iter(arr)]*3))
100000 loops, best of 3: 6.55 us per loop
Yep.Test 4: Larger data set, no interference
In [40]: arr = numpy.random.ranf(3000000)
In [41]: %timeit [(arr[3*x], arr[3*x+1], arr[3*x+2]) for x in range(len(arr)/3)]
1 loops, best of 3: 624 ms per loop
In [42]: %timeit numpy.reshape(arr, (-1, 3))
1000000 loops, best of 3: 1.06 us per loop
In [43]: %timeit zip(*([iter(arr)]*3))
1 loops, best of 3: 335 ms per loop
The numpy time doesn't change at all from test 2 despite the larger size, but the others suffer. Again, I suspect numpy is being intelligent here; my guess is that it doesn't actually apply the function and generate the real output, it just wraps the random generator in another one.Test 5: Larger data set, interfering with numpy
In [44]: arr = [x for x in numpy.random.ranf(3000000)]
In [45]: %timeit [(arr[3*x], arr[3*x+1], arr[3*x+2]) for x in range(len(arr)/3)]
1 loops, best of 3: 321 ms per loop
In [46]: %timeit numpy.reshape(arr, (-1, 3))
1 loops, best of 3: 354 ms per loop
In [47]: %timeit zip(*([iter(arr)]*3))
10 loops, best of 3: 83.6 ms per loop
There we go; we're back to roughly the original timing ratios.So, surprise! You always have to measure. Measure, measure measure. My bias is to write code first for legibility and modifiability, and then optimize hot spots if needed (and add comments, please, when you do so).
Without doing deeper analysis I'd say one moral of the Python story is, this shows the potential power of generators. But in real-world data sets this isn't always ideal -- is it faster to load up the whole data set in memory and blast through it, or load it from disk on demand with a generator? In really high performance scenarios, is it faster to preprocess the data to fit into the CPU's cache lines? You can't tell without measuring, and you have to measure in the environment you're deploying to, since the answer may be different on a machine with 1GB RAM vs. one with 128GB RAM, or 32KB L1 cache vs. 8KB.