I've had this experience when writing a simple algorithm for run length encoding, a problem which I struggled to vectorize. Compare the numpy algo with pure python:
def rle_encode_pure_python(sequence):
"""Run length encode array."""
previous = sequence[0]
count = 1
out = []
for element in sequence[1:]:
if element == previous:
count += 1
else:
out.append((count, previous))
previous = element
count = 1
out.append((count, previous))
return out
def rle_encode_numpy(sequence):
diffs = np.concatenate(( np.array((True, )), np.diff(sequence)!=0))
indices = np.concatenate((np.where(diffs)[0],np.array((sequence.size, ))))
counts = np.diff(indices).astype('uint16')
values = sequence[diffs].astype('uint8')
return np.rec.fromarrays((counts, values),names=('count','value'))
The numpy version is faster, but a lot less legible (and makes some assumptions about data types and array length). Luckily the pure python version can be jitted with numba and is faster than numpy.