PyCon 2011: How Dropbox Did It and How Python Helped
pycon.blip.tv
pycon.blip.tv
Also, "Optimizing CPU is easy, Optimizing Memory is Hard"
From my understanding, the any() version is faster is primarily because it avoids looking up .update() method each iteration. The usual Python way of optimizing such cases is:
# the original slow version
# 2.87753200531 sec
def run():
_md5 = hashlib.md5()
for i in itertools.repeat("foo", 10000000):
_md5.update(i)
# slight change
# 1.82029104233 sec
def run():
_md5 = hashlib.md5()
update = _md5.update
for i in itertools.repeat("foo", 10000000):
update(i)
# using any()
# 1.50683498383 sec
def run():
_md5 = hashlib.md5()
any(itertools.imap(_md5.update, itertools.repeat("foo", 10000000)))- imap returns an iterator, any is used to "force" the iterator run through all elements.
- the loop runs on C because of any. Instead of a for loop where jump and body are repeatedly evaluated by the interpreter, any consumes the iterator through a loop implemented in C.
http://j4.video3.blip.tv/0370005188920/Pycon-PyCon2011HowDro...
Once they integrate a way to sync google docs & docs.com then they are done!