IncPy: Automatic memoization for Python
stanford.edu
stanford.edu
There have been a few libraries posted here on HN recently about adding functional elements to python, so apparently it's not just me. Maybe it's time for a "functional fork" of python?
Yet even in Python3 where reduce has been moved around, there's still the spirit of functional programming present. The map, filter, reduce functions all return generators. This is a pretty good improvement as you can now use them on (theoretically) infinite sequences.
Yes, lambda is the lame, dead horse. It's just a syntactic issue. I'm sure most people in the Python world would be happy to receive a multi-line lambda that can return more than just an expression. It's just that no one has been happy with any of the syntax proposals to make it happen.
However, there are other things about lambda in Python that make it difficult to implement as well. But I think baby steps are important.
Is a fork necessary? Well... it would be nice to see some people experimenting with getting lambda to work. However I don't think a fork of the interpreter is necessary just to support a style of programming. Indeed Python prefers one way to do things, but functional programming has proven practical enough I think to be an exception to the rule and so it lives on. Sort of. :)
class memoized(object):
def __init__(self, func):
self.func = func
self.cache = {}
def __call__(self, *args):
try:
return self.cache[args]
except KeyError:
value = self.func(*args)
self.cache[args] = value
return valueyup agreed, but the programmer needs to figure out:
1.) when it's safe to memoize
2.) when it's worthwhile to memoize
also, your memoization decorator doesn't save data to disk. if you wrote a persistent memoizer, then you would need to also track all dependencies for the data you memoized, so that you can know when it's safe to invalidate on-disk cache entries.
IncPy takes care of all of this automatically ;)
If you'd design a declarative systems where you declare your datasets and how they transform into each other, then you could analyze the dependency chain and do the same thing as a library instead of a separate interpreter.
Sprinkle some transparent pickling, hashing and timestamping in there and you get all of the benefits but much more reusable.
Am I underestimating the problem?
That sounds like a lot of effort on the part of the programmer. The author's approach - which I like - is to require as little intervention from the programmer as possible.
Don't confuse the author's research implementation with how it should look in practice. I imagine the author implemented a light-weight interpreter that does the memoization on top of CPython. That's far easier than hacking CPython itself, which gives him a faster path to proof-of-concept implementation and publishing evaluations of the idea. If the research gives good results, then maybe this approach could get implemented as a VM optimization - you wouldn't know it's happening, your programs just run faster.