Specific ways to write better Python (2017)
github.com
github.com
Its been updated to be exclusive to 3.x, including 3.8 samples.
[1] https://learning.oreilly.com/videos/effective-python/9780134...
I blogged about this: https://opensourceconnections.com/blog/2020/06/10/python-gen...
So you can’t just replace lists with generators.
This break the uniform access principal. Client code IMO shouldn’t need to think about whether they got a generator or an interable container.
How would they not? The caller provides the file as an argument to the generator function. Further, the caller decides when to close the file, presumably via leaving a context and after the generator is consumed.
with open(...) as f:
lines = process(f)
# do stuff with lines
^ if `process` returns a generator, the file will be closed without having done anything.Do you `list(...)` literally everywhere you want a list, whether or not it's currently a list? Extreme-defensive-programming like that tends to be extremely rare in my experience. Not nonexistent, but I'd be willing to bet that if you took a random stack overflow or github line of code, it wouldn't do this.
But, there are plenty of instances of what you're talking about. Take a look at the standard library. Itertools often creates an iterator as the first operation, not knowing whether it received an iterator or some other kind of iterable.
https://github.com/python/cpython/blob/384621c42f9102e31ba2c...
If you search for `list` in the CPython code, you'll find many examples of what you're asking about.
Similarly, Pandas often makes a copy of a Series or DataFrame. It might be inefficient, but it keeps the API (more) simple.
Rather, if you need a list and the otherwise perfect library method `foo` returns an iterator then you as the caller have almost no downside in writing `list(foo())` instead.
On the flip-side, if you need an iterator (e.g. because of memory concerns) and the otherwise perfect library method `foo` returns a list then you as the caller have no options other than re-implementing `foo` or finding another tool to solve your problem.
If your project usually requires lists rather than iterators then the syntactic burden of wrapping everything in `list(...)` could be annoying, and I wouldn't be at all surprised to find that performance-critical code couldn't tolerate the iterator overhead (though I'd posit that most of the time just using a list instead probably wouldn't fix it), but using iterators rather than lists seems like a good default if you don't have a good reason not to.
E.g. lists are strictly more flexible than iterators since they are multi-pass and can index. If you take a return from lib func A and pass it to lib funcs B and C (possibly in a different lib), if A returns an iterator then you need to know what B and C are going to do with it. If A returns a list, you do not - B and C cannot affect each other in this way unless they directly mutate the argument, which is relatively rare and usually has very strong documentation and/or naming patterns to make it not surprising. And B and C and all their calls can change without you needing to change your code between them and A.
You can of course `list(A())` before passing it to B and C, but that's arguably defensive programming unless you know A returns an iterator.
tbh I'd say I see multi-iteration and indexing several times more often than I see iterator-only use in most code, and almost never see memory issues except in large or embedded systems. Gigabytes of RAM are standard now, and dumping a few thousand files into memory before processing will quite often out-perform a more memory-efficient streaming approach. In the rare cases where you do exceed those memory bounds, you're likely in a niche area (gigabytes of files, tiny memory space, significant computation cost) and are likely using niche-oriented libraries that benefit more from coordinating tightly than from being perfectly general.
In contrast, I see them all the time (damn you, Pandas). Beyond out-of-memory errors, generators can be vastly more compute-efficient because of the CPU cache.
How would you handle the problem? Or do you not see it as a problem?
with judgments_open('judgments.txt') as judgments
gather_features(judgments)
train_model(judgments) # <-- bug.
The context manager is still returning a single generator, not a list, so when you call train_model(judgments) it will have been exhausted, no?- - - -
Second, the bug is not really solved by the proposed solution.
This code is fine.
with open('judgments.txt') as f:
judgments = judgments_from_file(f)
for j in judgments:
process(j)
This code is buggy. with open('judgments.txt') as f:
judgments = judgments_from_file(f)
for j in judgments: # <-- bug.
process(j)
The generator must not be used after the file is closed at the end of the context manager's scope.If you did this you would still have the bug:
with judgments_open('judgments.txt') as judgments
foo(judgments) # okay
for j in judgments: # <-- same bug.
process(j)
In any event, in this case, I would just do: with open('judgments.txt') as f:
judgments = list(judgments_from_file(f))
gather_features(judgments)
train_model(judgments)On saving off the judgements var, I think I disagree that’s as much an issue. “With” is the languages feature explicitly designated for safely managing resources/variable lifetimes. One could save a file off created using with in the same way. The language basically tells you doing that with a var created with “with” is a bad idea. The articles solution makes lifetime of the generator more explicit.
Does it? What's the difference between these two snippets?
with judgments_open('judgments.txt') as judgments
gather_features(judgments)
with open('judgments.txt') as f:
gather_features(judgments_from_file(f))
You still have to complete all the work within the scope of the context manager either way, right?The code ehsankia presented combines the generator with a closure that includes the context manager, and automatically carries the scope of the context manager along with it, so that when the generator is exhausted the context manager's scope also ends.
def judgments_from_file(filename):
with open(filename) as f:
for line in f:
yield parse_row(line)
Why wouldn't you let the method that deals with the file also own the file pointer, and then the context manager from the with-statement close it whenever the function is done? Instead of having to do some magical stuff to make sure it closes at the right time? Is it to make unittest/mocking easier? for judgements in judgments_from_file(filename):
for j in judgements:
process(j)The only reason I can think of to do it the other way is if you want to make unit testing easier, allowing you to pass a fake iterable instead of a file pointer to `judgments_from_file` for testing.
What you’re describing is more or less what judgments_open does. I just did it with a context manager. But you’re right there are other ways to safely manage the files lifetime.
I also use this pattern for writing judgments, which has work to do when done judgments are all collected and the context exits.
Taken to the extreme, the "fit as much logic as possible into a single screen" approach leads naturally to code that looks a lot like something written in APL or one its descendants (J, K, etc.). I recognize this is not for everyone.
Personally, I think Jeremy Howard and the folks at fast.ai have struck a pretty good balance of readability and succinctness with their coding style for ML/AI code:
If you really need to pack a lot of math, just define it in a module (e.g. tensorflow does that nicely).
I've had an intern who came from MATLAB. Convincing them to stop packing numerical constants in the middle of every expression was unexpectedly difficult. "But it's shorter that way!"
I have used code from the fast.ai codebase and personally I find it horrendous to work with - I find the way they structure and name classes consfusing, they use wildcard imports everywhere, they use tiny variable names for everything - it's all extremely reader-unfriendly, which for an educational tool seems completely bizarre.
Not only that, but you can assign to a slice with a stride:
In [1]: r = list(range(10))
In [2]: r
Out[2]: [0, 1, 2, 3, 4, 5, 6, 7, 8, 9]
In [3]: r[1::3] = 10, 20, 30
In [4]: r
Out[4]: [0, 10, 2, 3, 20, 5, 6, 30, 8, 9]Is it even true that comprehensions are faster than reduce and map? I write a lot of python with reduce and map, and I know it's not considered a best practice for perf reasons, but whenever I go-ahead and rewrite it in the "more pythonic" way, I can't help but think that it ends up a lot less readable. I feel like this divide makes functional programming in python more of a hassle than it should be.
#8 is a little unclear as to which expressions it is counting (particularly, whether it counts the return expression, which I don't think it intends to.) I think it's referring to the two total “for” and “if” clauses, so either two “for” or one of each as the preferred limit, which seems to me to be a sensible guideline.
It think both #7 and #8, even with the additional clarification, need so to be considered in combination with #4: they set the rules as to what constitutes a complicated expression—either with comprehensions or map/filter/etc.—that calls for factoring part of it out into a helper function or named subexpression to avoid overly complex, code-golfy one-liners.
> Is it even true that comprehensions are faster than reduce and map?
While I'd love to see “accumulator comprehension” syntax* added to Python, in real current python comprehensions aren’t generally an alternative to reduce.
* something like:
(compute x from 0 as x+n for n in ns)