Python Tips and Traps
airpair.com
airpair.com
>>> from collections import namedtuple
>>> LightObject = namedtuple('LightObject', ['shortname', 'otherprop'])
>>> m = LightObject()
Traceback (most recent call last):
File "<stdin>", line 1, in <module>
TypeError: __new__() takes exactly 3 arguments (1 given)
>>> m = LightObject("first", "second")
>>> m.shortname
'first'
>>> m.shortname = 'athing'
Traceback (most recent call last):
File "<stdin>", line 1, in <module>
AttributeError: can't set attribute
Also, bare try/excepts as in: try:
# get API data
data = db.find(id='foo') # may raise exception
except:
# log the failure and bail out
log.warn("Could not retrieve FOO")
return
are really bad. The failure might be caused by a ^C or MemoryError, or even a SystemExit, should db.find() desire to do that.Instead, qualify it by catching Exception:
try:
# get API data
data = db.find(id='foo') # may raise exception
except Exception:
# log the failure and bail out
log.warn("Could not retrieve FOO")
return
It's also poor form to "return True" in the exit method of the context manager. If there is no exception then that's not needed at all, and if there is an exception ... well, that code will swallow AttributeError and NameError and ZeroDivisionError, and leave people confused as to the source of the error.E.g. when I want to carry a foo, a bar, and a baz around, but namedtuple isn't right (e.g. I need mutation). I'd prefer
config = Someclass()
config.foo = 1
config.bar = "quux"
config.baz = 3.3
and then using config.bar, etc.over:
config = { 'foo': 1, 'bar': "quux", 'baz': 3.3 }
and then using config['bar']However, I don't know the larger context. If this were at the top-level of a web services handler where a "return None" indicates a 400 - Internal Server Error, then logging the database failure and stopping is likely acceptable.
Even then, I wouldn't call it good code. However, the goal of this essay seemed to be to give the minimal example, and a more complete example would have required introducing a fake database module with its own exception type. I believe that would have obscured the intent.
I would have preferred real, working code. In this case, with sqlite3. That's fundamentally a pedagogical choice though.
[0] https://docs.python.org/3.1/library/collections.html#collect...
You can start by converting all tabs into 8 spaces. This can be tricky should some strings have tabs. That's a bad idea in the first place. Use "\t".
Don't mix tabs and spaces to get the same indentation level. Python 3 prohibits it. With Python 2 use "-t" or "-tt", which respectively warns and raises an exception if both spaces and tabs are used in the same block.
>> nested_dd = collections.defaultdict(lambda: nested_dd)
>> nested_dd['a']['b']['c']['d'] = 'hello' >>> NestedDD = lambda: collections.defaultdict(NestedDD)
>>> nested_dd = NestedDD()
>>> nested_dd['a']['b']['c']['d'] = 'hello'
does the proper thing. >>> nested_dd['d']
'hello' from collections import defaultdict
defaultobj = lambda: type('defaultobj', (defaultdict,), {
'__getattr__': lambda self, x: self.__getitem__(x),
'__setattr__': lambda self, x, v: self.__setitem__(x, v)
})(defaultobj)
names = defaultobj()
names.mammalia.primates.homo['H. Sapiens'] = 'Human being'
print(names)
The same but also allows dot notation.Also, I found the name and the tagline to be distasteful.
[1] https://github.com/mewwts/addict#addict---the-python-dict-th...
At least you won't forget.
pip install addictIt is fragile design to write code that depends on 1) the data is so large that you don't have room for a set() BUT 2) it is small enough for an in-memory sort. (IOW, the almost-out-of-memory case invariably degrades over time to flat-out-of-memory).
Another thought: people seem to place too much concern about about the size of various data structures rather than thinking about the data itself. Python containers don't contain anything, they just hold references. (Usually, the weight of a bucket of water is mostly the water, not the bucket itself).
Finally, if your task is to dedup a lot of data, it doesn't make sense to read it all into memory in the first place (which you would need for a sorting approach). It is better dedup it using a set as you read in the data:
# Only the unique lines are ever kept in memory
with open('hugefile.txt') as f:
uniq_lines = set(f)
Dude, sorry to go off like this, but the advice you gave is almost always the wrong way to do it.> rather than thinking about the data itself
…i.e, in this case, we need to realize that the set will require on the order of the size of how many unique items there are, and that in the worst case, that's O(n).
Of course, if you can guarantee that the size of unique set will fit into memory, then you're fine. But you might need a bit of knowledge about the data to do so.
But of course that will be unavoidable if you want fast deduplication. You could do fully lazy deduplication using a generator, but you'd have to avoid using sets.
A set by definition contains only unique items, that is all items in a set are different from each other.
What I meant is you could in theory process a generator and omit duplicates without any real memory usage (even with a file of millions of lines) by chaining generators together. This would be slower than a set but much more memory efficient.
ETA: See raymondh's post below.
* The with-statement only causes the file to be closed after use. It has nothing to do with iteration.
* Files themselves are self-iterable (anything that loops over a file object uses a line-at-a-time iterator). Even writing "for line in f: ..." causes you to loop a line at a time without the whole file being in memory all at once.
* Sets are just one of many objects that takes an iterable as an argument: min(f), max(f), list(f), tuple(f), etc.
I'm not sure I understand your third comment though. As I understand it, iterables have an __iter__ method. __iter__ methods return iterators. So an iterator for the iterable f traverses f and sends f's values to, in this case, set(). I believe we agree there. My initial concern was simply that the statement could be read "The iterator, set(), uses f.", and I don't see where the disagreement arises.
s = set()
for x in f: s.add(x)
Does not read everything into memory - by using iterators, it only reads one line at a time.