Python Default Dict
docs.python.org
docs.python.org
state2name2visited = defaultdict(lambda: defaultdict(list))
state2name2visited[“PA”][“Joe].append(“Pittsburgh”) def tree():
return defaultdict(tree)
>>> t = tree()
>>> t['a']['b']['c'] = 10
>>> t
defaultdict(<function tree at 0x10c40df28>, {'a':
defaultdict(<function tree at 0x10c40df28>, {'b':
defaultdict(<function tree at 0x10c40df28>, {'c': 10})})}) $M{hello}[4]{world}[0]++
gives you: {
'hello' => [
undef,
undef,
undef,
undef,
{
'world' => [
1
]
}
]
} trie_struct: Callable = lambda: defaultdict(trie_struct)
trie = trie_struct()
for word in words:
ref = trie
for char in word:
ref = ref[char]
It won't work with words that substring other words, but its interesting. >>> state2name2visited = {}
>>> state2name2visited.setdefault("PA", {}).setdefault("Joe", []).append("Pittsburgh")
>>> state2name2visited
{'PA': {'Joe': ['Pittsburgh']}}Also setdefault causes confusion for less experienced users in a way that the defaultdict format does not.
I don't see how you can declare what the most common use of a standard function is, so, uh, citation needed.
> But what if you want to use it to count the order in which we see keys
Don't bother. Since python 3.7 (really 3.6) you can just look at the key order after adding your strings to a normal dict comprehension because dicts are ordered now.
{s: 0 for s in strings}.keys() will give you the order. If you really want a number associated with each, you can wrap that in enumerate.
In a casual discussion you don't formally "declare", you just state something based on your observations and the patterns you see in the wild, that might or might not be true in general.
Similarly, if it's a casual discussion, one can just read "the most common use case" as "a common use case", and just move on to the substance of the parent comment.
Counters are for counting, not for determining order of discovery.
Yes but the question as framed ("first failed lookup initilizes to 1, the second to 2, etc") does not require counting the occurrences anymore, so using a counter does more work than necessary. If you want both counts _and_ order, then, yeah, counter is great.
> The most common use case for default dict is counting the number of occurrences of some string
I would assume very few people use defaultdict for this since Counter exists.
d = defaultdict(lambda: len(d))
However, it's probably better to avoid defaultdict altogether and implement __missing__ in a subclass of dict.```
import itertools, collections
cnt = itertools.count(1)
d = collections.defaultdict(lambda: next(cnt))
s = "abcad"
[d[c] for c in s]
```
d = defaultdict(count().__next__) days = {}
day = '2021-05-05'
if day not in days:
days[day] = []
days[day].append(event)
vs. just days = defaultdict(list)
day = '2021-05-05'
days[day].append(event)
what a blessing! Thanks for posting this. I do this so often.Now maybe add an `OrderedDefaultDict`.
This is the TXR Lisp interactive listener of TXR 257.
Quit with :quit or Ctrl-D on an empty line. Ctrl-X ? for cheatsheet.
Caution: objects in heap are farther from reality than they appear.
1> (defun pend (list item) (append list (list item)))
pend
2> (define-modify-macro pendto (item) pend)
pendto
3> (defvar x)
x
4> (pendto x 3)
(3)
5> (pendto x 4)
(3 4)
6> (defvar days (hash))
days
7> (pendto [days "2021-05-01"] 'event0)
(event0)
8> (pendto [days "2021-05-01"] 'event1)
(event0 event1)
9> (pendto [days "2021-05-01"] 'event2)
(event0 event1 event2)
10> days
#H(() ("2021-05-01" (event0 event1 event2)))
11> [days "2021-05-01"]
(event0 event1 event2)[0]: https://docs.python.org/3.6/whatsnew/3.6.html#new-dict-imple...
-if day not in days:
- days[day] = []
+days.setdefault(day, []) (push event (gethash day days))
does this with zero extra effort, since the default value for fetching from a hashtable is NIL. The magic of reasonable defaults... I feel like it's underappreciated sometimes.Another way of avoiding KeyError is using
dict.get(val, default_val)
I find it a bit cleaner, since you don't have to create a function or import collection.[1] https://docs.python.org/3/library/stdtypes.html?highlight=di...
Except that
my_dict = {}
my_dict[val] = my_dict.get(val, []).append('foo')
doesn't work because list.append() doesn't return anything, while my_dict = defaultdict(list)
my_dict[val].append('foo')
does the right thing.Defaultdict is sugar for dict.setdefault, not dict.get.
(my_dict[val] := my_dict.get(val, [])).append('foo')1) the walrus operator is stupid
2) the walrus operator is dangerous
3) `(my_dict[val] := my_dict.get(val, [])).append` is the opposite of simpler compared to `my_dict.setdefault(val, []).append` or (cleanestly) just `my_default_dict.append`
dict.setdefault(val, default_val)
[1] https://docs.python.org/3/library/stdtypes.html#dict.setdefa... d['foo']
whereas defaultdict handles the default value for you correctly independently of the access method.IMO defaultdict is ideal for this use case.
I use it quite a bit for metadata structures with optional elements so reads will still give something back, even if an empty string or a default value.
if `a = {'a': 1, 'b': 2}`, and I do `cval = a.get('c', 3)`, `cval` will contain 3.
If I do `cval = a['c']`, I get an exception
The issue here is that the clients have to know the particular default values for each parameter. For me, it was easier to return a defaultdict where the callers didn’t need to know what default value to pass into their gets.
I generally prefer to use .setdefault(key, default_value) with regular dicts, as it's much more explicit. If I do use a defaultdict for convenience, I will usually only use it within a limited scope, and if I'm returning it then I'll cast it back to a normal dict to avoid surprising the caller.
A defaultdict is a normal dict[0], its just a space-efficient way of expressing a large normal dict, most of whose keys won’t be accessed.
The default function itself should throw KeyError on values that are logically not in the dict (including, due to Python’s dynamically typed nature, those which are outside of the key domain because of type.) Though in some uses you can skip out on this because its used in a very lonited scope where you know its not going to be indexed improperly.
> I generally prefer to use .setdefault(key, default_value) with regular dicts, as it's much more explicit.
I’m not sure why one would prefer one of those over the other, as they have very different use cases; certainly defaultdict isn’t a great choice for places where .setdefault makes sense, but that’s true in reverse, too.
[0] well, except for the unfortunate .get() behavior; a ReallyDefaultDict where rdd.get(k, default) works more sensibly, returning rdd[k] unless that throws KeyError, and default otherwise, would be better.
That's not how `defaultdict` works - the key isn't passed to the default factory, so there's no opportunity for it to raise a `KeyError` if the key doesn't "logically belong" in the dict. It's possible to get that sort of behaviour by overriding `__missing__`, but I very rarely see this sort of thing in the wild.
A more typical use case is something like `defaultdict(list)` as a convenient way to build a dict of lists. This is fine within a limited scope where it's obvious to the reader that they are dealing with a defaultdict that has special `__getitem__` semantics, however it's a bad idea to return a defaultdict to a caller who might be expecting a normal dict, and would be surprised that missing keys don't result in KeyErrors.
With `.setdefault(key, default_value)` it's unambiguous what we're trying to achieve - the reader doesn't need to know whether they are dealing with a defaultdict.
You're right. I was thinking of the how Ruby Hash-with-default works (I forget that you have to override __missing__ to get that with Python because .get() not working consistently with __getitem__ leads me to avoid defaultdict in practice.)
> With `.setdefault(key, default_value)` it's unambiguous what we're trying to achieve
Sure, but that’s only useful where you know the key that is going ro be accessed, which isn’t the use case for passing a “dictionary that can generate an appropriate default for keys that aren’t stored” either up or down a call chain. So while I agree that setdefault is ideal foe that use case, I don’t think it substitutes for a well-designed dictionary-with-default-generator. (Unfortunately, defaultdict doesn’t either, but its at least-usable-but-dangerous in that role, whereas setdefault() fundamentally can’t fill it since it involves situations where thr access is nonlocal to the code that wants the special handling.)
``` In [22]: foo = defaultdict(lambda k: k)
In [23]: foo[1] --------------------------------------------------------------------------- TypeError Traceback (most recent call last) <ipython-input-23-5a39798fc62e> in <module> ----> 1 foo[1]
TypeError: <lambda>() missing 1 required positional argument: 'k' ```
I could envision this being useful for using a defaultdict as something like a local cache for something where requests are expensive. Below is how one might want to use it but it'll fail because the lambda is not given the key argument.
``` profiles = defaultdict(lambda p_id: get_profile_from_site(p_id)) ```