Box: Python dictionaries with recursive dot notation access
github.com
github.com
I think the real reason not to do this is one that as of this writing still hasn't been mentioned yet, which is that when using dot notation, your key names have to fit the Python grammar for identifiers, so you can't have a key "the thing" with a space, or anything else that isn't a Python identifier. So you can't just say "I can access this dictionary with dots", it has to be "I can access some keys of this dictionary with dots, but others I have to access another way". This greatly, greatly reduces its utility. There's a variety of ways of trying to address this; this package appears to try to normalize key names, which is very prone to surprising behaviors and makes it difficult to reason about what will go into what bucket, and also seriously mitigates the virtue of this entire approach because you, the programmer, must also run the normalization algorithm yourself in order to use the dot notation, which rapidly eats away the gains of typing a dot instead of two brackets and two quote characters. Rewriting keys like that is really icky.
The other problem is that the "dot" namespace, as it were, is as others have noted here used for method and property resolution, so you also end up with "dictionaries with random values they can't really contain because they're being used as method names" or "dictionaries that if you put the wrong key in them override a method" or something else like that. Also, consider the ability to subclass these things, making many of the things you might think to do to hack around this break in subclasses.
It's very superficially tempting but it has a looooot of issues that become evident over time. This isn't even a complete list.
I actually created the first iteration of this around 3 years ago, aimed primarily for sysadmins and others who use the console more than actual developing. As I find it extremely convenient to tab complete keys and a lot less typing.
From there it has obviously evolved to try and fit more peoples needs, and I use it regularly in my projects.
The worst thing about abstractions is when they break. I can always figure out how to fix or extend ugly but straightforward code. I may not even know where the manual is when a fancy wrapper dumps a trace.
I read somewhere that if it is written in self-documenting code, the problem and its solution will be obvious.
What counts as "self-documenting"?
Also, I seldom find that string based dicts are what I need. One of the best things about Python is being able to use virtually anything as a key.
I find that to be the thing one of the worst thing ever made into a programming language.
You are facing a program/API/library and you can have absolutely no clue of what it will give or expect without extensive reverse engineering.
Or reading the docs.
In [4]: %timeit d['b'] # dict
10000000 loops, best of 3: 42.7 ns per loop
In [6]: %timeit f.b # class using __slots__
The slowest run took 30.10 times longer than the fastest. This could mean that an intermediate result is being cached.
10000000 loops, best of 3: 44 ns per loop
In [9]: %timeit b.b # class using __dict__
10000000 loops, best of 3: 48.1 ns per loop
You're talking <6ns per access; except in the very exceptional case where you know you need this, in Python of all languages, it's an over optimization. The maintainability of having the stupid-simple class vastly outweighs the speed.In the past I have had to deal with very complex configuration systems; several thousands of lines of XML which were supposed to direct many applications.
And writing:
value['this']['that']['the_other']
vs value.this.that.the_other
Really starts to matter for code legibility. Of course if you change your XML configs your Python will break - but it would likely break in any case.The other case where I have found this applicable is when I have complicated JSON structures in the form of user input. Typically I augment this style by using a JSONSchema that ensures attributes do exist and have defaults (or None).
value['this.that.the_other']
Bonus: Make it support JSONPath and you can use it to extract lots of things at arbitrarily nested levels.Personally I'd prefer a helper functions like
deep_get(dict_, dotted_path[, default]) -> value
You stil have to check the docs/source what exactly happens, but it is just one simple function instead of magic methods. And i can keep using plain dicts everywhere.EDIT: To clarify a bit. An experienced Python developer should immediately recognize a call like `deep_get(value,'this.that.the_other')` as something project/framework specific and not built-in, while `value.this.that.the_other` is ambiguous.
NT = collections.namedtuple('NT', ['x', 'y'])
nt1 = NT(1, 2)
nt2 = NT(nt1, 'asdf')
assert nt2.x.y == 2
get_two = operator.attrgetter('x.y')
assert get_two(nt2) == 2
If you wrap the operator.attrgetter in a try/catch, you can force a default too: https://pastebin.com/bFxGr22E
_empty = object()
def deep_get(dct, dotted_path, default=_empty):
for key in dotted_path.split('.'):
try:
dct = dct[key]
except KeyError:
if default is _empty:
raise
return default
return dct toolz.get(value, ['this', 'that', 'the_other']
There's also a Cython implementation called `cytoolz`.To some extent this retreads ground covered by e.g. XPath, but the specific structure puts it closer to my day-to-day needs.
I've since learned to love the distinction between object attributes and items in a mapping and happily implement ['this style'] accessors without complaint now.
That being said, if the library just did difflib.get_close_matches() on the key lookup that would be neat for some use cases where you want fuzzy keep lookup.
I dislike these structures. Yes, Python lets you pull this trick. Everything, yes, can be represented as a dict, a list, or one of the primitive str/int/etc. types. But I find it's a lot clearer in the long run (and even in the short run), if you have a collection of heterogenous attributes, to define a class for them. Leave dicts (and the subscript notation) for homogenous collections of k/v pairs.
A class gives you the benefit of a type: you get a name, so you can recognize this bag-of-attributes from other, different bags-of-attributes, because it's been given a name. A class also — usually — gets you a list of attributes, and hopefully documentation about what those attribute's types and expected values are.
If you don't like typing (on a keyboard), the attrs package makes things easier[1]; it has the advantage, however, that you get a real class at the end.
The only place I've seen something like Box or Bunch work well is in config files, and even then, only at the uppermost layers of the config (some of the leaves, esp. when you start having a "list of X" in a config — X needs a type).
Python lets you do magic, but with great power and all. IMO, but this is one of those times: Explicit is better than implicit. Simple is better than complex.
Most of the time, this probably gives you exactly what you want, but then there are times where you discover bugs in production because data you assumed is your custom class is a plain old dict, and now you're raising AttributeError all over the place. Another wart is if you are unfortunate enough to have keys that match the name of one of dict's methods, then you have to resort to instance['items'], which defeats the purpose of using this in the first place.
This is a fun trick, but if someone one my team tried to introduce this, it won't make it through code review.
class Obj():
def __init__(self, d):
self.__dict__ = d
d = Obj({
'a': 1,
'b': 2,
})
print(d.a)Although there is this PyCon talk from 2012 advocating not to write classes if there is only one or two methods. [1]
from types import SimpleNamespace
d = SimpleNamespace(a=1, b=2)
print(d.a) my_dict = {
'a': 1,
'b': 2
}
d = SimpleNamespace(**my_dict)
print(d.a) class Box(UserDict):
def __getattr__(self, key):
return self.__getitem__(key)
def __setattr__(self, key, value):
return self.__setitem__(key, value)
I know a few folks who prefer this method of access, so I can't naysay against it too much, but personally I just prefer plain dictionaries.There's no reason to use UserDict, just extend `dict` directly, or implement `MutableMapping` instead. UserDict hasn't been useful since the types/class unification back in… Python 2.3 I think?
Dictionary lookup, attribute lookup and list-index lookups are implemented with a dot notation:
{{ my_dict.key }}
{{ my_object.attribute }}
{{ my_list.0 }}
If a variable resolves to a callable, the
template system will call it with no
arguments and use its result instead of the callable.
Which leads to some interesting and confusing errors if you start iterating over `.items` and you get a callable and not the list you expect. In [17]: a = {"a": 1, "items": {"b": {"c": {}}}}
...: a_box = Box(a)
...: a_box
...:
Out[17]: <Box: {'a': 1, 'items': {'b': {'c': {}}}}>
In [18]: a_box
Out[18]: <Box: {'a': 1, 'items': {'b': {'c': {}}}}>
In [19]: a_box.items
Out[19]: <function items>
In [20]: a_box.a
Out[20]: 1
[0] https://docs.djangoproject.com/en/1.11/topics/templates/#var...EDIT: This came up because our JSON commonly uses `items` as a key for a list of items, which I expect to be at `a_dict['items']`, and it has nothing to do with python's `a_dict.items`.
Pandas must use something similar under the hood to provide dot notation access to columns. I wish h5py did the same for hdf5 objects. In py3, I find myself needing to type list(X.items()) and then list(X['Y'].items()) and so on when I'm exploring a new dataset... fairly awkward for interactive use.
Probably not a huge issue if it's an isolated use-case. Not something I would want to see everywhere in my codebase.
Can you use 'self' as a key?
Ive used a class to provide this kind of dot notation fererencing of hierarchical data. If you could enclose the keys in quotes I might feel better about it but that's probably not possible.
Without looking at the code: yes. `self` in python is just convention.
Just yesterday I commented that if generators could refer to themselves in generator comprehension (recursion) I could express some algorithm as a single expression.
http://stackoverflow.com/a/41617394/180464
... anyway, this just turns out to be what looks like an "addict" clone.
It works similarity to addict, but does have important distinctions and IMO seniority (was in my 'reusables' pypi project named 'Namespace' before addict existed. Just finally spun it into it's own project).
Biggest differences:
* Box will convert items added to the object after creation, addict does not
* addict only acts as a defaultdict, Box can act as either regular or defaultdict
* Box updates it’s __dir__ so that attributes (keys) are tab completed in stuff like IPython
* Box repr clearly shows it is an object
* De/serializes JSON and YAML
* De/mangles, de/encodes keys
* Provides automatic, expensive hashcode
* Blacklists/transforms a bunch of likely keys because they conflict with reserved words
* Overlays attrs (__box_heritage)
* All in pretty complex code that obfuscates what you're really doing (especially to a maintainer)
For the ability to avoid importing json/PyYAML and use clear key lookups? The author must really like JavaScript syntax, or something, because I'm not seeing the point in this layer at all. It's the sort of library that you discover your inherited codebase is using, and just say, "fuckfuckfuck...", because it implies that the author cares more about pushing round pegs into square holes than writing the freaking application code; it reeks of inexperience.
I think the advantage over json/PyYAML is not having to supply conversion functions for Decimal, datetime, or user defined datatypes.
I like Python a lot, and I don't write much Javascript, but one thing I wish I could do in Python is the dot notation from a dictionary. I sometimes used namedtuple as a cheap (but "immutable") class, so I can simply use dot notation when I am passing my object around my functions, instead of always stuffing the data into a dictionary.
Why though?
foo['name']['attr1']['attr2']['morefuckingattr'] vs foo.name.atr1.attr2.morefuckingattr
More of a personal preference.
It'll enforce schema and destructure the nested dict/list that results from JSON decoding.
Depends on the usage. Sometimes you leave them as nested dictionaries and lists. Sometimes you transform them into tabular data. Sometimes you aggregate, etc.
I think this library has it's uses. I wouldn't use it uncritically but I wouldn't dismiss it either.
If that were true, there would be no type systems.
A) Giving off a signal that you're dealing with an object rather than a dict.
B) Making it cumbersome to swap out some of those selectors with variables.
C) Making it difficult to deal with the attribute not being there (with a dict you can say .get("attr2", {}) and it returns a default.
Quite annoying and I know some people considered what I wish is a bad practice.
x.y == x.__getattr__("y")
x["y"] == x.__getitem__("y")
assignment == set{attr,item}
del x.y, del x["y"], etc
len() and slices of items, not attrs
etc, etc. If you start messing with those semantics, it can become very confusing very quickly in Python and you probably want a data type. Remember that Python gets nervous about cleverness.One way I've approached your problem before is a recursive helper where I borrowed some concepts from jq:
val = fetch(result_dict, "foo.bar.baz.plonk[0]")
Because then you can also isolate the missing key handling and all that stuff into your helper, rather than changing the semantics of an untyped dict. Aside from typing the responses from APIs -- the way better option -- I found this a reasonably Pythonic approach toward dealing with the annoyance.a['b'] = 'c' # means i must use a dict a.b = 'c' # means must be an obj
if they were interchangeable then I could have more reusable code.
Many python REST API client libraries end up rolling their own one-off dict -> "dot dict" transformation helpers, for better or worse.
This sort of thing can be great for interactive use-cases, where terseness really helps (notebooks, REPL's etc). I'd definitely use something like this, there. But I would share the same reservations as you in using this library in a serious project.
Ugh, this really hits home.
I just started a new day-job where I'm essentially being asked to un-fuck a python codebase. Your statement neatly captures my feelings when I discovered the presence of hacked-together singleton classes littered throughout the thing.
Good luck unit-testing that...
box.might_exist.might_exist.might_exist.desired_key
where desired_key would return a default value if one of the keys doesn't exist?I find the bulk of my dict code is checking for keys before access, or using .get('', default) recursively. Gets really hairy for deeply nested dicts.
def nget(d, *ks, **kwargs):
for k in ks:
d = d.get(k)
if d is None:
return kwargs.get('default')
return d
>>> d = {'a': {'b': {'c': 12 }}}
>>> nget(d, 'a', 'd', 'c', default='Not Found!')
'Not Found!'
>>> nget(d, 'a', 'b', 'c')
12A better implementation would be:
def rget(d, *ks, **kwargs):
for k in ks:
if k not in d:
return kwargs.get('default')
d = d[k]
return dWrite a function. `def my_deep_get(dict, keys)`
Basically it's faster on standard soft creation as it converts on lookup. (And most use cases don't call for referencing every single key)
https://github.com/rcarmo/python-utils/blob/master/core.py#L...
...but this takes that notion much further.
class Box(dict):
def __init__(self, **kwargs)
super(Box, self).__init__(kwargs)
self.__dict__.update(kwargs)For all its faults, JavaScript has some pretty sensible defaults. Accessing a missing key returns what is basically its equivalent of nil. Extra parameters to function calls are simply ignored. Unsupplied parameters default to nil.