Glom – Restructured Data for Python
sedimental.org
sedimental.org
I prefer the `get_in()` method from Toolz: http://toolz.readthedocs.io/en/latest/api.html#toolz.dicttoo...
the very slight difference is that using T you must be explicit about attribute access vs key access whereas 'a.b.c' will try both
Here's a example comparison:
glom
glom(target, ('system.planets', ['name']))
# ['earth', 'jupiter']
toolz list(pluck('name', get_in(('system', 'planets'), target)))
# ['earth', 'jupiter']Check the post for more examples of where path-based access falls short of the mark!
(Sidenote: I find it very clunky that get_in() takes a "default" kwarg and _also_ a n"o_default" kwarg. Use a Sentinel object! http://boltons.readthedocs.io/en/latest/typeutils.html#bolto... )
DSL in strings = bad
DSL in native syntax = good
target = {'system': {'planets': [{'name': 'earth', 'moons': 1},
{'name': 'jupiter', 'moons': 69}]}}
glom(target, {'moon_count': ('system.planets', ['moons'], sum)})
# vs
def iter_moons(t):
for planet in target['system']['planets']:
yield planet['moons']
sum(iter_moons(target))
would have to combine with `defaultdict`s if your nested data is only sometimes there though>>> sum([x['moons'] for x in target['system']['planets']])
sum(planet['moons'] for planet in target['system']['planets']) sum(planet.get('moons', 0) for planet in target['system']['planets'])As a simple example, you get some syntax checking.
While mentioned in the post, it's not front-and-center because that's the sort of Python superpower that can look scary to less-experienced devs.
In my experience, it's almost impossible to dissuade people from generating strings if the API affords it. We can't even stop people from generating SQL by concatenation!
Musing: does Python have a way (Mypy?) to declare that a method can accept only constant expressions?
In addition to the string based lookup, it looks like there is an attempt at a pythonic approach:
from glom import T
spec = T['system']['planets'][-1].values()
glom(target, spec)
# ['jupiter', 69]
For me though, while I can understand what is going on, it doesn't feel pythonic.Here's what I would love to see:
from glom import nested
nested(target)['system']['planets'][-1].values()
And I would love (perhaps debatably) for that to be effectively equivalent to: nested(target).system.planets[-1].values()
Possible?--- edit: Ignore the above idea. I thought about this a bit more and the issue that your T object solves is that in my version:
nested(target)['system']
the result is ambiguous. Is this the end of the path query and should return original non-defaulting dict, or the middle of the path query and should return a defaulting dict? Unknown.The T object is a good solution for this.
def get_in(obj, lookup, default=None):
""" Walk obj via __getitem__ for each lookup,
returning the final value of the lookup or default.
"""
tmp = obj
for l in lookup:
try:
tmp = tmp[l]
except (KeyError, IndexError, TypeError):
return default
return tmp
data = {“foo”: {“bar”: [“spam”, “eggs”]}}
# find eggs
get_in(data, [“foo”, “bar”, 1])
By using __getitem__ you naturally work with anything in the python ecosystem.My interest stems from this issue[0] on the Ruby issue tracker to make a symmetrical method to Hash#dig (which does something similar to, but more limited than glom) called Hash#bury. The problem in the issue was that inserting a value at a given index in an array proved difficult and unnatural in Ruby, so I was wondering if there were other solutions out there.
Another question occurs to me - does glom only support string keys?
As for the data insertion, mutation may be in the future, but for now glom only transforms and returns new objects. Definitely something to think about though, bookmarked! :)
https://toolz.readthedocs.io/en/latest/api.html#toolz.dictto... https://toolz.readthedocs.io/en/latest/api.html#toolz.dictto...
I can't say how they compare, but they have some overlapping features.
foo = %{key: [[1, 2], [3, 4], [5, 6]]}
path = [:key, Access.all, Access.at(0)]
get_in foo, path
# => [1, 3, 5]
update_in foo, path, &(&1 * 10)
# => %{key: [[10, 2], [30, 4], [50, 6]]}
foo
|> put_in([:key, Access.all], “foo”)
|> put_in([:new_key], “bar”)
# => %{key: [“foo”, “foo”, “foo”], new_key: “bar”}
That third form is essentially the equivalent of building up a complex object through a series of mutations—but entirely functional.This seems like it usefully solves a problem, but the invocation pattern is suspect to me -- Instead of "glom" taking the target for picking-apart plus a magic little bit of DSL, what if "glom" took a single parameter, the aforementioned DSL, and returned a function that would perform the corresponding search when called on a target? Even if Python or this package optimises away repeatedly searching (by the same spec|in the same manner), the convention the package prescribes is odd to me, right after the first few paragraphs of intro.
The `T` object, which the article describes as its most powerful, can be a useful pattern in some situations, but it's worth pointing out it isn't new or unique to this project.
The author says in another thread here that he first started working on the "stuff leading up to glom" in 2013. One older example, which is virtually identical though less complete, is this Stack Overflow answer I posted in 2012: https://stackoverflow.com/a/9920723/500584
I'd seen the general pattern even before that post, if not the Pythonic syntax. I don't think that it's much of an improvement over defining a `lambda`, so again I would say the thing to focus on is the improved debugability and the simpler, dot-notation-as-generic-attribute-or-item-accessor syntax. I think `T` is largely a distraction, or should be reserved for advanced users.
Libraries like these shine only if they have brilliant tracing and debugging capabilities; otherwise are too easy to reduce to literally a single function.
affordances to add tracing prints, or drop into a pdb at any level
The Inspect specifier type provides a way to get visibility into glom’s evaluation of a specification, enabling debugging of those tricky problems that may arise with unexpected data.
Inspect can be inserted into an existing spec in one of two ways. First, as a wrapper around the spec in question, or second, as an argument-less placeholder wherever a spec could be.
Inspect supports several modes, controlled by keyword arguments. Its default, no-argument mode, simply echos the state of the glom at the point where it appears:
Declarative data transformation generates a lot of comparisons (almost all of them great, though!).
Striking a balance between ease of use / simplicity and powerful features is a tough exercise but you did well.
I can foresee the CLI being quite useful to do away with the run-of-the-mill sed / awk / grep [...] mess. Specifically for the less CLI inclined people out there.
it does not have the exact same feature set though. my focus was mostly on both reading and modifying nested structures in a type safe way.
I have a need to transform between pairs of structures, in both directions, and ever since I found JsonGrammar (https://github.com/MedeaMelana/JsonGrammar2) I've been pining for a Python version.
Lambdas and functions are always a safe fallback, but glom does its best to keep your specs readable and roundtrippable (gotta love a nice repr()).
It would be great to have clarification if this is JSON only, or supports other data structures, or parsers could be plugged in?
The CLI is in a pretty preliminary state, usable but not as robust as it will be in a few weeks. It only supports built-in parsers (JSON and Python literals) What formats are you thinking? YAML?
Does anyone know if something similar exists in Java/Scala land?
Nothing special, glom just reminded me of it.
> "as simple and powerful as glom"
> "big things come in small packages"
> "small API with big functionality"
> "power is only surpassed by its intuitiveness"
> "simplicity is only surpassed by its utility"
> "shortest-named feature may be its most powerful"
For heaven's sake, give it a rest!
It's a big red flag about your priorities that when I go looking for a precise specification, I can't find answers to simple questions and instead end up wading through incessant marketing phrases. I tried, and I finally gave up halfway through the API doc. It might even be the case that glom is a good idea—but you're making it really hard to trust you as a source of objective information about it.
Show, don't tell. My advice to you: you'll generate more interest if you delete every congratulatory word on those pages and focus entirely on helping your readers understand what glom does instead of trying to sell it to them.
How I wish one could publish a dry document and expect people to read all the way to the bottom. I've published enough libraries to know that's not the case. glom's free software so it's all there, as "shown" as can be.
But referring you to the code wouldn't be very considerate either. Instead, here's this literate code version that I prepared in advance. Hopefully this will be of more help to you: http://glom.readthedocs.io/en/latest/faq.html#how-does-glom-...
> After years of research and countless iterations, the glom team landed on this simple construct:
'years of research', 'countless iterations', 'glom team', really?
That said, I'm no liar. Kurt and I (as a team), really did write stuff leading up to glom in 2013 (years ago), and have written stuff like it enough times that I've lost count (countless :P). If this isn't research, I don't know what is. Heck, I'm even getting a fun little peer review!
I didn't find your writing insufferable. I've written tongue in cheek (or over the top) posts about my projects in the past. If they can't see the humor and the usefulness of the project, their loss.
Thank you for creating Glom and thank you for posting it on HN. Count me in as one of your users.
Personally speaking – my heuristic is that maintainers who put effort into marketing copy (even if it's awkwardly exaggerated) are the kind of people who really want their users to enjoy the project, and that often predicts a low-friction experience. Keep doing what you're doing!
see the python requests library documentation for a good example
this fallacy is called affirming the consequent. yes good libraries might not need marketing but that does not say anything about whether good libraries can have marketing.
Come on. kreitz holds the title of Python marketspeak tycoon for a reason. :P
Publicly deriding me as having an irreparable character flaw, based on something I said over ten years ago, which I can't possibly defend or apologize for because I have no idea who you are or what you're referring to—doesn't that seem a little low, though?
It sounds like this is still bothering you after all this time. Please consider reaching out to me (my e-mail address is my HN username at gmail); I'd be glad if we could sort this out in a private conversation. I can't promise that I'll take back what I said without knowing what it was, but I will do my best to understand what you experienced and why it was upsetting to you.