How to Use JSON Path
bump.sh
bump.sh
In addition to simple queries that allow you to select one (or multiple) matching nodes, it also provides some helper functions, such as arithmetic, comparisons, sorting, grouping, datetime manipulation, and aggregation (e.g. sum, max, min).
It's written in JS and can be used in Node or in the browser, and there's also a Python wrapper: https://pypi.org/project/pyjsonata/
I even had to click around the documentation link three times until I figured out how to get to the real docs. It's a really well thought out query language that is actually Turing complete. I really want to try it out on some real data in anger to see how well it holds up in practice. I'm also thinking about what could be taken over into PRQL which I'm somewhat involved in.
We see this kind of Nete’s data all over the place (json, yaml, python dictionaries, toml, etc, etc) and I’m thinking wouldn’t it be nice if we had a path language that worked across these structures, just like how we can regex any strings?
So we can have a pathql executable that we can feed yaml and json data to, but I can reuse the query in Python when I want to extract values from a json stream I just deserialized.
I think you could well apply jsonpath to yaml, except for the different data types, which is what makes you need xmlpath, jsonpath, file path, css and so on. If you're willing to do some automagic conversions, you could probably do that right now, if you write the code for it.
They vary in syntax and how they deal with scalars, objects and collections. Having to write 'myKey' is a bit unfortunate in json for example. Still, for any document larger than 20 lines yaml will fall apart easily.
Xml (dialects) mark the beginning and end of nodes explicitly, which deviates from json/yaml. Xml can represent nodes within the value, which is impossible for json/yaml. To convert such xml value to the latter, you have to break up such a node in fragments and represent them as a collection in json/yaml.
Are you talking about XML like `<text>Something something <para> inside </para> something else</text>`? I thought this would also presented as the text element having three children, the text "Something something", the <para> tag with its subtree, and the text "something else". Am I misremembering?
It's just when writing the document that you get the advantage, the data model is the same.
I think that is one reason why json took of for exchanging machine readable data, xml is too expressive, and leans towards document authoring.
Json got introduced because XMLHttpRequest was an IE-only component. Data exchange was done mostly between servers, back then the front-end was really dumb. Why xml got slowly replaced is because that for machine generated data, json proved to be sufficient. Any json library is easier to work with then their xml counterpart, exactly because of what we discuss here.
(Looking at you, jq, though to be fair JSON itself supports null...)
If you want this kind of semantics with JSON, take a look at https://www.jsoniq.org
torepr didn't quite work for me as I was dealing with objects containing large binary blobs and it was awkward.
fq is a great tool and I shouldn't have suggested this was a problem unique to it! I think this kind of "issue" is inevitable when dealing with so many types of input. And to be honest I struggle hard using jq as well for anything other than very basic paths, due to infrequent usage.
About blobs, if you want to change how (possibly large) binaries are represented as JSON you can use the bits_format options, see https://github.com/wader/fq/blob/master/doc/usage.md#options, so fq -o bits_format=md5 torepr ...
I can highly recommend to learn jq, it's what makes fq really useful, and as a bonus you will learn jq in general! :)
Speaking of this topic, I would be very interested in a CLI tool that translated between different dialects of regex. As opposed to general data or configuration, I find the case for cross format conversion very compelling for regexes; the objects that different regex languages deal with are functionally completely identical. I would be very happy to be able to craft a regex to search for something in vim, then convert the regex to grep or use another tool on the shell, then perhaps adapt it to some other scripting languages e.g. Javascript/Ruby/Perl.
Indeed I am sad that I am not finding any CLI tools (non web-based) for this beyond https://github.com/Anadian/regex-translator . Am hopeful for more suggestions.
But one problem is, that each format has slightly different ways it works. Some have nodes which support properties, some not. This makes building a proper query-language a bit more complicated if it should not end up ugly.
Yes, if you want to make the query language reject as incorrect queries that specify a property on a node that can't have properties, it is messy. But is this really necessary? You can write faulty queries anyway. And it's not a programming language for complex systems, where it really helps to prevent mistakes early. SQL works fine with no type safety and so on.
You generally don't see yaml or XML be called that, but there's some information on the net that you can find about it.
object graph - grove plus object references. eg serialization formats
knowledge graph - object graph where parent-child relations are explicit, using subject-verb-object clauses. like how Java Spring "flattens" object graphs.
document - grove plus random text nodes
- Integers (including 64-bit integers and longer)
- Non-finite floating point (Infinity, NaN)
- Keys of types other than strings
- Non-Unicode strings (e.g. byte sequences, TRON code, etc)
- Date/time
- Links
Additionally, they may differ of whether or not the order of keys should be retained.
The approach is to take the JavaScript object, convert it to XML DOM, run the query (either using standard XPath, or standard CSS selectors) and then either convert the DOM back into objects, or another way I've seen it done is to keep a register of the original objects and retrieve the original objects.
In this way, JSON, and any JavaScript object with non-circularity can be sifted and searched and filtered in reliable ways using already-standardized methods just by using those technologies together in a fun new way.
There is not necessarily a need for inventing a new custom syntax/DSL for querying unless you don't want to make use of CSS and XPath, or have very specific needs.
I wrote up my own detailed notes on that subset a while ago: https://til.simonwillison.net/sqlite/json-extract-path
In dev mode, our internal APIs return pretty printed JSON so one can inspect via view-source, more, or text editor.
https://news.ycombinator.com/item?id=40246089
(I'm being sarcastic, obviously. You are 100% right)
JSONPath can only pull data out, like XPath. jq can do much more, like perform transformations.
jq is also more concise:
.book[0].title
versus JSONPath: $..book[0].title
Here's a discussion with more comparisons: https://github.com/serverlessworkflow/specification/issues/2...This is about using JSONPath for OpenAPI Overlays and automated API Style Guides like Spectral.
You cannot use jq for either of those things.
Notbing against jq, just a different discussion.
There are a lot of factors at play so we can’t quite put our thumb on JSONPath, but it’s the current suspect and curious if others have run into anything similar.
Contrived example:
@localizedModel({
title: {},
quiz: { jsonPath: '$.[question,answerMd]' },
})
class MyQuiz {
title: string;
quiz: JSONObject;
}2. Ask [newest LLM] to write the proper json path to get to the element you want to reach.
1. Creat tree-structure document format that is flexible enough to handle all use cases.
2. Write a ton of content in this format.
3. Have to figure out a query pattern to accurately retrieve good info out of these structures.
Generally, I feel we’ve become good at querying normalized table data. But—-and maybe it’s just me being stupid—-wending through tree-structured data is still tricky. And I recently discovered LLMs are great at solving for it, if you ask clearly.
https://github.com/Tade0/permit-a38/tree/master
Ultimately the linting rules proved to be easier to write than read.
You can take my list comprehensions from my cold, dead hands.
There isn't a use case I've seen where these types of mini languages fit well. Ostensibly, you could give it to a user to write to query JSON in a domain-agnostic way in an app but I think it would just confuse most users as well as not being powerful enough for half of their use cases.
Sometimes it's better just to write code.
In some cases advantage is that you don't create new code and you just use some relatively standard tool. You just fetch some public package that handles various edge cases and you just prepare script that describes what you want to do with some program. This is useful, if you work in containerized environment and configuration exists as json or yaml. Often I just use jq or yq, instead of reinventing wheel to just read or write some values.
But there is a reason excel has a ceiling of maintainability that always turns it into a spaghetti mess once it's big enough.
$.store.book[?@.price < 10].title
Python: [x['title'] for x in data['store']['book'] if x['price'] < 10]
Javascript: data.store.book.filter(x=>x.price < 10).map(x=>x.title) data['store']['book'].filter(lambda x: x.price < 10).map(lambda x: x.title)
Yeah, not nearly as good. The syntax sugar of a.b instead of a['b'] and arrow functions instead of lambda really does make a pretty big difference. data.store.book.filter(_.price < 10).map(_.title)
Every language should just adapt it. data.store.book.filter(_.price < 10).map(_.title)
would be written as data.store.book.grep(*.price < 10).map(*.title) let possibleNow =
people
|> List.distinctBy _.Name
|> List.groupBy _.Age
|> List.map snd
|> List.map _.Head.Name
|> List.sortBy _.ToString() filter((x) => x.price < 10)
but why can’t we just write filter(x.price < 10)
and add a rule to the JS engine that says “when you encounter a ‘syntax error: undeclared identifier x’, rewrite the code to add `(x) => ` in front of where the syntax error occurred, if and only if this rewrite prevents the syntax error”.You might protest that reacting to syntax errors by inserting extra code and checking if the errors go away is an insane strategy, but I would note that JavaScript is actually a semicolon-terminated language in which most developers never write a semicolon, and the JavaScript engine is already using this insane strategy on nearly every line of modern JS to insert a semicolon whenever it encounters a syntax error, so it’s obviously practical.
filter(_.price < 10)
That may not work because plain underscore is already a valid identifier but another placeholder could potentially be used and no need for the parser backtracking / function insertion (which I don't like the idea of, there may be cases where an undeclared identifier was a bug and it shouldn't be turned into a function) grep(*.price < 10)
This is referred to as "Whatever-currying": https://docs.raku.org/type/Whatever*I think that the automatic semicolon insertion is a bad feature of JavaScript.
val numbers = listOf(20, 19, 7, 12)
val multiplied = numbers.map { 3 * it }
// [ 60, 57, 21, 36 ] data.store.book.grep(*.price < 10).map(*.title)
Although personally I would write that as: data.store.book.map: { .title if .price < 10 }
which combines the filter / map into a single operation.As an aside the js example above could be simplified to a reduce().
Can you write examples in Python and Javascript where you'd extract those titles from an arbitrary JSON structure? ;-)
$..book[?@.price<10].title
Yeah, I don't think javascript has that function in the standard library. Writing one is not super complicated, but having to put that into every file (or importing it) is not ideal. find_key = (data, key) => {
if(data instanceof Array){
return data.map(x=>find_key(x, key)).flat()
}
if(data instanceof Object){
let res = Object.keys(data).map(x=>find_key(data[x], key))
if(data.hasOwnProperty(key)){
res.push(data[key])
}
return res.flat()
}
return []
}
find_key(data, "book").filter(x=>x.price < 10).map(x=>x.title)How would one use JSONPath to extract all book titles from an arbitrary JSON structure?
Also, when might that use case apply? I can't think of when I've ever needed to do something like this, but I'm interested in learning. :)
I think you missed my point.
In realistic code you'd be using string interpolation to put the number 10 into this query language, and worrying about injection vulnerabilities while you did it.
Or even calling another arbitrary function to do the filtering. Which this query language can't handle at all.
[
for book in input.store.book:
if book.price < 10:
book.title
]You wouldnt use json path as replacement to any language. It might be marginally useful in configurations or passing queries between different services. But the complex syntax limits it in both cases, because you cannot easily automatically modify the query. In the case of configs it would be great to analyze hundreds of configs on different systems and change them automaticaly, same with queries exchanged between services which might even get stored in a database.
I do not understand why domain languages aren't designed with limited syntax in mind. In the style of lisp for instance. Because actually being able to programatically work with the language is a massive advantage that imo far outweights your own frustration with typing a paranthesis or two extra.
store.book[][$$.price lt 10].title
or the more verbose: for $book in store.book[]
where $book.price lt 10
return $book.titleWith XPath and JSONPath you're just dealing with DOM nodes and children rather than directories and children.
https://github.com/tomnomnom/gron
For 90% on my needs this is all I need.
What if the query is implemented in a lower level and more efficient programming language, or probably a completely separate DB engine.
The benefits of having a well designed Turing complete language available for the filtering criteria far outweigh the disadvantages.
The philosophy of language terseness uber alles ought to have died with Perl.
Regex has similar advantages, if one sticks to the subset of regex which is commonly understood between languages: so less so, for that very reason.
Another plus is the principle of least power. A JSON Path will halt, and it won't make syscalls. There are circumstances where that's useful.
More and more parts of the API ecosystem require JSONPath, and just saying "you should write code instead" doesn't actually help anyone write OpenAPI Overlays, so whats the point?
During parsing and manipulation of JSON data, the syntactical discrepancies/behaviours between various libraries might need a common specification, for interoperability.
features like type-aware queries or schema validation, may be very helpful.