Dasel: Select, put and delete data from JSON, TOML, YAML, XML and CSV
github.com
github.com
Or put another way, is there any data storage format that couldn’t be queried by SQL?
DuckDB is a good example of a (literally) serverless SQL-based tool for data processing. It is designed to be able to treat the common data serialization formats as though they are tables in a schema [1], and you can export to many of the same formats. With extensions, you can also connect to relational databases as foreign tables.
This connectivity is a big reason it has built a pretty avid following in the data science world.
[1] https://duckdb.org/docs/data/overview
[2] https://duckdb.org/docs/extensions/json#json-importexport
Is your SQL Turing-complete? If yes, then it could query anything. Whether or not you'd like the experience is another thing.
Queries are programs. Querying data from a fixed schema, is easy. Hell, you could make an "universal query language" by just concatenating together this dasel, with SQL and Cypher, so you'd use the relevant facet when querying a specific data source. The real problem starts when your query structure isn't fixed - where what data you need depends on what the data says. When you're dealing with indirection. Once you start doing joins or conditionals or `foo[bar['baz']] if bar.hasProperty('baz') else 42` kind of indirection, you quickly land in the Turing tarpit[0] - whatever your query language is, some shapes of data will be super painful for it to deal with. Painful, but still possible.
--
Sure there could be -- any turing-complete language (which SQL is) can query anything.
But the reason we have different programming languages* is because they have different affordances and make it easy to express certain things at the cost of being less convenient for other things. Thus APL/Prolog/Lisp/C/Python can all coexist.
SQL is great for relational databases, but it's like commuting to work in a tank when it comes to key-value stores.
* and of course because programmers love building tools, and a language is the ultimate tool.
> Or put another way, is there any data storage format that couldn’t be queried by SQL?
We created PLDB.io (a Programming Language DataBase) and have studied nearly every language ever created and thought about this question a lot.
Yes, there could be 1 language to query everything, but there will always be a better DSL more relevant for particular kinds of data than others. It's sort of like how with a magnifying glass you can magnify anything, but if you want to look at bacteria you're going to want a microscope (and you wouldn't want a microscope to study an elephant).
Now it may turn out that there is 1 universal syntax that works best for everything (I'm sure people can guess what I would say), but I can't think of a case where you wouldn't want to have a DSL with semantics evolved to match a particular domain.
Off the top of my head, SQL can't do lists as values, and doesn't have simple key-value storage. Json doesn't have tables, or primary keys / foreign keys, and can have nested data
Depends on how keen you are on pure SQL. For example, postgres and sqlite have json-extensions, but they also enhance the syntax for it. Simliar can be done for all other formats too, but this means you need to learn special syntax and be aware of the storage-format for every query. This is far off from a real universal language.
I suspect for most data structures you could construct an index to make querying faster. But think about querying something like a linked list: it is not going to be too efficient without an index but you should still be able to write an engine that will do so.
If you have something like a collection of arbitrary JSON objects without a set structure you should still be able to express what you are trying to do with SQL because Turing completeness means it can examine the object structure as well as contents before deciding what to do with it. But your SQL would look more like procedural code than you might be used to.
Even if SQL and/or another query language could be Turing-complete, that doesn't mean that you can have 1 universal language to perform all possible queries in an efficient way. In basic computer science terms that means that your data structure is linked with the queries, and efficiency you want to achieve, and ad-hoc changes should be created for specific problems.
User | Telephone Numbers
-----+------------------
A | 123, 456 <- not atomic; more than 1 number (i.e. a set)
B | 789
Now there are academic operators to convert to and from a purely relational system, but I don't think they are implemented/in the standard. I forgot what they are called, however.In general you don't want a universal query language. Depending on the shape of the data you want different things to be easily expressible. You can, for example express queries on tree-shaped data with SQL (see xPath-Accelerator), but it is quite cumbersome and its meaning is lost to the reader. I.e.: It's fine when computer-generated, but there is too much noise for a human to read/write themselves. I'd be glad to be proven wrong here, but as time has shown, there is no one size fits all for programming languages. The requirements for different applications just vary too much.
I would really like to find a good workflow for idempotent modifications to INI files, but haven't stumbled across one yet.
It is a bit perl-ish, but being pure and functional it is a little easier to reason about when you have to revisit your queries.
PS I am certainly bookmarking your tool as well =]
There is also amusing project jqjq that implements jq in jq itself that I love to point folks at to show how expressive the language is: https://github.com/wader/jqjq
Awaiting all the responses from people to show off or list what tool they've landed on to support their specific use cases; I always learn a lot from these.
I’m obviously going to biased here, but it’s definitely worth your time checking out some alt shells.
The last shell I was intrigued by was es-shell which despite being old is still being updated and uses functional semantics while still looking like a shell language. I had chatgpt generate a comparison of all these with Bash (take with a grain of salt, I already had to correct at least one thing):
https://gist.github.com/pmarreck/b7bd1c270cb77005205bf91f80c...
It looks really well done, I think I'm just failing to see how this is more beneficial than just opening a single file in the editor and making changes, or writing a quick functional script so you have the history of the changes that were made to a batch of files.
If someone could explain how I could (and why I should) add a new tool to my digital toolbelt, I'd greatly appreciate it.
Here’s a jq expression I used recently to turn a complete GitHub Issues thread into a single Markdown document:
curl -s "https://api.github.com/repos/simonw/shot-scraper/issues/1/comments" \
| jq -r '.[] | "## Comment by \(.user.login) on \(.created_at)\n\n\(.body)\n"'
I use this pattern a lot. Data often comes in slightly the wrong shape - being able to fix that with a one-liner terminal command is really useful.---
[0]: https://microsoft.com/powershell
[1]: https://learn.microsoft.com/powershell/module/microsoft.powe...
[2]: https://learn.microsoft.com/powershell/module/microsoft.powe...
And it's full of delightful tricks, like when you discover properly-throttled parallelism is easy (% -parallel {}) and you don't need to dive into yet another tool-specific abstruse sublanguage.
yq eval --input-format xml --output-format csv '[file_index, file_name, .project.parent.groupId, .project.parent.artifactId, .project.parent.version]' **/pom.xmlhttps://github.com/mikefarah/yq
| yq is a portable command-line YAML, JSON, XML, CSV, TOML and properties processor
People often say that they'd prefer to write their shell scripts in Python or even Go these days, but the problem there is that the elements of structured programming makes the overall steps difficult to follow. Typically, the paradigm with use cases adjacent with shell scripts is to be able to view what it is doing without any sort of abstractions.
It's not trivial to change a particular
version: 2.33
line in a JSON file, when it's possible that string is present in many places with different meanings.
You can just wing it and be right 90% of the time; but that thoughtlessness will bite you in time.
Multiply that by the number of different context you might want to make "the same" change in a project, and sed/awk get to be a poor fit. YAML and JSON are just plain not line oriented.
You really need something with a featureset like xpath to set the semantically correct node to the right value, and every few years the kids decide they need YET ANOTHER thing that's not XML, or Yaml, or JSON, or TOML, or..
Keeping a copy of the build number checked into your source tree seems like asking for trouble when the build number is so intrinsically tied to a single build pipeline run.
editor? when i pull up emacs, 50% of the time it's write emacs macros, and I do that because shell scripts don't easily go backward in the stream. (something rarely mentioned about teco was that it was a stream editor that would chew its way forward through files; you didn't need the memory to keep it all in core, and it could go backward within understandable limits)
writing an actual shellscript is only for when it's really hairy, you are going to be repeating it and/or you need the types of error handling that cloud up the clarity of the commandline
the commandline does provide rudimentary "records" in the saved history
It's exactly a quick functional script.
I have a hard time internalizing the jq query syntax, and am not overly excited to invest in learning all the quirks when it's not based on a widely-adopted open standard. Maybe `JMESPath` could be the way forward.
Sometimes `gron` can be a pretty great alternative approach, depending on your use case. At least it is very intuitive and plays nicely with other tools.
XPath, although it had some clunky artifacts for XML (which was the reason we moved from XML like namespaces... ugh), had basically the apex of expression/path/navigation capabilites. It would be really nice to see XPath ported to a general nav language that is supported by all programming environments and handled all the relevant formats.
I'd love to see gron/ungron implemented for all tree structures.
https://github.com/dbohdan/structured-text-tools
In fact it's already on it 6 times.
Sometimes we don’t actually want to parse yaml, we just want to mutate it without needing to module the underlying objects.
Being able to select and replace, add data to an existing yaml document is a huge win for automation.
I wrote about this a little in https://www.bbkane.com/blog/go-project-notes/#scripting-chan... and it's really helped me keepy GitHub workflows and various config files in sync across project repos
That's why wheb working with Json I Love Jsonata.org
And it's not surprise is that good the creator of jsonata is on the xpath and xquery commites.
- Is easier to learn
- Has most/best documentation
- Is faster to write in
Does anyone know of a good comparison article?(I still default to jq, I guess it has the momentum)
It's heavily used by AWS and Azure though.
It is missing a good "search" ability, though. If you don't know the full path down to the data, good luck.
I have tried jq, jmespath, yq and Jsonata was the best for me.
Please open a discussion if:
You have a question.
You're not sure how to achieve something with dasel.
You have an idea but don't quite know how you would like it to work.
You have achieved something cool with dasel and want to show it off.
Anything else!
---I really like dasal.
Can I pipe a .csv to dasal and have it spit it out in JSON? And is that the best way to do that? (arent there like a ton of ways to achieve this, or would dasal make it super simple?)
Also, what would be interesting would be to be able to pull and scrape text, to put into a structured JSON.
For example - I was talking about using a Discrenment Lattice to construct a profile for a PERON PLACE THING that one was doing research on, such that you can pull multiple sources/data-types for information on [SUBJECT] and have the knowledge dossier updated. Where, for example one could pull a lot of results that can be summarized by an GPT - then using Dasal to grab the relevant component-data-points and dasal-ize and feed them into the Discernment Lattice JSON File such as I described here:
https://i.imgur.com/vuuAtAL.png
So building out a structured lattice file for a senator would look like:
https://i.imgur.com/68WFiGA.png
So, using a crawlee txtai workflow --> dasal parse --> into lattice file.
Then the lattice file can be used to compare similar slices across all the different [SUBJECTS] -- such that further ties can be made.
So, in this example - we have the data being organized for all the various entanglements a congress person has - and we can use that as a constraint for searching for relations between [subjects] which share elements across ordinarily opaque threads.
The cool thing, is that one could then easily use it to ensure you scrub and manipulate the data into a more trainable lens for effectively fine tuning the data that you want to fine tune the model with/on - thus creating a hyper contextually focused lens - https://i.imgur.com/yngUwpr.png
https://github.com/sharkdp/hyperfine
Dasel MIT License Copyright (c) 2020 Tom Wright
desel: .data.all().filterOr(moreThan(.quantity,3),equal(.quantity,3)).mapOf(key,key,quantity,quantity)
yq: .data[] | select(.quantity >= 3) | {"key": .key, "quantity": .quantity}
jq: .data[] | select(.quantity >= 3) | {key, quantity}