Select, put and delete data from JSON, TOML, YAML, XML and CSV files
github.com
github.com
Also IMHO modules with user-defined (higher order?) functions are a big plus of a query language.
The developer wrote a comprehensive document explaining the rationale behind the porting that answered all my questions and a lot more: https://github.com/johnkerl/miller/blob/main/README-go-port.....
Thought other miller/mlr fans (that don't follow its development) might find this interesting as well.
(The dasel tool looks very cool, too -- looks like a good complement to mlr and similar tools!)
Combined with yj[2] you can also get TOML and HCL support, but not CSV.
It's certainly not because there's a lack of these tools :p
XPath was the best of the XML standards. Well, it helps that the language wasn't xml unlike XSLT and others.
Brackit is a retargetable query compiler and does a lot of optimizations at compile time as for instance optimizing joins and aggregations. It is useable as an in-memory processor or as a query processor of a database system.
The Ph.D. thesis of Sebastian:
Separating Key Concerns in Query Processing - Set Orientation, Physical Data Independence, and Parallelism
http://wwwlgis.informatik.uni-kl.de/cms/fileadmin/publicatio...
[2] https://sirix.io
What exactly is the "transformation" you envision?
let $array := [{"foo":0,"bar":"tztz"},{"foo":"hello","bar":null},{"foo":true,"bar":"yes"}]
let $value := for $object in $array
return
let $fields := bit:fields($object)
let $len := bit:len($fields)
for $field at $pos in $fields
return if ($pos < $len) then (
$object=>$field || ","
) else (
$object=>$field || "\n"
)
return string-join($value,"")
will output: 0,tztz
hello,null
true,yesthen just call to_csv() on the dataframe. (edited to add to comment on exporting to CSV as per the original question).
for instance:
from tablib import Dataset
json_array_of_objects = '[{"header": "data1"}, {"header": "data2"}]'
ds = Dataset()
ds.json = json_array_of_objects
ds.csv # data formatted as a csv
ds.xlsx # excel, only useful on a binary read or write
ds.dict # list of dictionaries
ds.json # list of dictionaries converted to json
ds.jira # table formatted for jiras markup
ds.html # html table
# and more
they used to vendorize dependencies, so everything worked out of the box, but now some features need to be installed specifically, or do pip install tablib[all], which is kind of annoying. I suspect they started doing it when they included support for pandas dataframes, because they didn't want to vendorize all of pandas. or force it to install as a requirement.There is no binary yet but there is a python CLI and library, even though it is written in rust.
It is the only tool that I know that deals with nested JSON and converts it into relational tables.
Here is a notebook of the python library usage.
https://deepnote.com/@david-raznick/Flatterer-Demo-FWeGccp_Q...
Pypi: https://pypi.org/project/flattentool/
Docs: https://flatten-tool.readthedocs.io/en/latest/
It's maintained by Open Data Services Coop, where we use it as a component in several of our web & data pipeline tools for working with data that is published in a Data Standard.
> echo '{"name": "Tom"}' | dasel -p json '.name'
1: https://augeas.net/docs/references/1.4.0/lenses/files/inifil...