List of command line tools for manipulating CSV, XML, HTML, JSON, INI, etc.
github.com
github.com
Of course, it would be even better if you could easily tell which of the dozen JSON query tools is the best choice for the task at hand, or which you should code if you only want to ever use one of them.
In fact I'd love if someone would like to share their set of tried-and-true tools. Personally I mostly go with the POSIX tools, plus jq or gawk on occasion (but I have to read their docs every single time...).
[1]: http://pubs.opengroup.org/onlinepubs/009695399/utilities/awk...
[1]: https://askubuntu.com/questions/1011414/gawk-is-crashing-for...
One thing I could suggest for the XML list is xmllint. It can be really useful for converting xml to canonical format so you can then use diff to compare it.
E.g. something like diff <(xmllint —c14 first.xml) <(xmllint —c14 second.xml)
I’d love to heat about more command line SOAP tools if anyone can recommend some.
tidy -xml -indent -wrap 0
or tidy -xml -indent -wrap 0 -quiet[1]: http://code.kx.com/q/ref/filenumbers/#load-csv
I personally prefer J to K in the APL family of languages. They also have a relatively cheap database, Jd [1]. Individual licenses are $600. Still a bit too much for my data mangling needs. :)
There's also a per-core/minute pricing which might be useful.
But a lot of the use cases these other tools are good for are small tasks every now and then. I feel kdb+ is in a different category.
For example, removing nonconsecutive, duplicate lines from a file, such as a CSV file:
exec echo "k).Q.fs[l:0::\`:$1];l:?:l;\`:$1 0:l"|exec q >&2;
where Q.fs is a function in a script thats bundled with the interpreter; the chunk size for reading the file into memory is adjustable by editing the function. l:0;.Q.fs[{if[x~l;:];-1 l::x}each]`:input
or if you have memory: -1 distinct read0`:input
...or if you want to use k: -1@?0:`:input k)a:{-1@?0:`:input;};a[] ;
What this does is return generic-null :: which .Q.s doesn't print.Sure, it is not as fast as many other formats, but on the other hand it integrates very well into Emacs an org-mode. I manage a large part of my different collections using a combination of both, and the Emacs integration means it is all less than 2 seconds away.
https://johnkerl.org/miller/doc/file-formats.html#Tabular_JS...
Edit: I’m looking for a command line tool that allows me to open an Excel file, make a few simple changes, and then save again as an Excel file.
This sheet was formatted like:
MEMBERS ...rows...
ADMINS ...rows...
EXECUTIVE COMMITTEE ...rows...
You could whip up any command line tool you need with that.
It won’t help if you need to retain anything Excel specific, but I find it very useful to deal with any Excel files that come my way.
That said, link to the manual in the bitbucket link not working.