Miller – tool for querying, shaping, reformatting data in CSV, TSV, and JSON
miller.readthedocs.io
miller.readthedocs.io
https://github.com/benbernard/RecordStream
It shows the unix model of many small, composable tools is very powerful, but also shows that POSIX is missing some essential pieces that everyone keeps trying to add/reinvent.
[0] https://github.com/multiprocessio/datastation/tree/main/runn...
Miller CLI – Like Awk, sed, cut, join, and sort for CSV, TSV and JSON - https://news.ycombinator.com/item?id=28298729 - Aug 2021 (66 comments)
Miller v5.0.0: Autodetected line-endings, in-place mode, user-defined functions - https://news.ycombinator.com/item?id=13751389 - Feb 2017 (20 comments)
Miller = sed, awk, cut, join, sort for CSV and tabular JSON - https://news.ycombinator.com/item?id=11674304 - May 2016 (1 comment)
Miller is like sed, awk, cut, join, and sort for name-indexed data such as CSV - https://news.ycombinator.com/item?id=10073040 - Aug 2015 (2 comments)
Miller is like sed, awk, cut, join, and sort for name-indexed data such as CSV - https://news.ycombinator.com/item?id=10066742 - Aug 2015 (75 comments)
This mentality is why modern software seems slower... Because it is.
So many technological advances, computers are exponentially faster in so many ways, yet we regress in performance.
It's a shame, really...
It's why you have user interfaces today that are often as slow as, or slower than the equivalent Windows 3.11 interfaces. In 1993 things took time because the processor was an actual potato and the data fetched off a floppy disk, today things are slow because the code being run is 1-10k times slower and messages need to be continuously sent to a datacenter half way around the world to inform an analytics server about your every input.
You're focusing on the developer, not the user, but this is a user tool :(
Miller has always been a user tool and will remain so.
I think the confusion arose because I did the port per se and feature-adds first (which has taken most of the time), and left analysis/optimization (even low-hanging fruit like output-buffering) until the very end -- https://github.com/johnkerl/miller/pull/786 for example merged just a couple days ago.
Another win (besides more features, and recent performance improvements) -- the Windows version is now a snap. Now `mlr.exe` takes it rightful place alongside the Linux and Mac executables (https://github.com/johnkerl/miller/releases/tag/v6.0.0-beta). Go+Windows is heavenly & I'll never need to tweak MSYS2/Appveyor/DLLs/etc again. :) See also https://miller.readthedocs.io/en/latest/miller-on-windows/
The most important release blocker (now resolved) is https://github.com/johnkerl/miller/pull/786 et al., thanks to which Miller 6 performance is now on par with Miller 5 for simple processing, and far better than Miller 5 for complex processing chains.
I saw in some other comments you mentioned some performance lift coming from thread level parallelism in Go. I wonder -- can you get oversubscription issues if the user is doing their own parallelism (I assume some folks will implement 'parallelism' by throwing a bunch of independent processes at a bunch of independent records).
I know in openMP (for example) this would be something where the user is expected to keep track of it, but maybe the Go runtime handles this stuff gracefully?
More specifically, Miller uses separate "goroutines" (more or less threads) -- 2 for input (one for raw byte stream ingest and one for forming records), 1 for each verb (sort then head would be two verbs), and 1 for output (records back to strings), all pipelined up. Some of the recent perf improvements came from splitting the record-reader into two concurrent goroutines like that. Some of the older perf improvements are from pipelining the verbs.
But yes, if you have say 16 CPUs and you launch 10, or 20, or 30 Miller executables -- in the C impl each would soak a single CPU and anything beyond 16, the OS would have to multitask things. With Miller 6, basically the same kind of thing except each executable will be trying to use more than one CPU if it can. If it can't, the Go runtime and the OS will multitask. And both do tend to handle this stuff gracefully and without tuning on the part of the user.
Also I'd point out that even with big files & deeper multi-verb processing chains, htop rarely shows over 250% CPU, maybe 350% for deeper chains -- the input and output processing typically take most of the time, and the verbs in between not as much.
When we run out of memory, systems crash; when we run out of CPU, things just take a bit longer. I.e. the oversubscription is a real but non-fatal issue, for C or for Go ...
The other day I ran across https://repology.org/project/miller/versions -- something autogenerated on the web; I don't maintain it. Anyway you can see most distros are on 5.10 but there are some farther behind ... I will probably start with Conda and Brew since these are more personal/interactive, then maybe Fedora and Ubuntu -- ?
https://github.com/capitalone/DataProfiler
Effectively loads anything into a dataframe. After that you can profile the data, merge, save, load, and take differences between profiles before generating a report.