I am wondering who really uses these tools and for what since there are R and python data science tools available?
I am wondering who really uses these tools and for what since there are R and python data science tools available?
To do a comparable amount of manipulation in Python takes a lot more boilerplate (imports, command line arguments, diety-can-we-default-to-Int64 already?, etc), plus you have to ensure you have a virtual environment with correct dependencies. Which is more or less standard numpy+pandas, but a single executable tool to do some data workup is always appreciated.
I am never performance constrained, but I have been told that miller is one of the slower tools in this space, but I still reach for it do to its wide format support.
I keep multiple little Python scripts around to do things like sum lists of numbers (think extracting a column with awk, then calculating a sum). Compiled vs an interpreted script really doesn’t matter. What matters is using the right algorithm for the job. R and Python data science libraries like to read in all of the data at once into one single data structure. That’s the anti-pattern to avoid if at all possible.
(But they are very handy for small datasets of complex calculations that require the entire dataset in memory. )