An introduction to data processing on the Linux command line
blog.robertelder.org
blog.robertelder.org
To visualize data coming in from a pipe, can pipe it to
https://github.com/dkogan/feedgnuplot
Very useful in conjunction with other tools to provide filtering and manipulation. For instance (the first one is mine):
https://github.com/dkogan/vnlog
https://www.gnu.org/software/datamash/
https://csvkit.readthedocs.io/
https://github.com/johnkerl/miller
https://github.com/eBay/tsv-utils-dlang
https://github.com/BatchLabs/charlatan
https://github.com/dinedal/textql
https://github.com/BurntSushi/xsv
https://github.com/dbohdan/sqawk
- The original award started in 1995. Even though pentium was already out, I think it is safe to say that was the era of 486 PCs. In 2019, for day-to-day shell work (meaning no GBs of file-processing or anything like that), isn't invoking UUOC and pointing out inefficiencies an example of premature optimization [1]?
- Isn't readability a matter of subjectivity, and that for some folks 'cat file' is more readable than '<file' or a direct use of a processing command (like grep, tail, head, etc) [2] ? (The whole stackoverflow page is fairly illuminating [3]).
[1] http://wiki.c2.com/?PrematureOptimization
[2] https://chat.stackoverflow.com/rooms/182573/discussion-on-an...
Shameless plug: https://github.com/csdvrx/sixel-gnuplot
mlterm works.
mintty had a regression, 3.1.0 may have fixed that
cat data.csv | sed 's/"//g'
can be simplified by doing this instead: cat data.csv | tr d '"'
This awk command: cat sales.csv | awk -F',' '{print $1}' | sort | uniq
Can be replaced with a simpler (IMO) cut instead: cat sales.csv | cut -d , -f 1 | sort | uniq
When using head or tail like this: head -n 3
You don't need the -n: head -3
Also shout out to jq, xsv, and zsh (extended glob), all nice complements to the typical command line utils.Ex: | sed -e step1 becomes | sed -e step1 -e step2 instead of adding another pipe and another "moving part" like tr
sort -u -t, sales.csv
However, those fail with quoted commas.
Also, head -3 is non-POSIX obsolete syntax.
Edit: I don't know why I didn't see other UUOC references initially.
cat sales.csv | awk -F',' '{print $1}' | sort | uniq
can even be further simplified to cut -d, -f1 sales.csv | sort -uGood introductory article here!
I was lucky that my first job was as a support engineer at a data-centric tech company, which is where I learned these. I've often thought about how to teach them to data analysts coming from a non-engineering background. This is comprehensive but clear and would be a perfect resource for training someone like that. Thank you!
[1] https://news.ycombinator.com/item?id=17324222
P.S.: Not essential, but it really becomes a joy when, as a touch typist, I have turned on vi mode in the shell (e.g., with 'set -o vi'). My fingers never have to leave the home row while I do my shell piping work from start to finish. (no mouse, no arrow keys, etc.)
The whole point of this article is to point out that a lot of common Linux tools can be used for Data Science like work (a significant part of which includes pre processing structured and unstructured text).
Even worst, most of the tools (cat, grep, awk) are Unix commands, redeveloped by the GNU project in most of the GNULinux distros.
Oh well, I guess a lot of people just think all Unix-like systems are called “Linux” now. Perhaps it’s become like the word “Kleenex”.
I find it more irritating when people try to score greybeard points by saying *nix (or Unix) when it's obvious that they're talking about a Linux-only mechanism and quite possibly haven't ever used Unix (or a direct derivative).
Also, the most popular Unix-like OS (far more than Linux) is macOS, basically the least “leet greybeard affectation” thing I can imagine. Your irritation is way off base.
Most popular Unix-like OS on MacBooks is MacOS.
To bring us back to the context of this post: I am quite willing to bet that “grep” and “cat” are used by humans more times per day on macOS than on any other OS.
Perhaps nothing? I was responding to the complaint in general terms.
> Your irritation is way off base.
Please allow me to feel irritated when people refer to obvious Linux things as something that's supposedly got something to do with Unix. It happens often enough.
I had thought that you were directly responding to the original poster.
While reading the idea that I know most of this, would that made me a data scientist? Jumped at me.
But then I quickly recovered from that thought that surely knowing some of the tools someone could use for a certain domain does not make you expert at that domain.
Might just be the case of same ingredients, different recipes.
awk -F, '$2 == "F" {$0=(($1-32)*5/9)",C"} {print}'
cat sales.csv | awk -F',' '{print $1}'
but I'd prefer cut -d, -f1 sales.csvRememeber, nearly all cases where you have:
cat file | some_command and its args ...
you can rewrite it as: <file some_command and its args ...
and in some cases, such as this one, you can move the filename to the arglist as in: some_command and its args ... file
— Randal L. Schwartz (http://porkmail.org/era/unix/award.html#cat)I actually prefer useless cat because when you're prototyping a pipeline it's very awkward to use non-useless cat. You'll probably start off with something like this to observe the content of the file:
cat something.txt
Using this doesn't work in bash: <something.txt
Then, continuing with useless cat to build on it you do cat something.txt | grep stuff
Which you can type easily from using 'up' in your terminal. But if you use non-useless cat you have to re-type the entire thing or move the cursor around: grep stuff < something.txt
With useless cat, you can keep adding things and check the result: cat something.txt | grep stuff | sed 's/"//g'
Or if you need to insert another filter before the last stage like this, you can just press "up" and insert it: cat something.txt | grep -v negmatch | grep stuff
I don't think there is any easily-typed equivalent workflow with non-useless cat.<file head -n50 | whatever
Can be the starting command. When you no longer need the head there, just get rid of “head |”.
Although I agree that the pointing out of “useless cat” is usually not particularly useful or constructive.
cat foo | bar
is useless use of cat, since it’s equivalent to < foo bar
which is both shorter and starts one less process. Why do you think that “cat” makes it “easier to author pipelines”?2). I often start with `bar < foo`, so if I need to add more arguments to `bar`, I need to always skip over the input.
3). If I just want to delete all processing and look at the input, I can’t just backspace away the processing because `< foo` is invalid.
tsort and comm were news to me.
Why not just do it all in Python or R? That way you also get something that will probably work on non-unix platforms.
I like to use command line tools for for one-off tasks that I'm unlikely to repeat. If there's a task I know I'll need to repeat or is too cumbersome to do in a couple of lines, I'll reach for Python.
For anyone who is interested in going a little deeper into data science, I’d also recommend the “Introduction to Data Science with R” series by David Langer:
My pet peeve is the "grep | awk" idiom. No, just use awk.
Awk does map/reduce, relational joins, associative memory, table lookup, and so on. Just use awk arrays, begin block, and end block.
But most people don't know awk. And awk requires more awareness. I break my awk when I fix things when tired.