Command-line tools for data science
jeroenjanssens.com
jeroenjanssens.com
[1] http://us-east.manta.joyent.com/mark.cavage/public/surge2013...
I guess the reason I ask is much of the "manipulate and check" that I do happens before I get things to where a one liner will work. That could very well be a programmer competency issue on my part though. :-)
It's faster than people think it is. Especially when you add in libraries like pandas, it's fantastic for data analysis.
Of course, by the time you get to using pandas, you have to have everything in memory.
This isn't true of python in general, though. For simpler tasks, you can easily write generators to read from stdin and write to stdout.
I'm not saying that it's better for things like log parsing, but for more complicated ascii formats, I'd far rather use python than awk.
That having been said, people who don't learn and use awk are missing out. It's a fantastic tool.
I've just seen one too many unreadable, 1000-line awk programs to do something that's a dozen lines of python.
In contrast, when you're building a custom pipeline in a high-level language, you're optimizing for simple solutions and are not likely to get better performance unless you hit an edge case where the standard tools do really poorly.
http://vkundeti.blogspot.com/2008/03/tech-algorithmic-detail...
It's usually quicker for me to iterate on building up a complex program using existing command-line tools -- up to a point. After that point, I switch to something like Node or Python.
One reason it's faster is that they're designed to be composable. They're flexible in just the right ways -- record separators, output formats, and key behaviors (like inverting the sense of a filter or whatever) -- to be able to perform a variety of tasks, but not so flexible that you need a lot of boilerplate, as with more general-purpose languages. They defer unrelated tasks (like sorting) to tools designed for that, keeping concerns separate.
Take an awk script that reads whitespace-separated fields as input and transforms that, adding a header at the top and a summary at the end. awk's got a really nice syntax for these common tasks, and at the end you're left with a program where nearly all of the code is part of the specific problem you're trying to solve, not junk around requiring modules, opening files, looping, parsing, and so on.
Yes, sometimes the full power of a programming or scripting language is what you need, and in cases it may execute faster (though you may well be surprised -- the shell utilities are often highly optimised), but if a one-liner, or even a few brief lines can accomplish the task, why bother with the heavier tool?
However, that being said, I do notice there is a disturbing trend of command line warriors trying to do absolutely everything on the command line resulting in spending 10 minutes to construct a perfect one-liner when they could have just wrote a python/perl script in 2 minutes.
Create a predictive model is as simple as: bigmler --train < data.csv
or create a predictive model for Bitcoin volume taking online data from Quandl and parsing it with jq in just one line.
curl --silent "http://www.quandl.com/api/v1/datasets/BITCOIN/BTCDEEUR.json" | jq -c ".data" | bigmler --train --field-attributes <(count=0; for i in `curl --silent "http://www.quandl.com/api/v1/datasets/BITCOIN/BTCDEEUR.json" | jq -c ".column_names[]"`; do echo "$count, $i"; count=$[$count+1]; done) --name bitcoin
More info here: http://blog.bigml.com/2013/01/31/fly-your-ml-cloud-like-a-ki...
You can actually do most of that with vanilla PowerShell and Excel believe it or not but it's much fuglier and you spend most of your way working around edge cases.
The only thing that scares me about this though is that JSON is a terrible format for storing numbers in. There is no way of specifying a decimal type for example so aggregations and calculations have no implied precision past floating point values.
Would be nice if some of these tools became standard at the command line. I don't know about "sample" though, since that can easily be implemented in awk:
awk 'rand() < 0.2' /etc/passwdI wrote a quick ruby script that converts one or more CSV files into an SQLITE db file, so you can easily query them.
"Doe, John", 1234 Pine St, Springfield
That would be imported as 4 fields, not 3.
Also supports parsing multiple CSV files at once, so you can easily do joins.
You can raise an issue (https://github.com/jehiah/json2csv) or just fork and add a flag to include a header line
Obligatory here-I-did-it
http://www.hdfgroup.org/tools5desc.html#1
The first thing I do when I deal with data which has more than --let's say-- 10,000 rows is to put it in HDF file format and work with that. Saves a ton of time while developing a script. I had a python script do a histogram and it ran ~15sec for a file with 100k rows. With converting it first to HDF it ran in ~0.5sec. The import in python is also much shorter (two lines).
HDF is made for high performance numerical I/O. It's great and you can query several structures and even do slices of arrays on the command line (with h5tools).
It's also widely used by Octave, Python, R, Matlab... And you don't have a drawback since you can just pipe it into existing command line tools with a h5dump.
HDF5:
http://www.hdfgroup.org/HDF5/RD100-2002/HDF5_Performance.pdf
Still, 15K rows is not much. Just did that 100K rows (read/sum) bit in Octave. Took about 1/4 second to extract and sum. I assume HDF rocks for much larger sets.
https://github.com/dbro/csvquote
And:
http://en.wikipedia.org/wiki/GNU_Core_Utilities (section "Text utilities")
http://directory.fsf.org/wiki/Textutils
Run "$ info coreutils"
More about Natural Language Processing.
https://github.com/benbernard/RecordStream/tree/master/doc
https://github.com/benbernard/RecordStream/blob/master/doc/R...
ghc -e 'interact ((++"\n") . show . length . lines)' < file
EDIT:
I know that it's the same like "wc -l" but using that approach I can solve also problems that may not be well suited for awk/sed. Or maybe I just have to see some convincing one-liners.
https://github.com/dkogan/feedgnuplot
It reads data on stdin, and makes plots. More or less anything that gnuplot can do is supported, which is quite a lot. Realtime plotting of streaming data is supported as well. This is also in Debian, and is apt-gettable.
Disclaimer: I am the author
Developing a Linux command-line utility:
I will try yet again to find it and edit.
Also interesting color scheme! Can you share this?
Here's section 5 from that article rewritten:
perl -MList::MoreUtils=zip -Mojo -E 'say g("http://en.wikipedia.org/wiki/List_of_countries_and_territori... > tr:not(:first-child)")->pluck(sub{ j { country=>$_->find("td")->[1]->text, border=>$_->find("td")->[2]->text, surface=>$_->find("td")->[3]->text, ratio=>$_->find("td")->[4]->text } })' | head
{"ratio":"7.2727273","surface":"0.44","country":"Vatican City","border":"3.2"}
{"ratio":"2.2000000","surface":"2","country":"Monaco","border":"4.4"}
{"ratio":"0.6393443","surface":"61","country":"San Marino","border":"39"}
{"ratio":"0.4750000","surface":"160","country":"Liechtenstein","border":"76"}
{"ratio":"0.3000000","surface":"34","country":"Sint Maarten (Netherlands)","border":"10.2"}
{"ratio":"0.2570513","surface":"468","country":"Andorra","border":"120.3"}
{"ratio":"0.2000000","surface":"6","country":"Gibraltar (United Kingdom)","border":"1.2"}
{"ratio":"0.1888889","surface":"54","country":"Saint Martin (France)","border":"10.2"}
{"ratio":"0.1388244","surface":"2586","country":"Luxembourg","border":"359"}
{"ratio":"0.0749196","surface":"6220","country":"Palestinian territories","border":"466"}
And here's the original command:curl -s 'http://en.wikipedia.org/wiki/List_of_countries_and_territori... | scrape -be 'table.wikitable > tr:not(:first-child)' | xml2json | jq -c '.html.body.tr[] | {country: .td[1][], border: .td[2][], surface: .td[3][], ratio: .td[4][]}' | head
Fairly comparable, but there's a whole world of Perl modules for me to pull in and use with the first one.
That's an odd thing to say as Python scripts are no more savable than shell scripts. Nor are Python scripts any easier to type than shell scripts. Also, as shell scripts tend to pipe commands together, it's multiple process and thus multiple core by default, unlike Python.
If I wrote a script for every command I executed to manipulate data I would have about a million small files lying around full of one off commands. Frequently I am just debugging or exploring data and I never want to see the command again - but if I do, it will be in my bash history.
One of Python's annoyances (which are few) is that it doesn't put enough in the default namespace to be convenient from the command line.