Show HN: Pythonpy – the swiss army knife of the command line
github.com
github.com
http://opensource.imageworks.com/?p=pyp
The Pyed Piper Tutorial: http://www.youtube.com/watch?v=eWtVWF0JSJA
Swapping pipe components at runtime with pipe_controller:
http://jugad2.blogspot.in/2012/10/swapping-pipe-components-a...
Using PipeController to run a pipe incrementally:
http://jugad2.blogspot.in/2012/09/using-pipecontroller-to-ru...
PipeController v.01 released - simulating UNIX-style pipes in Python:
http://jugad2.blogspot.in/2012/08/pipecontroller-v01-release...
Some ways of doing UNIX-style pipes in Python:
http://jugad2.blogspot.in/2011/09/some-ways-of-doing-unix-st...
Mostly out of frustration for PyP not being lazy (on large inputs it reads the entire file up-front, or at least used to).
But it was quite interesting to implement all the standard python idioms (like slicing) in a lazy way, and it's not that complicated a tool. I still use it a lot whenever I have a nontrivial pipeline to write.
py 'itertools.count(1)' | py 'itertools.islice(stdin, 0, 10, 2)'
However, the number of times that you need this are surprisingly rare. Most lazy operations don't require that each row be aware of the surrounding row context, and using the much simpler: py -x 'new_row_from_old_row(x)'
will get the job done in a lazy fashion. Usually, when you need rows to be context aware, as in: py -l 'sorted(l)'
or py -l 'set(l)'
it's just not possible to accomplish your task without reading in all of stdin.Some things can't be done without reading everything. But there are still a number of operations on "all of stdin" that can safely be done lazily. I'm particularly fond of "divide stdin into chunks of lines separated by <predicate>" [0]. Which does need context, but only enough to determine where the current chunk ends (typically a few lines).
`py` seems to be aimed at a single expression per invocation (nice and simple), while `piep` recreates pipelines internally (more complex but also means pipelines can produce arbitrary objects rather than single-line strings). So I'm not really sure how you'd do the above in `py` anyway.
[0] http://gfxmonk.net/dist/doc/piep/#piep.list.BaseList.divide
And Kickstarter/Indiegogo/Crowdfunding has shown genuine interest to help fund useful tools and libraries lately which is great.
It isn't about motivating people to do things; that creates abandonware. It is about solving a problem that needs solving.
ls | py -x '"mv `%s` `%s`" % (x,x)' | sh ls |py -l 'sum(1 for x in l)'
And also I can count the number of files whose names matche a certain extension: ls |grep .py |py -l 'sum(1 for x in l)'
I am not saying it was not possible before, but I don't know how to do it in bash.Looking forward to use that for some basic stats on file. The following should sum up the first column of a csv.
cat data.csv| py -x 'x[0]' |py -l 'sum(x for x in l)'To work with input split into columns, use awk. By default, it assumes the columns are separated with spaces. This can be customized with the '-F' option.
So your last example can be rewritten as
awk -F, '{s+=$1} END{print s}' data.csv ls | wc -l
ls | grep .py | wc -l ls | wc -l
ls | grep -c .py
(note the cat and pipe are not needed) awk -F',' '{sum+=$1}END{print sum}' data.csvIn the meantime feel free to use this python thing, its rather fun to be honest. Just wanted to demonstrate the unix-y way of doing it.
That won't work if there are files with names like foo.py.txt.
These will:
ls | grep -c '.py$'
or better, because less names to filter out:
ls -l *.py | wc -l
py -l 'len(l)'
That being said, I still use wc -l every time. In fact, I always prefer regular unix commands (e.g. grep, xargs, head) to pythonpy when possible. But there are some things, such as: py -l 'l[::2]'
which while possible with tools like sed, are just much better expressed with pythonpy. $ ruby -e 'puts (0...3).to_a' | ruby -ne 'p $_.to_i*7'
$ ruby -e 'puts (0...8).to_a' | ruby -ne 'puts $_ if $_.to_i % 2 == 0'