Faster Command Line Tools in D
dlang.org
dlang.org
I found this article interesting but to be honest I hate it when people do this. Usually it is the real-world considerations and error handling that cause the code to cease being an elegant demo and look more like the same stuff everyone else writes.
Sometimes I wonder if software should error first, then when you bounded the failure space, you iterate on the success space as you see fit.
The frequency command computes the frequency of values in 2 columns in < 2 seconds on my machine. Its a different task, but there are some similarities.
xsv frequency -n -s 1,2 -l 100 googlebooks-eng-all-1gram-20120701-0.tsv
5.19s user 0.06s system 351% cpu 1.490 total
[1] https://github.com/BurntSushi/xsvIn my opinion omit all of the discussion on python and just talk about "how to optimize a D program" b/c that's what this article is.
https://github.com/eBay/tsv-utils-dlang
They basically explored new languages to rewrite some perl script in and liked D enough to shift over. They have other tooling in other languages, my guess is they'll unify a good amount of it in D. Disclaimer: this is based on my own assumption that they like D so much they want to just use it all over. It wouldn't surprise me to see this confirmed by an eBay employee. Although seeing as he wrote TSV Utilities, it really wouldn't surprise me if he wants to rewrite all in-house tooling he uses in D as the repository states.
PyPy probably manages to optimize the hash map/method call lookups for these small programs, which explains the speedups. Removing memory allocations is still hard.
The D language provides finer mechanisms to control memory and data structures. This makes the language larger, but enables you to optimize if it becomes necessary.
Still, I agree and I would like to see a Python expert to optimize it.
Python is a highly dynamic language with an API (towards both Python and C) that is very invasive. These two things, taken together, make optimizing the interpreter extremely difficult, because practically all of it can be modified or introspected. CPython being implemented largely as a hashtable-interpreter is only one facet to its performance.
Perhaps a talk recommendation: https://www.youtube.com/watch?v=qCGofLIzX6g&list=PLRdS-n5seL...
When I switched to pypy the csv library actually made it nearly 2x slower than pypy using .split(delim)
Edit: do you mean the tsv utilities? Because they're working fine here on the Creator's Update.
Typically there are two modes in my computing: (1) Scripting / Command-line get stuff done and throw-away, and (2) Serious applications that are heavily used, need to process lots of data and be as fast as possible (e.g. processing millions of files like this one where the algorithm constant factor really matters).
In case 1: Hack something together with shell or python and get an answer. If it takes 100-1000x of the equivalent C program, than fine..
In case 2: Custom special purpose C or C++ code
I really don't understand the middle ground here. The equivalent C version that does the same job runs in ~250 millis on my slow Yoga 2 Pro laptop. Total line count 83 of pure C (no other libraries).
Is it "elegant"? Depends who you ask.. But, at the end of the day, "elegant" doesn't pay the bills..
People use Python all the time to manipulate data sets >= 100G in size despite its speed failings at that size. Why? Because Pandas is just so damn convenient. It would take me a grand total of 30 seconds to write Pandas code which read a TSV and gave me the sum of two multiplied together columns grouped by the day of a timestamp column. Doing that in C would take several orders of magnitude more time.
It's an optimization of people's time problem. You could probably spend several hours (or days) writing a C program for a specific problem. But if you can spend only 40% of the time writing the program and have it only 20% slower, then that's a definite win (these numbers are just an example).
http://tech.adroll.com/blog/data/2014/11/17/d-is-for-data-sc...
Edit:
To be fair a lot of the companies there look like they handle heavier loads. Also Garbage Collection is optional, and there are alternatives to the "standard library" for D that others have made that are probably usable without GC. Some people have done successful embedded systems programming without the GC, I remember one guy talking about it on the D irc channel.
Of course, it is also true that much of the standard library doesn't actually use it... but these objections are never actually about facts.
That's the point. There are already many popular languages with GC. Why would people switch from C++ (this was the original question) to D instead of one of those much more popular languages that you mentioned?
D's goto build system is called dub, it also handles packages hosted on dlang.org: I know no defects or problems with it, it's pretty good.
For simple (non-meta-programming) code. Dmd and Go are in the same league with respect to compilation speed.
Btw compilation speed is the main reason for me to put off Rust and C++ and Scala. I cannot stand slow compilers anymore after using D for a while.
By chance, I found (while looking for a rob pike quote) this article critiquing Go's design: It uses D to demonstrate it's arguments. http://nomad.so/2015/03/why-gos-design-is-a-disservice-to-in...
For production you want to omit things like debug from binary so your binary is smaller. You have to play with the build flags. I can get a simple TCP server in Go in just a dozen lines but the binary is around 5MB before optimization.
[1]: https://blog.filippo.io/shrink-your-go-binaries-with-this-on...
With D, in addition, you can link against C shared libraries (dynamic linking) AND you can write standalone libraries in D that even C can link against (see Mir/betterC).
AFAIK, you can do neither of these in Go.
Dlang: Traditional concurrency (thread based), great generics implementation, nice algorithm/container libraries and C++ interop
Still, Go is really hard to beat on that front because it makes it super easy and has the really brilliant feature of making almost everything that waits for IO interruptable inside a function when you call "go function()" without having to mark operations with "async".
That set of good internet sensibilities is what attracted me to Go in the first place, and why I still like working in it in my spare time. If D has a similarly strong story in that area, I'd love to know about it.
As a tiny nit, it's interesting to compare D's `out` keyword to Go's multiple return values. An `out` keyword seems like a prime example of "thinking in blub."
It's not to everyone's taste, but it was designed by people who've been programming for a long time to be a language they'd like to program in. It is my favorite language for many tasks (and I've been coding for a couple decades now).
For me what's impressive about it is it's very simple, reasonably expressive, and just a really well-designed cohesive whole that doesn't usually expose sharp edges.
The longer I program the less I care about fancy things or being elegant or writing the smallest possible code, and the more I care about eliminating bullshit problems and wtf moments. Go's good at that.
i am thought impressed with how fast pypy did
https://stackoverflow.com/questions/27801945/surprising-resu...
Another likely optimization would be to use the csv module (which would parse the rows in C).
$ cat sum.awk
{ a[$2] += $3 }
END {
for (i in a) {
if (a[i] > max) {
maxk = i
max=a[i]
}
}
print "max_key:", maxk, "sum:", max
}
$ time gawk -f sum.awk -O <ngrams.tsv
max_key: 2006 sum: 22569013
real 0m7.041s
user 0m3.797s
sys 0m3.156s
(This under Linux subsystem for windows)According to the gawk profiler, and strace -c (count) -- the awk program mainly spends its time reading the file (without the loop at the end, looking for the max value, the runtime is essentially the same).
In fact, on the surface, pypy and python3 are quite similar on the syscall/strace front - with roughly 23k "read" calls -- awk did 375k. And adding cat in front sped it up by about two seconds:
$ time (cat ngrams.tsv |awk -f sum.awk )
max_key: 2006 sum: 22569013
real 0m3.969s
user 0m3.719s
sys 0m0.516s
$ time awk -f sum.awk ngrams.tsv
max_key: 2006 sum: 22569013
real 0m6.465s
user 0m3.609s
sys 0m2.859s (cat ngrams.tsv \
|parallel --pipe awk -f map.awk \
|awk -f reduce.awk )
max_key: 2006 sum: 22569013
This now runs in 17 to 18 seconds... :-/ $ cat map.awk
{ a[$2] += $3 }
END {
for (i in a) {
print "ignore", i, a[i]
}
}
$ cat reduce.awk
{ a[$2] += $3 }
END {
for (i in a) {
if (a[i] > max) {
maxk = i
max=a[i]
}
}
print "max_key:", maxk, "sum:", max
}
[ed: However, there are faster awks than gawk: $ time mawk -f sum.awk ngrams.tsv
max_key: 2006 sum: 22569013
real 0m2.826s
user 0m2.391s
sys 0m0.422s
mawk is (a little) faster than pypy on my machine.]
Which is a 15000-line Perl script, so to be expected.
Try --pipe-part instead:
parallel -a ngrams.tsv --pipe-part --block -1 awk -f map.awk |
awk -f reduce.awkThe (old) version of parallel packaged with Ubuntu 16.04 (linux subsystem for windows) - doesn't have --pipe-part -- but running from upstream, the speed is more reasonable:
$ time (./parallel-20170522/src/parallel -a ngrams.tsv \
--pipe-part --block -1 -j4 mawk -f map.awk \
| mawk -f reduce.awk )
max_key: 2006 sum: 22569013
real 0m2.265s
user 0m4.672s
sys 0m1.672s
(Tried a few variants with/without -jN -- and this seems typical for the fast end of the spectrum). $ time (cat ngrams.tsv \
| mawk -f map.awk \
| mawk -f reduce.awk )
max_key: 2006 sum: 22569013
real 0m3.472s
user 0m2.891s
sys 0m2.406s
[ed: btw, did a double-take when I saw your Gnu Privacy Guard id: 0x88888888 :-) ]What is the smart way to do this in kdb+?
This is my naive, sloppy 15min approach.
Warning: Noob. May offend experienced k programmers.
k)`t insert+:`k`v!("CI";"\t")0:`:tsvfile
k)f:{select (*:k),(sum v) from t where k=x}
k)a:f["A"]
k)b:f["B"]
k)c:f["C"]
k)select k from a,b,c where v=(max v) 1#desc sum each group (!/) (" II";"\t") 0: `:tsvfile
Took about 3 seconds, 2.5 of which was reading the fileEDIT:
q)\ts d: (!/) (" II";"\t") 0: `:tsvfile
2489 134218576
q)\ts 1#desc sum each group d
486 253055104 A 4
B 5
B 8
C 9
A 6
How to solve with only a dict?Regarding the 1gram file at https://storage.googleapis.com/books/ngrams/books/googlebook...
This is the result I got
3| 1742563279
using q)\ts d:(!/)(" II";"\t")0:`:1gram
q)\ts 1#desc sum each group d
1897 134218176
371 238872864
or k)\ts d:(!/)(" II";"\t")0:`:1gram
k)\ts desc:{$[99h=@x;(!x)[i]!r i:>r:. x;0h>@x;'`rank;x@>x]}
k)\ts 1#desc (sum'=:d)
1897 134218176
0 3152
372 238872864
No doubt I must be doing some things wrong.With the reverse thrown in to switch the key/value around we get the correct answer
q) 1#desc sum each group (!/) reverse (" II";"\t")0:`:1gram
2006| 22569013
or k) {(&x=|/x)#x}@+/'=:!/|(" II";"\t")0:`:1gram
(,2006i)!,22569013i
Works the same for the simple example k)e: 4 5 8 9 6!"ABBCA"
k){(&x=|/x)#x}@+/'=:e
(,"B")!,13Certainly eye-catching but I wouldn't call it conclusive.
https://www.ibm.com/developerworks/community/blogs/jfp/entry...
I would speculate using numba or Cython would yield further performance gains over PyPy...but that's mostly just based on anecdotal comparisons:
https://cardinalpeak.com/blog/faster-python-with-cython-and-...
I just think it is a bit dishonest to try and make a claim as pointed as this article's in 2017 by stopping at simply running an un-optimized CPython script with PyPy.
http://cython.readthedocs.io/en/latest/src/tutorial/cython_t...
http://cython.readthedocs.io/en/latest/src/quickstart/cython...
But that is just my opinion.