This strikes me as an odd conclusion to come to if speed was the main motivator.
This strikes me as an odd conclusion to come to if speed was the main motivator.
Speed is the main motivation, but total time is TimeToWriteCode + TimeToRunCode.
Python has the lowest TimeToWriteCode, but very high TimeToRunCode. C++ has lowest TimeToRunCode, but high TimeTowWriteCode. Haskell is often a good compromise for me.
Also, with Haskell, it can be very easy to take advantage of 20 CPU cores, while I don't have as much familiarity with high-level C++ threading libraries.
If you provide me with a self-contained code example (with data required to run it) that is “too slow”, I’d be willing to try and optimise it to support my point above.
Also, have you tried Numba? It maybe a matter of just applying a “@jit” decorator and restructuring your code a bit in which case it may get magically boosted a few hundred times in speed.
[1] https://git.embl.de/costea/metaSNV/blob/master/metaSNV_post....
Here is an earlier version (intermediate speed): https://git.embl.de/costea/metaSNV/commit/ff44942f5f4e7c4d0e...
It's not so easy to post the data to reproduce a real use-case as it's a few Terabytes :)
*
Here's a simple easy code that is incredibly slow in Python:
interesting = set(line.strip() for line in open('interesting.txt'))
total = 0
for line in open('data.txt'):
id,val = line.split('\t')
if id in interesting:
total += int(val)
This is not unlike a lot of code I write, actually. interesting = set(line.strip() for line in open('interesting.txt'))
total=0
for c in chunks: # im lazy to actually write it
df = pd.read_csv('data.txt', sep='\t', skiprows=c.start, nrows=c.length, names=['id','val'])
total += df['val'][df['id'].isin(interesting)].sum()
I'm not exactly sure, but pretty sure that isin() doesn't use python set lookups, but some kind of internal implementation, and is thus really fast. I'd be quite surprised if disk IO wasn't the bottleneck in the above example.Reading in chunks is not bad (and you can just use `chunksize=...` as a parameter to `read_csv`), but pandas `read_csv` is not so efficient either. Furthemore, even replacing `isin` with something like `df['id'].map(interesting.__contains__)` still is pretty slow.
Btw, deleting `interesting` (when it goes out of scope) might take hours(!) and there is no way around that. That's a bona fides performance bug.
In my experience, disk IO (even when using network disks) is not the bottleneck for the above example.
There's also code in genetic_distance() function that IIUC is meant to handle the case when sample1 and sample2 are not similarly-indexed, however (a) you essentially never use it, since you only pass sample1 and sample2 that are columns of the same dataframe (what's the point then?), and (b) your code would actually throw an exception if you tried doing that.
P.S. I like the part where you've removed the comment "note that this is a slow computation" :)
scikit-allel: http://scikit-allel.readthedocs.io/en/latest/index.html
scikit-allel example: http://alimanfoo.github.io/2015/09/21/estimating-fst.html
with open('interesting.txt') as interesting_file:
interesting = {line.strip() for line in interesting_file}
with open('data.txt') in data_file:
total = sum(int(val) for id, val in map(lambda line: line.split('\t'), data_file) if id in interesting)Also, if the data you're reading is numeric only - or at least non-unicode / character data - you might be able to get a speed boost reading the data as binary not as python text strings.
Given his code you referenced, could you elaborate on what makes it look slow at a glance, and how you might speed it up? :)
if snp_taxID not in samples_of_interest.keys():#Check if Genome is of interest
Tracing through, looks like samples_of_interest is a dict. `snp_taxID not in samples_of_interest` would make membership check constant time.Numba does not support dictionaries and has limited support for pandas dataframes (only underlying arrays, when convertible to NumPy buffers, if I understand correctly). This limits usefulness for many non-array situations, as well as some existing code-bases (the dictionary is fundamental in Python and typically used everywhere -- often for performance).
"Despite the examples and docs, Numba is voodoo. Damned cool voodoo, but still voodoo"
I'm working on my first serious Python project right now, and I find it's super easy to throw together some code that more or less works; but for solid, readable, documented, properly unit-tested code I hope is production-ready, it's not any faster than Perl or Golang.
(Sure, if you're a Python expert it's faster for you than for me, but if it's about TimeForExpertsToWriteGoodCode I'm not any more convinced.)
Proper unit-testing is also going to take roughly the same time in any language, just because you have to think hard about sensible tests (although I still love mocking/patching in Python, so I'd give it an edge, plus pdb/ipdb for debugging tests is cool). Production-ready also includes deployment, which for anything non-trivial I'd say Golang > Python > Perl.
Finally, if we're talking "serious project", IMO tooling and how that tooling integrates into a CI pipeline are more important than development speed, because as a team or project goes, terrible CI will slow developers more than any language. Although again here I think Python does quite well with decent linting, unit test frameworks, and code coverage options, Golang's opinionated tools are simpler in this respect.
(I enjoyed C# for similar reasons, although I don't think it's kept up w.r.t. tooling - been ages since I used it though.)
One big point I would give to Golang, about which lots of people disagree with me, is the "opinionatedness" of it. It seems to me that Python, like Perl, has a "There's More Than One Way To Do It" mentality, and after many years of that I really appreciated Golang's emphasis on the "idiomatic." That goes for the tooling too.
I have also noticed that the Python ecosystem doesn't have a strong documentation culture, which I find annoying as a relative newbie. But that presumably matters less over time, and it seems to be part of the Python Way to use libraries that "just work" and not worry about the details.
In lots of areas, "good code" doesn't matter much, if at all.
Scientific computing is full of those cases -- you write code to run a few times, and don't care for maintaining it and running it ever again (as long as the results are correct).
It usually starts with "oh it's just a one-off thing" and then it turns out to be useful and the rest is messy history.
But sure, within that genre I could see Python being a faster language to write in than many others.
This is the received wisdom in biological science but I’m convinced that it’s trivially wrong. I’ve seen a lot of research code, most of it bad. I have no idea how many bugs are in this code, and I know for a fact that the original authors also don’t know. And it would be truly exceptional if these pieces of code were bug-free (in fact, there’s enough software engineering know-how to categorically conclude that a very high percentage of such code has bugs). How many of these bugs affect the correctness of the results?
… since the code quality is so bad, this is impossible to quantify. So, yes, code quality does matter in science, since it affects the probability of publishing wrong results.
Incidentally, there are cases of retractions of high-impact papers due to errors in code. Of course this will also happen with better code quality; but if conventional software engineering wisdom is right then it will happen substantially less.
The issue is when I cannot.
If you can't make the core pandas code decently fast, dask won't save you.
I've had success using it on non vanilla stuff (i.e. code that could not get converted to play natively with numpy/pandas structures)
As a bonus, the nice profiling tools (built within dask) have also helped me improve the performance of the code.
C++11 has all the nice features you might expect from python with the only drawback being the lack of a REPL.
If you know exactly what you need to write, you're just as quick in C++ as in Python, that's true. Programming is mostly about learning what to write, though, and here C++ loses.
EDIT: Not to mention, if you write your code as a lot of tiny functions you could just as well write it in C. Once you go for classes and templates, that's where C++ power is visible, but that's also where its compile times suck.
Then weigh in the hard realities of some engineering problems. It won't matter that it takes 1% of the time to implement a video decoder in python if it can't deliver decoded frames in a timely manner. It won't matter that the C solution will run 1000x faster if you need a month to develop what should be delivered on Friday.
I'm sorry if this is already covered in the article. I had a brief look before but it won't currently load.
I even wrote up a few utilities to make use of multiple threads while working at a high level: https://hackage.haskell.org/package/conduit-algorithms-0.0.7...
I have found that most managed languages generally come within 2-5 times slower than C and C++. Which is good enough for me.
cpp webserver + sqlite database + cygwin overhead? 9MB RAM lua worker + websocket client? 4MB RAM JVM server + static html page + websocket relay + kurento api? 500MB RAM
With the exception of the cpp webserver none of these tasks are CPU or memory intensive yet the JVM is still off by orders of magnitude.
It's ok if it's the only application running on a server with multiple users but there is just a single user and that's me.
Python itself is a slow language but it has a lot of fast packages, so it shows poorly when you actually write your benchmark in python.
Haskell is a faster language but because it is high level there are more pitfalls you’ll get into if you don’t know the ins and outs of getting fast code out of the compiler. The guys who write fast benchmark code aren’t ‘average’ developers.
So in Haskell it’s “the code is slow and I don’t know why” vs python “the code is slow because python is slow, import fast package someone wrote to speed it up.”
All that said, I think Haskell is the better language but you have to put in more effort to get experienced in it before you see returns on the investment. Python has a shallower learning curve and an easy way to get “good enough” performance (a bit slower than C).
The best criticism in the article is of the multi-core deficiency of python’s interpreter. But that’s only briefly touched on. It isn’t a friendly environment to write complicated multi core code.
The article we all reply to exactly claims that as soon as you don't use e.g. NumPy, it's not "good enough" anymore, and I agree with that. The article also argues that e.g. JavaScript isn't more in the same category with Python, but much faster, even if it's not less dynamic.
I think the reason for JavaScript's speed vs. Python's slowness is obvious: there were wealthy companies involved, which were, due to competition pressures, motivated to speed up their own JavaScript engines.
To get to the point where Python has similar speeds somebody would have to be motivated enough to invest heavily, and then it could happen. As far as I know, there aren't technical limitations against that.
So much Python has historically been tied to CPython's specific ideosyncracies that there is significantly more onus on the upstart VM developers to maintain compatibility with paralinguistic behaviour (things like expectations regarding object destruction sequencing).
Javascript also has the "benefit" of an appalling base library, while the base library that CPython provides is quite large, and growing.
C++ would probably increases his development time significantly compared to Haskell.
Things like C#, F#, Java, Kotlin, Nim, Lua would be more natural things to turn to when you want something "Easy" like python but faster, I think.
So no, it's not easy in a similar fashion that pointers or double pointers in C/C++ are not easy. Or understanding call by value vs. call by reference semantics are not easy. The list goes on. It's probably the largest barrier to learning the language.
But once you get a handle on the evaluation model it becomes a lot more natural. At least that was my experience, maybe it is not typical.