Why Python rocks for research
stat.washington.edu
stat.washington.edu
- Matrices are a pain. The r_[] and c_[] operators could be a reasonable replacement for Matlab's elegant matrix construction syntax, but they do not work as expected (as smart hstack and vstack), instead doing something completely different and inconsistent for vectors and matrices.
- Tensors are a bigger pain. Matlab has a very well-defined semantics for operations like permute and reshape; in NumPy these operations sometimes create just a view, at other times they reshuffle the memory contents. I know the idea was to "protect" the user from having to know the memory layout of data, but this idea is bad.
- Ipython is great in every way except when it comes to reloading parts of your program. After any tiny change to your code, the only safe thing to do is to quit ipython and start it again. All the other options (run, reset, reload...) make some secret and wrong assumptions on what you want to reload. In contrast, this works flawlessly in Matlab.
There's nothing intrinsically science-apt about Python/Perl, but Ruby and friends can't compete when it comes to the programming environment; that's what counts.
I think that making some tools available in a totally different language (maybe something functional) would be much more useful, because it would allow for a very different approach to the problem if needed.
In a perfect world, we could also have the option of using light wrappers around OpenCL matrix libraries, and push the linear algebra to GPUs that eat matrices for breakfast.
This is exactly why I'm focusing on Python lately. Ruby is a great language but I don't want to be pigeon-holed as a web guy forever. I've already done over 10 years of web dev and I'd like to try out a couple of new problem domains before I kick the bucket.
Are there plans to make Qt feel more natural on Gnome and OS X?
There are an awful lot of languages that provide iterators, a powerful set of data structures, extensive libraries and facilities for structuring and maintaining large codebases. .Net languages (maybe F# would be good for this?), Java or most of the emerging languages for the JVM stack, Ruby (which is generally considered to be "different but equivalent" to Python), and so forth.
Still, how many of them have a fast interactive interpreter ("command line") with a decent usability? How many of those provide good libraries for numerical as well as symbolic math? With an API that is easy to write, to understand and to extend?
Python may not be the only language with those qualities, but there aren't many languages (and ecosystems around them) which can compete on all those areas.
Python seems to be one one the few "best fits" for scientific applications.
Haskell seems a perfect fit for mathematical use and while I haven't used it in a couple years, I would hesitate to suggest it due to a lack of mature library options, difficulty of FFI and perhaps a steep initial curve.
Scala is a good language for an entire application but provides too much scaffolding for scientific applications.
R is fairly widely used but is also itself very quirky.
It counts on Freedom, Readability, Documentation System (including lhs2tex which will turn your "integrate f 0 a" into \Int_0^a{f dx}), High-level vs low-level, Standard library (including hackage/cabal), Data structures, Module system, Calling syntax, Default arguments (currying), Multiple programming paradigms (there was a saying that Haskell is the best imperative language). It partially counts on most other points.
Myself, I choose Haskell for my research project, as it was best language on (expressiveness times safety) scale. Strong type system certainly helps sweeping out errors.
I've shopped around a lot for alternatives to Python and looked at Ocaml, F#, Boo, Clojure, Common Lisp, Haskell and Scala.
There are a couple of things still keeping me with Python:
* The REPL, especially IPython,
* Numpy & Scipy,
* Networkx & igraph,
* jpype for fairly seamless JVM integration (this way, I can interact with Cytoscape), rpy for fairly seamless R integration,
* ZODB,
* IPython's parallel processing framework.
The .net environment probably comes the closest to providing everything here, but the REPLs need a lot of work and the graph libraries have more complicated interfaces which make them a pain to use on the command line.
Python is not without its warts but when I recently had to spend time with Matlab again after a few years, I was reminded of just how nice the Python ecosystem is in comparison.
Update: I should add that if you rely on one of Matlab's toolboxes, you might not find any decent alternatives outside of Matlab. You can always use a bridge like MLabWrap to access Matlab from Python.
In theory, .Net could displace Matlab or Python as the canonical platform for scientific researchers. And in theory Python could displace PHP as the canonical platform for classic CRUD web apps. In practice neither is likely to happen, no matter how much we might or might not wish it to.
(disclaimer - i work for Enthought).
Personally, I rather like Python XY ( http://www.pythonxy.com/ )though. It is totally free, open source, and for small little scripts I am a huge fan of the Spyder IDE. It lacks some of the features of bulkier IDEs so I also use Eclipse from time to time. But Spyder is light and much faster than Eclipse with every feature I would want when working on small projects that can be contained in just one or two files.
Java probably does and .NET may have something like this but I don't know of them or their amount of documentation.
But I think Python probably has the most active community as its been used as glue for a long time in this space.
IMHO, languages matter less than the available libraries, and in my experience only Java matches the depth of the Python ecosystem.
That Python is a nice language to work with, that's just a bonus.
But if I need to do scientific or linguistic programming, Python is absolutely amazing. SciPy, NumPy, Matplotlib and NLTK support a rich and deep ecosystem. And it's vastly nicer programming environment than Matlab and Octave.
(Of course, GNU R is also pretty useful if you're doing pure statistics.)
but I also get the sense that this is the first time he's seriously delved into a dynamic programming language. much of what he's saying about Python is exactly what bioinformaticists were saying about Perl in the late 90s / early naughts.
Context and corner cases maybe makes readability suffer. For example, what does `m///g` do in scalar vs. list context? I don't recall.
- it's a (very) good thing, in that there's -a- -lot- of expressive power hiding in those corner cases (esp list v. scalar context). in fact perl is nearly unique in the degree to which it embraces, rather than seeks to quash the potential for nuance and flexibility to be found in the various eddies and whirlpools that lie "between" non-whitespace elements of code. this is something that goes to the deepest intellectual roots of the language (to be more like the way the human brain thinks, rather than the way machines think).
- it's a (very) bad thing, in that the same "power" for expressiveness just represents little more than an endless stream of banana peels to most first time users (understandably sending many of them running into the arms of the other major dynamic programming languages). and that it just makes it too damn easy to write (barely) functional, pothole-laden frontline scripts (such that Perl is partly responsible for "scripting" being such a dirty word, in many quarters).
another aspect people don't like to talk about its nearly coprophilic fondness for not just nuanced and context-sensitive, but intentionally obtuse syntax (the impossible to remember "special" variables such as $[, $;, $' etc being probably the worst examples).
this also (sadly) has a lot to do with the community's deference to the aesthetics of its original authors (and is also plenty understandable, given the extent to which unabashed ugliness -- as personified by makefiles, shell languages and macro-laden C and C++ -- ruled the day at the time).
unfortunately it also blinds a lot of Perl folk to the fact that beloved muse just happens to look awful darn cluttery (or worse) to many reasonably intelligent people who come from more "modern" programming backgrounds (or who at least came onto the scene comparatively recently).
Also, you want a non compiled language. Most of the time you do interactive programming and change parameters on the fly, according of the result of the analysis.
Finally, matplotlib, one of python's most complete graphic library, is a breeze to use. Making graphs in an interactive way with java or .net is simply impossible.
I am a neuroscientist and most people in the field use Matlab. I use python (in fact I use ONLY open source software, by choice). It's amazing how many advantages python gave me on my daily life.
Yes I am a bit overreacting since the blog post is very well written and I actually agree 100% with the content. But please people: respect the PEP8 [1]. It makes your readers feel at home while reading your code. It is very important if you want to get new contributors to your project. See [2] for instance.
[1] http://www.python.org/dev/peps/pep-0008/ [2] http://www.dataists.com/2010/10/whats-the-use-of-sharing-cod...
there are some things in pep8 that are bad for science, the spaces around operations, and also the 80 chars to a line... scientific expressions are often long and complicated, yes you can do it while adhering to pep8, but its kind of a PITA
Furthermore having 80 chars is great to have vertically split editors with the code on one panel and the tests or the documentation on the other panel.
I disagree. There are many cases, especially after 2-3 levels of indentation, where 80 characters is an unreasonably narrow space. I don't have a strong preference for reading code within 80 characters. And I'd much rather comments use 80 characters plus indentation rather than worry about whether I've got my screens vertically split.
> Furthermore having 80 chars is great to have vertically split editors with the code on one panel and the tests or the documentation on the other panel.
Having lines here or there that go beyond 80 characters doesn't completely prevent you from doing this, and having an entire statement on a single line of code makes line-based tools like grep or kill-line more effective.
Limiting to 80 characters is a good idea, but it's easy to see that there are a few significant tradeoffs, and that someone trying to get something done is not going to want to bother.
> I disagree. There are many cases, especially after 2-3 levels of indentation, where 80 characters is an unreasonably narrow space.
Some people would say that after 2-3 levels of indentation you should be looking at refactoring your code. Probably to pull something into a separate function/method.
But in science, we have longer and more complex expressions in general.
Our priorities are different.
Code should not be Documentation.
Further nobody trusts anybody's code anyway unless it's just a couple of trivial calls to a pre-vetted software package like IRAF, AIPS (to name some astronomy related one), or LAPACK. So generally they don't want your code. the exception is grad students trying to apply your old work to new data because they aren't in a position to be trusted with completely original research yet.
Yes it'd be nice if every one had great readable code and handed over the 2 terabyte data sets that it needs without batting an eye. but in practice code quality is pretty low on the ladder of "things that get in the way of collaboration"
Hence code should be both published, well documented and readable.
Existing techniques in general really. Fields where the interest is the data and the implications of the data. Fields like ML and NLP where the algorithm/technique is the thing of interest then yeah sure the code is important.
Code is for humans to read, that it compiles/interprets to a program is a side effect. Otherwise we'd all be passing around binaries (or byte encoded files) with our thick stacks of documentation.
Well, we should be able to, but no, we can't, precisely because we don't get code - we get binaries.
Or maybe not just yet.
Why do you assume I buy any software?
The claim "code is for humans to read" does not logically lead to claim "code is the only thing for humans to read". There are different kinds of humans, programmers, maintainers, end-users, and idiots are some. You're a member of the later.
* lack of elitism in documentation (e.g. there are always plenty of examples)
* lack of elitism in conventions for code use: everything "just works", generally without any boilerplate
* installing libraries is a snap, and the whole module organization system is intuitive and elegant
* assumption that anything that's not a script is a library
* documentation conventions (doctests, e.g., are a nice stepping stone to good documentation _and_ code testing)
* the "there's only one way to do it" attitude
* large standard library
On the other topic: you are describing the way research works "today", which is actually pretty poorly (why, e.g., does all data need to be surrounded by so many words of introduction and discussion? why can't I just add something to someone else's work like I can add to an open source project?). This model of research will change, at one point or another, to resemble the much more efficient, effective, and fun, open source project model.
> The preferred place to break around a binary operator is after the operator, not before it.
I'd be interested in hearing the justification for this rule. I think that leading a continuation line with the binary operator makes it super-clear that it is a continuation line. What is the benefit of the preferred style? Compare:
if (the_result_of_this_function(on_this_arg) == 10
and this_overly_descriptive_boolean):
do_stuff()
if (the_result_of_this_function(on_this_arg) == 10 and
this_overly_descriptive_boolean):
do_stuff()
To me, the first one is quite clearly a continuation line (no statement can start with "and"). The second requires closer inspection.But I don't think anybody will complain if you use either of the them whereas 160 chars long expressions with no spacing between operators and funkyCamelCasing all over the code are just show-stoppers when I want to contribute a patch to a project.
if the_result_of_this_function(on_this_arg) == 10 \
and this_overly_descriptive_boolean:
do_stuff()
Indenting the second line of the if statement would, at first glance, indicate that it's part of the block instead. Then again, it depends. If it was the header of a def statement, I would follow the PEP, e.g. def __init__(self, width, height,
color='black', emphasis=None, highlight=0):
On a side note, I once did the "Art & Logic challenge" [http://www.artlogic.com/] and they use guidelines that apply to several languages, e.g. you would use the same formatting style for C++ and for Python, if at all possible. Much of it flies in the face of PEP 8.I did my whole phd in matlab.
EPD is much cheaper and is free for academics
even if it weren't free, I would use it anyways.
but it really isn't why is EPD better than matlab, it's why python is better than matlab. matlab is a domain specific application with a domain specific language. It doesn't work well with things outside of its domain.
python is a general purpose language (And as such, has good general purpose constructs) but it happens to have excellent scientific and mathematical libraries. This is useful when you actually have to apply your research and build an application.
numpy is also better for large data, because slicing arrays does not create copies of them (you can make it do so if you want to, but it doesn't by default) in matlab, slicing large arrays can cause you to run out of memory.
Cython makes it really easy to start out with python, and then optimize your code down into C.
with python you can run your calculations over a massive compute grid. Use messaging libraries like PyZMQ to distribute your data and result, and build real time GUIs to consume the final results.
- a matlab cluster is quite expensive
- chacko - another enthought python library which is free and open source is great for real time datavisualization, matlab does not have anything equivalent.
- python has a large number of messaging libraries, with matlab I think you're stuck with MPI.
Matlab always made me feel limited. I would work on a problem, and then reach a point where Matlab could not do what I needed to do.
That rarely happens to me with python.
use IPython, not just python shell for interactivity.
also checkout 3d datavisualization with mayavi, that stuff is really awesome.
there was a mayavi tutorial, and the files are available at the link
or bpython
Ah, you mean "Chaco" -- it's easier to find with the correct spelling. :) http://code.enthought.com/projects/chaco/
I use python, ipython, matplotlib, numpy and R. I call my R scripts directly from python using rpy.
A single example:
f(x):= x^2+3x+7;
Maxima provides: Symbolic computation, blas and laplack integration for numeric algebra, 500 pages manual in several languages, a complete library for statistics, differential equation, calculus, series. Graphics with matplotllib. Also maxima language is not much complicate that python:
for i in range(10):print ii versus for i:0 thru 9 do print ii;
[i2 for i in range(10)] versus makelist(i*2,i,0,9)
But Matlab libraries are greater than python and maxima.
Also see wxMaxima, which will (among many other things) produce LaTeX for you.
- installing all these packages on (any) system is painful. Different versions don't play together or don't work (yet) on some platform and or architecture. This stems from my own experience of getting a version of python to work with numpy, scipy, matplotlib, opencv and PIL on a windows, mac, and linux machine. No 100 percent success yet on any platform.
- central and consistent documentation. Even for very simple cases, I got a bit of a headache. I encounter a python print statement for the first time that obviously differs somewhat from its c printf cousin. I google "python print syntax" only to find that the first xx hits, including the official documentation, do not cover the full specification of this statement. I fear the moment I might actually need detailed information on something less trivial.
- Numerical integration is more accurate in Matlab.
- Visualization capabilities of matlab are more powerful. But who knows, perhaps there is yet another package floating around :-)
- Matlab may not have advanced data-structures, but it is a rapid prototyping tool, for testing ideas. If I need to write an actual application, I will use a tool and language geared for that task.
http://docs.python.org/reference/simple_stmts.html#the-print... lists eveything about `print` - it takes stuff in, converts 'em to strings, and sticks in on stdout. It doesn't refer to prinf-like formatting because that's for strings in general. If you weren't aware of this, you probably should have been going through a basic Python tutorial, rather than just jumping into the middle of things.
Python's documentation is the best I've encountered so far, and I find good docs to be an important value in the community, as well. I guess YMMV, though.
agreed on documentation
actually I think python's visualization capabilities are more powerful, have you looked at mlab? the 3d capabilities there are insane
I use python because I can do rapid prototyping, and turn it into a full application with the same code base.
did you ultimately go back to matlab?
I wanted to venture beyond Matlab because for what I am currently doing the environment and language is to limited, yet I do not wish to prototype in C++. Python together with some libraries seemed to be a deal in heaven. I also thought it would function as a better stepping stone towards an actual application.
Moving your development to a linux machine will clear up all of those issues.
Also 'easy_install' should get you all of the packages you want.
I completely disagree. Reusable Matlab code has been my holy grail for the last couple of years. The key is to break out specific functionality as subfunctions. When these are abstracted and generally useful elsewhere, then they become new tools for the toolbox. The subfunctions also make great starting points for repurposing code. This layout results in much less work.