Array Programming with NumPy
nature.com
nature.com
And because of this, numpy has to be much larger and less orthogonal. It needs to provide a wide range of functions that ensure that the user will never have to drop into regular Python and suffer a 100x performance hit. This is unfortunate, and prevents numpy by having the beauty and orthogonality of true array languages like apl, j, or kdb+/q.
This problem makes it difficult for new scientific users of Python to write fast code. Everything can be fast and then you add one if statement and the performance falls off a cliff. I've helped many people in various labs as this is a recurring problem that researchers face.
As such, the ease of use that Python supposedly has for scientific computing is a lie. The APIs are extremely complex, from numpy to Pandas to statsmodel and scikit. And they have to be, because scientific Python can't compose code due to performance reasons.
In contrast, Julia is much easier to learn, easier to use, and is pretty much better than Python for scientific applications in every way imaginable. In other words, it is strictly superior. The performance is consistently fast and array manipulation is built into the core language, unlike with Python. You are free to write Julia code instead of shelling out to C.
I'll probably get down voted, but Python should not be used for new scientific computing projects. I've had tons of experience with ml/financial Python and after trying Julia, I'll never go back willingly. The difference in speed, complexity, ergonomics, and expressivity is so stark it's hard to imagine why anyone would choose Python for a greenfield scientific computing project in 2020, especially since Julia can trivially access all of Python's ecosystem.
Once you move to the other end of the spectrum where the size of the problem requires a supercomputer, that is when the Python ecosystem becomes clunky and painful. I think that could be an area that Julia shines in. But even that, I'm not sure if Julia is a clear winner yet, as the field is still largely dominated by C, C++, and Fortran.
Also, check ref. 56 in the article. They did cite Julia, among others.
Python is so slow it is even unusable on trivially small data sets. Also it uses a massive amount of ram. I loaded a 2gb csv file the other day in Pandas, and it consumed over 40gb of ram. That's totally unacceptable in my view.
Try to make a plot. Last time I checked, the time to first plot was still a pain point of Julia.
> it's harder to use and less expressive than Julia.
I agree that Julia is more expressive because it's a Lisp in disguise. This is definitely a plus. But in practice, I don't find Python to be "harder to use." There is good consistency among the mainstream packages in terms of API. For example, if you have a NumPy function np.mean, you can assume that Pandas would have a method .mean() for the DataFrame, and Dask would have that for DaskArray as well. Not always, but things are moving toward that direction.
> I loaded a 2gb csv file the other day in Pandas, and it consumed over 40gb of ram.
Have you tried the `engine='c'` option when calling pandas.read_csv? Pretty sure there is also another option of chunking that may be useful.
But after paying that startup cost, the speed boost can be transformative in terms of the flexibility and dynamism that it buys. I've been using Pluto.jl lately, and interactive data analysis feels like way less of a burden than working in jupyter.
As to harder to use, writing fast vectorized numpy code can take a lot of mental effort. For one project I ended up forcing things into this 5-D array broadcasting mess, where in julia it's way simpler to just write intuitive code that's performant without bending over backwards to avoid explicit loops. For numpy code you can "just" drop down to cython, but for pytorch you can get stuck with thinking up some clever broadcasting solution.
Julia's GPU programming stack is fantastic and very advanced btw.
> New generation languages, interpreters and compilers, such as Rust [55], Julia [56] and LLVM [57], will create new concepts and data structures, and determine their viability.
I don't disagree with the statement, but there's not a lot of meat there either.
- It exposes an API and I want as big an audience as possible and widespread familiarity with Python means that there is a lower barrier to adoption;
- The interop works both ways - if really someone wants to use Julia they can;
- I felt the JIT startup times would lead to a poorer user experience in some circumstances.
Essentially, I accept it's not as elegant as Julia but I'm prepared to accept the extra complexity etc behind the scenes to try to help users.
Not saying that this is great (or even the right decision in this case) just that there can be straightforward reasons for choosing Python.
EDIT: You can write loops explicitly and have OpenMP style for-loops and they are fast.
Are there any comprehensive benchmarks that show Julia outperforming Pandas or PyTorch or SciKit?
Obviously pure Python is terrible. But the library algorithms written in C seem fairly competitive.
I’m a fairly boring user who doesn’t do new science, and is fine just composing existing boring algorithms to solve problems in my subject matter domain.
To my knowledge if you stick to things that calls optimized C and fortran code, it's a draw between the compiled code and Julia.
But even boring problems ends up doing things that are easily expressed in a loop, but ends up being a hard to read chain of pandas.
I agree that Julia code is aesthetically superior to a long chain of Pandas code. But at this point I’m used to reading a bunch of chained pandas code. Often I think of myself as more of a Pandas programmer than a Python programmer.
Numpy is not just a helping hand to python; instead, it is this efficient, giant, comprehensive engine, and python is merely the interface used to feed it data, operate it, and read from it.
Of course, under the hood, it's still just language and library. But the mental model is quite different.