Matlab vs. Julia vs. Python
tobydriscoll.net
tobydriscoll.net
It's just not designed to make good pipeline tools that are maintainable, easily tested, and easily refactored, and never was. Its ability to handle things like "easy and sensible string and path manipulations" are ... rudimentary, at best, and a weird pastiche of C, fortran, java, and whatever else language was faddish when that feature was added.
I have more or less one real annoyance with the technical content of matlab. (One-indexing i can live with):
size([1])
ans = [1 1]
which is flatly wrong, and numpy gets it right: In[0] np.array([1]).shape
Out[0] (1,)
Python is actually a general purpose language which has a mature scientific stack, and i feel more secure in my numerical computations there because i can have a robust test suite and command line entry points into my code that increase my confidence that my code's doing what i think it should be, and makes it easy to use.Packaging is more or less a coin-flip. Python packaging is a giant faff; matlab packaging is nonexistent (you have to vendor every dependency yourself) and expensive (your users have to shell out for the toolboxes you use, and/or have the MCR installed and can't edit your code).
I'll maintain matlab when i have to, but i don't enjoy it very much.
Part of when I help on-board newly graduated colleagues (typically engineers and scientists, not hired as programmers) is to reassure them that they can get a MATLAB license if will help them get productive quickly. This has happened once or twice. Many of them never bother, as they get busy enough with CAD and basic design work, that they don't really find a use for scientific computation. But to an increasing extent, they're willing to make the hop to Python, possibly just because it's a popular buzzword, but in any event, they are able to get themselves up to speed pretty quickly.
A single user license and a few toolboxes brings in more cash than all the student licenses my University probably used that year.
You would find it an impossible task to track down an automobile manufacturer or a platform level supplier which doesn't leverage Matlab & Simulink all over the place.
Why is it wrong?
Edit: typo
Amusingly this also leads to things like
x = 1
x(end+1) = 2
x == [1 2]matlab:
size(1)
ans = [1,1]
python: In[0]: np.array(1).shape
Out[0]: ()
0-tensors are not 2-tensors, either. You can read matlab's result as "scalars are matricies that have one row and one column", and that's nonsense -- matlab effectively lacks a scalar numeric type; under the covers, everything is a matrix.what numpy calls it is more or less moot.
This is not a CS data structures 101 thing, this is a "there are proper names for mathematical entities" thing. I think you're confusing representation and existence, maybe.
That's my point. Have you ever used a proper tensor? I mean, the mathematical entity: https://en.wikipedia.org/wiki/Tensor
[Edit: You may care about vector algebra but 99.999% of the users of numerical arrays do not. I just think that the proper term in a general discussion about data structures in numerical computing like the one here is “array”.]
[Edit2: in case you’re really not familiar with the abuse of the term, people do now often use the word tensor to talk about multidimensional arrays that have nothing to do with the mathematical entity called tensor. And I don’t think either matlab or python pretend that they have tensors.]
Matlab will try to optimize and avoid a copy if the function does not modify the argument. Matlab is also (sometimes) smart enough so that A=f(A) will modify A in place instead of making a copy.
This is what I expect from a math oriented language. Maintain the illusion of referential transparency but optimize under the hood if possible.
Also Matlab has a reasonable JIT compiler. And a good debugger.
I no longer use Matlab but it is a very productive environment for scientific computing (simulations, exploration).
but... what else is there, in life? Flipping big matrices around is nearly everything I do, and the python stuff seems too cumbersome for me to bother.
import numpy
import scipy.sparse
import scipy.sparse.linalg
just to begin writing something.You cannot create a literal array without calling a function. You cannot concatenate arrays without calling a function, and moreover this function has a different name depending on whether your arrays are full or sparse. The @ notation for matrix products is horrendous (and .dot is even worse). Arrays are not first-class objects of the language, and you have to use an external library. This is more infuriating by the fact that other, more complex data structures like dictionaries or strings are natively supported, even if they are mostly useless for numerical computation.
Compare the clean matrix flipping in octave
z = [ kron(speye(rows(y)), x) ; kron(y, speye(cols(x))) ]
to the python monstrosity z = scipy.sparse.vstack([
scipy.sparse.kron(scipy.sparse.eye(y.shape[0]), x),
scipy.sparse.kron(y, scipy.sparse.eye(x.shape[1]))
]) from scipy.sparse import eye, kron
from scipy.spare import vstack as vs
z = vs([
kron(eye(y.shape[0]), x),
kron(y, eye(x.shape[1]))
])The thing that kills my soul is the need for the "vs" function. Why aren't the symbols [] not enough ?
It has some key strengths though:
1. RStudio IDE as you note. It's a really great, focused IDE for doing most of the things people do with R.
2. Shiny. Such a well conceived and constructed toolkit for building interactive apps
3. The package ecosystem: lots of really good quality, high performance packages
4. RStudio the company, who contribute a lot to the community - both open source (RStudio IDE, Shiny, tidyverse, ...) and commercial (RSConnect, Package Manager).
From a language design pov I like Julia over Python over R. But for number-heavy computing I prefer the R ecosystem overall.
https://rstudio.github.io/reticulate/articles/rstudio_ide.ht...
Related tweets:
Also, it used electron.
So long live RStudio, I guess.
Tidyverse is massively overrated if you ask me. The good parts of it (dplyr and ggplot) are nice for interactive work. And that's about it - if you're deploying the code in an application, you're best off sticking to base R as much as possible.
MATLAB is adored in academia for a number of reasons. It is easy to make readable small scripts for in-class examples. The debugging feature/IDE is easy to navigate. The school pays for the licenses; there is no overhead work to compile or download packages (unless you want to do something 'fancy').
I took a ML course that was taught in Python. All my Numerical Analysis and Modeling courses relied on MATLAB for examples and homework. I (as a programmer outside of just the math world) picked Julia for research. Now I do much more theoretical research, as I did not enjoy mixing coding and mathematics.
A fellow student, developing PDE solvers in FORTRAN was told by a mentor to get it to work in MATLAB first and then move on to faster languages.
Happy to answer any questions :)
The other thing I found weird was the complaint about matrixes objects being deprecated. After the addition of the '@' operator it is really the same as matlab, except their default is a matrix, while for numpy it's an array. As a side note the author complains about the '@', what about matlabs stupid decision to use the most easily overlooked ascii character for distinguishing between matrix and element wise operations. In my experience, almost all the time a matlab calculation returns weird or garbage results, the bug search is a "find the missing '.'
Python 3 has supported unicode variable names for 12 years. Not all of unicode is permitted, but all the useful bits are.
In [1]: α₁ = 1
File "<ipython-input-1-3c2973844bb9>", line 1
α₁ = 1
^
SyntaxError: invalid character in identifier α₁ \alpha<TAB>\_1<TAB>
vₓ v\_x<TAB>
H₀ H\_0<TAB>
χ² \chi<TAB>\^2<TAB>
Aᵀ A\^T<TAB>
When I went to do the python demo above, IPython also tab-completed the `\alpha` above... but to get the ₁ the quickest and easiest way for me to get it was in my Julia editor.Fair enough in some contexts, but I don’t find that to be the nicest general-audience feature. YMMV.
help?> α
"α" can be typed by \alpha<tab>
Code is for reading much more than writing — especially for a new person. I maintain that matching the canonical form of
the algorithm I'm implementing will help them gain their feet faster.Good example:
https://github.com/JuliaStats/Distances.jl/blob/c21aab0fae30...
vs.
https://www.npmjs.com/package/haversine-geolocation#introduc... (note the mathematical formula on that page and the JS code compared to Julia's)
Sure someone can pattern-match the code to the already-derived line in a reference book somewhere, but that doesn’t help at all with reasoning about the geometrical relationships or developing new algorithms, following control flow, building abstractions, improving precision, handling edge cases, ...
It’s mostly helpful when you want to treat your formulas as an externally defined black box.
* * *
Aside: The problem in this particular example is that spherical coordinates and spherical trigonometry are just not a very good formalism for calculating anything, in either theory or practice. Unfortunately cartography, geodesy, etc. are bound by tradition and there hasn’t been much effort to switch them to better tools.
I’ve been trying to read a spherical trigonometry textbook (Todhunter, 1878) the last few days and following along is a huge pain.
Much better is to switch to cartesian coordinates or stereographically projected coordinates, and then use vector methods (and skip writing explicit coordinates in your code to the extent possible). All of the proofs and derivations get nicer, with geometrically meaningful steps and conclusions. Now your points can just be called p and q or a and b (or if you have a lot of them and aren’t pressed for space in each expression, points[i] and points[i+1]), and the coordinates stay internal.
In addition to using clearer code, the calculations will also be faster, more precise, with better numerical stability, take up less memory/bandwidth, ...
Here’s some general math for an arbitrary-dimensional sphere; for just the 2-sphere everything is simpler. http://geocalc.clas.asu.edu/pdf/CompGeom-ch3.pdf
The point isn't the algorithm itself. The point is just how using unicode allows you to match the style of an arbitrary algorithm out of a textbook.
In general I see code or explanations relying on Greek letters as no better than ones with English words for names.
That's a lot of information for such a compact representation. And that's mostly unconscious, which is great. I really found the Julia example above to be far, far easier to understand than the linked JS (though I think that JS was not particularly great).
To me, it seems like it'd be nice-to-have, but alone, I don't think it would be enough to make me switch languages.
I mean everyone who uses my python code is going to have to be trained in pip, virtualenv, mypy, sphinx, and git. Anyone editing code written by and for mathematicians or scientists is going to know LaTeX, it's like the html of academia.
And as pointed out, you don't need any IDEs, this is even supported in the REPL. And using a specialized editor for a language/environment is hardly unusual (not that it's needed).
Besides that: how do you actually input a significant number of greek letter w/o consulting a layout-picture or wikipedia?
I learned Mathematica and MATLAB in my math and science courses (Physics) and was going to learn Java before I dropped Computer Science. Interesting I could probably replace all those with Julia now.
Flipping through a few of his notebooks, I guess he's from the part of academia that was core Matlab-land, engineering-style number-crunching. Mathematica is mostly people doing symbolic things, and people doing lots of stats are yet another story.
The vast majority of Matlab's vaunted numerics performance comes from using MKL instead of OpenBLAS. However Intel has made MKL free software. Meaning that you can easily build NumPY on top of it. Numpy+MKL will compile down to virtually identical assembly as Matlab.
There's very little reason in this day and age to pay Mathworks such an insane licensing fee.
If I recall, a big chunk of the code that controls Toyota engines is generated from a massive Simulink model.
Not at all. At Lockheed (probably one of Mathwork's biggest customers) the big use case is the toolboxes. Scientific/engineering packages like signal processing, radar, phased array, embedded/VHDL are 2nd to none and are used daily. Some sites also used a lot more of the Simulink side for modeling & simulation.
[0] https://www.mathworks.com/help/signal/ref/signalanalyzer-app...
[1] https://www.mathworks.com/help/comm/examples/rf-satellite-li...
[2] https://www.mathworks.com/help/fusion/examples/multi-patform...
[I agree about the marketing aspects and the huge waste of university resources buying into that rather than funding development of free software.]
I'm an AI researcher / practitioner. For me code accompanying papers is very useful and usually this code is in Python. Occasionally it's Matlab but let's be honest, who cares about those papers :). I'd love to use Julia but the package support just isn't there. Ironically people like me are supposed to be writing this code but with a demanding job and a family it's not likely I will be improving their DataFrame effort anytime soon.
Anyway the MAIN reason I use open source software is because if it isn't working correctly I simply fix the code myself. This isn't possible in the proprietary world. Why would you trust your research or production work with code you can't see and edit?
There's been a lot of talk about documentation. Docs are secondary sources, like WIRED, read the code if you're serious about being correct. Even (especially) hired hands make mistakes and fail to write good tests.
This article reminded me of the fictional Simpson's news article "Old Man Yells at Cloud". It's funny, and he may have a point, but it has no relevance.
I'll never forget when I took a controls class and we were given an option to use Python on our own or matlab with guidance and support from the professor and TAs. I chose Python since my background was slightly different that most of the students, but most everyone chose matlab. It was highly amusing to watch the whole lab suffer for days because no one understood matlab's semantics. (The base framework had been written by some expert who retired and made extensive use of both handle and value classes. Problem was, no one still involved with the clas knew what the difference was.
Meanwhile, I spent about 6 seconds longer writing out np.dot a few times.
Matlab is good for math. Most software (even math heavy stuff, EM simulations, etc) has little math (in terms of source, no execution time).
mmmmhmmmm...
The car analogies in the blogpost are not particularly useful... Why do people feel the need to dumb down a topic with off-the-wall analogies? Talking about Julia like it's Tesla is laughable. Tesla is a huge innovator, Julia is another tool that does similar things to the other tools. The apt analogy for Julia would be a new ICE company, not a new EV company.
As an example, most physics programs don't even have an introductory statistics class in their curriculum.
Some engineering disciplines use it more, and it is often relied a lot in industry. But most engineering research makes little to no use of it.
And publication-wise, it's a bit of a chicken and the egg problem: since most reviewers are also not well-versed in statistics, they are less likely to accept a paper that uses statistical methods to confirm or reject a hypothesis. Thus, there's no institutional pressure for a physicist to learn statistics.
(And a graduate course in statistical mechanics does not count, saying that as someone who used to use that excuse)
Not to mention that if experimentalists actually used statistics, they would have to describe their experiment in more detail when publishing for their statistical analysis to make sense. And in my experience, they really prefer to list as little about their experimental setup as they can get away with.
>And a graduate course in statistical mechanics does not count, saying that as someone who used to use that excuse
Heh heh. Most physicists I know would invoke "But we use probability in quantum mechanics all the time!"
I have absolutely heard that one before :-)
Some can't even work: take a look at this one[1], which is part of the current release of BioC: it will never work because the hardcoded host there is no longer functioning and the author wrote that for their master's thesis and have since moved on.
In addition, there's no centralized bug reporting, and some packages use GH issues while for others you may need to email the author and hope for the best (sometimes works, sometimes doesn't).
I use BioC regularly every day because I have no other alternatives, but its main advantage is the sheer numbers of packages available for almost every bioinformatic niche, rather than the quality of the packages themselves.
[1] https://github.com/AllenTiTaiWang/anamiR/blob/7a7a133c553f8c...
And an important aspect of scientific computation is data visualization, an area in which R is decidedly more advanced than other languages.
[HPC clusters on which you'd run CFD typically don't let you run for more than days, let alone months. There's probably no reason you couldn't use R the same way as Python for such things.]
I don't know that that's strictly true but it gives you a general idea for why things are the way they are.
Nothing beats Matlab/Gnu Octave for me when I just want to make some quick calculations.
Though you can also use 'import' to keep the namespaces separate.
h2o.ai has been working on data.table for python (https://github.com/h2oai/datatable). Hope it matures quickly
http://www.math.udel.edu/~driscoll/SC/ or on github https://github.com/tobydriscoll/sc-toolbox
But I can’t say I agree with much of this.
> One function per disk file in a flat namespace was refreshingly simple for a small project, but a headache for a large one.
This is an absolutely horrible bonkers limitation for a “small project”. It’s “refreshing” like someone constantly dumping buckets of ice water over your head. Being able to define functions in the repl, export multiple functions from a single file, etc. are things I pretty much can’t live without in an interactive programming language. Needing to make every tiny utility function into its own file adds SO MUCH FRICTION to basic prototyping workflows.
The result is that when working with Matlab I try to make as many functions as possible into lambdas, e.g. square = @(x) x∗x. But these are very limited in practice, and the workarounds to make a function work as a lambda often compromise readability, performance, correctness, and functionality.
+ + +
Is an occasional V.conj().T @ D∗∗3 @ V really that much worse than V' ∗ D^3 ∗ V?
I find that in even small programs, syntax complexity and clarity is dominated by logic flow rather than matrix multiplication. My Python programs end up dramatically easier to read than my Matlab programs, including the almost-pure-numerics parts of them.
If concise built-in syntax for numerical operations were our highest ideal, we’d all be using APL.
> exists a matrix class, and yet its use is discouraged and will be deprecated
I think Driscoll doesn’t understand what’s going on here. There is plenty of discussion for someone who searches about the problems with the matrix class. The matrix class dated from a time when there was no @ operator, so ∗ was used for matrix multiplication instead of elementwise multiplication. This made it inconsistent with everything else and artificially limiting in obnoxious ways. Now that Python has an @ operator it is no longer useful or necessary. This has nothing to do with matrices being important or not.
> Matplotlib package is an amazing piece of work, and for a while it looked better than MATLAB, but I find it quite lacking in 3D still
YMMV, and maybe I’m spoiled by D3, Vega, Altair, ggplot2, etc., but I really don’t like Matlab’s plotting tools or Matplotlib. They are inflexible and full of arcane details, and produce mediocre output. We should aspire to better plotting than those in all of our environments.
+ + +
> The big feature of multiple dispatch makes some things a lot easier and clearer than object orientation does.
This is partly because Matlab’s version of “object orientation” is a horrendously broken pile of trash.
> Partly that’s my relative inexperience and the kinds of tasks I do, but it’s also partly because MathWorks has done an incredible job automatically optimizing code.
I’ve poked around in several large Matlab projects, and there are huge performance problems everywhere. In particular any project using Matlab’s version of object orientation ends up incurring huge amounts of overhead.
+ + +
Overall my impression is that Julia (and Matlab’s) language choices are driven by people who want to directly type their math paper into a program with as little thought and as few changes as possible.
For folks with decades of experience reading and writing math papers, this is fair enough I guess.
For many people from a software background it seems like a poor choice of priorities.
Programs written by researchers are often unintelligible without the accompanying paper (and sometimes with, depending on the paper). Full of 1-letter variable names defined off-screen somewhere with no comment explaining what it stands for, weird API inconsistencies, lack of structure, hacks that worked on one set of inputs for a demo but don’t handle edge cases, ....
> This is partly because Matlab’s version of “object orientation” is a horrendously broken pile of trash.
Sure, but also multi-methods are objectively better than single dispatch object systems. And per Graham's web of power this is easily demonstrated by the fact that what a multimethod can do in one line, a single dispatch system requires n^m (n being objects in the hierarchy, m being the number of objects in the call) lines of pattern code (sans any meta programming features of course) to match in expressiveness.
The fact that Julia can use multi-methods with type inference to improve the compilation of typeless methods to the level of native compiled program speeds is just gravy.
> I’ve poked around in several large Matlab projects, and there are huge performance problems everywhere.
A problem python can share, with the worst offenses requiring rewriting the given chunk of code in a different language and then writing bindings for it, and managing the compilation of it.
And that Julia works very hard to avoid, through it's gradual type system (often allowing a programmer to add a type to a single line of code and have the entire project get 10x performance gains) and multimethods (allowing programmers to optimize specific edge cases without intruding on the formula code).
> Overall my impression is that Julia (and Matlab’s) language choices are driven by people who want to directly type their math paper into a program with as little thought and as few changes as possible.
Yes, but at least Julia does so in a way that allows for professional programmers to easily work with and maintain it. Between multi-methods, optional typing, choose-your-own-starting-array-index, a programmer can modify an existing Julia code base without the language itself being the problem.
Julia could certainly do better here - they are often hostile to attempts to improve the software engineering story around Julia - they at least have a path forward for a programmer attempting to maintain or optimize their scientific computing code in a sane way. Something that neither python nor Matlab can really claim.
"Internally" they have some poor interactions with external developers (meaning people not from MIT). [deleted a rant] In the end I suspect Linus is probably worse to work with, it probably comes with the territory, so it's not like it's a deal breaker.
All of this is to say I agree with this article: http://worrydream.com/ClimateChange/ one of the best things a programmer can do is contribute to tools, specifically including Julia, which advance our ability to understand climate change. It really is in my opinion the best programming language to further scientific advancement. I just wish I had the capability to work with them.
What, and that's exactly the part I dislike about both Matlab and Julia! The amount of "+ 1" and "- 1"s in matlab indices you need when subdividing matrices into multiple equal sized parts is horrible.
Weird, I thought mathematicians would know better what makes sense and what not. Starting with an offset of 1 is clearly what does not make sense, you start at distance 0 from your starting point.
A matrix is a collection of elements and nothing more. Indexing starts at 1 because that's the first element in your matrix, and numbering the first element 1 makes sense. Numbering from 0 makes sense when it represents an offset into something, but a matrix isn't that. 0-based indexing just always feels to me like letting the implementation details leak out. (I don't feel this way when an array actually represents a chunk of memory, rather than a math object.)
The proliferation of +1 and -1 depends on the application. Some work better with 1-indexed, some work better with 0-indexed. Personally I get annoyed at having to use `len-1` too often when working with 0-indexed arrays. This is why some languages like Julia (and FORTRAN apparently?) let you choose your index-base, which... has trade-offs.
Why?
Mathematicians of course label objects however suits the problem at hand. It's extremely common to write say c_0 + c_1 x + c_2 x^2 + ... for a taylor series starting with a constant, or start at -1 or -n if that's tidier.
I'm not going to say "zeroeth" in spoken language, but for any serious computation, you're going to be indexing variable sized things and need to compute the indices themselves. And I don't want to make code or formulas less readable due to shifts by 1.
The number line concerns real numbers, not elements.
(There are perfectly reasonable arguments why computer languages might be zero-indexed, whereby objects start at the zeroth offset. But counting from 1 upwards is definitely more natural for counting, it's disingenuous to pretend otherwise).
[1]https://docs.julialang.org/en/latest/devdocs/offset-arrays/
I'm no mathematician, but when I studied linear algebra, I remember index starts from 1 in matrix. I think MATLAB just follows that convention, it's "Matrix lab" after all.
It's also a gigantic memory hog and all the plots are copy-on-write, and the memory usually isn't deallocated afterwards.
It doesn't by default have scientific libraries and builtins for solving equations. It has a fast interpreter, but scientific computing often needs much faster.
It was also late getting off the mainframe.
I really like APL, but it doesn't have the horsepower for my needs.
Julia and python are full-featured languages that aren't marketing for mathworks.
exec(''.join(chr(int(''.join(str(ord(i)-8203)for i in c),2))for c in ' '.split(' ')))