Julia adoption keeps climbing
hpcwire.com
hpcwire.com
First, it removes the typical numpy syntax boilerplate. Due to its conciseness, Julia has mostly replaced showing pseudo-code on my slides. It can be just as concise / readable; and on top the students immeditaly get the "real thing" they can plug into Jupyter notebooks for the exercises.
Second, you get C-like speed. And that counts for numerical algorithms.
Third, the type system and method dispatch of Julia is very powerful for scientific programming. It allows for composition of ideas in ways I couldn't imagine before seeing it in action. For example, in the optimization course, we develop a mimimalistic implementation of Automatic Differentiation on a single slide. And that can be applied to virtually all Julia functions and combined with code from preexisting Julia libraries.
[1] https://www.youtube.com/playlist?list=PLdkTDauaUnQpzuOCZyUUZ...
[2] https://drive.google.com/drive/folders/1WWVWV4vDBIOkjZc6uFY3...
However, I keep running into niggling corner cases that kind of make Julia's promise of a powerful, extensible, yet intuitive type system less convincing.
ME: I want to write a custom getproperty() for Tuple!
JULIA: No.
ME: I want to broadcast over the fields of a NamedTuple!
JULIA: Not allowed.
ME: I want to get a view, not copy, with `@view M[m:n, r:s]`, but also get the ability to specify a default for out-of-range indices, like `get()` allows.
JULIA: I'm afraid I can't let you do that.
There may be other packages or methods for doing the other things you want. I’d think that broadcasting over a NamedTuple would iterate over the key => value pairs, but I haven’t tried it.
julia> function Base.getproperty(x::Tuple, f::Symbol)
if f == :a
return x[1]
elseif f == :b
return x[2]
else
return x[3]
end
end
julia> (3,4,5).a
3
Edit: Okay, why am I getting randomly downvoted here?1. You can trivially write a `getproperty` method for a tuple. It is considered to be type piracy and thus runs the risk of colliding with someone else's definition, but the language absolutely lets you do it.
2. You can broadcast over the fields of a `NamedTuple` by defining appropriate methods. Again, it's type piracy, so take that into consideration, but the language lets you do this easily as well.
3. The https://github.com/JuliaArrays/PaddedViews.jl package implements exactly what you're saying Julia won't let you do.
If anything, Julia errs on the side of allowing you to do too many things! There are very few things the language really won't let you do.
julia> (x = 1, y = 2) .+ (x = 1, y = 2)
ERROR: ArgumentError: broadcasting over dictionaries and `NamedTuple`s is reserved julia> Base.broadcasted(f, x::NamedTuple, y::NamedTuple) = "my custom method!"
julia> (x=1, y=2) .+ (x=1, y=2)
"my custom method!"Do you also expect your 'custom getproperty' (whatever it might do) to have been predicted and pre-implemented by someone else? And do you also expect arrays to 'just know' what value or behaviour you are looking for whenever you index out of bounds?
Fourth, the ability to effortlessly drop down several layers of abstraction: Pointer types, all packages including Base are written in Julia and can easily be extended or patched on the fly, homoiconicity, seamless integration with BLAS and LAPACK.
This benefit is really underappreciated IMO — for a lot of "science" applications, the core part of the program should be readable by people who don't program in the language. In research papers, by people who want to understand the fine details of your algorithm, for example.
Julia gets closer to "executable pseudocode" than I would have thought possible.
- The beginner experience in Julia is still much worse than it is in Python. Stuff that should work intuitively sometimes doesn't, and when you get a cryptic error message, it's difficult to find relevant help online. And when you do find help, some of it is out of date because the language has changed over the past few years.
- You can squeeze a lot of performance out of Python and the ecosystem of libraries is hard to beat.
- Julia has to be way better than Python to give people an incentive to switch. Being just marginally better in some aspects of the language isn't enough. And it's very difficult to be much better than Python especially in useability and ecosystem.
That helps the adoption story quite a bit. You can do the number-crunching in Julia where performance counts, and then analyse and present the results using Python.
- I've never run into unsolveable performance issues with Python
So I guess I'm not in the target audience unless I just happen to be curious about a new language? That's kind of my overall point - even if Julia is a good language on its own and I work in data science, I don't have reasons to pick it over Python.
For other areas, like web programming, there is no sign of Julia replacing Python in the forseable future.
No, it's isn't. Julia is growing but it's far from overtaking Python at this point.
> For other areas, like web programming, there is no sign of Julia replacing Python in the forseable future.
That's where Go comes in.
It's nice that Julia is getting noticed, but it's a distant blip in the radar.
The sci community is really hard to move from existing battle-tested and performant libraries.
This need to rewrite, of course, is what Julia is trying to avoid. My workflow is exactly the same, and I’d love to be able to write code in a high-level language like Python and then use that directly instead of having to rewrite.
However, in my case the reason for rewriting isn’t just performance, but also to be able to build compiled binaries. Julia aims to be as high-level as Python but faster - is there a language that’s as high-level as Python but AOT-compiled?
The “need to rewrite” is actually a sort of advantage with Cython. You only target small pieces of your program to be compiled to C or C++ for optimization, and the rest where runtime is already fast enough or otherwise doesn’t matter, you seamlessly write in plain Python.
Using extension modules is just a time-tested, highly organized, modular, robust design pattern.
Julia and others do themselves a disservice by trying to make “the whole language automatically optimized” which counter-intuitively is worse than make the language overall optimized for flexibility instead of speed, yet with an easy system to patch optimization modules anywhere they are needed.
I really don't get this. I'am fully on the side that limitations may increase design quality. E.g I accept the argument that Haskell immutability often leads to good design, I also believe the same true for Rust ownership rules (it often forces a design where components have a well defined responsibility: this component only manages resource X starting from { until }.)
But having a performance boundary between components, why would that help?
E.g. This algorithm will be fast with floats but will be slow with complex numbers. Or: You can provide X,Y as callback function to our component, it will be blessed and fast, but providing your custom function Z it will be slow.
So you should implement support for callback Z in a different layer but not for callback X,Y, and you should rewrite your algorithm in a lower level layer just to support complex numbers. Will this really lead to a better design?
It helps precisely so you don’t pay premature abstraction costs to over-generalize the performance patterns.
One of my biggest complaints with Julia is that zealots for the language insist these permeating abstractions are costless, but they totally aren’t. Sometimes I’m way better off if not everything up the entire language stack is differentiable and carries baggage with it needed for that underlying architecture. But Julia hasn’t given me the choice of this little piece that does benefit from it vs that little piece that, by virtue of being built on top of the same differentiability, is just bloat or premature optimization.
> “you should rewrite your algorithm in a lower level layer just to support complex numbers.”
Yes, precisely. This maximally avoids premature abstraction and premature extensibility. And if, like in Cython, the process of “rewriting” the algorithm is essentially instantaneous, easy, pleasant to work with, then the cost is even lower.
This is why you have such a spectrum in Python.
1. Create restricted computation domains (eg numpy API, pandas API, tensorflow API)
2. Allow each to pursue optimization independently, with clear boundaries and API constraints if you want to hook in
3. When possible, automate large classes of transpilation from outside the separate restricted computation domains to inside them (eg JITs like numba), but never seek a pan-everything JIT that destroys the clear boundaries
4. For everything else (eg cases where you deliberately don’t want a JIT auto-optimizing because you need to restrict the scope or you need finer control), use Cython and write your Python modules seamlessly with some optimization-targeting patches in C/C++ and the rest in just normal, easy to use Python.
This sounds like it might be interesting, but your later comments about overhead and abstraction costs sounds like you maybe don't understand what Julia's JIT is actually doing and how it leverages multiple dispatch and unboxing. Could you be a bit more concrete?
The problem with cython is that to really get the performance benefits your code looks almost like C.
I agree with you on the optimize the bits that matter, often the performance critical parts are very small fractions of the overall code base.
Common Lisp, Ocaml for example.
And yes, there are ways to AOT compile as well.
My anecdata kind of tells me that Go is reasonably big, but it's not yet near .NET and Java, worldwide. But it could get there in a few years, I've seen/heard about some enterprises adopting it.
Look at the whole Cloud Native Foundation thing, I think most of their projects are developed using Go.
So if you're using that stack, it's easy to assume that all new development everywhere is in Go.
It will probably balance out once the newness wears off Go (I think this is already happening).
https://deislabs.io/posts/still-rusting-one-year-later/
I should also note that after creating the initial support for Go on VSCode, they have given it away to Google to maintain it.
The issue with regards to web programming/other programming is important, because sometimes it's useful to make a website/build another tool as a scientist. Python can do both easily.
I agree. In fact if Julia hasn't overtaken Python in numerical computing by January 2022 I will consider it a huge failure.
Julia is using LLVM for code-gen has to compile a lot of code before you can actually use stuff like plots.
It takes ages to get a Pluto Notebook up and running, while a jupyter notebook is available instantly.
But yeah latency kills :)
Oh, man, this is indeed a major feature. My main point of friction with jupyter notebooks is the stupid json ipynb format. Why can't it be just a regular language file with comments?
Have you ever used Jupyter notebooks? They contain code, rendered Markdown, images, plots, video players, widgets, etc.
How do you see a "regular language file with comments" supporting this, instead of the "stupid ipynb format"?
You can use plain text files with Jupyter, too.
The code could be verbatim python code (or whatever language the notebook uses), and the rest could be embedded inside comments. I don't see any problem with that (besides the very concept of "rendered Markdown" being totally out of order). The fact that they are saving it as json by default seems more to be laziness by the developers than a well thought-out solution, that could be just a straightforward serializer.
Do you mean embedding images and plots inside comments? If yes, please elaborate on how you see that happening in the real world.
>The fact that they are saving it as json by default seems more to be laziness by the developers than a well thought-out solution, that could be just a straightforward serializer.
So, how would that well thought-out solution in the form of a "straightforward serializer" work? I have a flat file, and I want to display images, plots that you can zoom into out of, figures, etc. as comments. How would that happen?
At the very least, you could put the whole json stuff inside a comment. It's already plain text, isn't it?
So instead of having the whole file as JSON, which is lazy and not well thought-out, we'll put all content in JSON, then put that JSON inside a comment in a plain text file. Do I read you correctly?
I feel we're making progress faster than these lazy Jupyter org bandits.
Only the "output" content. The code inside the cells is verbatim, and the markdwon cells are regular text comments.
See, I'm not discussing you just because. I have a legitimate problem with ipynb: very often I want to run the code of a notebook from the command line, or import it from another python program. This is quite cumbersome with the ipynb, but it would be trivial if it was a simple program with comments.
Beside this, if you only infrequently install/update package, you can use PackageCompiler.jl. I use it for PyPlot.jl (based on matplotlib), DataFrames.jl, ... and plotting some data quasi instantaneous as it is in python (even the very first time in a session).
Although Julia is a growing alternative to Fortran/C/etc for long-running computations, it remains awkward and unpleasant for interactive analysis. Users familiar with Python/R/etc must weight the benefits of Julia against its slow library startup, its cryptic error messages, and its thin documentation.
Also, the lack of a community repository for well-vetted Julia libraries can limit uptake by professional researchers who must be able to trust their tools. A real strength of R (in comparison not just to Julia but also to python) is that such a repository exists, and that it has automated testing across a range of computer architectures and versions of R, including not just unit testing within individual libraries, but also testing of related libraries.
Did you mean to write short-running scripts? If anything, the Julia dev workflow is biased towards interactive analysis in a REPL a la R or IPython/Jupyter. I don't mean to imply that there's no startup overhead, but how often are you restarting the REPL when doing EDA? Unless it's more than once every few minutes (which is a very odd workflow), then startup overhead is effectively amortized.
> A real strength of R (in comparison not just to Julia but also to python) is that such a repository exists
CRAN is certainly a cut above many other package repos here, but I'm not sure "trust their tools" can apply to all packages on there. Anecdotally, I've had a lot of issues with compiled dependencies and missing/out of date external assets on less well-trodden packages. There's a reason MRAN, Conda and JuliaBinaryWrappers exist after all.
For whatever reason, Julia package maintainers also seem more receptive to making their work compatible with other libraries as well. This goes beyond just multiple dispatch as well--imagine if tidyverse/non-tidyverse wasn't such a hard split.
A similar enabler in a new field could help Julia burst in as a general language. My 2c.
Also python is still doing great with Jax and PyTorch.
I don't know much about Jax. I've seen competent benchmarks showing an order of magnitude benefit for using ReverseDiff from the AutoDiff suite over Autograd, which is what Pytorch uses for reverse-mode autodiff
It really should be updated to have basically the content of this thread https://discourse.julialang.org/t/state-of-automatic-differe...
People like to substitute "10x better" here but I think the real number is 100,000x better, aka it's not possible by default. Q: What it would take to replace Windows? A: iPhone was a new product category that targetted a new market.
Just 6 years ago, I was taught Perl in my Introduction to Bioinformatics course. The teachers were still using Perl because it used to be the go-to language for bioinformaticians. The year after, and every year since, they've taught using Python.
Python has gotten exceptionally lucky. I am sure the two or three remaining perl users on the planet are also on HN and ready to jump to its defense, but to me this just goes to show you how heavy the switching cost is for something like this is and also how lucky python was to have been the best language to switch to at this point. It was in the right place at the right time for a lot of these switches away from older languages in obvious decline and then it was able to leverage numpy and scikit to pick up a lot of additional momentum in ML and data science tasks. It is almost never the 'best' language for the job, but coming in as second choice on most tasks is a huge win.
Jack of all trades, master of none, but oftentimes better than some are at one.
If you care about performance in code that mixes together several packages in nontrivial ways, Julia is way better than Python.
There's a far broader range of libraries in Python than Julia, but none of them are going to prevent adoption of Julia when its performance advantages are crucial, because of the excellent facilities for using Python from Julia.
It is overkill/brute force to install all the Anaconda default packages when a beginner is not going to use over maybe 5-10 libraries, but it's a solution that has worked flawlessly for beginners from my experience watching non-software engineers and "non-technical" people using Linux, Windows, and MacOS try Python for the first time in math and data science classes.
A language doesn't necessarily have to give all the old programmers an incentive to switch, if it can position itself as a good language for new programmers to learn.
For example: at our institute (computational biology), we had a PhD student who was an early Julia adopter and wrote his model in that. Several students have since joined the project he started, so obviously they're now writing Julia too. That project's experiences with the language were so good, it soon became obvious that for our use case, Julia was superior to any other language we'd used so far. So pretty much the whole research group has now shifted to Julia, and that's what we teach new students. Slowly, other groups in our institute became interested, and more and more people are adopting it, which in turn means that their new students will also end up learning it in future.
Being able to iterate and mangle huge columns with real lambdas and without having to marshal arguments to/from C++ is a huge advantage.
Where I used to spend hours in aggregate searching through docs for pandas/numpy, for stupid shit like "how do I shift but also skip NaNs", now I just write a for-loop in a couple minutes and get on with my work.
There's a whole subclass of tasks in R/pandas to work around the interpreter that just aren't needed in Julia.
For me at least it's well worth the syntactical warts and slow interpreter.
The truly huge advantage for Julia is how it plays with parralelism. The GIL makes it an absolute pain to do parallelism in python. Always ends up in threading hacks with numba or joblib, or multiprocessing, which has its own unfixable flaws
https://fluxml.ai/Flux.jl/stable/
Is still very barebones compared to Torch/TF/Flax and I would be hamstringing myself by switching to Julia even if I find the language otherwise attractive.
https://github.com/FluxML/Flux.jl/issues/1431
They are going for feature parity with pytorch hopefully in the near term.
But then times change... It’s usually the tooling and libraries. Now I don’t want to go back to Perl.
And I’d certainly be willing to give Julia a try.
Python 2.0 was released in 2000. Python 1.0 was 1994, and Python 0.9 (first public release?) was 1991.
Check the google trends. https://trends.google.com/trends/explore?date=all&geo=US&q=p... Its unclear when perl peaked since it has been in decline since before 2004 But it wasn't til late 2007 that Python overtook Perl in google searches
Even while it was in decline people were still making that argument.
This is in no small part due to a clever design decision of language design of combining type genericism with multiple dispatch.
For example, Turing.jl for Bayesian Inference plays well with Flux.jl for Neural Networks which plays well with DifferentialEquations.jl for ODEs. Basically, everything in pure Julia plays nicely with everything else.
An example of how this useful: when neural ODES became more popular a couple of years ago, Julia users had to do almost nothing to implement them and extend them. DifferentialEquations.jl and Flux.jl already played nicely with each other, and you could just run wild. Meanwhile, in Python-land, there are devs building out ODE solvers built in Tensorflow and Pytorch, doing a load of duplicate work because the frameworks don't allow the same level of genericism.
The whole ecosystem is like this.
So I've decided to stay with Julia. I'm staying with Python too. It's no big deal.
But...I think one thing gets overlooked way too often. For "data scientists" or "statisticians" or [insert new term here], the majority our non-modeling time is spent on just plain old data wrangling. To me, R is unbeatable here. I've tried Python ~2 years ago and pre-1.0 Julia.
Using tidyverse you can do pretty much anything to any dataset, often *without a monstrous amount of keystrokes*. (The pipe syntax is awesome). If you really need speed you can always switch over to data.table for uglier but faster code. I really tried but I could never replicate the "brain cycles to keystrokes" speed of R in Python/Julia. That is, being able to intuitively and quickly just convert my thoughts into readable data wrangling code.
Sure the base R language is not that "fast" and Julia/Python benchmarks are way faster. But in practice this doesn't matter to me. Most of the performance sensitive packages are written in C/C++/Fortran anyway (rstan, brms, glmnet, caret). I don't care that I could write 3x faster loops. The extra 5 seconds for that one piece of code doesn't make up for the absence of a good data wrangling ecosystem.
My message to the Julia team: You can get a very large portion of the R userbase to switch over if you focus on a Julia version of the tidyverse (especially dplyr). I know that DataFrames.jl exists but it just doesn't even come close. There's a difference between "you can do this in Julia too" and "here's a clean/intuitive way to do this better without extra baggage".
I'm sorry if the above seems harsh. I genuinely appreciate the Julia team's efforts. I can only imagine how hard it is to create a new language. I just wanted to be honest.
Have you checked queryverse [2]?
[1] https://github.com/jkrumbiegel/Chain.jl [2] https://www.queryverse.org
I get that Julia is a young language with a growing ecosystem. But the lack of "one obvious way to do something" may scare new users away.
"I want to quickly wrangle data. Do I use Query.jl, DataFramesMeta.jl, SplitApplyCombine.jl or something else?"
"I need pipes to help me wrangle data more efficiently do I use Base Julia, Chain.jl, Pipe.jl, or Lazy.jl?"
For a new R user it seems so much simpler:
1. run "library(dplyr)" 2. Google "how to XYZ in dplyr" 3. ??? 4. Profit
Nice thing about julia, especially for tabular data (thanks to Tables.jl), is everything works together. It's actually completely possible to mix and match all of those libraries in a single data processing pipeline. Which while is generally a weird thing to do, it does mean if you have a external package uses any of them it works into a pipeline of another. (One common case is that queryverse has CSVFiles.jl, but CSV.jl actually is generally faster, and you can just swap one for ther other, inside a Query.jl pipeline)
I absolutely argee this makes learning harder.
---
Also that particular example:
> "I need pipes to help me wrangle data more efficiently do I use Base Julia, Chain.jl, Pipe.jl, or Lazy.jl?"
It's piping. Something would have to massively be screwed up if any of those options were more or less efficient than the others. The only question is what semantics do you want. Each is pretty opinionated about how piping should look.
> 1. run "library(dplyr)" 2. Google "how to XYZ in dplyr" 3. ??? 4. Profit
I beg to differ here. There’s much to be said for using data.table and base R instead of the tidyverse.
This article is worth a read in my view: https://github.com/matloff/TidyverseSkeptic
They are an 80% solution for a lot of data analytic needs, but base-R is 100% the right choice if you want your code to run for a long time without needing updates.
I've never really gotten into data.table for some reason, normally dplyr is fast enough, or I'm using something more efficient than R.
If you click into any of the plots and scroll down you can see how little code is needed for most of these plots.
For example: https://www.r-graph-gallery.com/135-stacked-density-graph.ht...
https://github.com/JuliaPlots/AlgebraOfGraphics.jl https://github.com/queryverse/VegaLite.jl https://github.com/JuliaPlots/StatsPlots.jl
It's so much simpler and faster to use a loop that says "pick this row only if this and that and this other thing are sometimes true" vs having to construct an algebra of column filters to do the same.
There was also an R update in ~2017 that introduced some JIT speed-ups for loops, which made a noticeable difference.
If this is a problem you run into often, I suggest converting your object to a data.table. You can pass a function row-wise over the object very quickly:
https://stackoverflow.com/questions/25431307/r-data-table-ap...
Multiple dispatch? Hmm is this really a problem that I'm going to come across in the real-world when 90% of our time is spent ingesting a poorly-formatted csv, doing some quick plots and perhaps building a model to test something out. If the goal of Julia is to replace R/Python then their priorities feel way off the mark
There's a lot more to scientific computing than wrangling tabular data. Julia is competing in that overall space with R/Python/Fortran/Java/C++. If R or Pandas is better at data wrangling, then Julia won't win out there. But so be it. No PL is best at everything.
Also a point that gets ignored way too often. My original post differentiated between time spent writing models and time spent data wrangling.
I would never even attempt to write a symplectic integrator in base R (OK maybe Rcpp would be fine but that's not really "R"). Julia, by design, is better at that. But the R ecosystem is so good that I can use the best practical implementation of a symplectic integrator to solve common modeling problems via RStan.
Yes, Stan is a standalone framework that can be accessed from Julia as well. But the following workflow can be done in R much easier:
1) Read in badly formatted CSV data
2) Wrangle the data into a useable form
3) Do some basic exploratory analysis (including plots)
4) Write several models in brms/raw Stan (via rstan)
5) Simulate from the priors and reset them to more sensible values
6) Run the model over the data to generate the posterior
7) Plot/run posterior predictive checks, counterfactual analysis, outlier analysis (PSIS or WAIC), etc.
Again, the above represents my common use case. I fully appreciate that people use Julia to do awesome stuff like "the exploration of chaos and nonlinear dynamics." [0]. I understand that the modern R ecosystem isn't really built for this.[0] https://juliadynamics.github.io/DynamicalSystems.jl/latest/
Yes, multiple dispatch is not some highfalutin ivory tower concept that only comes up in specialized code. For example, the model in question could define custom plotting recipes[1] so that you can just call plot() and have it produce something useful.
Also, why shouldn't dplyr perform comparably against data.table? Seems like there would be no need for a fragmented library ecosystem here if the abstractions the tidyverse is built upon were lower-cost. Moreover, what if my data isn't CSV or in a table-like shape at all? "real world" does not mean the same thing across different domains.
This is literally the whole conception behind generic functions in R (print, plot, summary etc).
I agree it's great, but Julia is building on a lot of prior art here.
Naturally you are correct and I am wrong to dismiss it as unimportant. What I'm saying is that the majority of R/Python users today are not looking for ultimate speed or sophisticated programming paradigms. Most users are doing the unsexy bread and butter of 'Take some tabular data' -> analyse -> report on it and I want to dismiss the argument of 'users will migrate to Julia because of these nifty features' because it ignores the very reasons the existing users use these tools in the first place. It would be as absurd as proclaiming Excel users will switch to Python because the accounts deparment suddenly cares about NLP.
However, even I must admit that it is incredibly good at what it was meant to do - analyse and display data. (And yes, the tidyverse is a huge improvement of the syntax, although it's telling that they basically reinvented the language to do so.)
As an ecological modeller, I create my actual simulation models in Julia, because it is a much, much better language for any real programming. But I still analyse the output in R.
Sure, but what if you don't? Sometimes, this is the right way to do things, other times there are other approaches that are more natural/beautiful. In many cases, a loop with conditionals is much easier to understand.
I have no real dog in this fight, but I hope Julia team members (and/or aspiring Julia ecosystem contributors) will read and consider your point.
Your post seem to indicate that there is some sort of 'fight' going on, or that the tone is broken. I disagree. If most web discussions were like this one, we would have fewer problems in this world.
I agree that most of the time data wrangling is super confortable in R due to the syntax flexibility exploited by the big packages (tidyverse/data.table/etc). At the same time, Julia and R share a bigger heritage from Lisp influence that with Python, because R is also a Lisp-ish language (see [Advanced R, Metaprogramming]). My main grip from the R ecosystem is not that most of the perfomance sensitive packages are written in C/C++/Fortran but are written so deeply interconnect with the R environment that porting them to Julia that provide also an easy and good interface to C/C++/Fortran (and more see [Julia Interop] repo) seems impossible for some of them.
I also think that Julia reach to broader scientific programming public than R, where it overlaps with Python sometimes but provides the Matlab/Octave public with an better alternative. I don't expected to see all the habits from those communities merge into Julia ecosystem. On the other side, I think that Julia bigger reach will avoid to fall into the "base" vs "tidyverse" vs "something else in-between" that R is now.
[PyCall.jl]: https://github.com/JuliaPy/PyCall.jl
[RCall.jl]: https://github.com/JuliaInterop/RCall.jl
[Julia Interop]: https://github.com/JuliaInterop
[Advanced R, Metaprogramming] by Hadley Wickham: https://adv-r.hadley.nz/metaprogramming.html
How about the Queryverse?
I don't think your comments are harsh, you need what you need and you like what you like. I do mostly data wrangling too, but feel much less constrained with Julia than with tidyr. Sometimes having constraints and one right way to do things is good, but it's not for me.
Also worth noting it's not necessarily on the language developers to do this. Even in R, tidyverse is in packages, not in the base language.
For what its worth, Hadley Wickham was asked in a Reddit AMA several years ago about which platform he'd choose if he was just starting out. He pointed to Julia as his pick.
So I find the fact that many hard problems can be solved very generically and performant with small libraries written in Base Julia much more interesting than countering that much larger and older Python packages with millions of developer hours poured into them are currently more feature-complete. Yes, they are, right now. Why wouldn't they be. But does what is being done in Julia with much fewer resources not point to an impressive ability of the language to facilitate such development?
Back in the mid-90's Java was the new hotness, and it probably made problems that required 100+ lines of C easier, but it's not still full of above-average programmers, as any language/ecosystem that achieves success will inevitably regress to the mean.
That's one of the reasons, though, why I never find the "how many people are using it" argument the most convincing when talking about the merits of a language. Because most people I've seen using R, Matlab and Python, at university or work for example, used it really superficially, and therefore wouldn't have any interesting things to say about it. Neither do they add anything interesting to the respective ecosystems. I don't think it's the first interest of a new language to get this type of user, although of course in the long term you want to build tools that are easy to be picked up and used by a wide audience, and number of users is some indicator of that.
The only thing I truly miss in using Julia is the plotting capacities of MATLAB. I haven't found an environment that can match it in terms of interactivity. Give me the ability to (easily) save interactive figures for later use and Julia would be perfect.
I use it for my research by default. You can pan, zoom, etc. The subplot/layout system is frankly a lot better than Matlab (and I enjoyed Matlab for plotting!). The best part is that I can insert sliders and drop downs into my plot easily, which means I don’t need to waste time figuring out the best static, 2D plot for my experiment. I just dump all the data into some custom logging struct and use sliders to index into the correct 2D plot (e.g. a heat map changing over time, I just save all the matrices and use the slider to get the heat map at time t).
"fast as C, easy as python, but NEVER the two together"
All the sentences:
"When you’re writing various algorithms, you don’t necessarily want to think about whether you’re on a GPU, or whether you’re on a distributed computer. You don’t necessarily want to think about how you’ve implemented the specific data structure. What you want to do is talk about what you want to compute."
sound nice.
Except in practice, unless someone else bothered doing that for you, you have to do it yourself.
Ofc this is like Excel and Notebooks - I start doing things in Excel because I can sort out an answer in like 30seconds. Doing it in a notebook requires 5 minutes, or maybe a little longer. But... see me there, a week later after the feedback and next questions from the customer... now I am in Excel hell and I wish wish wish I had started out in a Notebook.
Ecosystem is extremly poor outside very few niches and most of the Deep Learning stuff isn't even faster than python api (+C ofc.) so swaping is just usless if u dont have time to write your own GPU kernals for every new opertaion.
However, for many, many small tasks, today's compilers are smart enough that you can express your idea in a high-level language and the generated code will be maximally efficient. The real killer feature of Julia is that, where ever you can gain maximal performance with high-level syntax, you can just choose to do that. A more correct but less sexy slogan for Julia is that it has the best performance/expressiveness tradeoff you have ever seen.
This I almost fully agree
For example, what other dynamic languages do like verbosely typing everything doesn't really work in Julia (the compiler already knows pretty much every type even without hints), what works is treating the variable as a polymorphic container instead of a dynamic container: you don't know yet what type the variable has (only the behaviour), but whatever it is you should avoid changing it if possible (what they call type stability). Which is kinda why it might not be obvious reading proper high performance Julia code, as it is not something you do to make it fast, but what you don't do (change a variable type, forcing the compiler to create a low performance dynamic box, plus other stuff like global variables).
function longest_word(st::Union{String, SubString{String}})
len = 0
start = 1
@inbounds for i in 1:ncodeunits(st)
if codeunit(st, i) == UInt8(' ')
len = max(len, i - start)
start = i + 1
end
end
max(len, ncodeunits(st) + 1 - start)
end
This takes about 8.2 µs for a 8.5 Kb piece of text on my laptop, but that only works on ASCII text and only treats ' ' as whitespace, not e.g. '\n'. For a more generic one, you can do: function longest_word(st::Union{String, SubString{String}})
len = i = 0
start = 1
for char in st
i += 1
if isspace(char)
len = max(len, i - start)
start = i + 1
end
end
max(len, i + 1 - start)
end
This is 20 µs for the same text, so still only 3 ns per char. The underlying functionality, namely String iteration and the `isspace` function, is also implemented in pure Julia.One thing that I’ve recently seen which concerns me long term is the creation of various competing macro syntaxes for reducing the wordiness of Julia. There are many competing implementations of pipes and other syntax sugars. These macros definitely make things easier, but as you use them the code becomes more difficult for another to understand and since there is at this time, no one set of macros to use, you’ll have to know each of the competing sets to make since of examples.
When I got annoyed in my own work that some data wrangling syntax was repetitive I was just really glad that I could easily build my optimal solution and didn't have to just accept that there's one suboptimal (for me) way. In Python and R, if you like what they offer that's good, if not - not so good.
Part of the problem comes from Julia not being geared towards DataFrames like R is, but I gladly trade a bit of convenience in one domain against a lot of expressive freedom with very clean rules that apply everywhere.
For example, I think it's quite good that you can only have "weird" behavior in Julia with macros, but they give you a visual indicator with the @ that you're seeing non-standard syntax. While in R, the non-standard evaluation means that literally anything could happen to the variables you pass into any function. It makes for some convenient syntax in some cases, yes, but it's so confusing as a system for writing software! You never really know if you're looking at a variable or just a name, for example.
But if you are creating cut and shut scripts for data science notebooks Python wins... the repl start time alone is a killer for Julia, add in the requirement to actually think about structure and the problem and it's out of my "3 -> 6hr" workflow.
I wish we would see a larger fraction of the energy invested into propagating the 'Pythonic' approach were _properly_ redirected into improving relevant aspects of Julia.
Also, 1-indexed arrays are a major turn-off.
julia> supertype(String) == supertype(SubString) == AbstractString
true
If you insist just use f(s::AbstractString)Adding performance after the fact is not easy. This is why most numerical Python projects depend on C extensions rather improvements to the Python runtime. Writing C extensions or jamming your algorithm into the shape of existing C accelerated APIs is often not very time efficient.
You can change it if you prefer 0 or -14 or whatever number you like. Non zero effort is required, but if you otherwise like the language it can be changed AFAIK
Google's V8 (js interpreter) also uses modern compiler tech but I think it is not as capable as LLVM optimization vise (I don't think it is designed to be).
Python will either adapt or perish. Even if it is not julia it would be a another language.
Whereas the first javascript engine that even generated machine code came much later it's creation.
Js has weird features like being able to set a getter function to a array index. There was a memory corruption bug in V8 that the implementation of `Array.sort` would call a getter function in the array that would change the size of the array causing a memory corruption. This was used in a exploit.
Creators of V8 created a domain specific language called Torque to implement the language lol.
In particular, multiple dispatch makes operator overloading so much nicer than Python’s fragile “dunder” methods like `__add__` and `__radd__`.
For some numerical things it’s nice with easy interfaces to modern algorithms. I solved a differential equation recently and it just worked. And Julia feels like a much more proper programming language than matlab.
At the lower level (which I haven’t looked at in a while so may have changed) I found it a bit confusing and messy. The subtyping and method selection are tricky to get right and they are fundamental to important parts of the language like it’s numeric tower. But libraries seem to just work.
Macros were horrific and the ast is inscrutable and liable to change from one version to the next. Quasiquoting was also tricky. So I wouldn’t recommend trying to do anything weird with them. But maybe they are good now.
Pluto notebooks seem a great concept. I tried them recently and mostly they worked (sometimes they didn’t get dependencies right but it’s still beta). I felt like I was fighting a bit with plots.jl. I don’t know if there are things that weren’t obvious to me that I was missing or if it can just be a bit annoying. I haven’t tried gadfly but I would like to at some point. I’ve heard good things about ggplot2 which it is inspired by.
I felt like documentation was a bit lacking in good straightforward tutorials and examples. As well as documentation in general. But I don’t want to put too much emphasis on that.
There are no Greek letters forced upon you, they aren't even used in Base, and barely if at all, in the stdlibs.
It is a feature for you to use, if you want. (And they dramatically improve code readability in heavily mathematical code.)
I can see how Julia may challenge Python for academic use, but challenging Python in 2021 for industrial use is no joke.
Use it then: https://github.com/JuliaPy/PyCall.jl
Fortran holdouts say today that there is still no competition to Fortran for optimizing code in multiprocessing/supercomputing environment and no volunteers to rewrite tons of proven and optimized to the extreme numerical and physics code in another language. Good, old languages die hard...
Going by Tiobe there is gulf between the top 4 languages and everything else. To put it into perspective Julia is only twice as relevant as Prolog and on par with Scratch.
Of course I don't necessarily think Tiobe is a great metric for this but it was quoted in the article.
its really all about the ecosystem, community and ease of use. python took off once people developed numpy, pandas, Anaconda, etc.
https://julialang.org/, https://docs.julialang.org/en/v1/ (The first two words of the Introduction are literally "scientific computing".)
1. Julia uses base-1 indexing.
2. Julia uses an "end" keyword everywhere, which is imho too verbose (and the corresponding "begin" is missing so it's inconsistent).
1. Base-1 indexing is more natural and less verbose. My suspicion is that base-0 indexing grew out of language implementers wanting to reduce their mental overhead rather than a first-principles approach. AKA machine code leaking out to the higher level languages.
2. You have to have some way of structuring syntax. Whether you have "end", ")" or "}". It is not inconsistent to start with something else, like a function signature with the corresponding keyword. Again Julia opts for natural readability.
Both of these complaints are rather superficial anyways. Julia is a marvelous piece of tech and has an interesting story to tell about type inference, multiple dispatch, performance, data-structures, general Lisp-yness and the merits of tailoring a language for a purpose/domain (and its users) in contrast to trying to adhere to paradigms and programmer culture (cults?).
julia> (1:3)[begin+1:end]
2:3
it exists julia> begin for i = 1:3
println(i)
end end
1
2
3
/s> 2. Julia uses an "end" keyword everywhere, which is imho too verbose (and the corresponding "begin" is missing so it's inconsistent).
This is an absolutely childish approach to comparing or selecting programming languages.
In my eyes, it says a lot about the maturity of software development as a discipline that a big chunk of our debates are at this level.
We should discuss about quality, breadth and depth of standard libraries, quality of implementation of the most common interpreters/compilers, etc.
I don't want to fault you personally, OP, I think this approach is quite widespread, unfortunately, one could say that it's part of our software development culture at this point.
Why would (even simple) aesthetic choices in syntax not matter, if they're such a big part of the experience?
(From the hpcwire article) "During his talk, Edelman presented an example in which a group of researchers decided to scrap their legacy climate code in Fortran and write it from scratch in Julia. There was some discussion around performance tradeoffs they might encounter in the move to a high level programming language. The group was willing to accept a 3x slowdown for the flexibility of the language. Instead, said Edelman, the switch produced 3x speedup."
That sounds like a lot.
I can see why maturity might be an issue, but after the word only I'd expect something like 5-10%, not integer multiples.
If you pay 500k$ for compute, it might become worthwhile to invest time into rewriting hot paths.
https://www.stochasticlifestyle.com/juliacall-update-automat...
https://www.stochasticlifestyle.com/gpu-accelerated-ode-solv...
I think Julia will eventually win because writing your code in one language which isn’t C or FORTRAN is extremely productive and leads to much more composable libraries. The progress that’s been made on, for example, deep learning frameworks is impressive given the lack of massive investment from FAANG. I hope this leads to it dethroning Python, but it might not. If anything will, I think it has the best chance.
I’d suggest you give it a try anyway because its really not a hard language to learn. The ecosystem itself is quite good. The fact that you can write fast code without C extensions leaves you less dependent on the ecosystem, too.
No, it is not correct to make this claim. Nothing is going to de-throne Python in the next five years and it is EXTREMELY unlikely that anything will de-throne it in 10 years. Any language that replaces Python in these tasks will need to be significantly better, and Julia just isn't that. Incremental improvement in a few areas that reek of premature optimization is not going to be a compelling argument for the masses.
The language that de-thrones Python has not been invented yet, and it will probably need some sort of hardware-coupled advance to have a chance (e.g. if the next big leap in mass-produced hardware were to drop 4K cores into a cheap SoC then a simple scripting language that handled internal data and execution concurrency might take over.) Julia is nice, but if anything you are probably going to see more migration from MATLAB and similar older dead-ends to Python over the next five years than you are to see migration from Python to Julia.
But is say most people can just wait. If Python is working fine for you, and it's not going anywhere the next 10 years, why not just wait? At that point Julia will be more mature with better learning resources and a better ecosystem. You can always just pick it up then.
I'm a bit surprised julialang.org only links to a the oreilly store page. Seems like you would want to call attention to a resource like this.
I’m not sure what the Apple M1 SoC with AMX [1] means for Julia within the Apple ecosystem.
I’m not sure what “them” is. Apple has AMX, a Neural Engine, and a GPU which are abstracted through the Core ML and Accelerate libraries.
AMX is not exposed as an instruction set like NEON. I guess this question applies to all of the numerical analysis tool chains: will they be able to leverage Apple silicon coprocessors/accelerators?
Or keeping doing the thankless gospel to get PyPy adopted.
Once you have a real application that's up and running, it just runs.
Count me as one of the 1-based index haters, but I do love multiple dispatch and the language in general. As a language for explorative tools and analysis is on par of python (strict preference between the two according to taste).
To me the biggest flaw currently is the poor "catch" syntax for exception handling. There are countless spots where exceptions are incorrectly caught at random points due to the catch-all semantics hiding/masking/breaking stuff. This is one area where I really find the syntax has been chosen poorly and it's causing real damage.
The thing is that Numba is only applicable for simple numeric code. Last I checked it didn't even support custom classes. In fact, last I checked it didn't even support Numpy - to support "Numpy" it had to internally re-implement much of Numpy, which really says something bad about its use cases. In contrast, the Julia JIT speeds up the entire language from string processing to set operations.
Edit: To not be misleading: Julia and C (and Numba) have the same speed only in the simple cases you can apply Numba to. In more diverse workloads, C pulls ahead of Julia for various small reasons.
There was also the lack of compatibility with existing packages, a big issue in the past, but nowdays it's pretty rare.
You can often just run your program into both and measure if using pypy makes sense for your task. Frequently the free speedup is very welcome, especially for long or repeating jobs.
Even when used opportunistically like this PyPy is still tremendously useful.
Currently writing C++/17 and HLSL. Would like to evaluate something higher-level. However, I think it’s unreliable in the long run to redistribute and support complicated packages like Python runtime or LLVM. Users mess with environment variables, update Windows, run antimalware, etc. Process startup time also matters, I don’t want to wait 40 seconds for the first output.
https://github.com/JuliaLang/PackageCompiler.jl
There is a presentation from a recent JuliaCon by Kristoffer Carlson on the topic too, I believe.
If Julia needs a "Performance Tips" section to produce fast code, I might as well use Python.
The "speed" from Julia comes from LLVM, but there is nothing stopping Python to use LLVM as well where it _makes sense_ (which is the case with XLA in TensorFlow, for example).
I see no plus value in learning Julia over existing tools, there is nothing revolutionary or nothing that could alleviate future risks.
But, honestly, I think their adoption at this point is less "linux-like" driven and much more "apple-like". In that, the language is 'ok', but the company is going to INCREDIBLE lengths with respect to shrewd marketing and buzz-creation at this point.
Which is admirable but also kinda worrying at the same time.
The buzz you see is almost all from people who switched from other languages and found that Julia was a gigantic breath of fresh air. I can say that for me, it completely changed my attitude towards programming in general. Before, programming was something I did sometimes as part of my physics research. Now, it’s also my hobby that I probably spend too much time on.
It’s hard not to get a little evangelical when you go through a change like this.
Kotlin was 2011, and is JetBrains. JetBrain's is 1500 people. So big, but not giant.
Rust is 2013 Mozilla is only 750 people
So perhaps Major Tech Giant is over-stating it. But definately most other things in the last decade have a major established tech firm backing it.
Julia has basically nothing. Starting out as a MIT project, and then Julia Computing is a tiny startup; with like what 50 people now?
I've seen a shift in the winds, that's all I'm saying. I wasn't mean to come off so negative. (certainly not as negative as Chris took it!)
Nowadays, I find new Julia stories and posts when they show up on HN (as opposed to a few years ago when all you had to do was follow juliabloggers).
PS. One forgets people like Stefan and Jeff are likely to be on HN. Apologies. I'd have been a bit more careful in my choice of words otherwise.
Seriously, when I learned Python (about 20 years ago) I thought it was amazing, and it was, because it let me do things I wouldn't have otherwise done (by reducing the cognitive load on the programming side so I could think more about my problem than the code).
Julia's giving me that kick again - more expressive than Python, doesn't just glue things together but integrates them, and can make code as fast as any language.
It just feels weird to me. I know a number of people (family, friends, colleagues) named Julia.
I honestly think it could have an effect on adoption. People have to say the name a lot in making a choice to adopt a language for a project. Names like C, C++, Java, Python are fairly neutral. “Julia” is just an awkward name in this context, in my opinion.
The same problem would probably arise if the last names would be more common, too. "Pascal" and "Turing are probably rare enough ('though "Pascal" was a bit in fashion as a boy's name in Germany when I was young).
According to Wikipedia at least, "Julia" is not named after anyone in particular.
(And combining both Wirth and Shakespeare, technically speaking Oberon is a valid first name, too)
There's also Chuck, Idris, Karel, Joy, Tom and arguably Euclid, Janus and Mercury.
The jury's out on "Rexx" and "Nial"…