Julia v1.3
github.com
github.com
In one of my CS classes we wrote a LISP using Julia. Our professor wrote the parser for us, so we just had to focus on the actual interpreter bit. The pattern matching/multi-dispatch mechanisms were really nice. We didn't get to play too much with the parallelism mechanisms, but I liked what I did see.
I feel like Julia is a much richer language than, say, Python, and it would be a nice drop-in replacement for numerical computing tasks: it's faster, in some ways more ergonomic, and afaik can call out to Python libraries if needed. Has anyone here switched from Python to Julia for scientific/ML/AI purposes?
On that train of thought, how interoperable is preexisting numpy/pandas code with Julia?
On the other hand, it should work very well as a replacement for Numba. Numba already comes with a one-off jit cost, and it is much more limited in the sort of code it can speed up.
I personsally don't mind the startup too much. But I do get where people are coming from
Yes! this is the biggest roadblock for me. I find python startup time unbearably long. I do not want to imagine even using julia. And this is sad because the language is orders of magnitude better (for math, at least).
I think Julia is meant to be used like Lisp:
- Step 1: Open the REPL.
- Step 2: Don't close it until you are finished (both writing and running your program).
My best guess is that people that complain about startup speed of Julia (and Emacs too for that matter) ignore step 2.
Library ecosystems have a human workforce value of billions of dollars, available for free. Any modern language that want to go mainstream must offer transparent and complete interoperability with an existing ecosystem. TypeScript -> js Kotlin -> jvm C++ -> C Elixir -> Erlang Hack -> php Crystal -> ruby
As we can see there are still Big ecosystems lacking a new modern language™ The python ecosystem The C# (CLR) ecosystem
It could be said that's mainly because those languages are already mostly "modern"
There are also dying big ecosystem and there would be economic value to make their old libraries accessible. Fortran Pascal Perl Etc
F#, obviously.
Modern is not some "PL-community favorite with small following" a la Haskell, it's what programmers in the trenches are using. The same way a modern car is (pick a popular current model) and not some limited edition car with limited sales.
C# "keeps up with the Joneses", brings in ideas from functional programming (bringing in F# features incrementally), etc.
LINQ, pattern matching, dynamic, value types, tuples, destructuring, async/await, null-coalescing operators/assignment, streams, and much more...
Not only is up there with any popular language in the same domain (Swift, Kotlin, etc), and way ahead of Java, but compared to something like Go that gets people excited this is space technology.
(Plus is not like multiple dispatch is a new development, or gradual/optional typing is the "new" thing - we had those things for decades, and I don't see them catching on except the latter where necessary - e.g. adding gradual typing in JS/Python/etc variants like TS, where they can't have strict static typing for compatibility/historical reasons. The only language with gradual typing I remember in recent times was not very cutting-edge itself, Dart, and didn't manage to go anywhere).
What do you mean by memory hopping, reference types? C# has had value types since the beginning unlike Java.
1. https://docs.microsoft.com/en-us/dotnet/csharp/whats-new/csh...
And here is the kicker: Packages is Julia are far more interoperable than in any other eco system I have seen.
Than means 10 packages in Julia can quickly do more than 50 packages in other language. Say somebody makes a GPU processing package and somebody else makes a Machine learning package. They don't know about it each other.
Yet with Julia we frequently see that these guy can use each others package. The ML package can with little effort run on a GPU by including a GPU package even if it was never designed for it.
Julia just moves way faster than the competition and that is why Julia is NOT doomed.
That's because Python already has tons, while Julia lacks tons.
Are the new ones added to Julia as good or better though?
>That means 10 packages in Julia can quickly do more than 50 packages in other language. Say somebody makes a GPU processing package and somebody else makes a Machine learning package. They don't know about it each other. Yet with Julia we frequently see that these guy can use each others package. The ML package can with little effort run on a GPU by including a GPU package even if it was never designed for it.
This sounds like magic bullet / handwaving.
Perhaps Julia's multiple dispatch helps here (?), but why would the above be the case and wouldn't be achievable in other existing languages as is?
Yes, ubiquitous, zero-overhead multiple dispatch is the reason, as explained in depth here:
It's not handwaving when it already exists. CUDAnative.jl gives native compilation of Julia code to GPUs, which allows CuArrays.jl to do efficient loop fusion among other tricks. This is then enough for DifferentialEquations.jl, Flux.jl (the big neural net library, think PyTorch), CLIMA (a MIT-CalTech next-gen climate model), etc. (too many packages to list) to all utilize the same underlying GPU framework. It's as if someone got rid of the GPU functionality of Numba, PyTorch, TensorFlow, etc. and just made one good enough Python library for everyone to share. As you can imagine, that reduces the amount of work to do package development but gets the same functionality in the end with a lot less maintenance issues. You can also say the same for how many JITs and transpilers to C there are in Python libraries vs the single Julia JIT.
The real danger of the Python monoliths is that when something as simple as a neural network package starts to include a JIT, a GPU compiler, and automatic differentiation library, and it's all written in C++, it gets really hard to attract developers because the code base is so hard to understand. This means that instead of getting new developers and more maintenance, people just spawn new libraries all of the time in Python, starting from scratch to write a new JIT, new GPU kernels, etc. I just cannot see how that's a good thing.
> Than means 10 packages in Julia can quickly do more than 50 packages in other language.
Maybe you guys would like to compare your text editors now?
One of them is Julia! - https://github.com/JuliaPy/PyCall.jl .
The other is Swift - https://github.com/tensorflow/swift/blob/master/docs/PythonI... .
The main bottleneck was computing the Mittag-Leffler function. First we tried in Python, then we get a ~10x improvement by outsourcing that bit to Fortran. The Julia implementation of the Mittag-Leffler function is comparable to Fortran but writing the entire software package in Julia led to ~5x speedup due to avoiding language interop issues and taking advantage of other Julia-specific benefits.
And for anyone wondering about the state of the Julia package ecosystem compared to Python, someone had already written the main two packages we needed in Julia (for Mittag-Leffler computation and Inverse Laplace computation). At that time there was no Mittag-Leffler implemented in Python or SciPy that I could find.
I could also add that we were computing integral equations with Mittag-Leffler function in the Kernels so in worst cases we needed O(n^2) calls to Mittag-Leffler function where n could be O(10^5).
> No one should use global scope of Julia for calculations: [...] which is said that takes 35 seconds. > If you wrap it in a function and not had a type instability, You'd get 4.704 seconds...
Half-decently written Julia ought not to be much slower than C.
But it does also allow you to write quick and dirty scripts with no care for speed, which is also useful. It would be nice if people wouldn't benchmark the latter. The manual has a pretty helpful section for avoiding this: https://docs.julialang.org/en/v1/manual/performance-tips/
Even more importantly, from a cursory look online I cannot find any Lua implementation of either the Mittag-Leffler function or Inverse Laplace routines. And there is no reference to any wider ecosystem of scientific computing packages.
Sure I could call some C functions but then I'm back to facing a 2-language problem which I don't want.
Furthermore, it looks like there has been no new releases of LuaJIT since 2017. Development on it looks to have stopped.
- I want to build an new (yet another) distributed data processing framework, say, Hadoop or Spark Next, generic or for a particular industry, and I do not want to use Java or C++.
Can Julia be a feasible choice for a distributed computation like Spark, for distributed file system like HDFS, for resource allocator like Yarn? Can it be a good choice for a database engine, or its place is in layers above the engine - say a layer to support aggregation and computations?
Or it's not a right choice comparing with Rust, C++, or Java for core systems, and it's better to stick with it just for computations on the top of those core systems?
If I am a Chief Engineer / CTO of a company, and need to choose a language for Yet Another data platform, I would definitely look first at Java / Kotlin / Scala for implementation - they are a proven pragmatic choice for such tasks. If there are memory and garbage collection concerns of the former, I would look at C++ or Rust (still risky for staffing).
What I like in Julia is its expressiveness, and that it was built for writing maths and computations. So it's a great candidate for top level libraries above core components - it may be a reasonable choice for a business now.
But Julia promises good performance as well. Therefore, if one does not have very strict requirements to avoid GC it may look like an interesting idea to stick with the same language for the core components. Here was my question above :)
JuliaDB is about storing and retrieving Julia data end to end, although it works with e.g. CSV files etc as far as I know.
The benefit of using Julia over something like Java or C++ is that user defined function can easily be added and they are JIT compiled for maximum performance.
The main consideration with respect to using Julia is about whether your domain has problems with JIT compilation or not. Anything that is started and shut down frequently and only runs for a short time will not work well with a JIT based system like Julia.
Also places where you need fine control over latency such as computer games or real time systems may not be suitable for Julia. Then I suspect Rust or Swift would be better.
Other than that Julia is good for almost anything. I has great support for parallelism, concurrency and crunch numbers really fast.
While C++ and Rust may beat Julia in performance I think you should be able to outperform Java, because Julia is designed much more with performance in mind than Java. E.g. you have better control over memory layout, cache misses etc in Julia than in Java.
However that will be eventually fixed, and then we won't be able to bash Java any longer for lack of value types.
In what GC languages are concerned, it would be more interesting to compare Julia's generated code against .NET Core languages, D (specially ldc), Nim.
You can play around with value types experimental releases already.
https://jdk.java.net/valhalla/
Adopting value types, while keeping 20+ year old jars working without any kind of changes is an engineering feat.
I bet that Java will get value types before Go gets its generics, if ever.
It looks like Julia integrates quite seamlessly with python, so I'm hoping that we can start using it to easily speed up exploratory research code without having to spend a lot of time.
Only thing I'd wish for is better control of memory layout for mutable structs, but I know that's unlikely since that's one thing that the Julia folks want to keep abstracted from the user.
That's great news, it drives me crazy when languages or libraries get that wrong.
Edit: Dumb mistake on my part. I was looking at the video linked in the top comment (https://www.youtube.com/watch?v=YdiZa0Y3F3c) and then I came back to this thread and thought it was the topic.
https://keno.github.io/julia-wasm/website/repl.htm
Here's the repo:
https://github.com/Keno/julia-wasm
It's still a work in progress to get the full package manager on there, but there's funding from Mozilla to get it done. I'm excited to see how that turns out.
Upcoming release: v1.3.0-rc5 (Nov 17, 2019)
We're currently testing release candidates for Julia v1.3.0
Still when I built julia this morning, from the 1.3 banch, I can confirm the -RC5 label was gone.
...if you do anything towards ML or datascience, typescript and kotlin are the poor relatives. For most tasks 90% of the ecosystem is useless to you and you can spend more time wrangling the dependencies of a nasty nodejs package than quickly coding smth from scratch.
That "richness" is also... "garbage".
The best chance Julia has isn’t attracting average devs alone, but the library developers. Writing new python libraries for research purposes often requires writing C/C++ code. There’s already a number of Julia libraries focusing on interesting research topics in scientific, mathematics, and AI. Those lead to better and more useful libraries over time which can interact over time. Plus Julia Computing seems to be going towards building a successful commercial side of Julia. Personally I don’t find the software poverty thing much of an issue, outside a few core areas (database drivers, basic http, c interop, networking, text editors).
What exactly would you like Julia to be compatible with?
Python? https://github.com/JuliaPy/PyCall.jl
JS? https://github.com/SimonDanisch/JSCall.jl
C++? https://github.com/JuliaInterop/Cxx.jl
Java? https://github.com/JuliaInterop/JavaCall.jl
R? https://github.com/JuliaInterop/RCall.jl
Not all of these are fully polished, but they're all being worked on (plus there are a bunch of others that aren't being actively worked on that I haven't put here).
C/Fortran anything else that compiles to C-API supporting shared libraries is supported in built by `ccall`
JuliaLang is amazing at of FFI
I wrote a custom data pipeline for huge HDF5 files for some ai thing, which was a headache in linux but doable. Then it had to run on windows and I just wanted to kill myself.
Doing the same in Julia or even Matlab was easy, and faster.
And python is supposed to be easy to use :<
Even among static languages, it’s a relatively rare capability: Go has excellent support (of course), C extensions like Cilk and TBB support it, Rust has something similar with tokio, and I’ve heard people claim that Haskell can do something similar with appropriate extensions.
Could you elaborate on how you got that impression? Do you have any benchmarks that the Raku core developers could look at?
Threading can make a difference on non-IO bound tasks if you're interested in wallclock rather than CPU: for instance, `say (1..Inf).grep( { .is-prime } )[9999]` (showing the 10000th prime number) runs 22 seconds on my machine, but the threaded version of that: `say (1..Inf).hyper.grep( *.is-prime )[9999]` runs in 8 seconds. That's just by adding the `.hyper` method to the chain!
julia> using Primes, Lazy
julia> @time drop(9999, filter(isprime, Lazy.range(1)))[1]
0.119811 seconds (1.25 M allocations: 23.502 MiB)
104729
The Lazy package doesn't support using threads yet since 1.3 was just released, so doing the threaded comparison isn't simple at the moment. However, this illustrates what I was getting at: the sequential Julia code is 185x faster than the the sequential Raku code and 67x faster than the threaded Raku code. It's great that Raku threading gets some scaling here (not sure how many cores you have, so it's unclear if 2.75x scaling is good or not, but it's not nothing). But it's considerably easier to scale if there's already a lot of performance on the table. The faster each operation is, the harder it is to make a threading implementation that has low enough overhead to make threads worthwhile. From the performance perspective, any effort spent on threading in Raku would be better spent on sequential speed until that has been maxed out.I should also note that this is not how one would actually find the 10000th prime efficiently. For that you'd use the `nextprime` function also provided by the Primes package, like so:
julia> @time nextprime(1, 10000)
0.010456 seconds (17.00 k allocations: 265.562 KiB)
104729
That's another 10x faster than the lazy sequence approach. Which mostly tells me that the lazy code is impressively efficient—I would have expected a dedicated function to have more of an edge. Lest anyone cry foul about using C or whatever, this function is implemented fairly straightforwardly in Julia:https://github.com/JuliaMath/Primes.jl/blob/ce0c1e388e1fd375....
Just for the record: yes, there are faster ways of finding the 10000th prime, but I use this example often as an example of a CPU-intensive task that can be spread over multiple threads easily.
Also, please note that all features of the example are built-in into Raku: no external module loading needed.