Some fun with π in Julia
julialang.org
julialang.org
I write this as a huge Julia fan; I use Julia daily, and it is both my favorite language and the language I know best. So I already think Julia is great. But reading many Julia blogs, especially those from julialang.org, would make me think Julia is only useful for very narrow scientific applications if I didn't already know better.
I like Julia because it's extremely fast and extremely expressive - and I don't just mean "expressive" as in "can be written like a dynamic/scripting language," though it can. I mostly mean that the combination of its type system and multiple dispatch allows some really elegant abstractions.
Or you know.. it's Pi day.
On a 4-core CPU I have the environment variable JULIA_NUM_THREADS set to 3 so that when Julia is doing threaded worked there's still a core free for web browsing.
One potential gotcha: the Julia docs will mention that you can access shell commands from the REPL by starting your line with a semicolon, e.g. ;ls. For me this only works if you start Julia from something like Git Bash, not cmd.
Lots of people seem to like Juno (junolab.org) as an IDE. The team behind it has made incredible progress recently, and sometimes I use it for its GUI around the debugger, but for the most part I tend to stick to Notepad++.
...compared to Python, under some circumstances, disregarding startup time.
Don't oversell it. Don't confuse what Julia aspires to be with what it is, or you'll just turn people off when they feel they've been misled. There are extremely fast ahead-of-time-compiled languages that would have finished their computation while Julia is still JIT-ing its kernel.
Disregarding JIT time, Julia is as fast as compiled languages – we benchmark against C and Fortran and don't grade on a curve. This is not an aspiration, it's a well-established performance characteristic in both benchmarks [1] and real-world applications [2].
Regarding JIT time, it's important to understand that this is a fixed cost – for any computation that takes long enough to matter, the time taken by JIT will be a negligible fraction of the time. The reason to exclude JIT time when benchmarking is because most benchmarks don't take much time – tens of milliseconds, compared to which JIT is significant. But we don't actually care about the speed of things that only take tens of milliseconds – we're using benchmarks to extrapolate to things that take much longer, which is why you exclude JIT.
The time to parse the data I needed was many minutes, slower than Python. Restarting a Jupyter kernel would break my flow every time.
The benchmarks are all for numerical computing, where the hard work is typically offloaded to a BLAS library anyway, and don't involve strings or dictionaries. Let's see a JSON parser in there and not just a Mandelbrot set.
How is string performance doing?
You should seriously try out Julia again. I think it really works well now. I used to have a lot of issues with it before. Stuff would break, slow startup times etc. Today I don't really have any complaints.
Ok I got a few small ones. The limits on the ability to redefine stuff as you develop in the REPL can be a bit cumbersome.
The way I wrote the current code in Python was abusing sets and dicts a lot to take advantage of those fast data structs. Rewriting in Julia was fun because it was different and because multiple dispatch is fun.
However, it was roughly the same speed as Python. Ended up sprinkling some Cython on top of Python and resulted in 10x speedup. Didn't take much time to add types/pass pointers instead of strings to functions. I am not at all familiar with C++/Cython.
I think even if they speed strings/dicts up by a lot, there seem to be lots of breaking changes between releases so I wouldn't try it for something big.
I think if Julia is to succeed in the near future in the same way Python is successful for data science, it needs to be more usable for general tasks. Things like web servers, fast JSON parsing, maybe static binaries or easy parallelism. Basically, some more selling points. So far, for me personally Python is faster and easier to read for most of the things I write. At least given comparable amount of work.
I built an in-memory relational query compiler in Julia which keeps up with Postgres on the Join Order Benchmark.
http://scattered-thoughts.net/blog/2016/10/11/a-practical-re...
Last I checked there were still some limitations around stack-allocation of types containing pointers, which is occasionally painful, but other than that it gave me all the tools I could possibly want.
https://github.com/JuliaLang/julia/pull/18632
The startup time is annoying, but I typically use Julia from Juno and I restart at most a couple of times per week. Ctrl-shift-enter recompiles the current module, which is usually all I want.
Startup time isn't the biggest problem, although it is nowhere near "comparable to the JVM". The JVM will start cold, load a Hello World application from a JAR, run it and shut down in 60-80ms. A Julia Hello World app would take at least a second, even when precompiled into a native executable.
The biggest problem, IMO, is packaging and deployment. I used BuildExecutable, which works OK most of the time (although startup times are still slow), but it results in so many shared libraries, with no easy way to cull them. Without BuildExecutable, I couldn't even find documentation on how to deploy a Julia app as a self-contained program (even not precompiled, but in a way that doesn't require Julia to be installed on the user's machine).
I'm curious because I'm trying to find an example of a "non-scalar" programming language, a la this attempt here[1].
Thanks for the help!
I translated some of my code from one of those general purpose languages to Julia 0.2, years ago. It was both more readable, and 20x faster. Never looked back.
In Python, I can do the same thing if I install the Pandas and Statsmodels libraries first.
Try that in Ruby, Perl, C++, C, Java, Rust, Haskell, Common Lisp, or just about any language you can think of. Good luck.
This recent post gives some more motivation: https://discourse.julialang.org/t/julia-motivation-why-weren....
"A lot of Julia's design fell directly out of making it easy to generate LLVM IR – in some ways, Julia is just a really great DSL for LLVM's JIT."
Excited to learn more
Quasiquoting means that codegen is as easy as string interpolation is in other languages.
quote
let
$(index_inits...)
$(results_inits...)
if $(reduce((a,b) -> :($a && $b), true, index_checks))
let
$(var_inits...) # declare vars local in here so they can't shadow relation names
$body
end
end
tuple($(results...))
end
end
Great introspection into the inference and compilation pipeline, directly from the repl. julia> function double(xs)
[2*x for x in xs]
end
double (generic function with 1 method)
julia> double([1,2,3])
3-element Array{Int64,1}:
2
4
6
julia> @code_warntype double([1,2,3])
Variables:
xs::Array{Int64,1}
#s1::Int64
#s2::Int64
#s3::Int64
x::Int64
#s4::Int64
Body:
begin # none, line 2:
GenSym(1) = (Base.arraylen)(xs::Array{Int64,1})::Int64
0:
GenSym(3) = (top(ccall))(:jl_alloc_array_1d,(top(apply_type))(Base.Array,Int64,1)::Type{Array{Int64,1}},(top(svec))(Base.Any,Base.Int)::SimpleVector,Array{Int64,1},0,GenSym(1),0)::Array{Int64,1}
#s1 = 1
#s2 = 1
#s3 = 0
unless (Base.box)(Base.Bool,(Base.not_int)(#s3::Int64 === GenSym(1)::Bool)) goto 2
3:
#s3 = (Base.box)(Base.Int,(Base.add_int)(#s3::Int64,1))
GenSym(9) = (Base.arrayref)(xs::Array{Int64,1},#s2::Int64)::Int64
GenSym(10) = (Base.box)(Base.Int,(Base.add_int)(#s2::Int64,1))
#s4 = 1
GenSym(11) = GenSym(9)
GenSym(12) = (Base.box)(Base.Int,(Base.add_int)(1,1))
x = GenSym(11)
#s4 = GenSym(12)
GenSym(13) = GenSym(10)
GenSym(14) = (Base.box)(Base.Int,(Base.add_int)(2,1))
#s2 = GenSym(13)
#s4 = GenSym(14)
GenSym(4) = (Base.box)(Int64,(Base.mul_int)(2,x::Int64))
$(Expr(:type_goto, 0, GenSym(4)))
$(Expr(:boundscheck, false))
(Base.arrayset)(GenSym(3),GenSym(4),#s1::Int64)::Array{Int64,1}
$(Expr(:boundscheck, :(Main.pop)))
#s1 = (Base.box)(Base.Int,(Base.add_int)(#s1::Int64,1))
4:
unless (Base.box)(Base.Bool,(Base.not_int)((Base.box)(Base.Bool,(Base.not_int)(#s3::Int64 === GenSym(1)::Bool)))) goto 3
2:
1:
return GenSym(3)
end::Array{Int64,1}
julia> @code_llvm double([1,2,3])
define %jl_value_t* @julia_double_21481(%jl_value_t*, %jl_value_t**, i32) {
top:
%3 = alloca [4 x %jl_value_t*], align 8
%.sub = getelementptr inbounds [4 x %jl_value_t*], [4 x %jl_value_t*]* %3, i64 0, i64 0
%4 = getelementptr [4 x %jl_value_t*], [4 x %jl_value_t*]* %3, i64 0, i64 2
%5 = getelementptr [4 x %jl_value_t*], [4 x %jl_value_t*]* %3, i64 0, i64 3
%6 = bitcast [4 x %jl_value_t*]* %3 to i64*
store i64 4, i64* %6, align 8
%7 = getelementptr [4 x %jl_value_t*], [4 x %jl_value_t*]* %3, i64 0, i64 1
%8 = load i64, i64* bitcast (%jl_value_t*** @jl_pgcstack to i64*), align 8
%9 = bitcast %jl_value_t** %7 to i64*
store i64 %8, i64* %9, align 8
store %jl_value_t** %.sub, %jl_value_t*** @jl_pgcstack, align 8
store %jl_value_t* null, %jl_value_t** %4, align 8
store %jl_value_t* null, %jl_value_t** %5, align 8
%10 = load %jl_value_t*, %jl_value_t** %1, align 8
%11 = getelementptr inbounds %jl_value_t, %jl_value_t* %10, i64 1
%12 = bitcast %jl_value_t* %11 to i64*
%13 = load i64, i64* %12, align 8
store %jl_value_t* inttoptr (i64 140140655870352 to %jl_value_t*), %jl_value_t** %5, align 8
%14 = call %jl_value_t* inttoptr (i64 140149385306688 to %jl_value_t* (%jl_value_t*, i64)*)(%jl_value_t* inttoptr (i64 140140655870352 to %jl_value_t*), i64 inreg %13)
store %jl_value_t* %14, %jl_value_t** %4, align 8
%15 = icmp eq i64 %13, 0
br i1 %15, label %L4, label %L1.preheader
L1.preheader: ; preds = %top
%16 = load i64, i64* %12, align 8
%17 = bitcast %jl_value_t* %10 to i64**
%18 = bitcast %jl_value_t* %14 to i64**
br label %L1
L1: ; preds = %idxend, %L1.preheader
%"#s3.0" = phi i64 [ %22, %idxend ], [ 0, %L1.preheader ]
%"#s2.0" = phi i64 [ %26, %idxend ], [ 1, %L1.preheader ]
%19 = add i64 %"#s2.0", -1
%20 = icmp ult i64 %19, %16
br i1 %20, label %idxend, label %oob
oob: ; preds = %L1
%21 = alloca i64, align 8
store i64 %"#s2.0", i64* %21, align 8
call void @jl_bounds_error_ints(%jl_value_t* %10, i64* nonnull %21, i64 1)
unreachable
idxend: ; preds = %L1
%22 = add i64 %"#s3.0", 1
%23 = load i64*, i64** %17, align 8
%24 = getelementptr i64, i64* %23, i64 %19
%25 = load i64, i64* %24, align 8
%26 = add i64 %"#s2.0", 1
%27 = shl i64 %25, 1
%28 = load i64*, i64** %18, align 8
%29 = getelementptr i64, i64* %28, i64 %19
store i64 %27, i64* %29, align 8
%30 = icmp eq i64 %22, %13
br i1 %30, label %L4.loopexit, label %L1
L4.loopexit: ; preds = %idxend
br label %L4
L4: ; preds = %L4.loopexit, %top
%31 = load i64, i64* %9, align 8
store i64 %31, i64* bitcast (%jl_value_t*** @jl_pgcstack to i64*), align 8
ret %jl_value_t* %14
}
julia> @code_native double([1,2,3])
L128:L159:L177: .text
pushq %rbp
movq %rsp, %rbp
pushq %r15
pushq %r14
pushq %rbx
subq $40, %rsp
movabsq $jl_alloc_array_1d, %rax
movabsq $140140655870352, %rdi # imm = 0x7F750A030190
movq $4, -56(%rbp)
movabsq $jl_pgcstack, %r15
movq (%r15), %rcx
movq %rcx, -48(%rbp)
leaq -56(%rbp), %rcx
movq %rcx, (%r15)
movq $0, -40(%rbp)
movq $0, -32(%rbp)
movq (%rsi), %r14
movq 8(%r14), %rbx
movq %rdi, -32(%rbp)
movq %rbx, %rsi
callq *%rax
movq %rax, -40(%rbp)
cmpq $0, %rbx
je L159
xorl %ecx, %ecx
movq 8(%r14), %rdx
nopw %cs:(%rax,%rax)
cmpq %rdx, %rcx
jae L177
movq (%r14), %rsi
movq (%rsi,%rcx,8), %rsi
shlq $1, %rsi
movq (%rax), %rdi
movq %rsi, (%rdi,%rcx,8)
incq %rcx
cmpq %rcx, %rbx
jne L128
movq -48(%rbp), %rcx
movq %rcx, (%r15)
leaq -24(%rbp), %rsp
popq %rbx
popq %r14
popq %r15
popq %rbp
retq
movq %rsp, %rsi
addq $-16, %rsi
movq %rsi, %rsp
addq $1, %rcx
movq %rcx, (%rsi)
movabsq $jl_bounds_error_ints, %rax
movl $1, %edx
movq %r14, %rdi
callq *%rax
It catches type errors early, thanks to the typed multiple dispatch. julia> xs = []
0-element Array{Any,1}
julia> push!(xs, 42)
1-element Array{Any,1}:
42
julia> push!(xs, "foo")
2-element Array{Any,1}:
42
"foo"
julia> ys = Int64[]
0-element Array{Int64,1}
julia> push!(ys, 42)
1-element Array{Int64,1}:
42
julia> push!(ys, "foo")
ERROR: MethodError: `convert` has no method matching convert(::Type{Int64}, ::ASCIIString)
This may have arisen from a call to the constructor Int64(...),
since type constructors fall back to convert methods.
Closest candidates are:
call{T}(::Type{T}, ::Any)
convert(::Type{Int64}, ::Int8)
convert(::Type{Int64}, ::UInt8)
...
in push! at ./array.jl:432
Plus, I only have to think in one language, but I can write sloppy dynamic heap-allocating-everywhere code in the compiler and with just a bit of thinking emit zero-allocation statically-dispatched code in the output.When I'm building software, I tend to like things to have very tight interfaces -- box everything up into components, understand how they talk to each other, etc. Define interfaces, implementations, types, etc -- In general, optimizing for long-term maintainability.
When I'm exploring data, my thought process is much more "scatter everything about on the desk" and "let me run these 5 lines of code again within the current context". "What does this thing look like", etc. Having "a table of data" as a first-class citizen in the language, with all the libraries assuming that as input and everything optimized to work around / display / visualize such is incredibly useful.
https://github.com/interplanetary-robot/Verilog.jl
This is very important if you're planning to build hardware to do specific math - because Julia is incredibly good at mathematical modeling, and you can very rapidly set up comprehensive tests for your hardware designs with confidence.
The closest alternative is chisel, which is written in scala. Although it's more professionally maintained and more fully-featured, it's hard to call the verilog chisel emits "human-readable", and it's harder to set up comprehensive tests - berkeley hardfloat, which is a very impressive project in chisel, had several critical bugs in its implementation (that I found, using julia).
Drop a line. My contact info is in my profile.
So it understand the concept of missing value and works with logical operation. Also it have factor concept into the language too.
General Language may not have the tool to deal with missing value and having a library doesn't necessary means it's good as a built-in concept NA.
Python panda and such does not handle NA as beautifully as R. (http://pandas.pydata.org/pandas-docs/stable/missing_data.htm...)
Python Panda default to Null and before that it uses something else to represent missing data. Null value in general is the last thing you do when you don't know how to represent a certain type. If you have a type language with sum type and pattern matching then you don't even need null. And yes R have its own Null type so NA is separate from it.
Erlang have PID as a primitive.
I think any domain specific languages are smaller in rules and syntax, and it makes it very very easy to learn for experts and people of those domains.
In addition, scientific languages will have libraries mostly in that domain. You can't beat R's library when it comes to statistics. They will also have good plotting libraries.
What you also get with Julia and Fortran is code designed to be optimized for numerical computing. Python attempts to offer this kind of performance via Numpy, which is Python wrapper on C or Fortran BLAS library. Or by JIT compiling with Numba, which is something Julia does automatically the first time you call a function.
Because you need special math support, easy access to suitable operators (and/or overloading), etc.
At any rate - a lisp-like language with llvm backend, package manager, and a (small, but not trivial) group of users writing real software - what's not to like?
I (still) recommend the talk: "Julia - to lisp or not to lisp?", that gives a quick overview of some of the design choices:
For me, I like that it has a real story trying to balance "actual integer math" (not this silly machine constrained twos-compliment hack" and "you could conceivably write a ray tracer that wasn't unusably slow" :)
If you're asking "why Julia versus other languages" it's that, well, Julia is fighting to answer that question for itself. As far as I can tell:
- Versus R and Octave: performance, coherent syntax, and more features for writing "programs" instead of just "scripts" - Versus Python + the Scipy stack: its scientific features are built into the language (instead of being an awkward layer on top of it) - Versus any proprietary platform (SAS, Matlab, etc): it's open source and free-as-in-beer, and therefore not confined to legacy/enterprise applications
I'm a data scientist and I currently use R and Python. I've been wanting to give Julia a try for months, and now that the ecosystem is starting to mature (plotting and data frames are must-haves for me) it's making more sense to spend some time with the language.
Applause to the Julia contributors for their work on this innovative language with great out of the box support for modern computer chip architectures. However, they have fibbed to build up momentum, particularly their performance benchmarks. The tests are of compiled Julia with OpenBLAS for the benchmarks against out of the box versions of the other languages. Also, the benchmark code in other languages is with a style that is among the slowest implementations for each language. No seasoned programmer in any of those languages would write code in such a way.
It does seem that there is co-ordination to get Julia posts most attention, timing and upvoting.
Nonetheless, credit where it's due. It should become an awesome language, marketing hacks notwithstanding.
As for co-ordination on Julia posts - there is none. We submit all our blog posts to HN, and while some do reach the front page, many others do not.
Glad you like Julia and hope this helps.
Wonderful.
Insightful :)
It all depends on what you want to measure. The whole point of the micro-benchmark suite is to test very specific language primitives. I'd argue that the current set of benchmarks are more valuable for that than an "expert" implementation would be — that may end up simply testing the cleverness or resourcefulness of the expert.
1. https://github.com/JuliaLang/julia/blob/64409a0cae8b52d3f795...
Of course, such a benchmark it might not have the same marketing hue as claiming that Julia is 553 times faster than Matlab at parsing an integer, for example.
edit: looked up "call R from Julia", it looks like there's a n "RJulia" library for this. Assuming there's a Python equivalent? How does this compare to, say, using RPy2 in Python (which is nice but kind of annoying)?
You want RCall.jl and PyCall.jl. Both allow you to directly work with and manipulate native objects. RCall even allows you to bring up an R REPL.
As for your broader question, Julia definitely fills a gap that I have been pained by for years. I've used R since it was in beta, and it is slow, which is a pain when you are discussing numerical needs. Yes, you can program in something like C/C++, but that is painful because of its overhead and dependency complexity (although it's surprisingly become less painful over time). Python could be used too, and probably is better at this point in that regard, but it has many of the same problems as R.
Julia is open-source, fast, and well-thought out with regard to modern numerical programming problems. I can write something in Julia and it performs essentially as well as something in C, which is a huge time saver in multiple respects.
I do wish Julia were more general-purpose in its orientation, or that the solutions it offers were coming from a more general-purpose language, but at the moment that doesn't seem to be in the cards. Maybe as it grows it will find use as a more general-purpose language, which is possible; maybe as languages like Rust or Go grow they will occupy this niche as well. Rust is interesting to me in this way, but currently it has little to offer in terms of simplification over C++ for numerics, and Go is not friendly to numerics. I personally like Stanza, but it's in its infancy, and no one probably even knows what I'm talking about.
For whatever reason, my experience has been that numerical programming has been a kind of isolate in programming. Numerical computing has always seemed slightly neglected in programming languages, and languages that have targeted numerical computing have often never been able to shake the "domain specific" label. I've just sort of come to see it as part of the territory.
There's nothing wrong with Python, C, or R. Also, languages change rapidly, so who knows what will happen. At the moment, though, Julia offers the best of all three and the only big downside is lack of libraries, which is becoming less and less of an issue every day (I wouldn't say there's a lack of libraries, more that there's fewer libraries). So I think it's deserving of its current attention.
I guess the question is, why would a systems programmer use C, or a web programmer use Javascript, or a network infrastructure programmer use Erlang, etc. etc. etc.?
I'll mention multiple areas I think Julia excels. If you want to write quickly high performance numerical software then I don't see what the alternatives are.
It is the best language I've encountered for writing shell scripts. Bash is terrible due to the difficulty of factoring the code into functions and the frequency in which you forget to quote variables properly. Many end up thus using Python or Ruby. Ruby seems very popular, but Julia really works better. It has tighter integration with the shell and handles chaining, reading and writing to shell commands much nicer.
It also has all sorts of useful stuff ready to use out of the box and sane function names etc. I find it much more cumbersome to write shell scripts with python. Awkward to call processes. Got to always remember what every little module you got to import. Because of multiple dispatch, Julia allows much better naming of functions.
And as a computer language geek I love powerful LISP macros, but I can't get used to LISP syntax. Julia is the only language I know of which gives access to powerful macros and code generation similar to LISP.