Julia 0.5 Highlights
julialang.org
julialang.org
In my experience there are still some rough edges as compared to the Python ecosystem (of course!), which together with the 0.x status make it impractical for many production situations. However it is fantastic for prototyping numerical code, the type system is a pleasure, and the JuMP mathematical optimization library is a gem. Being able to have fast code be "first class," as opposed to the impedance mismatch of dealing with numba/Cython, feels great and is a real boon for trying new things. Then, it is fairly straightforward to port the final solution to whatever production language you use (e.g., Python with a sprinkle of numba).
- Needing to interact with an existing codebase, and an existing developer base. If everyone knows and uses Python and only a few use Julia, it is too early to put Julia in production. If there are proprietary libraries, now may not be the best time to commit to porting them to Julia.
- Language and ecosystem stability. I started something on 0.4, and with 0.5 there were a raft of deprecations. If the code will live several years, that's a support commitment with unclear value.
- Library maturity. If I need to build a web app, read an Excel, read a CSV with dates quickly, consume a SOAP endpoint, etc etc in Python -- no problem. With Julia I will mostly be fine, but am likely to run into some cases that are not yet 100% there.
- Most code does not need the extra performance, so once you have a fast prototype as a performance target it is often not that hard to hit similar performance with Python + numba/Cython.
Note for that last point: there is a lot of value to not worrying about this in the exploratory stage, and getting a performance target (for later optimization) as a nice byproduct.
The big breaking change is the array indexing one – that will very much require adjusting multidimensional array code.
I'm interested in people's everyday use of Julia and how it has impacted your workflow. I don't work with "Big Data" most of my data sets are bellow 100k in size. Anyone using Julia for medium and small data sets?
Fortunately the DataFrames ecosystem and related packages (like Query.jl and StructuredQueries.jl) seem to be headed in the right direction.
Could you elaborate? Which new tools?
RStudio has also been a God send for me. RMarkdown and now R Notebooks just amaze me. Then I work in a very MS Office environment that ReportRs (http://davidgohel.github.io/ReporteRs/) is the most under valued library right now.
I didn't have these things 5 years ago. (RStudio was initially released Feb 2011 but didn't start with it till 2013.
If you're comfortable using R and the tooling and performance it offers, you should probably keep using it. Julia isn't just about big data, but it does tend to appeal to people who are struggling with problems that are sufficiently hard that their old tools left them in a world of pain. That pain may comes from size, complexity, CPU-intensiveness, or need for more language expressiveness (a hard thing to define). In particular, Julia offers a unique combination of productivity and speed for numerical work that can't be found anywhere else.
I guess I am seeking the one "pet project" that would get me to jump in and give Julia a test drive.
Thanks for all your work even though I don't directly benefit from your work.
I generally find Julia to be a more expressive and fun language to program in than R or Python. It may be my background -- I've done a fair bit of work in Scheme, and in many ways Julia has lots of what I liked about Lisp in it. I don't need all the numerics and stats libraries that R or Python have; what little I need is easy to cobble together very quickly in Julia or has already been implemented. And the FFI is easy to use; calling C is pretty straightforward.
I have used Julia with a 50GB dataset for feature extraction while I was waiting for R to process the same dataset (10 minutes to an hour depending on the function). I actually learned some Julia while waiting for R. For what I was doing, Julia felt roughly 100 times faster (most delays under a minute).
However, I did not know about dplyr at the time and I will certainly try it in the next project. Although, I am skeptical that using fast libraries is a silver bullet since, at some point, the data will pass through some slow custom R code that I wrote myself.
I prefer to use Julia - the API feels cleaner and more manageable, I don't suffer from subtle type errors, and there is also the possibility of running things in parallel from the same repl. I also had no memory leaks -- I suffered a lot from the bloat caused by my use of R closures.
That said, the power and output quality of ggplot means that I will probably stay with R for plotting. There will also be times when CRAN has the only library that does something we need.
data.table is fast for your project I would imagine.
> Julia’s LLVM version was upgraded from 3.3 to 3.7.1.
> [...] we’re very happy to be back to using current
> versions of our favorite compiler framework.
LLVM's current version is 3.9, is this a typo or are there problems that prevent Julia from being used with any release newer than 3.7.1?Either way, looking forward to reading about where the optimizations themselves came from.
Disclaimer: I am not affiliated with the LLVM project.
I have tried to use it several times for toy compilers, and each time is a pain for do simple things. I can't imagine how much pain is for depend your project on it.
Certainly worth the effort, but as expected of a C++ tech: You need a army of specialist to make it work.
Julia uses its type system for:
- self-documentation
- reduction of boilerplate manual type checking that litters libraries in dynamic languages
- all those times you need to express the type of something, which happens especially frequently in numerical code
- performance
In the future, some "type linting" could be built into the standard library since we can infer types for so much code, but it isn't the top priority.
The only confusing edge case I've encountered is in biology, when manipulating a genomic region of zero length. Such a region is used, for example, when representing the location of an insertion of additional sequence. Before the insertion, the region has zero length. So, if the insertion occurs after the 500th element, then the zero-length range, in 1-based indexing, starts at 501 and ends at 500. However, it's rarely necessary to deal with this case directly, as such a range can be constructed as "starting at 501 with length 0".
Could you elaborate on that? What about statistics do you feel makes it necessary to have 1-based indexing?
I've used both 0 based (python, C++), and 1 based (R, matlab, mathematica), but I really belive 1 based is the right choice for Julia. It isn't necessary, but it feels better.
It's mostly about expectations. Users of statistical software, who may or may not be programmers by inclination, use 1-based indexing. The program they came from uses 1-based indexing (R, SAS, Stata, etc.). When you talk about data, you rarely talk about the 0th observation. It's just a recipe for errors to have to code switch between 0 and 1-based indexing.
I am totally fine with zero base in Python and Lisp and such. BUT when I am doing statistics and the math is 1 based I think the ability to make a mistake is to big. Also for subset in R df[1, 1] would be the 2nd row and 2nd column would just throw most users of R.
The REPL is on par with ipython (perhaps even better). Also are some IDEs for it but they aren't quite there yet.
Would that be covered by https://github.com/JuliaLang/julia/pull/18632 or is it a separate issue? I have some really ugly code that ought to be built out of nested generators, but I can't afford the millions of heap allocations.
addendum: read 'static' for 'strong', please
Someday, Haskell will get a good matrix library, and then I'll be happy.
julia> xs = ["foo", "bar"]
2-element Array{String,1}:
"foo"
"bar"
julia> push!(xs, 42)
ERROR: MethodError: Cannot `convert` an object of type Int64 to an object of type String
This may have arisen from a call to the constructor String(...),
since type constructors fall back to convert methods.
in push!(::Array{String,1}, ::Int64) at ./array.jl:479
julia> xs = Any["foo", "bar"]
2-element Array{Any,1}:
"foo"
"bar"
julia> push!(xs, 42)
3-element Array{Any,1}:
"foo"
"bar"
42
If what you are looking for is a statically typed language, then Rust has a very similar approach to data, types and polymorphism but with much more emphasis on catching errors at compile-time, at the cost of poor interactive development.Although I'm not sure if it's dead or just pining for the fjords
As for being functional, strictly speaking it is not, but still has some features like closures and higher order functions, immutable variables, the usual suspects (map, filter, ...) can be inlined with zero overhead, side effect tracking and so on (in addition to a very good macro system)