I Love Julia
technology.stitchfix.com
technology.stitchfix.com
Yes, there are a number of workarounds. You can use python or R for plotting, you can keep a single interactive session going and reload your source file in it, you can precompile Gadfly into Julia's global precompiled blob, etc, but the solutions take time and effort that significantly offset Julia's value proposition (for me, at least). If you haven't looked at Julia yet you might want to hold off for a while until they get it sorted out. It looks like it might happen in the next version, 0.4, which hopefully will support pre-compiled libraries.
I feel like a bit of a dick for complaining about hiccups in a beta version, but I really like Julia and I really don't want to see it go the Clojure route and simply accept several-second delays in oft-repeated tasks. Drake, I'm looking at you -- Make can build my entire C++ codebase in less time than it takes your "make-replacement" to launch! I love your features, but the stuck-in-molassas feeling I get when using your program was enough to scare me away.
Base.require("Distributions.jl")
Base.require("Optim.jl")
Base.require("DataFrames.jl")
Base.require("Gadfly.jl”)
You'll need to create a new build after adding this userimg.jl file. Also be aware that changes to these packages will stop showing up when you type "using Gadfly" since you'll always access your precompiled version.
(and PyPy does break some python APIs and expectations, speed does come at a small cost, after all)
but yeah, for... probably 90%+ of use-cases, it is the PyPy<=>CPython differences that are more notable
The libraries, on the other hand... I know it's new, but the DataFrames.jl package in particular gave me fits. Data frames are essential tools for statistics, and there are many problems. When I last used it, it took several minutes to load modest 10MB TSV matrices, and segfaulted entirely on slightly larger ones. It doesn't support indexes on both axes, and the developers made the extremely questionable decision to require that index names be valid symbols. I could go on.
I think the core developers should exercise more control over the library ecosystem, at least for the packages that are crucial to the type of workflow they're building the language for.
If you have any ideas about how we should modify the basic data types and functions defined in DataFrames, those ideas would go a long way to making Julia a better language.
I fully appreciate that the type system imposes constraints that don't exist in Python or R. For my purposes in particular, and I think many people, I don't actually need a full-fledged data frame with heterogeneous types. What I actually want is a numeric matrix with labels on both axes and good methods for querying, group-by operations, etc. (And an equivalent numeric Series type). Big bonus for memory mapping and/or fast I/O.
I think this is an easier problem to solve, especially since factors and ordinals can be considered as a special type of numeric.
It has been too long since I've looked at the internal code structure of DataFrames.jl, but I think the biggest design flaws at the time were the requirements of index names to be symbols (probably should either be a flat String, or a choice between String and Int64), and axes on columns only. I can only assume the symbol decision was made for performance but you surely have worked with datasets given by investigators that use all kinds of random conventions for index names that don't fit the constraints of a symbol. Not to mention the very common case of numeric index names. I find it very annoying to read such a file in R and get "X1000" or whatever as my index names.
I actually tried briefly to dive in and fix the I/O problems, but the code style was daunting -- a few, very huge functions. If it hasn't been done, I would suggest breaking it up a little.
Anyway, I didn't mean to be overly critical -- I think you're doing a very important task -- but as an honest assessment of why I, as a busy scientist, found Julia to be more trouble than it was worth.
[0]: http://junolab.org
Is it different from pip install git+...? See https://pip.pypa.io/en/latest/reference/pip_install.html#vcs...
Put another way, does Julia sacrifice something relative to Python? Are the objects less flexible?
Could I write Python that compiles to Julia, without losing features?
tl;dr of the talk: Syntax has little to no effect on how a language performs. What distinguishes Julia from Python is that Julia's semantics were designed to be amenable to type inference. The results of type inference allow Julia's compiler to generate very efficient machine code.
* It looks like packages are installed in a global namespace that is shared by all projects. This seems like it will get messy when you try to run older and newer projects.
* The default way to add packages is to just Pkg.add("package-name"), this makes reproducible builds difficult. This is especially an issue with a language used in scientific contexts where reproducibility is extremely important.
Are there solutions to these issues that I can't see? I'm aware that Julia is a young language so I don't expect them to solve everything at once.
http://heike.github.io/stat590f/gadfly/carson-knitr
IJulia[1] is probably closest to the community supported equivalent, but it's based on IPython
Julia has roots with MIT which immediately lends it some cachet, and probably made it easier to grow a vibrant community.
Lush (which is an unfortunate name btw) used Lisp syntax, which has never been favored by math, science and engineering types. It seems obvious that equations as expressed in the language should closely resemble those used in actual math - thus the syntax of languages like Fortran, Matlab and Julia.
Writing (or reading):
(setq vx (+ vx (* ax deltat))) ; update velocity
is much more awkward than:
vx += ax * deltat # update velocity
This issue gets worse with more complex examples.Lush also allowed inline C, which was probably a bad design choice. Julia allows painless C library calls, which is much cleaner.
Those are just a few thoughts off the top of my head...
(incf vx (* ax deltat))
The equivalent of += is INCF.HN thread about it: https://news.ycombinator.com/item?id=7109982
Parts of Julia's library and compiler are implemented in C, but this actually isn't very relevant to the speed of the generated machine code that actually runs.
Statements about Julia being "on par" with C mean that if you write code in a straightforward way to solve some problem, e.g. "find the three largest even integers in a collection," then Julia is capable of generating machine code that executes with efficiency "on par" with the machine code that C generates.
The "straightforward" part in the last paragraph is actually important. You could in principle solve this problem in any language by writing your own machine code generator in that language, and then the distinction between efficiency of different languages breaks down. But usually you won't do that, and so usually the distinction does have some meaning.
However, numerical Python can be nearly as fast as C as well with very, very little additional work (using Numba means adding @jit on top of a function). The downside is that Numba only works on the 'numpy' subset of Python, basically.
But if you are suspicious, there are introspection utilities that let you see the generated native code. Give it a try.
It's using LLVM on the backed, so it should be quite possible.
Everytime I see this BS I think of this:
https://www.youtube.com/watch?v=lrp57IAlh84
It's fine to say something more reasonable like "I have high hopes for this very early stage language in the future", but this kind of fanfare is the reason why stuff like Java got so big.
If you don't like it, don't use, but don't piss and moan about the fact that some people enjoy it.
Out of principle I don't engage with products like these because you either become part of them or will never get truly involved.