Switching from Common Lisp to Julia
tpapp.github.io
tpapp.github.io
Most implementations do it (CCL SBCL CMUCL ECL LW ACL SCL ...), except for a few unpopular implementations. Seven do what you want, but a couple don’t. What gives? Julia has neither a standard nor more than one implementation.
2. No efficient parametric/ad hoc polymorphism.
You can get far with inlining. You can also use libraries specifically built with this in mind. But you’re right, it’s not a built-in feature. But scientific codes are generally not super polymorphic anyway. Look at all the FORTRAN.
But besides that, Lisp let’s you build the language you want to use, and have it interoperate for the rest of the ecosystem. If you are surprised that building a language is work, then Lisp may not be the choice for you. If you’re willing to design the language you want for numerical computing, Lisp has the essential data types and primitives to allow you to do that, even portably.
If you want off-the-shelf features of Julia, then using Julia is more optimal than Lisp. If you want a language that will readily adapt without changing under your feet, without risking your programs not working 2, 10, or 25 years from now, then Lisp might be a good choice.
I totally get the reason why he would go to Julia for doing data jobs. Data structures and already done libraries are very valuable.
Fortran and C conventionally deal with this by adding the relevant type signature to the function name.
Scientific programming is also, regularly, polymorphic in 2, (often parametric) types. This heavily motivated the design choices in Julia esp the multiple dispatch infrastructure that is descended from Dylan.
You can make them generic, but it’s not often you gain anything at all.
They do this, not because it's inherently a good thing, but because the advantages of using Fortran or C have outweighed the disadvantages. But where that's not true, for example in end user computing, Fortran has not been popular for decades.
What the Julia devs are attempting to do is to design a language that has the ergonomics of a Matlab or R to make scientific programmers productive whilst having the performance of a Fortran or C. They have a sophisticated type and dispatch system to help bring these goals together.
Now, their current focus appears to be more on the R/Matlab replacement than on becoming the language of choice for HPC, say. So the steady state performance is unlikely to be competitive at scale as yet.
Maybe they won't get there. Maybe the decisions they've made mean that it's effectively impossible to eek out those last few percent so we'll still have a 2 language model. I don't know but I applaud them for trying.
More, I applaud the way they have gone about their development: focus on usage, actively learning from other languages, very little linguistic dogma. I hope they succeed.
The fact that it can be used in an important use case is a necessary but not sufficient condition that it can replace the incumbents. Same goes for end user computing.
Now, it gives me great confidence that, without the apparent short term focus of the core devs in this field, there has been such progress as Celeste. This is a good sign that the fundamentals are right.
But I'm probably more encouraged that the community are focused on user experience and delivering an environment that scientists want to use. This is essential, IMO, and appears to be working well.
1. Allowing arbitrary array types means your algorithm can already work on the GPU via GPUArrays.jl (https://github.com/JuliaGPU/GPUArrays.jl). This means that all differential equation, optimization, numerical linear algebra, etc. routines which internally utilize the functions which GPUArrays implements will automatically compile a special version when encountering a GPUArray that will do all computations on the GPU without any data transfer back and forth.
2. Matrix types allow the user to specify algorithms for linear solving. Numerical linear algebra is an entire field based around coming up with good algorithms for solving linear systems (\), and the whole point of the field is to specialize on the type of matrix you have. With Julia, you can set it up so that way the user passes what type of matrix they have (Tridiagonal, sparse, PETSc for multi-node parallelism, or a special matrix type with \ overloaded to do multigrid) and the optimization, differential equation, etc. routine will use the fast linear solver routine specific to the problem. This makes it easy to expose very large amounts of performance gains and specialization opportunities to the user in ways that are usually automatic (i.e. sometimes that user doesn't even have to know!).
3. Complex numbers. Large fields of scientific computing (physicists and engineers) use complex numbers. Many libraries have pitiful support for complex numbers. Generic typing helps a lot here.
Those are 3 off the top of my head, and there are many more diffeq-specific examples that have come up as users have asked to have resizing control models with discrete variables mixed in with continuous variables, etc., and this can all be handled efficiently via the type system without having to restructure the core algorithms to handle these extra features.
The key thing here is that users can add new features specific to their problem into your algorithm as they need to, without having to modify your code, just by smart uses of the type system.
I suspect this strength of Lisp could have undermined improvement of Lisp compilers in long term. That is, the fact that programmers can customize the language to run efficiently for their applications put less pressure to make the existing compiler better, compared to the languages that don't give programmers such flexibility.
I've worked on performance-sensitive commercial Common Lisp applications. We employed heavy macrology so that optimal instructions were generated in the performance critical regions. Effectively, it was reimplementing part of the compiler to deal with out domain-specific meta information such as parameterized types. It worked, but came with the cost of maintenance--as if we were maintaining another layer of the compiler. It was a burden. I'm sure there have been such effort spent in many other places and eventually abandoned.
Meanwhile, languages that give less power to the programmers, put lots of efforts to improving the compiler and/or the language that work more effectively with the improved compiler, and over a few decades, they have quite sophisticated compilers.
I still mostly use Lisp-family languages at work, but sometimes wonder if the power could have adverse effect in long term.
Perhaps, then, Julia makes sense in those cases.
Or, it's time for an enhanced Common Lisp implementation targeted specifically for scientific computing. I agree that one thing is "extending the language" and other is "implementing things that the compiler should have given me. "
Thus in this case Julia is a sane choice, unless some valiant, progressive Lispers want to fork one of the implementations and create a new Common Lisp implementation focused on high performance scientific computing...
This implementation of parametric types for Common Lisp looks nice:
https://github.com/cosmos72/cl-parametric-types
It allows you to do templating which additionally declares the (parametrized) types, so the compiler can optimize accordingly.
Worth a look.
Whereas to ignore CL & stick to SBCL will be working to fracture the community-- perhaps necessary, but CL is already a niche community
The CL community knows one uses the implementation one needs. Lispers are using ECL whenever they need to embed with C code, ABCL to interact with JVM, CCL when they're with macs, this is not fragmentation at all, they are all compatible with the CL stanfard.
Which major incompatibilities are you thinking of between CPython and PyPy, which are the two I use most often? Are they more severe than the incompatibilities between the different Common Lisp, which the author describes?
(I regard C extensions as implementation-specific features outside of the Python language. The different Common Lisp implementations also have implementation-specific features which are not portable.)
MicroPython is deliberately not Python compatible (differences at http://docs.micropython.org/en/latest/pyboard/genrst/index.h... ) but the others have a goal of not having major incompatibilities.
Yes: Common Lisp, Java, C and C++, to give four examples. Javascript as well. All these languages have formal standards, which helps a lot.
>Which major incompatibilities are you thinking of between CPython and PyPy, which are the two I use most often? Are they more severe than the incompatibilities between the different Common Lisp, which the author describes?
PyPy attempts to implement the full Python 2.7 language but there are things missing. You argue that Python's CFFI features are "implementation-specific", let's just agree with this; but the Python 2.7 language also defines -as part of the spec- many modules, and PyPy does not implement all of those modules, so it's not implementing the full set of features defined by that spec.
Code written in Common Lisp should run (and often runs) correctly with no change in all mature Common Lisp implementations (and there are many of them), unless the code uses implementation-dependent features. Those implementations implement the full set of features defined by the ANSI Common Lisp standard.
Regarding missing modules, which ones? And don't C compilers have similar issues? https://en.wikipedia.org/wiki/C99#Implementations helpfully points out Clang "Supports all features except C99 floating-point pragmas", and Microsoft's compiler doesn't implement tgmath.h.
I thought also that most Python code runs correctly with no change in CPython and PyPy, where "correctly" includes not using implementation-dependent features like third-party C extensions or garbage collection behavior.
That was one of the ways how Sun made money with it, by certifying implementations for the trademark symbol.
Some examples you might know:
HotSpot, the base for Java SE and OpenJDK
OpenJ9, now run by Eclipse.
DoppioJVM, a JVM for JavaScript.
[0] https://en.wikipedia.org/wiki/List_of_Java_virtual_machines#...
And you really want to use PHP as an example of a good situation?
The first language you name in your "single implementation"-list is Java, which is standardized and has multiple implementations.
Maybe it's "quality" or "lack of bloat"? That can't be, because one of the largest clusterfucks in those departments, PHP, is on the left.
Maybe you're talking about popularity? The list of "standard clusterfucks" includes C, C++ and (though you don't mention it) JavaScript, which are some of the most popular programming languages in the world.
What are you trying to say?
Wouldn't sticking to say SBCL (that does give you specialized arrays) be equivalent to using a one-implementation language that does the same?
Julia sounds quite appealing as a replacement for MATLAB or R with some Dylan-like semantics plus types and really efficient code generation on LLVM. I wish the Racket - Chez merger lead to something that targeted LLVM to be able to do front-end and back-end stuff using Scheme.
A JVM-centric stack is one of my preferred alternatives. Clojure is great for data preprocessing and manipulation, plus ClojureScript for coding all front-end. Datomic, core.spec, core.logic, anglican, plumatic.plumbing just to name a few are a joy to use. Then there's Scala, which is also a great asset, and tons of fantastic Java / Scala libraries like Stanford NLP, Markov Logic Networks (Tuffy, RockIT...), Factorie, Deeplearning4j, etc. Sadly, Scala-Clojure interop is not very good.
Python is the other obvious option, with tons of good libraries, including a fantastic data analysis ecosystem built around NumPy, SciPy, Matplotlib and Pandas. Plus most deep learning libraries targeting Python first. I just feel the language doesn't scale that well, although things like Numba or Cython help.
The tradeoffs compared to Python: LuaJIT is much faster, which should ease your scaling worries, but the ecosystem is not as developed.
If you compare LuaJIT with PyPy or Numba, the “much faster” argument will simply not be true. What’s different is that both of the above only cover a part of the language, but for specialised (e.g. numeric) application that’s often not a problem.
Source? AFAIK it does hold true. Last time I benchmarked those, LuaJIT was significantly faster, with PyPy almost being on pair with vanilla Lua, due to Python being a much more complex language and harder to optimise.
I would actually recommend against lua because of the lack of libraries for doing every day data science tasks.
I mean look at what facebook had to do to justify using lua, it had to invent the notebook for it.
My team builds deeplearning4j. We're aware of the massive demand for python and built a bridge to our tensor library: https://github.com/deeplearning4j/jumpy
This library does direct pointer mapping between our JNI based tensor library and cython (no network!)
So you could off load some of your work to the JVM using pyjnius (which this library uses underneath)
It's not a full solution yet but it's definitely a start to something promising!
We also import python models. We only support keras right now but our new autodiff library (samediff which will also be usable from python!) will handle onnx and tensorflow.
For visualization we tend to use zeppelin which has worked well enough in practice. If you have any specific suggestions or use cases I'm more than glad to take input though. We would love to build a python friendly JVM backend.
Other notable work in this space is what wes is doing with arrow. We are looking at using their tensor interop (it's still kinda green field yet..but it holds promise!) to do zero copy ETL between python and java. That should help as well.
However, I really don't understand how the author complains that CL doesn't "have" certain features that in fact are available by choosing a suitable CL implementation. But no, he prefers switching to a language with no standard (yet) and only one implementation...
Well, many Common Lisp implementations, which includes most of the famous ones like SBCL, CCL, and ABCL, will give you exactly...
... an array of double-float!!
So where is the problem?
>"However, this gets worse: while you can tell a function that operates on arrays that these arrays have element type double-float, you cannot dispatch on this, as Common Lisp does not have parametric types."
Well, there are many options. First, let me reiterate that an array of a certain element type, stays of that element type. Example:
CL-USER> (defparameter *a* (make-array 0 :element-type 'double-float ))
*A*
CL-USER> (type-of *a*)
(SIMPLE-ARRAY DOUBLE-FLOAT (0))
I'm going to give the simplest, quick&dirty options:If you're using arrays of different element-types, option A is write your function and just use the ETYPECASE function to do the dispatch according to the element type, so you can ensure that you select the code correct to the element type. Use the DECLARE declaration specifier so the Lisp compiler knows which type is your array /array elements and thus the machine language code produced is optimal.
Option B is simple, just define classes or structs, one for each array type you intend to use; and then take advantage of CLOS dispatch so invoking the generic function dispatches to the correct code for each array.
These are two options which don't even need any macro solutions; there are many more options as well, i'm just proposing two.
Option C, more elegant, could be perhaps using Fare's "LIL" (Lisp Interface Library)? https://common-lisp.net/~frideau/lil-ilc2012/lil-ilc2012.htm...
I really think Julia is a nicely designed language, and is a language I often tell people to take a look at. However, if the author already has a program in Common Lisp, and has good experience of Common Lisp (which, by the way, means the author might be a quite skilled programmer), why not dig deeper into the facilities that Common Lisp brings to solve those problems?
NOTE: This is an edited version of the comment i left on the page.
- 1) lisp 1 vs lisp 2
- 2) Matlab syntax (y tho?), infix, and sygils everywhere
- 3) Clos and metaobject protocol and ability to adjust these as easily at run-time
You do appear to gain some things though in all honesty. I'm trying to program up an dataframe/analytics type program/library in Common Lisp as we speak, and I'll admit, getting that stuff right and efficient is hard work. Gives you immediate respect for anyone who has done the same in other languages.
Additionally:
- 1) Parametric types and dispatch on them
- 2) Potentially inlining and compiler optimisations and integration around their generic functions. My gut says SBCL would need some compiler magic and metaobject protocol type stuff to do the same, and it would no longer be standard Common Lisp. That's not necessarily a bad thing, there are some warts around in-builts/objects/generic functions in Common Lisp, but it would be a fracturing of the community to update/change them.
[0] https://github.com/guicho271828/inlined-generic-function
Eh. Emacs vs Vim.
2) Matlab syntax (y tho?), infix, and sigils everywhere
The target audience is scientific and mathematical users, not necessarily "programmers", and certainly not "software engineers". Matlab syntax, infix, and sygils read like math notation. The target use-case is mathematical and scientific programming. This is a benefit in my opinion.
Also, with respect to sigils, I can think of at least one popular sigil in Common Lisp: the #' reader macro for accessing functions! So much for Lisp-1 vs Lisp-2.
3) Clos and metaobject protocol and ability to adjust these as easily at run-time
Fair, but do you really need CLOS with a delightful type system and multiple dispatch?
Structures, types, generics, etc, do cover a lot of what people actually want in practice...
For scientific computing, no. However it is very useful in other areas such as web, agents, gui, where you have many methods to combine in many ways and OOP really suits in.
Maybe Julia should be seen as an infix typed Lisp DSL for math. It'd make a good stack (see my other comment) with Clasp or any other Common Lisp that offers good bindings to LLVM.
[0] https://github.com/rigetticomputing/cmu-infix/blob/master/RE...
I am also working on a data frame like library, using Tamas's experimental cl-dataframes as a starting point.
I'd be interested in learning a bit more about your effort - whats the motivation, why something like clml was not appropriate and so on. There might be an opportunity to collaborate?
Cheers
Lisps tend to love the "you can build anything in it! You can make all of your own syntax in it!" and I think that goes too far. Julia lets you do this with macros, but then there's a good convention to not overuse this. There has to be a balance between being too structured (which hurts innovation) and too dynamic (which makes it hard to learn someone's new library because everything is too different), and I think Julia strikes the right balance.