JuliaCon2020: Julia Is Production Ready
bkamins.github.io
bkamins.github.io
The size of our Julia codebase is about 400,000 lines of code, spread over numerous internal and open source packages. You can cite the "we're hiring" slide from my Juliacon talk for that https://raw.githack.com/oxinabox/ChainRulesJuliaCon2020/main...
As far as I know, all ISO's in the USA use either ABB, Siemens, or GE for their market optimization software (huge systems) or basically those same vendors plus a few more for their EMS systems. There are also plenty of other amazing products for transmission planning, short term studies, and transient stability studies. More specifically, how are you linked to the system operators? When you say you optimize the grid, that could mean any one of a thousand different things. Are you providing various inputs (forecasts, risk factors, etc) that are used by some markets? Please note that I'm not questioning that y'all work with some of these entities in some fashion, I'm just trying to get a better handle on the specifics of where y'all are plugged in.
At the moment there's no Julia episodes yet. If you want to come on, head over to https://runninginproduction.com/ and click the "become a guest" button near the top right to get the ball rolling.
Writing servers in Julia is really pleasant thanks to the clean coroutine-based task model. Under the hood, Julia uses libuv to get efficient high-concurrency I/O with an event loop. But from the user's perspective, you just write simple blocking code and use `@sync` and `@async` to spawn concurrent tasks. If you want to use multiple threads, just use `@spawn` instead of `@async`. This is very similar to Go's highly successful I/O and concurrency model.
Several Julia Computing customers also run long running Julia jobs in production, and some are discussed in our Case Studies (https://juliacomputing.com/case-studies/).
One of the things that still keeps me at bay is the JIT and the pre-compilation in general. It seems like it's still not very easy to actually compile a Julia library or an executable. There is PackageCompiler.jl, but my understanding is that it necessitates pulling in sys.so, which is some 130 MB in size, making it prohibitive for many projects.
I'm curious to understand why the 130MB is prohibitive. Are you looking at embedded or resource constrained environments?
I was just trying to do a quick test app the other day to see if I could build a Hello World app, and (this may have been my fault) but I hit the threshold of cognitive overhead before deciding to revisit later. This isn’t a criticism so much as a recognition that compiling binaries may take a bit more care / attention / learning in Julia than something like Go, Rust, or even Racket.
As for binary sizes, for some serverless deployments smaller binaries can open up a little easier path.
I’ll try and take another run at PackageCompiler and give it a more fair shake. If in that process I find anything constructive to report or submit documentation PRs, I will do so.
Thank you for your work on Julia. Julia was on my radar a few years ago and then dropped off again. I’m excited to see where it has come in that time.
This pattern can be seen in several places and frameworks like Spark & company, and it's also pretty much the foundational idea behind processing data with Joyent's Manta[0]. It always pretty much boils down to "serialize a function, and send it somewhere, to run on the data locally".
Some of these frameworks do it only through native functions (Spark runs serialized JVM procedures, and for Python functions I believe it shells out to an actualy Python interpretes), others are completely agnostic to this and work with binaries (Manta).
It'd be pretty hard to do, and have a lot of overhead, if all of those "functions" that I'm shuffling around are, at the lower bound, 130MB in size.
I understand that a lot of this can be done directly at the Julia layer for Julia things (much like Spark does with JVM things), but it kinda limits the applications of this nonetheless.
https://docs.julialang.org/en/v1/manual/distributed-computin...
However, this is not related to the system image (sys.so) that is used to start the Julia process, which has a compiled version of base, stdlibs and whatever else you choose to build in.
This is a nice alternative to the AOT compilation offered by the likes of PackageCompiler.jl
Hopefully we'll have more / better AOT compilation and deployment in the near future. There's a lot of rumblings of exciting things starting in Julia AOT compilation, and I think it'd make a killer combination with our package manager's binary artifact deployment story.
In practice, it may often inline more aggressively as there aren't any boundaries that will prevent inlining.
If you write function `foo` calling `buz` from package `Buz.jl`, which calls `bar` from package `Bar.jl`, `bar` can be inlined into `foo`.
You can think of it as always have `-lto`, because the source of the Julia functions you're calling will always be available to the compiler and thus subject to (potentially aggressive) optimization.
I legit think that even if you are using say pytorch. You are better off writing your code in julia and using the python interop.
Though I much prefer Flux.jl
In other more complex cases beyond big fat matmuls and such, it's much faster, even comparing to torchscript (sciml, neural ODEs etc).
Still, the codegen is getting there with the funded work on the GPU compiler and memory allocation and there eventually won't be any sort of tradeoff.
In the meantime, Yes, there is a package at allows calling pytorch kernels without modifying your code. This is also only possible with such finesse due to Julia's power and composability. See here for the package: https://fluxml.ai/2020/06/29/acclerating-flux-torch.html
https://github.com/FluxML/Torch.jl https://fluxml.ai/2020/06/29/acclerating-flux-torch.html
Also the gap between native GPU codegen in Julia and the hand-optimized kernels in PyTorch is lesser than you'd imagine.
In fact, see this talk at JuliaCon, which compares Julia on a 1000 GPUs against a CUDA/C implementation, and finds them to be extremely close.
https://www.youtube.com/watch?v=vPsfZUqI4_0 https://github.com/eth-cscs/ImplicitGlobalGrid.jl
However, we are integrating a lot of the underlying infrastructure to the point that both Knet and Flux will only differ in API approach.
[1] https://gist.github.com/ChrisRackauckas/cc6ac746e2dfd285c28e... [2] https://gist.github.com/ChrisRackauckas/4a4d526c15cc4170ce37...
I think the ecosystem could be a bit better, but that's more of a personal opinion rather than a serious concern as, like you alluded to, you the interop story is insanely good.
TimeArray is like a DataFrame, but with a lot less features. You can't really do that much with them, and they only have limited support in an auxiliary package for resampling.
It would be nice if a new DataFrame library could be created that was inspired by Pandas.
In general, I think Julia suffers from the problem of too many small packages. For example, there's 3 libraries I have to use for time series: TimeSeries.jl, TimeSeries Resample.jl,and TimeSeriesIO.jl. None of the libraries are particularly large and should all be merged into one package.
Part of the reason for the dirth of small packages is because, unlike Python, Julia's strong support for function polymorphism allows for seamless extension of existing functionality with new packages. Unfortunately, I think this has created a cultural problem in which new packages are created that only slightly modify existing functionality.
It's a pretty good scripting language, and it's a "good enough" interface that implementers of C packages can expose powerful features, and that can only go so far. User's can't easily compose the powerful features that library designers didn't intent.
Julia, on the other hand, was designed for the explicit purpose of high-performance scientific computing, with a syntax informed by the tools people already knew (python, MATLAB). It's no wonder it's so much easier to use.
For scientific computing and advanced machine learning domains, Julia is certainly a breath of fresh air and a no-brainer decision. However, as a general-purpose language as well as technology stack, I think Julia has still a long way to go. This is mostly due to relative immaturity of that part of the ecosystem, spotty - and sometimes even non-existent! - documentation (of course, except for the language itself and quite a limited number of core and popular packages) as well as some other issues, including tooling, development/compilation performance and limited pool of skilled developers.
So, from a startup founder perspective [who has to select the optimal, risk-minimizing, platform stack], despite Julia's many attractive features (including powerful meta-programming facilities - in my case, for potential DSL development), I'm now leaning toward Python and .NET ecosystems. Both offer very mature and large package ecosystems along with comprehensive tooling support and incomparably larger pool of skilled developers. Additionally, .NET offers stability of consistent development improvements, backing of a major corporation [no acquisition risk] along with support for modern enterprise-focused architectural practices and patterns, including DDD and multi-tenancy.
Not trolling. I don't really know why people like to shoehorn a purpose-oriented language into all domains. Eg: Julia for scientific computing, Rust for low level programming. They would be great if they focused on those "strong zones".
And Julia creators do focus on it's strong zone, with all the works on TPU, HPC and parallel computing, and so does the community with the stuff presented in juliacon like interactive reproducible notebooks (Pluto.jl) or data dashboards (Stipple.jl), both using the web domain to improve the scientific computing domain.
1. Intro section (italics emphasis mine).
"... delivering complex enterprise projects:
Julia is fast, and has a very nice syntax, but its ecosystem is not mature enough for use in serious production projects.
For many years I would agree with it, but after JuliaCon 2020 I believe we can confidently announce that Julia is production ready!"
2. Section "Building microservices and applications". Even more so, practically all sections in the post, except just one ("Managing ML workflows"), describe various general-purpose and enterprise computing aspects.Therefore, I don't see how one could view the post and author's conclusion purely from scientific computing perspective.
And while right now I wouldn't recommend Julia for that web domain unless you also need it's number crunching features (which is more than just scientific computing, I work with large finance systems and there is definitely a lot of that stuff), I do think it's more of a library/community problem than a language problem. The language is well equipped to handle it (easy to use from a scripting perspective, easy parallelism, fast after warm-up, allows for clean abstractions to create sophisticated web frameworks, the mentioned easy FFI), what it lacks is the coverage and maturity of the tools and support (so I don't feel like I'm at a risk at any point that I need to do some integration, such as integrating with Kafka or any other systems, without having to write my own solution or integrating multiple languages at every step).
That's different from systems programming, embedded, game development and even GUI development (at least until smaller binaries with less warm-up) right now. Even with good library it will not be competitive with the languages that already claim those domains. Julia is a general purpose language already and more than a matlab substitute, it does not need to be good at everything (and it shouldn't try), but I don't think it should restrict the scope too much either.
People find the need to identify themselves as gopher or rustacean or something like that. Apparently Programmer is not hip anymore.
But then for my applicationsspace there are problems:
- compilation seems possible now, but one has to be careful of GPL (like in the fft)
- no way to turn off GC (so byebye real-time possibilities).
- what about GUI design? Heard some mixed messages about it
While Julia does not offer semantic guarantees that you can avoid the GC, it;s actually quite possible and easy to write code where you're manually managing all your memory and the GC is never invoked.
> - what about GUI design? Heard some mixed messages about it
There's a lot of promising things happening with Makie.jl, Stipple.jl, Dash.jl, Pluto.jl and others. This space is still a little immature in Julia, but it's progressing fast.
But I like your question, lets call it 'which languages are all purpose production ready'; probably only C, C++, Java (there is a real-time version), LabView.
I'd say that if people have other reasons to use julia in a realtime low-latency production system -- like needing it's scientific stack -- then you can certainly make it work, you just need to do some testing to make sure you're not hitting GC.
Docker has some nice infrastructure that we like, such as the ability to set memory limits, automatically restart processes on crash/if a healthcheck goes sour/on reboot, easily spin up different versions of the same service (based on which image you tag to a name), etc... Containerization has done a lot for our ability to bring servers up onto heterogenous hardware easily and quickly.
As an example, this repository [0] is deployed on 10 machines around the world, providing the pkg servers that users connect to in geographically-disparate regions. The README on that repository is a small rant that I wrote a while back that walks through some of the decisions made in that repository and why I feel they're good ones for deployment.
To answer your specific questions:
* How to deploy Julia: we use docker containers and watchtower [1] to automatically deploy new versions of Julia.
* Keep dependencies sane: it's all in the Pkg Manifests. We never do `Pkg.update()` on our deployments, we only ever `Pkg.instantiate()`.
* Pin memory usage: We use docker to apply cgroup limits to kill processes with runaway memory usage [2]. This is only triggered by bugs (normal program execution does not expect to ever hit this) but bugs happen, so it's good to not have to restart your VM because you can't start a new ssh session without your shell triggering the OOM killer. ;)
[0] https://github.com/JuliaPackaging/PkgServerS3Mirror [1] https://github.com/containrrr/watchtower [2] https://github.com/JuliaPackaging/PkgServerS3Mirror/blob/c6a...
Everything is available on their YouTube channel:
https://www.youtube.com/playlist?list=PLP8iPy9hna6Tl2UHTrm4j...
Now, 8 years later, I'd say essentially everything there has been achieved. The most important difference between Julia now and Julia then is the gigantic, vibrant and interacting ecosystem that's sprung up (though share are other important technical differences too!)
_______________________________________________
Here's the pitch I'd give:
Julia's value proposition is mostly geared towards people trying to do things that are just too difficult, awkward or slow in other languages. It's not just about speed, dynamism and friendly syntax. It's also about solving the expression problem and providing unprecedented composability.
So while it may not necessarily have great appeal to 'end users' (who are people just calling numpy/scipy/pandas/whatever functions the way they were meant to be called by the library writers), it does appeal to people who make the sort of things end users want so the community is currently swelling with very talented people making state of the art research libraries. The language is designed for 'composability', so that it's very easy to combine packages together in ways the authors didn't necessarily intend or forsee.
This approach is really starting to bear fruit. Our package ecosystem has some real, state of the art stuff not easily available in other languages, and it's often built in less time and with less resources. For 'end users', whether or not Julia makes sense for them really just depends on if the things they need are already available.
Another huge advantage is that even if you're not some super researcher developer, Julia substantially lowers the bar for mere mortals such as myself to do package development. The composability I mentioned above means that I can stand on the shoulders of giants as I implement my own stuff, plus the fact that the majority of libraries are written in pure julia means that I can easily "peek under the hood" to figure out how they work and learn from them. This is really special and can't be replicated in python, because 'peeking under the hood' invariably means reading very specialized C code, rather than Python code.
I think Julia will someday eat Python for scientific computing, but it may be a while before that becomes obvious to Python users. In the meantime, we have an excellent package for calling and running Python code from withing Julia, called PyCall.jl [1], so that can make the burden of switching even easier.
So here is an example and my worries ... My project uses GDAL, PIL, SentinelSat libraries to collect Sentinel 1 imagery and process it, pandas, numpy, scipy extensively. Would I find a bunch of those modules missing? I assume so ...
I guess then I might be better of to trying a new project in Julia. Something not too easy but not too hard.
Worst case scenario, you might need to do the GDAL, PIL and SentinelSat stuff with PyCall and then the Pandas, Numpy and Scipy stuff should all be quite well supported from the Julia side. I can ask around to see if there's anyone who can answer this better though.
Julia more than covers all the functionality of those packages.
>PIL
https://github.com/JuliaImages/Images.jl
> GDAL
https://github.com/yeesian/ArchGDAL.jl
>SentinelSat
I'd use PyCall.jl for that. It's very easy and does automatic conversion of types etc
> I think Julia will someday eat Python for scientific computing, but it may be a while before that becomes obvious to Python users.
I can’t say I understand this rather pervasive attitude. If you have well tested code in a language which you can easily interface with, why bother rewriting it except for reasons of performance or compositionality? Or perhaps you mean that we’ll all be writing Julia even if some algorithms still sit in Python?
On the other hand Julia probably has a very strong long game, given that the Federal reserve bank uses it.
The containerization I used was singularity. I don't really love them much anymore, for various non-technical reasons, but it got the job done.
I mostly mean that I think Julia will someday have more mindshare in scientific computing than Python.
However, I'd say that the reason you see the "rewrite it in pure julia" attitude everywhere is that we enjoy a lot of benefits from having more and more things in pure julia, namely that our very powerful metaprogramming and code introspection tools can more easily 'get inside' Julia code, and that the compiler is able to perform interproceedural optimizations (IPO). IPO are not always important, but when you need them and don't have them, you really miss them, so at least for numerical kernels and such, it's really nice to have as much pure julia infrastructure as possible.
In the meantime though, as I daydream about a world where everything is written in Julia, there's nothing stopping me from being pragmatic and wrapping / bridging to and from other languages.
Julia's syntax and language design is brilliant, but the biggest reason Julia has been my favorite language since I first learned to program is how astonishingly easy it is to use tools other people have made, and also to create tools for others to use.
The packaging system is just great for development, both as a user, and as a maintainer. I've been surprised by how easy it is to take stuff I've made and get it to work with a myriad of packages for the purposes I need. Multiple dispatch does wonders here too.
In the recent survey they said that 30% of respondents had contributed to a package. 30%! It's a super active ecosystem and the barrier of entry is low enough where I don't feel "imposter syndrome" as strongly as I have elsewhere--that I need to pass some threshold in order to properly contribute.
It's a hard problem to combat, because fundamentally, package code must be able to do arbitrary things. One advantage julia has though is that packages tend to be written in pure julia, which makes the duty of checking and detecting malicious code much easier.
> lowers the bar for mere mortals such as myself to do package development
That's a powerful combination. I've been close to contributing to Python packages a few times, but the overhead is pretty high. At this point, I also feel like it would be a waste of effort not to write stuff in Julia instead.
> Julia will someday eat Python for scientific computing
It's well on the way. Julia users have a reputation for being evangelists. Heck, I made a convert even before I actually started using Julia in earnest: my labmate had a problem he couldn't solve in his favorite language, and I encouraged him to try Julia -- it turned out great!
I also think writing out math is much more concise in Julia, no more np.whatever(np.sin(x.detach()) + cos(np.sqrt(y)).mean().
Yes, far exceeding python actually. - https://github.com/jump-dev/JuMP.jl - https://github.com/JuliaNLSolvers
Are just some.
> graphing
Excellent support.
- https://github.com/JuliaPlots - https://github.com/JuliaPlots/Makie.jl
>Munging xml/Json etc
- https://github.com/quinnj/JSON3.jl - https://github.com/JuliaData - https://github.com/Algocircle/Cascadia.jl - https://github.com/queryverse/Query.jl
>oauth2 etc
- https://github.com/JuliaCloud - https://github.com/JuliaWeb/HTTP.jl
for anything not available yet, you can call it from python using PyCall.jjl
The question I am currently researching is not what Julia is good at but what are the pitfalls and drawbacks of particular technical choices. I know them by heart now for Python, and my team relies on knowing what doesn’t work well. For Julia it’s an unknown unknown. Understandably all the blogs are about why Julia is great so it makes seems to go slowly and coexist instead of moving.
Mind elaborating?
Julia's equivalent to xarray: https://github.com/JuliaArrays/AxisArrays.jl
Original issue explaining the problem: https://github.com/pydata/xarray/issues/525
Latest issue detailing the work: https://github.com/pydata/xarray/issues/3594
Edit: This looks like a fantastic resource! Full link: https://juliaacademy.com
https://julialang.github.io/PackageCompiler.jl/dev/devdocs/b...
If you can ensure your code and dependencies does not have dynamic things that need compilation at runtime, it is fairly easy to remove the dependency on bundling LLVM. Speed is the same, since you are then just running cached native code.
Edit: it’s a bit more into the weeds than I thought. Honestly the docs give a better picture of DataFrames.jl
Row indices are super powerful and things are annoying without them. I've actually been thinking about programming a new and improved data frame library inspired by Pandas (but not so ugly and complicated), but the amount of work to write and then maintain it would be prohibitive.
Please somebody enlighten me