Julia 1.4
github.com
github.com
The bar was low enough for Julia to sound good at the start, but it gives the impression of being "designed" "on-the-fly".
IDE friendly? IDE support is only good because you absolutely need an IDE for this language.
It's an unwieldy, verbose mess of a language.
Julia works really well for power users. There are no huge libraries full of C code like pandas or scipy. Instead there are dozens of small, well-tested packages that fill the same role, all hosted on github. That makes fixing issues so much easier.
Granted: outside of numerical/technical computing, the libraries can be lacking (eg. web development) compared to other languages. We're doing it anyway, but it's a more difficult decision.
Could you elaborate? What was the problem and what was the solution?
Memory leak. As explained, it's not really Julia's fault.
This I also found amazing: https://github.com/JuliaLang/julia/issues/28726 . It took 4 hours to go from bug report about bad compiler code generation to a patch.
I have commented with some details there:
https://github.com/JuliaLang/julia/issues/30653#issuecomment...
Isn't Julia's LinearAlgebra a wrapper around BLAS/LAPACK implementations (OpenBLAS/MKL etc.)?
This is not my experience, at least for numeric code. It generates faster code than Golang, because it uses an LLVM backend (and actually supports macros and parametric polymorphism, so doesn't need to do the work at runtime), faster numeric code than Rust via @inbounds and @simd annotations (way more work to disable bounds checks in Rust), and faster than Swift because it doesn't have pervasive reference counting that can sneak in and destroy performance.
How many go or swift or Nim issues discuss matmul vs the Julia repo?
Swift have built in simd same about Rust https://github.com/apple/swift-evolution/blob/master/proposa...
Nim have arraymancer https://github.com/mratsim/Arraymancer
@inbounds it is just compiler option u can use it in any language with LLVM backend i would be suprised if Julia will be faster then Rust/Swift or Nim in this regard. But True about GO in that particular case.
I agree it wouldn't necessarily be faster, but it also wouldn't be slower. Plus in Rust at least disabling bounds checks requires marking code as unsafe, which really gets the community's hackles up.
Here are my early experiments at making a pure-julia multi-threaded BLAS using LoopVectorization.jl https://github.com/MasonProtter/Gaius.jl. It absolutely blows a naive triple for loop out of the water and is quite competitive against OpenBLAS until you get to very big sizes.
MArrays will be stack allocated if they don't escape.
One of my in development packages also uses it's own "stack" (MMap a chunk of memory), so that it can have pointers to fast "stack-alocated" arrays.
I played around with LLVM's alloca a bit, but it seems like I could only ever use a single alloca at a time; if I ever used more than one, LLVM would just return the same pointer each time instead of incrementing it. If I have to manage incrementing the pointers myself anyway, I may as well use my own stack, too.
For (the problems I have tested and tuned it on), LoopVectorization produces faster code than C/Fortran, e.g.: https://chriselrod.github.io/LoopVectorization.jl/latest/exa... But it may be more fair to compare it with plutocc. In my early tests (which involved much larger problem sizes), plutocc does a lot better, because (unlike LoopVectorization) it seems to consider memory/caches rather than just registers allocation and instruction costs.
However, in response to discussion about Julia having an integrated set of abstractions that provide high performance - you are discussing a collection of features in Rust, Go, Swift, Nim, TensorFlow, Jax, PyTorch, etc. For any single feature, there can always be some other thing that does it better. But it is unclear how that makes one system better than another.
I wish I could tell the compiler to compile a function only if it can do it without using memory allocations. Often I would rather see a compilation error than the GC making performance of a function unpredictable.
Of course borrow checker would be helpful, but I guess that's too hard to implement, that's why something like this would be a good balance.
I also think Julia would benefit a lot from a list of "we want these but haven't had time for them". People like contributing things they know are desired, and with Julia it's hard to know whether people want or will accept something until you make the PR.
My employer, Invenia, uses Julia for what I would say is at least a medium sized project.
~200k LOC, ~30 people concurrently working on that code base, literally the core product/system that makes us money.
Julia's not perfect, but I would not blanket discount it out of hand for medium projects. Like all technologies the tradeoffs need to be investigated with reference to the task (and team) at hand.
It's a beautifully designed language with incredibly responsive and wise developers and 1.4.0 is a great release I've been on 1.4 release candidates for over a month and haven't had a single issue, and love the new features and improvements.
If I were in the position of making or advising someone making the decision of whether or not to use Julia, an unsubstantiated, untestable assertion like that makes virtually no difference in my decision or advice. I hope it's clear this is not an insult, it's just that if I have no way to evaluate your assertion or its applicability, there's no way for me to incorporate it into my decision or advice.
By contrast, if you could point to even one concrete example, I could file it away in my head as "someone on HN pointed out that Julia has X, Y, and Z problem". If I were in a position of making or advising a decision on Julia, I would be able to evaluate whether it applies the use case in question, whether it points to some deeper design problems with Julia, or whether I even think it's a problem at all.
Until then - I concede, please take it with a grain of salt and use your own judgement.
Rather than rewriting code in Python, we took an approach of using PyCall. It turns out that the overhead of calling Python is very small... at least not to an extent that I had to worry about. For my production system, we used this strategy to access Oracle and Apache Kafka. The code still looks very clean as the calls to the Python packages are just like normal Julia function calls.
1. Get rid of global dependencies. Store all of your dependencies in a project local deps directory with a project lock file.
2. Opinionated file system structure. Maybe you don't need this in one-off scripts but definitely for some sort of "project" layout
3. Force all packages in the package manager to obey these constraint.
The ship may have already sailed on 3), sadly.
I agree that an opinionated file system structure is nice. Julia already wants you to put source code inside `src` but then underneath you are free to do whatever you want. We generally manage this via coding convention since people can never agree on a standard.
1. Every package must declare it's dependencies (and to register must declare compat bounds on them, and good devs do that always anyway) 2. Every package must have Project.toml in the base directory, source code goesin `src`, test code goes in `test`, documentation goes in. `docs` (and more structure there if using Documenter.jl) What more could one want? If a package needs more folder structure within `src` then it should be multiple packages. 3. These are the rules, done.
Further you are describing every sensibly done julia application.
Basically no seriously Julia developer uses the global environment for anything but dev-tools, like BenchmarkTools or ProfileView. Certainly one does not depend on the content of them for any reused code -- that is what Project.toml and Manifest.toml is for.
Can you point me to some documentation on those best practices. My girlfriend - architectural acoustics consultant - is working on a project in Julia and is having a hell of a time managing dependency shift underneath her program (also she barely knows how to use git).
The Pkg manual doesn't have a tutorial on standard practice. But it's worth reading the compat section and making sure to always set your compats https://julialang.github.io/Pkg.jl/v1/compatibility/
And the section on creating one's own project. So as not to need to use the global environment. And alternative to using `activate` after starting Julia (and my preferred way) is to start Julia with the `--project=.`
https://julialang.github.io/Pkg.jl/v1/environments/#Creating...
And can go further and create projects via PkgTemplates which is was "everyone" does. Because a good project looks just like a package (one might as well consider them synonyms when looking for thus kind of advice) https://github.com/invenia/PkgTemplates.jl/tree/v0.6.3
To be honest, you can incur technical debt with any language. Software must be designed and maintained properly if you need a long-lasting solution. Research projects are also different than production code. Proper training and involvement with the developer community could help a lot.
My suggestion to everyone reading this thread - if you are new to Julia programming and need to work on a production project, do talk to Julia Computing folks. Their consultants can lead you to the right track. In addition, join the Julia Slack and Discourse community. You can almost get instant answers to anything you ask there.
Lastly, here's a selfish plug to my book: Hands-on Design Patterns and Best Practices with Julia https://www.amazon.com/Hands-Design-Patterns-Julia-comprehen...
BTW, this kind of discussion is healthy. The Julia core developers are already taking notes and I'm certain they will continue improving the language or ecosystem.
In my experience, one of the biggest problems is that people approach julia as "python or matlab but faster", which is going to be a recipe for failure. Julia is a very different language with its own idioms and very different style and if people try to just write Python with Julia syntax I can see how problems arise.
To be clear, I do definitely believe there are some big pain points in julia at scale, but depending on the usage domain and needs, my impression is that julia is an appropriate tool for at least some niches.
If my job was just that then I would certainly be heavily burned by it and I wouldn't want to touch the language. But the better solution would be to instead learn the language properly and create a serious guideline that everyone has to follow. If people can't use macros properly (or any of Julia's most powerful features), then every PR with a macro requires a very convincing justification and a review from a more experienced developer. If you're adding a new functionality then you have to be sure there isn't cyclic dependencies with other functionalities, and possibly making it a small independent library instead. And more importantly code is not done until it's properly documented, reviewed and unit tested.
I'm working with a computer science undergrad at the moment who put multiple full versions of his code to into the Github repository. He didn't use branches, he didn't use pull requests for updates, the diff log is just a new file that is the previous one with changes, all of which sort of defeats the purpose of using source control like git. But, as you say, it's something that takes some guidance and experience to understand. I do find that it makes me lean towards tools that encourage good coding practices by design and which make untying knots easier, which it sounds like Julia might not make super easy (I don't have any production experience with it). Obviously, the practices you laid out definitely help to make sure things don't slip through the cracks, though!
And until then the best practices for larger project using Julia's paradigm (Multiple Dispatch, which is by itself not as well understood as OOP or functional) will probably become more largely known. I feel like Julia projects should not grow large not because it's not suitable for large scale projects, but because even large scale Julia projects should emerge from the composition of many small and maintainable Julia projects.
The NEWS file may not be designed for that, but IMO, it still is way better than this text, which, for casual followers, doesn’t say much more than “1.4 replaces 1.3, no breaking changes, a few new features”, without even mentioning those features, or why they were added.
Maybe, combining the two in a single document, and copy-pasting NEWS.md in the announcement, would decrease the amount of work and improve things?
Looks like Julia still haven't managed to handle updates well.
I remember being so amazed that they'd blown their 1.0 announcement (went from 0.7 to 1.0 over a Juliacon).
And because there were so many changes (most of which i thought were good), it was pretty broken for new users at that time, which I firmly believe limited their adoption.
And to be clear, I love the idea of Julia, and that first document made me fall in love. And I think that the design for statistical computing is really, really good.
It's just a shame that this ops/packaging stuff is holding them back.
Otherwise it really is a pleasure to use, and I have found that the LTS install has zero compatibility issues thus far.
I hope you did file an issue with the Plots package. The reason is that we generally like to make sure that regressions like these are not language level regressions. Sometimes package use undocumented internals - but we generally chase each and every single one of these.
For those who haven't seen it yet, we have described our release process in great detail in this blog post: https://julialang.org/blog/2019/08/release-process/
IMO, 1.0 was rough not for new users but people who had invested a bunch of time in Julia codes pre 1.0. We anticipated this and therefore had a very carefully planned release strategy to ease the transition with depreciation warnings and preparing 0.7 as a migration aiding release for 0.6 users to 1.0. In fact our release announcement discussed all of this at great length.
That sounds as if the communication was of the “1.0 will be released at this date one year in the future” kind. But it was more like saying for a year “it will be finished and released someday” and then, according to the wikipedia, “the release candidate for Julia 1.0 was released on 7 August 2018, and the final version a day later”.
Really?
v0.7.0-rc1 - Jul 31, 2018
v0.7.0-rc2 - Aug 2, 2018
v0.7.0-rc3 - Aug 7, 2018
v1.0.0-rc1 - Aug 7, 2018
v0.7.0 - Aug 8, 2018
v1.0.0 - Aug 9, 2018
When I tried to install packages, I got errors after error as a result of deprecation warnings from 0.7 becoming errors at 1.0.
Again, I really like Julia, and want it to succeed. But the 1.0 situation put me massively off, and killed my plans to start evangelising Julia at my company.
It's just a shame, that's all.
That said, the tooling is still frustrating. Generating compiled binaries is a slow, painful process.
Swift can get you 90% of the way there, and that extra 10% can be more than made up the by the efforts of apple, google and other companies (including money, network/clout, kaggle which is owned by google etc). Despite predictions to the contrary, Chris Lattner's departure doesn't seem to have slowed down the project, and more team members from google have been added since.
Swift is rapidly approaching usability on windows with investment from google.
Further, at some point google will facilitate Swift's use for android apps, and then Swift's popularity will skyrocket, and all those developers will be naturally inclined to check out the ML stuff. Even facebook is getting in on the party: https://twitter.com/nadavrot/status/1241150682104606720
In addition, Swift has its own benefits over julia for production and large codebases, such as compilation to small binaries and static typing. Julia doesn't have a good story for either of these (yet?), and chasing down type instabilities in larger code isn't fun.
Julia will be just fine even if s4tf somehow steals all the mindshare (especially since a mature differentiable programming library will inevitably serve as inspiration for Flux itself) as the language and target audience is not very similar to Swift's and as such Flux and s4tf will also find different niches (for example one can be more used on high performance scientific research thanks to Julia's ecosystem and focus while the other can focus on mobile deployment of ML models).
I don't see the ML ecosystem developing in isolation because there's going to be overlap, especially as more and more code can be differentiated.
And in this marathon Julia is a very late runner, and Swift didn't even start properly running in this direction.
The ML stuff is just a start.
Machine learning programmers aren't about remake DifferentialEquations.jl or scipy in Swift. I've yet to meet a single scientist from a field outside of machine learning who was seriously excited for swift. This sort of machinery is hard to make and takes deep expertise, I really doubt it'll be made in Swift any time soon. Does swift even have plotting libraries yet?
Swift has a good automatic differentiation story, mostly because it is very focused on machine learning use-cases, has corprate backing and all efforts are on one implementation. However, having only one automatic differentiation implementation has drawbacks. It won't be suitable for everyone.
Julia on the other hand has a gigantic basket of different automatic differentiation tools all of which have strengths and weaknesses. This allows people to choose the right tool for the job and explore a very wide design space, allowing us to find which approaches work best for different circumstances. Our AD machinery is still evolving and definitely has problems, but progress has been fast and really encouraging.
Even if Swift becomes the next Python and eats scientific computing, I strongly doubt this will seriously hamper Julia's community. We've been doing great living in Python's shadow. Julia doesn't need to be the most popular language in the world to be useful or successful.
Google and apple. Apple already is working on a swift-numerics package.
Look at TF python and jax. They've re-implemented chunks of scipy and numpy twice, hired people to work on plotting (altair) etc
And that's with python. Their engineering time will go much further with swift, obviously.
Numpy is not the same thing as scipy.
DifferentialEquations.jl in julia is a great example of what it actually takes to make a real, competitive differentiation equation library. The sort of stuff that was built there requires a deep connection to the scientific and mathematics literature. Cash won't cut it.
Another great example that'll resonate with physicists at least is things like ITensors.jl https://github.com/ITensor/ITensors.jl. Apple and Google are not going to make something like that.
In particular, do you have another example aside from DifferentialEquations.jl ?
Neural ODEs are hot enough that something like that could easily pop up in swift.
ITensors is interesting, but that's only one.
Maybe, but I'm doubtful. Scientific domain experts flock to languages like Python, Julia, Matlab, R, etc. because they're interactive and allow them to quickly iterate on ideas, query data, produce plots, etc. Swift is not much of an interactive language and is not built around that kind of repl driven experience.
> In particular, do you have another example aside from DifferentialEquations.jl ?
Sure, here's a smattering of high quality packages made by and for research scientists:
https://github.com/JuliaApproximation/ApproxFun.jl
https://github.com/BioJulia
https://github.com/JuliaDiffEq/ModelingToolkit.jl
https://github.com/crstnbr/MonteCarlo.jl
https://github.com/chriselrod/LoopVectorization.jl
https://github.com/JuliaNLSolvers/Optim.jl
https://github.com/PainterQubits/Unitful.jl
https://github.com/mcabbott/TensorCast.jl
https://github.com/JuliaPhysics/Measurements.jl
https://github.com/Jutho/TensorOperations.jl
There are many many more, these are just the first that came to mind.Unless they show to be as committed Microsoft is with .NET Core, Swift will have the same support as Objective-C has had since NeXT days.
As for Google, I still don't believe that they are that committed as well, Chris Lattner left, and SF4T appears more on Reddit and HN comments than on any kind of Tensorflow official communication channels.
If it wasn't for the valiant effort of one single developer not employed from either of them, Swift for Windows wouldn't even exist.
This is how committed both companies are.
"As for Google, I still don't believe that they are that committed as well, "
They have any approximately 10+ person team, many of whom are new hires from apple etc. Even for Google that's a decent chunk of change.
On my book being committed, is paying for the race at the start.
When TF4S delivers even the half of what ML.NET or Julia allows for on Windows including IDE, libraries and tooling support, today, and they actually start doing a bit more information then what gets given to Python, C++ and JavaScript then they are actually committed.
Swift is mostly useless on Windows, with Windows users being told to use Google's cloud infrastructure for doing anything serious.
Even Apple apparently is hiring Rust developers for server side development on Linux, instead of improving Swift's history on Linux.
Keep hoping for Swift on Android, it will never happen, if that depends on Android team, Kotlin/Native might have a place on the NDK, Swift never will get one.
It was truncated to one day when they made it virtual. There were swift announcements slated for day 2, which was cancelled entirely.
Regarding windows: https://forums.swift.org/t/new-swift-installer-for-windows/3...
There are also frequent commits to swift master and projects like swift NIO working on windows compat, from google. Also the S4tf stated they had future plans for swift on android.
So it has an installer now.
Pity that it lacks the eco-system of libraries available to .NET, Java, Julia, and respective IDE support, on Windows.
So now besides import glibc on Linux, with platform ifdef, should we do import mvsc as well?
ST4F team having future plans for Android is meaningless, regarding what Android team actually puts on the SDK, NDK and Android Studio.
When I see Swift listed on https://developer.android.com/guide/platform, I will consider that Google actually wants it on Android.
Otherwise it is just like Flutter and Go, third party tooling, with various headaches and Android integration issues.
This way I can just immediately query and test things using the language itself without having to deal with any sort of constraints or quirks within the editors. And the Julia REPL allows for quickly switching to shell (just pressing ;) when needed. I would hope they keep improving the REPL (and the editor support features like linters) over focusing on a completely new IDE (there is even more potential when considering Lisp languages REPL, plus general usability improvements like first call latency), but I'm biased and that's not the workflow most developers use.
What's wrong with the IntelliJ plugin for Julia? https://github.com/JuliaEditorSupport/julia-intellij#readme
I don't know exactly what 'OP' means, but there are other ways to do ML in for example C++. I have some good experience with dlib, PyTorch's libtorch is on my todo list.
I think Flux really excels when you're trying to do ML as "differentiable programming" (e.g. model-based reinforcement learning), because it can differentiate so much control flow via Zygote, which hooks into the compiler and differentiates the AST: https://github.com/FluxML/Zygote.jl
As for features, I think this is because people coming from TF or PyTorch are used one monolithic package that does everything. That’s intentionally not how Flux or the Julia ecosystem is designed. I’ll admit that there are a lot of preprocessing utility functions that could be better in the larger Julia ML community. But for the most part, the preprocessing required for ML research is available. This is mostly the fault of the community of not having a single document explaining to new users how all the packages work together.
Where the difference between Flux and other ML frameworks is apparent is when you try to do anything other than a vanilla deep learning model. Flux is extensible in a way that the other frameworks are just not. A simple example is with the same lab mate and I trying to recreate a baseline from a paper that involved drawing from a distribution at inference time based on a layer’s output then applying a function to that layer based on the samples drawn. I literally implemented the pseudo code from the paper because in Flux everything is just a function, and chains of models can be looped in a for loop like an array. Dumb pseudo code like statements where you just write for loops are just as fast in Julia. And it was! Meanwhile my friends code came to a grinding halt. He had to resort to numerical approximations for drawing from the distribution because he was forced to only use samplers that “worked well” in PyTorch. This is the disadvantage of a monolithic ML library. I didn’t use “Flux distributions,” I just used the standard distributions package in Julia.
This disadvantage to TF and PyTorch will become even more apparent when you do model-based RL. Flux was designed to be simple and extensible from the start. TF and PyTorch were not.
But for production ready models PyTorch and TF is miles ahead first of all: NLP, audio and vision based packages building frameworks, (attention layers, vocoders etc.) then u have option to compile models using XLA and use TPU (about 2/3 times cheaper then gpu for most of our models [audio and nlp])
Next inference performance (dunno about now maybe this change but about ~8 months ago flux was about 15-20% time slower [tested on VGG and Resnet's]then pytorch 1.0 without XLA)
Time to make it to production: Sure maybe writing model from scratch can take a bit longer on PyTorch then Flux (if u not using build in torch layers) but getting in into production is a lot faster, first of all u can compile model (something not possible in Flux) and u can just use it anywhere from Azure and AWS to GCP and Alibaba Cloud make a rest api using Flask/Fast-api etc. or just using ONNX.
Dont get me wrong i love Julia and Flux but there is still a LONG way before most people can even consider using Flux on production enviroment not for reasearch or some MVP stuff.
It's worth knowing it is more like R's DataFrames than like Pandas.
It is getting pretty close to a 1.0 release. Probably a few months out (One more minor, then if all goes well 1.0 a month or so later)
Further one should know that there are many tabular data packages in Julia and they all use the interface defined by Tables.jl, and all interop very well.
Query.jl (which is something like Linq or TidyR) works with all of them, and do packages for loadingand saving (CSV.jl, LibPQ.jl etc)