My favorite things that are coming with Julia 1.0
white.ucc.asn.au
white.ucc.asn.au
EDIT: grammar
We all know that having a null value in your language is a mistake, but i have to admit that it had not occurred to me that the fix was to have two null values.
So Julia is typed ( I think) but it's not a mistake in any way to have null in a dynamic language by this argument.
Null is a thing and if you don't have it you reinvent it, it's the typesystem breaking that's the problem, maybe monads effectively have nulls it's just that they then get protected by the typesystem.
In future we'll likely have shortcuts for `T?` as `Union{T,Nothing}`; effectively a dynamic version of optional/maybe types, as in Swift and others.
Javascript is even better as it has 3 (null, undefined and NaN, and while NaN is available in most every language it is pretty common in JS as its conversion APIs tend to return NaN instead of error).
From Swift for Tensorflow doc:
Julia is another great language with an open and active community. They are currently investing in machine learning techniques, and even have good interoperability with Python APIs. The Julia community shares many common values as with our project, which they published in a very like-minded blog post after our project was well underway. We are not experts in Julia, but since its compilation approach is based on type specialization, it may have enough of a representation and infrastructure to host the Graph Program Extraction techniques we rely on.
I think the sort of pacakges showcased here is a good indication of the sorts of niches julia is worming its way into.
[0] https://discourse.julialang.org/t/what-package-s-are-state-o...
Edit: It's just amazing to have genuinely lightweight abstractions in your language. I could make Python code fast, but it precluded using much of the language. Switching to Julia has vastly simplified some of our code. If not for Julia we might have ended up going for C++ (or Python+Rust) but I can take a scientist with little knowledge of programming and get them productive in Julia in a fraction of the time it would take for C++ (or Rust).
We've been limiting our research to make it fit the programming paradigm.
Even if Julia wraps a library like Tensorflow, its API is looking really nice compared to Python [1]:
using TensorFlow
sess = TensorFlow.Session()
x = TensorFlow.constant(Float64[1,2])
y = TensorFlow.Variable(Float64[3,4])
z = TensorFlow.placeholder(Float64)
w = exp(x + z + -y)
run(sess, TensorFlow.global_variables_initializer())
res = run(sess, w, Dict(z=>Float64[1,2]))
Base.Test.@test res[1] ≈ exp(-1)
[1] https://github.com/malmaud/TensorFlow.jlIt looks like there are the beginnings of both of these in Julia:
Part of this is why we did JuliaDB: http://juliadb.org/ and continue to try push the boundaries on parallelism, missing data, OnlineStats.jl and making data manipulation and modeling that much easier.
I've joked before, about how there is no such thing as a "One Language To Replace Them All", however, I feel Julia is the best candidate for the "One Language To Rule Them All", since while it solves the "two-language" problem in many cases, you can use it bind code written in many languages together (hopefully in a bit nicer fashion than the "One Ring" bound the other rings and their users!)
I could see something like Julia earning a lot of mindshare if it had a really polished solution for the space between, “my data is hundreds of megabytes”, and, “my data is hundreds of gigabytes”.
Moving to a new language has more friction than basically anything else unless there's a real language feature missing or the budget doesn't allow for more compute hardware. Hundreds of gigabytes is well below where academic and industry labs will start having to think about these problems. It's going to be really tough to displace Python with anything equally as general purpose.
This is all to say that I buy that Julia can shine more than Python for I/O bound HPC, but it really shouldn't be I/O bound until you have terabytes of data (and likely tens of terabytes). And aside from that, the Python numerical computing ecosystem includes a lot more than just Numpy and Pandas. As other commenters have mentioned, you can use Dask if your hot data has grown into the terabyte range. Anaconda includes a lot of libraries which can bail you out of situations once you've left the familiar world of Pandas data frames.
https://en.m.wikipedia.org/wiki/Gustafson%27s_law
For instance.
Even highly tuned c++ code spends most of it's time on the CPU waiting on cache misses. Its a pretty exceptional usecase where your code is not IO bound
I think it's useful to differentiate between an operation which is purely CPU bound (i.e. it's just constantly calculating without a reference) and an operation which is cache bound (faster than memory but still bound over the CPU). But calling operations I/O bound when they're sufficiently optimized that they live in the cache and don't even hit memory, let alone disk is an abuse of terminology. In the context of what I'm talking about, most HPC is absolutely not I/O bound unless it's using SATA/SAS drives instead of cache and memory.
And circling back to my original point, most research labs which can afford it will sufficiently optimize their code and hardware so that they don't hit the disk unless they're working with terabytes of data. Python, C++ and R provide numerous packages between the three of them for numerical computing across each of these bottlenecks, so I don't think Julia can rely on differentiating itself by shining in an I/O bound setting (i.e., waiting on disk). And if it does, "hundreds of gigabytes" isn't really the data size in which people are (in my opinion) going to overcome the friction of a new language and ecosystem just to harvest those benefits.
Operations waiting on the disk are "disk bound", IO bound is the generic term for data access, instead of cpu processing, being the bottleneck (which is usually the case, just a question of where).
I agree that Julia probably doesn't have a huge advantage, that said, having been stuck with slow Python code before, you're often stuck with rewriting large parts of the system in c++ or another low level language. That's the reality of Python, but at least it's not the most painful thing to do.
In a compute-bound workload the CPU spends the bulk of its time actually retiring instructions, not stalled waiting on data.
Think about it from the perspective of what the FPU sees -- once it has done that FMA operation, does it have the data it needs to do the next one, or does it need to sit on its hands for a while?
The cache hierarchy, cache-friendly data structures and algorithms -- they all aim to reduce time spent waiting on IO.
Theoretically speaking we can model any process as one which has to wait and one which doesn't have to wait. But in modern usage we have a variety of types and speeds for reads and writes. When I/O simply means reads and writes, you lose all the practical granularity you'd otherwise get by decomposing the reads/writes into different bottlenecks. It's philosophically elegant, but practically unhelpful for optimizing HPC and distributed systems when, as the responder said, it's rare to be CPU bound.
I also think that the context of my original comment is pretty clearly using I/O in the modern sense of disk usage. Responding with a correction that everything is I/O bound is vacuous, not insightful.
Not when solving (partial) differential equations, which is what I am using Julia for.
Used to be you could count cycles, now that's only really true in the simplest of cases with trivial memory access patterns. Now high-information branches are much more expensive than cycle counting would have us believe, ditto pointer-chasing.
Then because putting all those dots can be unwieldy, there's the `@.` macro which puts a dot on all your function calls.
There's also Dask [1], a native Python framework for distributed computations (by Anaconda). Irina Truong gave an excellent talk at PyCon 2018 about it [2]. I had never thought to look into Dask because Spark worked well for my use cases, but it has a lot of advantages over Spark (e.g. speed -- it's faster and more lightweight than PySpark and has no JVM serialization overhead) if you're using Python. Dask also runs on Kubernetes clusters, so scaling is not an issue.
And yeah, a huge amount of important data analysis work will continue to be done on data that fits in memory. Data analysis on distributed datasets is important, but from what I can tell, outside of certain domains it's certainly not the majority of the data analysis work out there.
The example I usually use is allowing integers to overflow, instead of automatically promoting to arbitrary precision (Python), or converting to a sentinel value (R). Integers are used in a _lot_ of places, so inserting these checks (or worse, access to heap-allocated memory) makes it difficult to optimise. (throwing an error might be a reasonable alternative in some cases).
Another is that you make it easier for the compiler to figure out things about an object, such as its size (e.g. you can declare the types of the fields of a Julia struct) and whether or not it can be mutated (immutable objects are easier to optimise).
IIRC Julia used to automatically promote integers, is this the main reason why this was dropped?
If an operation involves two different integer types, they do promote to the larger one (i.e. an Int64 + a BigInt will give a BigInt).
And my (very crude) understanding is that the stronger type system makes this much easier than in Python. The compiled version of any function is specific to the types of its inputs, and thus need not contain any further checks: simple functions often end up with literally the same assembly as C would produce.
https://www.youtube.com/watch?v=qCGofLIzX6g
https://www.youtube.com/watch?v=HStF1RJOyxI
Take everything mentioned in these videos that make Python and R really hard to optimize and don't do those things :D
https://julialang.org/blog/2012/02/why-we-created-julia
https://discourse.julialang.org/t/julia-motivation-why-weren...
I don't expect anyone who has not spent a lot of time with Matlab to "get" it.
https://cheatsheets.quantecon.org/
Python works, but it's far from elegant in this domain.
[A, B]
concatenates two matrices/vectors, whereas in Python it "wraps" them in an `n+1`-dimensional "matrix".That said, you get most of what you want with libraries. In NumPy you won't write
[aRow + bRow for aRow, bRow in zip(aMat, bMat)]
because you'll just call `numpy.concatenate`.Also, Python is mostly not used for mathematical work, and programmers tend to assume matrices are scary or only useful for mathematical work, so it has a "boring" syntax more suited for operating on single items at a time, with lots of loops.
I think that is a necessary trade-off for Python as a "general-purpose" programming language. I had used MATLAB and IDL quite intensively before I moved on to Python and R. When writing MATLAB, I felt like a scientist and did not have to bother with programming practices, like coding style, unit tests, writing functions instead of scripts, etc. But Python forces me to think like both a scientist and a programmer. (For example, every time you write `np.array([1, 2, 3])` instead of `[1, 2, 3]` it reminds you that array operation is not a free lunch offered by the language; it comes from the NumPy library. Also, it keeps the namespace pure.) I personally like this way better. But I also agree that not everyone likes it. (In my institution, researchers are kinda split half-and-half between Python and MATLAB.)
Has the language improved dramatically in the last few decade?
https://cheatsheets.quantecon.org/
You can go as far as to say that Python is verbose in many cases, and non-intuitive in others (A @ B?).
That said, you would never want to write a webapp in MATLAB, so as people expand from "math scripting" to "programming" they run into MATLAB issues which have absolutely no good solution. This is where the Python pickup comes from: it's still decent for scientific computing, but it also is an actual programming language. However, Julia keeps the nice syntax of MATLAB in the mathematical domain, keeps the technical computing focus of its community, adds some speed, and also is a general-purpose languages where webservers etc. are being written. In that sense, Julia is a really good fit for people looking to ditch MATLAB.
(A lot of MATLAB's bad syntax was bolted on later. It started as MATrix LAB, and later became a programming language. You can easily see the elegance of its initial design, and the terrible choices when extending it.)
I don't think it is likely for a software engineer to understand Matlab's niche and effectiveness. It comes from a different direction.
A ball point pen, a paint brush and a piece of drafting graphite all have their uses.
It also has the drawback that most of its users generate write-only code, so everyone that learns it also learns to write code that way.
Also, if I'm just restricted to using Pandas on a laptop or small server instance, then loading in a several gigabyte csv file can really tax the memory.
What killed me with Go was lack of generics (a lot of copied code to implement exactly the same integration approach for real and complex functions; different accuracies, etc.); clumsy way to express formulas (no operator overloading to provide a clean syntax for matrix operations) was a pain as well.
I have seen many times that the use cases of scientific computing would benefit from the functional programming paradigms, efficiency with basic math operations and ability to define operations that match standard scientific notation (matrix algebra, convolutions, etc.), but there are few languages that do well in all of those.
https://juliacomputing.com/case-studies/celeste.html
It's tough to crack the old school high-performance computing world, which is very slow to change programming tools, but Julia is the first and only high-level dynamic programming language to do it. Now that it's been shown to be possible to exceed a petaflop/second in Julia without having to write any gnarly, low-level C++ code, I suspect we'll start to see it happening more in the future.
I am hoping for a nice web assembly[1] emitted from Julia
[0] https://docs.julialang.org/en/stable/manual/calling-c-and-fo...
> codes
I think there is a large overlap between users of that language and users of that word.
Except that this is not always true. R, Scala, and spark are heavily used in bioinformatics. Physicists and quantum chemists nowadays use C++ more often than Fortran when they build new models and libraries. Astronomers use Python. The only scientific fields that get stuck forever with Fortran seem to be climate modeling and weather forecasting (and they continue to use Fortran 90 and pretend the most recent Fortran 2008 standard does not exist).
Edit: I'm not saying Fortran isn't good. It's the best for bit twiddling. Offloading the performance critical part from Python or other high-level languages to Fortran is great. The problem is, in an academic setting, a lot of scientists untrained in programming choose to write programs from scratch in Fortran, which is too taxing on the debugging, testing, and maintenance time.
For anything else that requires long-term maintainability and has a business or someone's job depending on it, Julia is a risky choice to make. Few, if any, large organizations have done so yet, and usually don't want to go first.
Well after 1.0, if Julia Computing (who employ almost all of the core developers) is on sustainable financial footing and any large organizations have made investments into adopting and supporting Julia, it might make some sense for non-academic use cases.
It like the getfield overloading and a few other things really deserve a post of their own, so I wasn't going to try and push them into this one.
https://medium.com/@sdanisch/compiling-julia-binaries-ddd6d4... -> notice that's a third party package.
I think it's just a problem of manpower on the Julia developer side.
However, unless I'm missing something, that argument isn't that much stronger.
1. Is the package actually promoted and actively maintained by the Julia developers?
Or even better:
2. Is the package part of the default Julia installation?
And even:
3. Is the package documented on the Julia site, as part of an officially recommended process?
Yes.
>2. Is the package part of the default Julia installation?
No, mostly because it's not done yet and only experimental right now.
>3. Is the package documented on the Julia site, as part of an officially recommended process?
No, mostly because it's not done yet and only experimental right now.
However, go to my original comment and replace "third party" with "it's not done yet and only experimental right now" and the meaning stays the same... :)
The performance is of course also awesome.
So I hope we get more traction for Julia as a general purpose language. But for now I'm sticking with Node.
I was excited about v0.6 but then when I learned it took >15s to do whos(), I completely gave up...
That said, once you're up and running, plots (and everything else) are super snappy. When I'm doing plotting stuff, I'm usually doing it interactively, and when I'm scripting it, it's because I'm plotting hundreds or thousands of things (and then the startup time is vanishingly small).
For context, I come from using MATLAB interactively. When doing that kind of work, the MATLAB 'whos' command is like /bin/ls, it's hard to do much without it.
Is there a different command in julia for showing all numerical arrays/matrices/tensors in memory and their dimensions from the repl? whos() seems like another direct analog of MATLAB, but maybe there's something better that isn't unreasonably slow?
[names(Main) typeof.(eval.(names(Main)))]I know web frameworks exist for Julia, but I’m wondering how practical it is to actually use this language for that purpose?
But maybe I'm wrong here. It seems that "use the right tool" turns into use Python or JS or Excel for all the things.
Python had a long history of web server development. I think, aiohttp may be considered state of the art now. Let's measure its performance:
from aiohttp import web
async def handle(request):
return web.Response(text="Hello")
app = web.Application()
app.add_routes([web.get('/', handle),])
web.run_app(app)
using `wrk` for testing: $ wrk -t1 -c1000 -d30s http://127.0.0.1:8080/
Running 30s test @ http://127.0.0.1:8080/
1 threads and 1000 connections
Thread Stats Avg Stdev Max +/- Stdev
Latency 164.59ms 17.12ms 537.35ms 83.97%
Req/Sec 6.05k 0.95k 7.74k 74.33%
180749 requests in 30.08s, 26.72MB read
Requests/sec: 6008.75
Transfer/sec: 0.89MB
So we have ~6k rps on a single CPU core. As far as I remember, Tornado has ~4k rps, while built-in Flask server can process only about ~1k requests per second . Yes, you are unlikely to use Flask dev server in production, but for aiohttp it's indeed a recommended way.Now let's measure Julia's HTTP.jl server:
using HTTP
HTTP.listen() do request::HTTP.Request
return HTTP.Response("Hello")
end
which gives: $ wrk -t1 -c1000 -d30s http://127.0.0.1:8081/
Running 30s test @ http://127.0.0.1:8081/
1 threads and 1000 connections
Thread Stats Avg Stdev Max +/- Stdev
Latency 105.95ms 108.35ms 1.99s 98.40%
Req/Sec 9.65k 1.69k 13.66k 80.21%
271917 requests in 30.09s, 16.10MB read
Socket errors: connect 0, read 0, write 0, timeout 274
Requests/sec: 9035.73
Transfer/sec: 547.67KB
So it's 9k (with a few failed requests though).This doesn't include any routing, input data parsing, header or cookie processing, etc., but it amazes me how good the server is given that web development is NOT considered a strong part of the language.
The downside of Julia web programming is the number of libraries and tools (e.g. routers, DB connectors, template engines, etc.) - they exist, but are quite behind Python equivalents, so gotchas are expected. Yet I'm quite positive about future of web programming in Julia.
That said, my personal favourite thing about julia is not actually its speed but the fact that it's an incredibly expressive, flexible and composable language (I'd say julia learned all the mot important lessons lisp had to teach). If I had to guess, I'd think that building the infrastructure for devs to work on web servers from the ground up may very well be easier in julia than in most other langauges, including python. I may be off the mark though.
One other note since it was mentioned elsewhere, there actually is work being done right now to have julia compile to WebAssembly code which could be pretty cool!
However, it was pretty actively discouraged. I suspect they didn't want a lot of voices influencing the direction of the language while it was still being formed. A lot of what make R a poor general purpose language help make it effective in its specific domain for its specific users.
It might be worth people taking another look at it after 1.0 is released, as long as they have low expectations about influencing the direction of the language's development.
What is the specific context for this? Serving concurrent users? Or raw processing?
If it's the former, there are ways around this. If you'd like to replace Python for this, I'd look at something like Go rather than Julia.
If it's the latter, Julia might help, but there are a bunch of tradeoffs to consider. Julia is a language that was primarily designed for numerical computation.
> "Julia is the fastest modern open-source language for data science, machine learning and scientific computing...with the speed, capacity and performance of C, C++..." [4]
That's a bold claim! And there is much evidence to the contrary. I often find that the language falls into a similar trap of many other languages where the creators are evangelists not quite sharing the whole picture. Writing fast Julia code is not always a pleasant experience as you often need to fight the easier idioms. It is marketed as fast, but really, how fast is it?
[0] http://www.zverovich.net/2016/05/13/giving-up-on-julia.html
"What’s disappointing is the striking difference between the claimed performance and the observed one. For example, a trivial hello world program in Julia runs ~27x slower than Python’s version and ~187x slower than the one in C."
[1] https://www.ibm.com/developerworks/community/blogs/jfp/entry...
"We can code in Julia in a way similar to Python's code. However, that code is slower than it should. One way to speed Julia is to take into account the Fortran ordering it uses by looping on j before looping on i in the second loop. We also add decorators to speed up the code."
[2] https://www.codementor.io/zhuojiadai/julia-vs-r-vs-python-st...
"...comparing R's sorting speeds to Julia's is not the complete story, even though on the surface R appears faster, and from a users' perspective, (once the data is loaded) R is still the king of speed."
Stackoverflow has quite a few user posts [3] from folks trying to get their code to perform at the advertised speed vs popular alternative languages. It's not so easy.
[3] https://stackoverflow.com/questions/20613817/julia-julia-lan...
I used to call Fortran code from Python for a particularly complex function (Mittag-Leffler function) from Python, I then rewrote everything in Julia and it runs approx 10-20x faster. So there you go, I'm a primary source.
Edit: just for clarity: The Mittag-Leffler function itself runs just ever so slightly faster in Julia than Fortran-called-from-Python (due to the probably inefficient way I was sending and receiving the data to/from a Python subprocess, which itself is an example of the difficulty of combining two languages). It is the program as a whole which runs 10-20x faster.
The 10% figure comes from a comparison of the performance of the codes whose logic has been left virtually unchanged, of course.
What is awesome is the ability to run codes within Jupyter (like Python), and the ability to run "for" loops on large arrays fast and with little memory consumption (unlike Python). The 10% increase in computational time is fully absorbed by the reduced development/validation time!
About the second quote: this is not about Julia being slower than Python; it's about avoiding cache misses. Obviously you should exploit contiguity of the data in memory. Secondly, the 'decorators' like "@inbounds" just ensure that Julia does not redundantly checks bounds, otherwise the inner loop has branches and cannot exploit SIMD.
[1] https://tk3369.wordpress.com/2018/02/04/an-updated-analysis-...
[1] https://docs.julialang.org/en/latest/manual/performance-tips...
https://news.ycombinator.com/item?id=17164737
I guess the claim is that you could write all that in Julia, and hope to be competitive, while you could never do this in (pure) Python. It would still be a lot of work, and would require you to think about the cache, there's no way around that.
The selling point here (it seems to me) is that you can easily doodle up the most naive version of whatever algorithm you're thinking about, and then start optimising to the degree needed, where needed, without having to start over in C or something.
Julia in comparison is a fully featured language, with a lot to offer - you can write highly generic + fast code in a very high-level way. E.g. have a look at https://medium.com/@Jernfrost/defining-custom-units-in-julia... You can see, that Julia can be more elegant than python while being a lot faster.
I think this works best when you have some loops pushing arrays of Float64 around. And worst if you'd like to pass some library a million little functions to optimise which depend on some strange data type you just defined... Cython is quite a limited sub-language. But certainly many current Julians made a living this way in the past.
Especially when you consider what would result from that same user spending that same amount of effort trying to learn C.
What matters to me is that it is actually really nice to write. The syntax is elegant, the community is excellent, and the core of the language's semantics being multiple dispatch is a game changer.
I think that the ML community's decision to write core operations in C or C++ and provide Python wrappers is the way to go about it if you want the flexibility of scripting.
Julia's performance optimizations have to far been mostly focused on intensive numerical computations where the startup time and compilation time are merely a constant overhead that are irrelevant to high performance numerics.
I don't have to recompile a julia function every time I run it so long as I don't close my julia session. If I do need to close and reopen my julia session a lot for some reason, I'd just statically compile the function. In practice, one rarely needs to close a julia session and recompile functions and if one does do it occassionally, the compile time while a little annoying is not too bad.
Compile time will also be actively worked on to reduce it post 1.0 once all the breaking changes are done and pressing bugs are squashed.
The biggest issue with the standard Julia benchmarks is that they're not using what would be optimal C code to compare against.
In addition, if you look at the Julia issues on GitHub, you'll find hundreds of performance regressions where code performs more than 10 times as slow as what they expected/claimed at one point. It's not reliably fast, even when written by Julia experts/developers.
Yes, the C code benchmarks are not optimal. Neither are the julia benchmarks or any of the other languages for the matter. The julia devs made a hard decision with those benchmarks and decided that if they allowed arbitrary optimization, the benchmarks would become more a measure of who spent the most time and knowhow writing the benchmarks for ______ language. Instead, they tried to keep the code for all the languages at a reasonable level and avoided super specialized magic. That may make some uncomfortable, but keep in mind that the julia code used in the benchmarks also has a lot of room for improvement. Some julia devs are absolute wizards and getting performance out of julia code if you let them go crazy.
> In addition, if you look at the Julia issues on GitHub, you'll find hundreds of performance regressions where code performs more than 10 times as slow as what they expected/claimed at one point. It's not reliably fast, even when written by Julia experts/developers.
Julia 0.7 (which is still in its alpha build by the way) included a ground-up replacement of julias iteration protocol and a reworking of a ton of code optimization routines. If you think you or anyone else in the world could take a codebase as large as julia's and replace fundamental parts of it without seeing performance regressions anywhere you're delusional.
There are a number of performance regressions, some mysterious and some not mysterious and they will all be worked on because the julia devs take regressions very seriously. It won't be instantaneous, but I do not doubt that the bulk of them will be eliminated promptly.
There are also orders of magnitude more performance improvements in 0.7-alpha than there are regressions, they just aren't filed as issues. Some of these performance improvements, especially with broadcasting, are state of the art and not seen in other languages.
Calling julia a "language of failed promises" because the alpha build of a pre 1.0 version has some performance regressions is sensationalist and disingenuous.
When you're doing that much typing, you might as well write C or C++.
If I am reading you correctly, you are a library builder, an infrastructure creator, and comfortable caring about the bottom line when it comes to performance. You most likely prototype an algorithm in some high-level language, ensuring that it works, then push it down into C++ in order to make it scale to cool problems, lastly you may create bindings to a higher level programming language so that others without your C++ acumen can benefit from your labour. A lot of great code is written this way, some that I rely on in my day-to-day work would be OpenBLAS, TensorFlow, and PyTorch.
I can only really speak for myself, but I know that there are many like me, we are researchers and to us code is only incidental. We are judged based on our ability to churn out as many papers and results as possible, in as little time as is humanly possible (ever wondered why academic code can be absolutely awful?). We rarely know the structure of the solution a-priori, rather, we start throwing techniques at things and try to make the experiments run. At some point, a (wild?) performance bottleneck appears and we just want to get around it as soon as possible. Now, some like myself have been in the Python/Cython world for years and you can get around a lot of bottlenecks this way. However, it comes at the expense of additional boilerplate and mastering which parts of the Python programming model you must throw overboard, not to mention how you make your Cython code interact with pure Python code from libraries that others have written. This is where Julia shines, it allows you to much easier go between this “productive” and “performance” mode, to me, this is worth its weight in gold. It is not for everyone, but if I ever have a law named after me I would be happy if it was “Nothing is for everyone”.
Yet that's how Julia looks to us: Python + a sensible Type system.
You can improve much more gradually as needed, retain all the full language features and have much less menta overhead, needing only to keep Julia in your head.
Julia devs make performance regressions when replacing critical central code components just like everyone else does. The advantage is that julia makes it easier to reason about and fix those regressions
Personally, I still find C++ to be effectively a black box, meaning that if I want to understand it or make changes, I'm at the mercy of the maintainers or willing colleagues.
Isn’t Google’s “Swift for TensorFlow” move in direct opposition to this statement? I think you are right that it gets you 95% of the way, but that the final 5% of performance and portability simply will not be there.
https://github.com/tensorflow/swift/blob/master/docs/DesignO...
I might be interested in trying cxx.jl, thank you for the pointer.
https://github.com/JuliaLang/julia/pull/24473
which are making this:
https://github.com/JuliaLang/julia/blob/master/src/interpret...
a compilation-free version of Julia. At the same people there is discussion about how to cache native compilation results after precompilation, so that way packages can store all of their compiled code and users can directly use it without compilation through the interpreter. This second part isn't done yet, but it'll be interesting to see what happens when it has a truly dynamic + precompiled form.
I know several efforts that are transitioning Python + C Libraries to Julia. It's simply much much nicer and simpler to write fast code in Julia than in the Python/C paradigm.
This is not necessarily a bad thing. Just an observation from me translating my Python code to Julia.
Thats a deceptive and unhelpful 'benchmark', especially when the C++ compilation time was not included in its benchmark. Julia compilation and statup times are slow if you are expecting the feel of the Python interpreter, however, julia's compilation and statup time add a finite, constant overhead to its preformance and so are completely irrelevant for high performance numrical computing.
If you were wanting to spawn many julia instances and make them execute a single command then exit, yes julia would be a terrible choice of that (unless you turn of compilation and use its interpreter or statically compile your program) but for the most part measuring startup time just isn't relevant unless it takes a minute or something absurd.
In julia the only libraries I know of that take a minute to load are plotting libraries (once you make your first plot the library is blazing fast again) and thats been considered such a big issue that a new plotting library Makie is well under way which is supposed to be fully statically compiled in julia in order to slash the first time to plot.
exp.(v.^2) .+ (10 .* u)
can be written as
@. exp(v^2) + (10 * u)
with no added runtime cost.
Update: this post was edited a bit to remove a sentence about how friendly
the Julia community is since that no longer seemed appropriate in light of
recent private and semi-private communications from one of the co-creators
of Julia. They were, by far, the nastiest and most dishonest responses I've
ever gotten to any blog post. Some of those responses were on a private
discussion channel; multiple people later talked to me about how shocked
they were at the sheer meanness and dishonesty of the responses. Oh, and
there's also the public mailing list. The responses there weren't in the
same league, but even so, I didn't stick around long since I unsubscribed
when one the Julia co-creators responded with something bad enough that it
prompted someone else to to suggest sticking to the facts and avoiding
attacks. That wasn't the first attack, or even the first one to prompt
someone to respond and ask that people stay on topic; it just happened to
be the one that made me think that we weren't going to have a productive
discussion. I extended an olive branch before leaving, but who knows what
happened there?
Update 2, 1 year later: The same person who previously attacked me in
private is now posting heavily edited and misleading excerpts in an attempt
to discredit this post. I'm not going to post the full content in part
because it's extremely long, but mostly because it's a gross violation of
that community's norms to post internal content publicly. If you know
anyone in the RC community who was there for the discussion before the
edits and you want the truth, ask your RC buddy for their take. If you
don't know any RC folks, consider that my debate partner's behavior was so
egregious that multiple people asked him to stop, and many more people
messaged me privately to talk about how inappropriate his behavior was. If
you compare that to what's been publicly dredged up, you can get an idea of
both how representative the public excerpts are and of how honest the other
person is being._Edit: To add, Julia is now my go-to language for any scientific/numeric programming and would be my number 1 suggestion to anyone else doing similar work.
Unfortunately, it doesn't take much negative energy to spoil a community, even if it comes from just one person. That it was a language co-creator is troubling but not necessarily a reason to avoid the language if the rest of the community is nice (which maybe it is, maybe it isn't).
However one thing I have noticed is that most language communities do tend to follow an attitude set by the language creator(s). This seems to have played out quite a bit for Clojure, Python, Elm, Elixir, and other languages where, at least to me, the overall shared perspective of the community is closely aligned to the personal attitudes and opinions of the language author(s).
I'm not sure if there's a fair way to quantify drama in a community, but it hasn't been a big issue for me relative to all the positive interactions I've had with other julians.