This amuses me. I hadn't really considered the author's perspective, and now I think it aligns with my take on it pretty well.
This amuses me. I hadn't really considered the author's perspective, and now I think it aligns with my take on it pretty well.
Source: my current and previous job were basically data viz programming jobs which were all about optimizing said iterative loop for scientists. Going from minute-long to sub-second rendering speeds is a game-changer for many.
EDIT: having said that, this kind of reminds me of what the biggest difference between analog and digital photography is for me, namely whether or not you get instant feedback. I do remember from my art school days that in my experience film was a much better option for training the skill of observation and composition than digital, because it forces you to essentially picture the photograph before you take it. However, once you get somewhat decent at that... I'd switch to digital and reap all the benefits it has ;). The same logic might apply to learning how to plot your data.
Second, third, forth etc plots in Julia are fast. Likely faster than any of the competition as it is running highly optimized native code at that point.
On the long term? I think Julia is a better language than Python, Matlab or R for several reasons like modularity, package management and performance. But these are things that require at least 10 hours of use (to say something) instead of 10 minutes. With so many languages promising enlightent and transcendence to a next power level you cannot expect people to make that kind of investment.
Also, note that for people to whom plotting is really important, it's quite easy nowadays to just AOT compile your plotting library to your sysimage with PackageCompiler.jl [2] for instant plots.
At the language level Julia separates files from modules. Modules are just a language construct, another object. They can span several files, have several them in the same file or even declare them in the REPL. This last point is important for me since I can (re)evaluate code in the REPL without polluting it. Usually Python feels more interactive but this is a point where Julia wins at the REPL.
Julia's package manager is just another library and you have support integrated for it inside the REPL.
I'm not devops, I don't think I can judge the relative merits of Julia's package manager design relative to Python's but for me "it feels" better. You will have to judge for yourself reading the docs and this things about federated package management.
I have not tried it but Julia allows to build a sysimage with all dependencies included. I know there are similar things to build standalone Python apps but last time I tried (long ago) it was a pain.
Finally, Python depends a lot on external C libraries to achieve performance. This obviously complicates deployment and is a reason why I use Ubuntu: I have binary wheels for almost everything. Julia provides performance without external tools and packages are usually pure Julia code.
I think Julia's is vastly superior concerning packaging but of course for 99.9% of people this is not enough to make a switch yet (me included, for the moment I use it for hobby projects).
Regarding modularity is amazing and not obvious at first why Julia's performance increases modularity. The reason is that in Python since you need to use C/C++ for performance your data structures need to be shaped appropriately when they cross this interface. This rigidity propagates through your program and makes you build big frameworks. I have in mind for example PyTorch or Tensorflow. So you have Numpy arrays, PyTorch tensor... you have of course almost transparent conversion between them since numpy is a standard. But all of this is achieved because a behemoth like Facebook or Google are behind injecting money and manpower. It's pyramid building: some engineering but a lot of work. Even then you are stuck with arrays and contorting your program to vectorized operations.
There are of course some dark spots for Julia but I have the feeling they will be solved. I don't consider myself a fan boy or an early adopter but I think it has a future. So I thought for Python 20 years ago :)
On one hand you have people that say "thinking takes a lot longer than waiting a bit for compilation or actually editing source code" (I'm in this camp) and people that go "I don't want to wait for compilation and I want my editing to be hyper-efficient even if I have to invest hundreds and thousands of hours into it, so that I'm always in the zone/in the flow".
People are just different but every camp thinks They're Right and The Others Are Dumb and Stupid And Dangerous.
My personal guess is that besides the split in personalities/workflows, there's also a difference in projects. People who work on existing projects tend to read a ton more code/docs/team comms/architectural diagrams and edit/compile less so they care less about these issues. People who constantly create tons of mini-projects with short lifecycles care more about them.
What I mean in practice is that if you take Rust for instance, the compile times can be fairly long but the type checking occurs early on and is quite fast. Therefore once I know that this step succeeded I can usually let the compilation continue in the background while I focus my attention elsewhere.
If on the other hand if I need to wait a lot longer to confirm that my code is actually valid I find myself just staring at the output window, not willing to let go of my short term memory until I get a confirmation that my code was accepted.
The problem is not adding 20s to your overall dev time, it's to have a 20s interruption while you're "in the zone".
Seems to me like people are making a mountain out of a molehill.
Seeing these answers makes me think Julia will never be fixed. I forecast Julia will be back in a niche within five years if they don't get their act together. And it's sad, because the alternatives are fundamentally broken. Julia isn't fundamentally broken, but the devs and the community seem to insist on superficial breakage.
Could you elaborate on this? I'd say repl based interactive programming is one of julia's greatest strengths, and avoiding the repl is probably setting yourself up for pain.
That said, if you do find yourself running lots of scripts and paying this penalty all the time, I'd suggest https://github.com/dmolina/DaemonMode.jl as a great way around these pains.
With REPL you have an invisible global state, can't reproduce what you have done, changes earlier in code path don't propagate to results, you don't have documentation of what you did.
It's for me really like trying to write a book by dictating. Except that you're dictating to somebody who's gonna give an independent summary of it to a third party and never gonna write down what you dictated. It boggles my mind how people can work like this, but they probably get hooked to REPL from the first tutorials and just don't know better.
I'll look into DaemonMode.jl. Not a fan of using a daemon (and I'm guessing there will be problems with e.g. interactive plots), but in the short term I'll take anything that could make Julia programming tolerable.
Mhm, that's fair. I think Pluto.jl has a really neat approach to this, using reactivity (and technically even more state) to actually eliminate that experienced state.
If I could use it from emacs it might even be my goto way to interact with julia, but I also don't mind the statefulnes and find it manageable.
For me, the most important thing is that when I'm writing serious code, I create a local package. Then, in the REPL I load that package and have Revise.jl active so that it can watch the the package source ode and constantly do hot code reloading for me so that I'm never stuck with old versions of code running.
Then I do all my interactive analysis in the REPL, and plumbing in the package module. This eliminates a lot of statefulness, but keeps restarts to a minimum.
I create a "local package", meaning a file from which I relatively import. During development/analysis it's hard to foresee what the package structure is gonna be, so it's quite pointless to go through the whole packaging ceremony at this point. FromFile works fine for this.
As a temporary hack I could use REPL to call my "main" function and let Revise.jl update automatically. (In long term this is bad for interoperability with rest of the system). But in my experience Revise.jl tends to break a lot. Julia breakage is hard to analyze by itself, and Revise.jl often makes this more or less impossible.
I have to repeat that I really don't see how caching of the compilation results is even close the complications that Revise.jl or Pluto.jl have to do.
Not quite. It builds a dependancy graph of your code and can figure out what definitions depend on others. So depending on what you change, maybe only one or two cells need to be rerun. Or in other circumstances, the whole notebook will have to re-run. It just depends on what changes.
> I have to repeat that I really don't see how caching of the compilation results is even close the complications that Revise.jl or Pluto.jl have to do.
I think the main trouble with the caching is that the native code you cache can depend very strongly on the exact combination of packages you have loaded. This means you can hit a combinatorial explosion of different methods to cache pretty quickly, so you'd need to find a very clever way to find the right methods to keep and which ones to delete once the cache gets too big.
I think there's also other potential issues that I understand less. This is being actively worked on though.
But the effect is still that any changes up-file will be always reflected down-file? If so, I don't care how it's implemented (given it's fast enough and doesn't break), the semantics is the point.
> I think the main trouble with the caching is that the native code you cache can depend very strongly on the exact combination of packages you have loaded. This means you can hit a combinatorial explosion of different methods to cache pretty quickly, so you'd need to find a very clever way to find the right methods to keep and which ones to delete once the cache gets too big.
Yes, I think this is a problem for a clean solution. But for a big fat ugly hack that isn't too picky on wasting disk space or occasionally recompiling stuff needlessly it's probably less so.
For a lot of cases very rough invalidation would probably suffice. E.g. invalidate all definitions from all files that are changed from the last run (i.e. like Make does). And invalidate all definitions for any name that gets any definition. I'd guess accomplishing this would cut the startup time greatly; the end-user code rarely redefines (at least intentionally) anything that's in the packages, and vast majority of time is spent (re)compiling the packages themselves.
I'm sure there are complications with type inference. But I'd be willing to pepper some explicit typing in my code if it means I don't have to recompile it every time I run it. Binary of a method with concrete types should at least be trivially cacheable (given no library changes between runs).
> This is being actively worked on though.
It's been worked on for as long as I've known of Julia. AFAIK there's still absolutely zero logic on caching compilations of "end-user-stuff" (as opposed to stuff like package precompilation). I don't think this is necessarily due to technical issues, but because the community says that REPL (or notebook) is the only way of using Julia, and those don't suffer from the problem that much (Revise.jl breakage notwithstanding).
Technically it's probably very difficult to do "perfectly", and I'm thinking this is how the compiler devs want to do it. I'm not sure they even mean persisting-between-runs caching when they say "caching" in compiler related discussions. It may well be just some run-time caching of some compilation artefacts that are now compiled multiple times. And that would probably not have that dramatic performance gains for the re-run case.
For an AOT compiler Julia is clearly fast enough. There are probably no easy tricks left to make it a lot faster. But re-run performance doesn't need faster AOT, it just needs the compiler not to recompile the same identical stuff every time.
Yes, I was just bringing this up because it means that various things can be significantly faster, causing you to experience less latency than you normally would by re-running a whole file.
As to the rest of your most, I agree it'd be interesting to see a more quick and dirty solution. It appears that everyone who has the know-how to do this wants to 'do it right', so on the public facing side there's very little visible progress.
> It's been worked on for as long as I've known of Julia. AFAIK there's still absolutely zero logic on caching compilations of "end-user-stuff" (as opposed to stuff like package precompilation)
This is not really true. E.g. there's PackageCompiler.jl which does sysimage based caching and works quite well (at the expense of slow compilation and large binaries), and briefly there was StaticCompiler.jl which did good small binary compilation but then bitrotted quite fast.
All of our CPU compliation stuff is built using a small binary, static, AOT compiler (currently hosted in GPUCompiler.jl) and it's quite reliable. There's active work being done to make this work on the CPU again (basically a modern version of StaticCompiler.jl). So while I feel your frustration that this has been 'coming soon!' for a long time, progress has been made. The new compiler hooks for version 1.6 are partially designed to make this whole process less hacky and easier to iterate on.
It would be huge if Julia could be compiled to shared objects with e.g. C interface. I don't even care if they are bloaty or hacky. Any way of accomplishing this would be an instant boost for using Julia in production. And would go beyond anything even close to Julia's productivity.
I think Julia people may underestimate the potential Julia has as a general purpose language, and overestimate the short term efforts to make it happen. Just add some hacks like AOT caching and any way to call with CFFI and it would go like wildfire.
I understand that most of Julia's community is about crunching data, and that's what I do most of the time too. But with that background it's probably not that clear how dire the situation in more general development is. An expressive, reasonably performant and interoperable language would be revolutionary.
I don't see why REPLing it is a more elegant solution. With that solution I can call the function form the REPL if I want, but also from the shell if I want. With shell I get the elegance of having a persistent, complete and reproducible description of the state all the time, which can be e.g. version controlled.
- The other side is "these people" that "won't stop"
- They're striving for "hyper-efficiency" at the cost of "hundreds and thousands of hours"
- People who complain about this issue don't (or do significantly less of) reading documentation/code/etc
People talk about TTFP because it is a real issue that is off-putting for many programmers that would otherwise love to use Julia. Julia is roughly an order of magnitude slower than python in this instance on my computer, and that's not a good first impression.
That doesn't mean that everyone is going to be impacted by this issue (obvious ex: you aren't), but this isn't akin to a flamewar because unlike editor choice (which is opinion), Julia would be better for everyone if TTFP was improved. Whether or not it should be prioritized as a development goal is up to the Julia team, but it's not just some difference of opinion like interpreted vs compiled or functional vs OO.
I personally do care for my concrete usage pattern.
This is my use case: A long shell script that does a lot of things. At some point, inside a loop that runs hundreds of times, it needs to solve a couple of small linear systems and plot a simple graph. There's hundreds of png graphs, that are then combined into a video sequence. Right now, the computation is done by calling octave (inside the loop) and then gnuplot (to create the actual graph from the octave computed data points). I would like to replace each call to octave+gnuplot to a single call to julia. Yet, this would make my script run in a few hours instead of a few seconds, because for this usage pattern all plots are first plots
Before you suggest that I should rewrite the whole thing in julia, maybe you are right but
1) it would take me a few weeks that I don't have
2) that's not my point. A good tool is a tool that can be used for purposes that it was not intended to, like this. If the time to first plot in julia was a millisecond instead of 10 seconds, then julia would be a much better tool.
Are there any downsides to this? You never care how long does a system update take. You always care how long do your programs run.
Well I love julia the language. It's the interpreter quirks that I find annoying. If julia had something lean and superfast like luajit it would be incredible!
I do all my plotting with Gnuplot.jl, as gnuplot is fast, and I can save .gpt files which reproduce the plots for later reference and making publication-quality.
And not that long before that you would print the results out and manually plot them on graph paper.
That said, there is some very interesting work happening on getting around these restrictions. There was a PR from Tim Holy a while ago that could have allowed it, but there were some problems with the PR, and there were also some associated costs that were deemed too steep to pay.
That said, there's other great work on other ways around this. For instance, you can dynamically redefine structs all you want in Pluto.jl notebooks and there's no performance penalty!
This is something that kinda just fell out as a natural consequence of it's reactive design.
It really doesn't if you use a module.
It's not perfect but it's something.
That's not quite right, I think. People are looking for excuses to not use Julia (r new technology X) because it serves as confirmation bias that their choice of <blub language> is still good and there is no need to start thinking of their extensive training and investment in blub is sunk.
As a language and technology it's way better than the alternatives, but usability of Julia as a programming language is broken by the ridiculous startup latency. I'm sure it's not even really hard to fix (at least by some caching hacks), but for some reason the Julia community is actively resisting such fixes.
And don't give me REPL. REPL is a fundamentally broken approach to programming, and REPL people just keep looking for excuses to keep using it.
Edit: And if you give me REPL, I can answer that I have tried it too. It's broken as well. Revise.jl breaks constantly with anything non-trivial. And with the effort going to horrible hacks like Revise.jl, I'm sure a simple caching of compilation results between calls would be nothing. It seems to be something ideological.
Aside from compiler improvements, most caching related optimizations are happening on the module level, because that's where namespaces are seperated.
I tend to refactor code out of there to separate files, and then somehow import it. An ugly way is include, and I've tried Revise.jl with includet.
But I think the least ugly approach is the @from macro from here: https://github.com/Roger-luo/FromFile.jl Judging from some opinion in bug trackers, this is probably gonna get totally shunned by core devs and they'll keep on bikeshedding about the import stuff forever.
With this setup I have about 400 lines of code in three files. It compiles for 15 seconds. After every single change, and actually without any changes too.
I think performance wise this should be equivalent to using modules, but saving some pointless ceremony.
It's not equivalent no - include doesn't introduce a namespace and neither does includet. Compiled stuff from packages (=modules with a Project.toml) is cached between runs, scripts just don't have that luxury of seperation. @from doesn't look into the files you're including and (somewhat simplified) verbatim pastes the code into your "main" file.
I don't think it's a lot of "pointless ceremony", especially since it keeps dependency management on a per project basis easy, is just a `]generate MyPkg` away and allows compiled code to be cached
If you don't want to use projects, that's fine - but please do so in a constructive manner and don't be surprised that the most common workflow (wrapping things in a package) gets more attention sooner. That just signals some disregard for other peoples' needs & wants, even if that's not intended.
If a "package" is used only by me and only from files controlled by me, Project.toml is clearly pointless ceremony. And `]generate MyPkg` too, and assumes REPL on top. Python manages this (albeit with some stupid arbitrary restrictions) fine, Node manages this fine. The compiler doesn't need that stuff for anything.
I didn't look into the implementation of @from, but I picked it up from a huge bikeshedding bug (still open, from 2013...) about local module imports, and assumed it's doing imports instead of including, as it also has a separate namespace. From the code [2] it's not clear to me exactly what it does when, but one branch seems to be generating a module with the code imported on the fly. Not sure this should be any different than any other module for the compiler. Are you sure you're not talking out of your ass on this one?
I don't care if people for some reason want to write their pointless ceremony, but what I don't understand is that people are so jealous of it that they insist of pushing it on everybody else too. I just want to somehow get access to those symbols defined in another file, why does this need more than the path of the file? I'm sure using just files-as-modules would probably be less work for the compiler, and it's easy to have a byzantine package ceremony on top if you want (Python has dozen or so available, so lots to draw from).
[1] https://github.com/JuliaLang/julia/issues/4600 [2] https://github.com/Roger-luo/FromFile.jl/blob/master/src/Fro...
Feel free to show me the PR’s that have have been rejected that would have solved the problem. Or if you think it’s easy, feel free to make that PR yourself.
As the quote goes, “There are two hard problems in programming: cache invalidation, naming things, and off-by-one errors”. This is cache invalidation.
Cache invalidation is not always that hard. For example pure functions are more or less trivial to cache, and you don't even have to do any explicit invalidation. Perhaps do some LRU type pruning if the disk starts to fill up.
I know next to nothing about Julia's internals, but given packages like Revise.jl, PackageCompiler.jl and SnoopCompile.jl are even possible, I don't think it can be that hard. Dumb caching should be a lot easier than any of these.
I may well be wrong, and that nobody has done this yet is a hint to me being wrong. But I think another scenario may be that Julia ecosystem is so hung up on REPLs and notebooks that this case just gets no attention. And very few non REPL-or-notebook people hang around long enough to get to know the internals at all. Maybe I'm just desperate enough?
There's also another possibility, which may sound bizarre but I think is possible. At some level people who come from scripting language background think that long compile times is a sign of a "real language". This is somewhat prevalent in e.g. Javascript scene, where more and more byzantine compilation systems are introduced for a language (or platform) that works just fine without compilation (or can do very fast on-the-fly "AOT" if needed).
What seems broken to me about REPLs is how text-centric they usually are. But I want to be able to easily introspect and play with my programs, and REPLs are one way to do that. Really good debuggers and environments for static languages are "another" way.
I agree that this is probably not for everybody, i.e. if you're not used to the CLI workflow. And admittedly it would be sometimes nice to have e.g. embedded graphics, but unfortunately the troubles usually outweigh the benefits (looking here at emulating damn 70's terminals too...).
I mostly do the introspection with print, dir and help straight in the code. Not ideal, but rarely fails you, and I've yet to find a debugger GUI or IDE that isn't more trouble than it's worth.
Something like autoupdating RMarkdown/Sweave/Pweave/etc would probably work often as well. I sometimes do use Pweave, although it tends to be a bit buggy as well, and doesn't have any caching logic (although it's easy to use your own).
Sadly most efforts seem to go to Jupyter notebooks and such, whose state/code inconsistency are simply a non-starter if one wants to keep some sanity.
This means you have to change your code in order to debug it. And you have to know what you're debugging before you change your code. This is really, in my opinion, much less than ideal, because you have to iteratively instrument your code while you figure out what is wrong. It's a cycle of You print out the first suspect thing, then that produces 5 potential suspects, and you have to decide which one to print next, or print all of them.
You're absolutely right that it rarely fails you, and so I surely want that facility to be at my finger tips. But getting a text representation of a value is literally 33% of what a REPL is for.
I'm not sure how REPL helps you out of this. You still have to somehow change the state of the program, but if you "monkeypatch" it using REPL, you now have to keep in your head what the state is.
In the end code is just description of how to bring the program to some state. I like to have that description on file so I don't have to keep it in my head.
But sure it may not fit your particular preference or it may be that you have simply not learned to use it effectively. It takes some time to work effectively in a REPL style. It took me some years.
Not every language is suited for REPL development. Julia, LISP and Haskell seem quite well suited.
I don’t have quite the same good experience with Python e.g.
How do you persist and document your code with REPL-development? Do you log the commands to some separate file? How do you recreate the REPL state if it crashes or you have to reboot? How do you make sure the REPLs state is what you think it is?
These are (some of) the concrete problems that I see with REPL, and that don't exist for program based workflow. And I think these are fundamentally impossible to solve for REPL, and very important for e.g. reproducibility (and IMHO sanity).
Other times I have the package I am editing loaded with Revise.jl active and I am calling in and trying out the methods I am concurrently writing in my text editor.
It's more like TDD than anything else. It's got that same quick back and for of run, write, run write. But a but more interactive. (Note I am not saying that is TDD -- it isn't -- tests are not nesc written or saved. Though I do often use this while doing TDD to run a test I have written)
For more complex things, I can even have a stub for a function I am developing and breakpoint into it and then I prototype the functionality in the REPL and that is wayyy faster and less bug prone than trying to go at it blindly in your IDE without being able to experiment and verify your code as you develop it.
What I don't see is the benefit of the REPL. I effectively use my editor and the shell as a "REPL", but instead prefer to have the code in a program structure all the time. This means I don't have to have extra discipline for not accidentally polluting the state. Plus I can use version control, which means I can quite easily try out quite deep changes and still revert back to any state I had before. This is difficult with REPL. And with this workflow I don't have to do any extra "graduating" step; the code usually cleans up during the process.
A big additional benefit is that I can use the shell. I can e.g. pipe stuff from other programs, and I make a CLI for the program on the fly.
The main problem perhaps is when there are some longer computations, as there often are in analyses. For these I prefer memoization to the disk. The usual way to do this across REPL-sessions is to write ad-hoc temporary result files. This gets hairy really fast when you have to update these when the codepath before the dumps change.
I don't see much benefits of REPL over my workflow. Maybe some completions are nicer in REPL and you may get a bit nicer formatted output, but these are quite trivial. Perhaps people who are not accustomed to shell think that REPL is the only way to "rapidly iterate"?
The problem is that I can't. The startup latency makes using scripts practically impossible. I think it could be relatively easy to fix (hack). I'll shut up about this on the very second there's some way to get same magnitude of latency with scripts as there's with REPL.