121 karma · joined March 23, 2022
Sometimes I get a lot of enjoyment of buying a thing that I know I will love and considering all of the alternatives. Other times, I just defer to what worked in the past.
I see what you mean by obfuscation, but I think that it's one of those things that feels really hard and stupid until you start being able to do it really quickly. When you learn a foreign language, you first read letters, then words, then sentences because you become accustomed to larger pieces of the language that you can predict what's coming next without reading it. A similar sort of thing happens with APL/BQN, you read letters (primitives), then you begin to recognise words (small, commonly used groups of primitives), then you see larger patterns which look like magical incantations to an inexperienced user.
These "words" are (typically) tacit phrases, many of them only existing due to specific primitives like swap. Once I used BQN to golf, I started wishing Julia had a swap for operators i.e.
-(3, 5) = -2
swap(-)(3, 5) = 2
I won't defend these languages to the death, but they are fun to puzzle your brain with in codegolf. Maybe Dex[4] will go somewhere too.[1]: https://mlochbaum.github.io/BQN/spec/system.html#operation-p...
[2]: https://saltysylvi.github.io/blog/bqn-macros.html
Unfortunately, as with Plots, the documentation is lacking. The basic tutorial does a good job introducing the aspects of the package at a high level, but the fact that some parts of the documentation uses functions/structs that don't have docstrings in examples makes it very hard to build on the examples in these cases.
I get it, I can do anything with Makie, and most things that I want to do work amazingly. But my code for a single figure can get huge because it's all so low level. See, for example, the Legend documentation[1].
[1]: https://docs.makie.org/stable/examples/blocks/legend/index.h...
I have not tried it. I like that the project makes broadcasting invisible, I dislike that it tries to completely replicate R's semantics and Tidyverse's syntax. Two examples: firstly, the tuples vs scalars thing doesn't seem very Julia to me. Secondly, I love that DF.jl has :column_name and variable_name as separate syntax. Tidier.jl drops this convention (from what I see in the readme).
> I'm not sure if someone is looking directly at the data.table parts
I believe there was some effort to make an i-j-by syntax in Julia but it fell through or stopped getting worked on. By this syntax I mean something like:
# An example of using i, j, and by
@dt flights [
carrier == "AA",
(mean(:arr_delay), mean(:dep_delay)),
by = (:origin, :dest, :month)]
# An example of expressions in by
@dt flights [_, nrows, by = (:dep_delay > 0, :arr_delay > 0)]
The idea of ijby (as I understand it) is that it has a consistent structure: row selection/filtering comes before column selection/filtering, and is optionally followed by "by" and then other keyword arguments which augment the data that the core "ij" operations act upon.data.table also has some nifty syntax like
data[, x := x + 1] # update in place
data[, x := x/nrows(.SD), by = y] # .SD = references data subset currently being worked on
which make it more concise than dplyr.The conciseness and structure that comes from data.table and its tendency to be much less code than comparable tidyverse transformations through some well-informed choices and reservations of syntax make it nicer for me to use.
According to these benchmarks: https://h2oai.github.io/db-benchmark/, DF.jl is the fastest library for some things, data.table for others, polars for others. Which is fastest depends on the query and whether it takes advantage of the features/properties of each.
For what it's worth, data.table is my favourite to use and I believe it has the nicest ergonomics of the three I spoke about.
What I always tell people is the following:
If you are writing code using existing libraries then use whichever language has those languages. The NN stack(s) in Python are great, the statistical ML stack(s) in R are simple and include SOTA techniques.
If you are writing a package yourself, then I assume you know the core of the idea well enough to be able to write your code from the "top down" i.e. you're not experimenting with how to solve the problem at hand, you're implementing something concretely defined.
In this case, and tailored to your use, I would argue that Julia has more advantages than disadvantages, especially compared to R or Python. Here are a few comments:
1. Environments, dependencies, and distribution can all be handled by Pkg.jl, the built in package manager. There is no 3rd party tool involved, there is no disagreement in the community on which is better. This is my biggest pain point with Python.
2. Julia's type system both exists and is more powerful than that of Python (types or classes) and R (even Hadley's new S7(?) system). By powerful I mean generics/parametric types and overloading/dispatch built in. You can code without them, but certain problems are solved elegantly by them. Since working heavily with types in recent years, I find this to be my biggest pain point in R and I wouldn't want to write a package in R, although I like to use it as an end user.
3. New developments in scientific programming, programming ergonomics, hardware generic code (as in this post), and other cool features happen in Julia. New developments in statistics happen in R (and increasingly Julia), new developments funded by big companies happen in Python.
4. The Python and R interpreter start up faster than Julia. The biggest problem here is when you are redefining types, which is the only thing in Julia that can't currently be "hot reloaded" i.e. you need to restart Julia to redefine types.
5. Working with tabular data is (currently) far more ergonomic and effortless in R than Python and Julia.
6. Plotting is not a solved problem in Julia. Plots.jl is pretty easy and pretty powerful, Makie.jl is powerful but very manual. Time to first plot is longer than R or Python.
7. Julia has almost zero technical debt, R and Python have a lot. Backwards compatibility is guaranteed for Julia code written in >v1.0 and Pkg.jl handles package compatibility. If I send you code I wrote 4 years ago along with a Project.toml containing [compat] information then you could run the code with zero effort. (This is the theory, in practice Julia programmers are typically scientists first and coders second, ymmv.)
8. You can choose how low level you want your code to be. Prototyping can be done in Julia, rewriting to be faster can be done in Julia, production code can be done in Julia. Translating Python to C++ production might mean thinking about types for the first time in the dev process. In Julia, going to production just means making sure your code is type stable.
From my perspective, however, DataFrames.jl's power is what makes it quite unergonomic for me. As an example, take the `args => transformations => result` syntax for doing pretty much anything in DataFrames. It versatile, but the lack of rank polymorphism in Julia i.e. broadcasting/mapping has to be explicit (which is usually a good thing given that type polymorphism is Julia's whole schtick) means that the transformation syntax feels cumbersome.
It's not that I want everything rowwise by default, an option provided by DataFramesMacros.jl, it's that I want things to be rank polymorphic when it makes sense. Base R got this right, hell S got this right, and so the Tidyverse inherited it and it makes the package so much more ergonomic than it would otherwise be.
I cannot overstate how impressive DataFrames.jl is, but I have to caveat this with "but I really try to avoid using it if possible". It's a shame, but I just think R's laissez-faire hackability, which in many cases results in spaghetti code, works really well in the tabular programming world where ergonomics are king and performance is easy.
In particular, as you mention, plotting is one of the evolving parts of the ecosystem. Plots.jl is fine, Makie is powerful but very DIY, and AoG is slick but unwieldy. ggplot2 is far from perfect, but it works so well due to its maturity and integration with the rest of the Tidyverse.
In my ideal world, there would be a DataFrames.jl wrapper to provide nice (not just nicer like the two DFM.jl packages) syntax, and a powerful high-level plotting package (Makie is powerful but syntax is low level, Plots is mid on both) which is heavily integrated with the wrapper package.
Admittedly, I'm not a data scientist (anymore) so I don't follow the new developments in the dataviz scene much. If something like this exists then I would love to find it.
I wonder what my ideal syntax would look like anyway. Maybe something close to Tidyverse but with symbols as column names `:col_name` for ambiguity reasons.
For context, I very often find myself building small libraries from scratch to solve very specific scientific problems which is somewhat performance critical. Julia has worked very well for me for this, but I recognise that moving outside of your comfort zone is the best way to become a better programmer.
RStudio is the perfect IDE. REPL/command-line + Scripts + Plots. I could not be happier using it and I wish I could get VSCode to be half as good. Julia for VSCode is pretty good, but the Python science tooling goes 100% towards notebook environments which I'm not a huge fan of so the Python Science VScode experience is subpar.
Without defining more precisely what succeeding and the timeframe means it is difficult to know what OP was after but the fed example stands well on its own.
Writing .jl files in VSCode and sending code to the integrated REPL with shift-return for playing around is great and then when I want to run code properly I just comment out my testing code, wrap my important code in a main() function (which is sometimes everything apart from the imports) and then just make sure that I'm not referring to any variables or function methods defined that are now commented. This is made relatively easy as VSCode will immediately complain about possible method errors.
It's not a perfect substitution for having fast startup and when I go back to Python or R for things like the holy grail called Tidyverse then it feels like a weight off my back, but the benefit of doing things the Julia way is that your code has a clear starting point and it helps readability as you can see what every script is supposed to do.
The first impression of how you are as a firm is the job ad, and if you don't post a salary the first impression is worse.
Just goes to show how important it can be to read code in completely different fields than the one you're working in. Some problems are common for others so they've already been solved. Knowing that a solution exists is half the battle.
How good is my Python out of 5? In what areas? Compared to who, someone with 30 years of Python experience or my peers who've never coded before?
Furthermore, with natural languages, there are commonly used scores (CEFR in Europe at least) that are more or less objective and don't depend person to person. If someone says they're B1, you know what to expect. If someone says they're 3/5, what does that mean?
I would like the option to not justify how skilled I am at something, or have more flexibility to qualify/quantify my skills such as, but not limited to, number of years of professional use.