105 karma · joined December 1, 2019
For my personal workflow (others may differ), compiling to html is only done once at the end of a session, and the latency wouldn't matter if it could execute like a script. Weave.jl^[1] has a great feature called `include_weave` which has the features I like.
But take my feedback with a grain of salt. I generally just save things in folders and compile a pdf separately with many tables and figures.
[1] https://weavejl.mpastell.com/stable/usage/#include_weave
1. Has scoping rules that make it difficult to debug 2. Has low latency, making it frustrating for debugging.
It's actually really nice. Works great over ondemand and is way more responsive than X-based apps.
Jupyterhub is a way to coordinate compute space / config across many users of notebooks or jupyterlab.
We made it illegal to build housing, so we don't have enough housing. We could fix our housing crisis by re-legalizing housing _without_ fixing all of the (very real) problems you describe. You are simply engaging in whataboutism which will _not_ solve our housing crisis.
But sympathy makes a bit of sense when you think about it economically. The businesses that do make it in have more market share and are selected to be those that are politically connected. The regulations serve to keep other businesses out, and we should try to help people that aren't as connected start businesses as well.
Yeah allocation seems like the biggest hangup here. I would rather have a function stick to a "no allocating" contract and allow for some undefined behavior than have a function unexpectedly allocate to preserve safety.
The histogram errors seem annoying though. Hopefully they can get fixed.
I agree it's unlikely that a user will name their column `.data`. But it certainly saves developer effort from thinking about these issues.
The larger concern, really, is that Julia needs to know which things are columns and which things are variables in an expression at parse time in order to generate fast code for a DataFrame. It needs to do this without inspecting the data frame, since the data frame's contents aren't known at parse time.
One option would be to make all literals columns. But then you run into issues with things like `missing`, which would have to be escaped or not recognized as a column. Its hard to predict all the problems there, and any escaping rules would definitely have to be more complicated than R's. So we require `:` and take the easy way out, which has the added benefit for new users who might get confused about the variable-column distinction.
.data[[a_var]]
? df = tibble(a = c(1, 2))
and you want to use a dplyr verb to modify it mutate(df, b = a + 1)
the `a` in the above expression refers to the column in `df`, but this means it's hard to reference a variable in the outer scope named `a`. Furthermore, if you have a string referring to the column name `"a"`, you can't simply write mutate(df, b = a_var + 1)
Contrast this with DataFramesMeta.jl, which is a dply-like library for Julia, written with macros. df = DataFrame(a = [1, 2])
@transform df :b = :a .+ 1
Because of the use of Symbols, there is no ambiguity about scopes. To work with a variable referring to column `a` you can write a_str = "a"
@transform df :b = $a_str .+ 1
I won't pretend this isn't more complicated or harder to learn. Some of the complexity is due to Julia's high performance limiting non-standard evaluation in subtle ways. But a core strength of Julia's macros is that it's easy to inspect these expressions and understand exactly what's going on, with `@macroexpand` as shown in the blog post.DataFramesMeta.jl repo: https://github.com/JuliaData/DataFramesMeta.jl
Here is a tutorial for those familiar with dplyr: https://juliadata.github.io/DataFramesMeta.jl/stable/dplyr/
I want to push back on this a bit, which I acknowledge is very ironic. In the past few years people consistently post on Discourse asking for fundamental changes to the language to make it more resemble python, C++, or whatever their preferred language is.
People often say Go is great because there is "only one way of doing things", yet people are very resistant to being told "the way" to do something in Julia. This has happened enough that it's prompted a pinned PSA on discourse: https://discourse.julialang.org/t/psa-julia-is-not-at-that-s...
It gets tiring! And i'm not sure how the community should handle these requests, but I don't think it's fair to blame all of the negativity on the Julia community when these somewhat misinformed, or even bad-faith posts are so frequent.
It's a very simple package that also displays some of the core strengths of Julia as a language: multiple dispatch and interactivity.
I can imagine someone slowly transitioning a complicated workbook to a script with this package.
One pain point is that Rmarkdown uses a different pandoc installation when executed by Rstudio than from the terminal.
sysimages are great, as is daemonmode. But really just do Revise at the REPL.
But if your point is the inability to do `julia script.jl` , yeah thats a pain point. Fortunately there has been some tooling to make running many jobs in a row easier: https://github.com/dmolina/DaemonMode.jl
That's not quite correct. The major `source => fun => dest` API as part of DataFrames.jl was designed specifically to get around the non-typed container problem. And it definitely works. That's not the cause of slow performance.
I think the reason is that, as you mentioned, DataFrames has a big API and a lot of development effort is put towards finalizing the API in preparation for 1.0. After that there will be much more focus on performance.
In particular, some changes to optimize grouping may have recently been merged but didn't make it into the release by the time this test suite was run, as well as multi-threaded operations, which havent been finished yet, should speed things up a lot.
That said, this new Polars library looks seriously impressive. Congrats to the developer.
According to Propublica, part of the reason for the Fitzgerald crash was bad interface design. https://features.propublica.org/navy-uss-mccain-crash/navy-i...
This might come with sacrificed performance. If you want a different implementation for a different data type to get the best performance, writing one method makes that hard.
> The `Measurements` author very specifically (if I understand this right) implemented a `plot` functionality.
This is a good point. But in that example, note that `Measurements` also works with all of differential equations, and they didn't implement any differential equations specific code, unlike as you pointed out with `plot`. The fact that uncertainty propagates through a differential equation solver is pretty impressive, imo.
With legalized drugs there isn't this trusted arbiter telling you what to do, so you can consent easier when you buy and take them.
I agree that it's fraught and there are tons of ethical parallels that need to be worked out but that's one reason.
foo(x::Baz)
then you will never enter the method that has those if else statements. This is the benefit of multiple dispatch. You can't add more if else statements inside a method that is already written but you can add new methods.It's great to have the flexibility.