Is it Jupyter envy? Why is it not possible to keep one good product and stay with it?
I wish MatLab licenses weren't so expensive, at this point I'd just buy one and sit all this churn out.
Is it Jupyter envy? Why is it not possible to keep one good product and stay with it?
I wish MatLab licenses weren't so expensive, at this point I'd just buy one and sit all this churn out.
Also, a lot of the Posit team are fully "bilingual", it's not like the old guard of academic R contributors. My impression is they appreciate both languages for what they have to offer.
For me, I'm apathetic to the languages, the only thing I care about is the output.
This is absolutely the case. Dplyr syntax is much more intuitive for many use cases than Pandas or Polars equivalents.
One thing I miss from RStudio is the Rmarkdown documents with inline outputs. Jupyter notebooks, even in VSCode, are so needlessly over-engineered and under-featured compared to the elegance of RMarkdown. So I am excited to see what Posit can do to bring that experience to python. My git repos will be thankfull anyway.
It’s basically RMarkdown for SQL
One can see that in the JVM world with java vs scala: people attracted to scala tend to like "cute" DSL, java people tend to be more careful with shiny new features. (This is an oversimplification, of course)
Specifically for dplyr: it looks cute and tends to be easier to use in a REPL setting (you can build your pipeline step by step by running your command, looking at the output, get the command from history, add a step, run again; and at the end you get a single line to copy paste in your script). But if you want to wrap it in a function, it tends to create issues.
It also provides guardrails and encourages best practices which I find a bit to paternalistic and annoying but again I can see the value.
I think most R users would be surprised and just how much tidyverse functionality is hidden in base R but majority of the dplyr versions of functions have at least some intended improvement over the base R versions, and some are a massive improvement in functionality.
For example in a typical script the only tidyverse package I may load besides ggplot2 is tidyr, because the pivot_ wider/longer() functions really do solve a problem that was not fun in base R.
This is already in Quarto! https://quarto.org/docs/computations/inline-code.html#:~:tex....
RMarkdown (Rmd) was recently developed into “Quarto” (Qmd), precisely because they now support Python as well. I’ve used it a bit and it’s excellent.
I don't remember invoking Python from RMarkdowm (maybe you already could in RStudio but I never did), so this will be a welcome addition in this new Posit program.
The funny thing is that the R in Jupyter actually stands for R (the language). It's Julia, Python and R. No need for envy.
Of course, RStudio/Posit != R (at least in theory)
If you had no legacy or compliance requirements, are you going to start a new data project in SAS, R, or Python? Where are you going to find the most talent?
Of course, ML projects and other things that need to result in production-grade models are almost always done in Python. This is currently the most visible form of "data project" due to all the ML/AI hype, but it is far from the only data work going on.
Also try:
gsub('serious', 'hyped', x)
That's a normal follow-up question that you should be able to answer. Otherwise, why are you even commenting?
Any criticism brought up you'd dismiss. Heres one: lack of native 64 bit integers.
I know it's not always easy to extricate oneself but it's helpful to remember that the only way to 'win' is to stop.
If you wouldn't mind reviewing https://news.ycombinator.com/newsguidelines.html and taking the intended spirit of the site more to heart, we'd be grateful.
I know it's not always easy to extricate oneself but it's helpful to remember that the only way to 'win' is to stop.
If you wouldn't mind reviewing https://news.ycombinator.com/newsguidelines.html and taking the intended spirit of the site more to heart, we'd be grateful.
"Does what it intends to do reasonably well" is going to be widely subjective, depending on whether the user's use-case is statistical/life-sciences vs more general purpose coding and relying on many packages; prototyping/experimentation vs production code; whether the user uses base-R, or tidyverse/data.table, etc.
Here are two of those many posts:
* An opinionated view of the Tidyverse “dialect” of the R language (July 5, 2019) https://news.ycombinator.com/item?id=20362626
* The R programming language: The good, the bad, and the ugly (epatters.org, 2018) https://news.ycombinator.com/item?id=35571659 -> https://www.epatters.org/post/r-lang/
If you check e.g. the journal of open source software (which does not have much ML/AI bias), most of the papers are python, with an occasional R and julia submission.
Python is the first language many people are exposed to today. It has a library and tooling for every use case.
I always found that was the group who used R - kind of a use what you are used to until it gets out of step with the remaining workflow.
I also would say that the amount of R I see is far less than python.
1. R is 100% better for analytics work and statistical modelling. There's just no contest.
2. Python is much, much better for data getting (APIs/scraping etc) and dealing with non table-like data. Again, there's basically no contest here.
3. Software engineers hate R (in most cases), which means that it's easier to hand over work for production in Python.
This leads to a situation where it looks like most of the prod-level work is being done in Python, but if you look under the covers you'll discover that most prototyping/analysis/exploration is done in R and then ported to Python if it works.
Like, Python is a great language for lots of things, but it's pretty terrible for exploratory DS work (pandas is like the worst features of base R and base Python mashed together in an unholy hybrid).
There's also the fact that all the NN stuff is predominantly Python, so lots of companies believe that they need Python people, which reinforces the stereotype.
And finally, while I love R, Python has more guardrails, and it's harder to make an unmaintainable mess with it (relative to R). Particularly when people use all the various lazy evaluation packages that the tidyverse has used over the past decade (I once maintained a codebase that used all of these in different places, it was not a fun experience).
Apropos this idea of a vs code competitor, I wish they would spend more effort on existing products. I find quarto frustratingly buggy and meanwhile see no reason to move my workflow from vscode to this new thing. Ymmv
Oh definitely, but at least Python's stdlib is relatively consistent, which helps packages be a little more so.
My favourite example is t.test, which is not a t method for the test class, unlike summary.lm which is.
And there's like 4 different styles of function naming in base & stats alone.
Python has problems (for gods sake, why isn't len a method?) but it's a little more consistent.
I used to think that R was responsible for a lot more of the mess than I now do, having seen the same kind of DS code (and I am a DS) written in both Python and R.
And it would be sweet if R had a pytest equivalent, if I never have to write self.assertEqual again, it'll be too soon.
For a "big data" project, people will probably use Python (though Google apparently retreats from it except in machine learning).
Why could you not use R and C++ or Java though? For example, Arrow has bindings for both, so the argument that Python is needed to shovel/scrape/steal data becomes less and less valid.
In practice we have nearly ZERO development for desktop apps, simply because modern desktops are still widget based stuff who was sold a "what we need", against the complexity of classic DocUIs, and then we migrate to modern web witch is a bad DocUI, so developing for the desktop is simply a nightmare. To add features we can't much "use the environment" so all apps tend to try doing anything inside evolving toward unmaintainable monsters no one can handle their codebase and at a certain point in time they became a kind of a framework where "features" became "ideas of someone" without a coherent vision or a target "for the application" like Eclipse from an IDE to a platform some use to code, some others to pay taxes (yes, Italian gov. have made Desktop Telematico witch is a custom Eclipse to fill taxes).
As the classic Greenspun's tenth rule we witness the same: we damn need an OS as a single user-programmable application like Emacs or doing ANYTHING is a nightmare and there is no long lasting solution.
A data science focused version of VS Code with some kind of notebook sounds rather awesome to me.