Maker of RStudio launches new R and Python IDE
infoworld.com
infoworld.com
Is it Jupyter envy? Why is it not possible to keep one good product and stay with it?
I wish MatLab licenses weren't so expensive, at this point I'd just buy one and sit all this churn out.
If you had no legacy or compliance requirements, are you going to start a new data project in SAS, R, or Python? Where are you going to find the most talent?
Of course, ML projects and other things that need to result in production-grade models are almost always done in Python. This is currently the most visible form of "data project" due to all the ML/AI hype, but it is far from the only data work going on.
Also try:
gsub('serious', 'hyped', x)
That's a normal follow-up question that you should be able to answer. Otherwise, why are you even commenting?
Any criticism brought up you'd dismiss. Heres one: lack of native 64 bit integers.
I know it's not always easy to extricate oneself but it's helpful to remember that the only way to 'win' is to stop.
If you wouldn't mind reviewing https://news.ycombinator.com/newsguidelines.html and taking the intended spirit of the site more to heart, we'd be grateful.
I know it's not always easy to extricate oneself but it's helpful to remember that the only way to 'win' is to stop.
If you wouldn't mind reviewing https://news.ycombinator.com/newsguidelines.html and taking the intended spirit of the site more to heart, we'd be grateful.
"Does what it intends to do reasonably well" is going to be widely subjective, depending on whether the user's use-case is statistical/life-sciences vs more general purpose coding and relying on many packages; prototyping/experimentation vs production code; whether the user uses base-R, or tidyverse/data.table, etc.
Here are two of those many posts:
* An opinionated view of the Tidyverse “dialect” of the R language (July 5, 2019) https://news.ycombinator.com/item?id=20362626
* The R programming language: The good, the bad, and the ugly (epatters.org, 2018) https://news.ycombinator.com/item?id=35571659 -> https://www.epatters.org/post/r-lang/
If you check e.g. the journal of open source software (which does not have much ML/AI bias), most of the papers are python, with an occasional R and julia submission.
Python is the first language many people are exposed to today. It has a library and tooling for every use case.
I always found that was the group who used R - kind of a use what you are used to until it gets out of step with the remaining workflow.
I also would say that the amount of R I see is far less than python.
1. R is 100% better for analytics work and statistical modelling. There's just no contest.
2. Python is much, much better for data getting (APIs/scraping etc) and dealing with non table-like data. Again, there's basically no contest here.
3. Software engineers hate R (in most cases), which means that it's easier to hand over work for production in Python.
This leads to a situation where it looks like most of the prod-level work is being done in Python, but if you look under the covers you'll discover that most prototyping/analysis/exploration is done in R and then ported to Python if it works.
Like, Python is a great language for lots of things, but it's pretty terrible for exploratory DS work (pandas is like the worst features of base R and base Python mashed together in an unholy hybrid).
There's also the fact that all the NN stuff is predominantly Python, so lots of companies believe that they need Python people, which reinforces the stereotype.
And finally, while I love R, Python has more guardrails, and it's harder to make an unmaintainable mess with it (relative to R). Particularly when people use all the various lazy evaluation packages that the tidyverse has used over the past decade (I once maintained a codebase that used all of these in different places, it was not a fun experience).
Apropos this idea of a vs code competitor, I wish they would spend more effort on existing products. I find quarto frustratingly buggy and meanwhile see no reason to move my workflow from vscode to this new thing. Ymmv
Oh definitely, but at least Python's stdlib is relatively consistent, which helps packages be a little more so.
My favourite example is t.test, which is not a t method for the test class, unlike summary.lm which is.
And there's like 4 different styles of function naming in base & stats alone.
Python has problems (for gods sake, why isn't len a method?) but it's a little more consistent.
I used to think that R was responsible for a lot more of the mess than I now do, having seen the same kind of DS code (and I am a DS) written in both Python and R.
And it would be sweet if R had a pytest equivalent, if I never have to write self.assertEqual again, it'll be too soon.
For a "big data" project, people will probably use Python (though Google apparently retreats from it except in machine learning).
Why could you not use R and C++ or Java though? For example, Arrow has bindings for both, so the argument that Python is needed to shovel/scrape/steal data becomes less and less valid.
Also, a lot of the Posit team are fully "bilingual", it's not like the old guard of academic R contributors. My impression is they appreciate both languages for what they have to offer.
For me, I'm apathetic to the languages, the only thing I care about is the output.
This is absolutely the case. Dplyr syntax is much more intuitive for many use cases than Pandas or Polars equivalents.
One thing I miss from RStudio is the Rmarkdown documents with inline outputs. Jupyter notebooks, even in VSCode, are so needlessly over-engineered and under-featured compared to the elegance of RMarkdown. So I am excited to see what Posit can do to bring that experience to python. My git repos will be thankfull anyway.
It’s basically RMarkdown for SQL
One can see that in the JVM world with java vs scala: people attracted to scala tend to like "cute" DSL, java people tend to be more careful with shiny new features. (This is an oversimplification, of course)
Specifically for dplyr: it looks cute and tends to be easier to use in a REPL setting (you can build your pipeline step by step by running your command, looking at the output, get the command from history, add a step, run again; and at the end you get a single line to copy paste in your script). But if you want to wrap it in a function, it tends to create issues.
It also provides guardrails and encourages best practices which I find a bit to paternalistic and annoying but again I can see the value.
I think most R users would be surprised and just how much tidyverse functionality is hidden in base R but majority of the dplyr versions of functions have at least some intended improvement over the base R versions, and some are a massive improvement in functionality.
For example in a typical script the only tidyverse package I may load besides ggplot2 is tidyr, because the pivot_ wider/longer() functions really do solve a problem that was not fun in base R.
This is already in Quarto! https://quarto.org/docs/computations/inline-code.html#:~:tex....
RMarkdown (Rmd) was recently developed into “Quarto” (Qmd), precisely because they now support Python as well. I’ve used it a bit and it’s excellent.
I don't remember invoking Python from RMarkdowm (maybe you already could in RStudio but I never did), so this will be a welcome addition in this new Posit program.
A data science focused version of VS Code with some kind of notebook sounds rather awesome to me.
The funny thing is that the R in Jupyter actually stands for R (the language). It's Julia, Python and R. No need for envy.
Of course, RStudio/Posit != R (at least in theory)
In practice we have nearly ZERO development for desktop apps, simply because modern desktops are still widget based stuff who was sold a "what we need", against the complexity of classic DocUIs, and then we migrate to modern web witch is a bad DocUI, so developing for the desktop is simply a nightmare. To add features we can't much "use the environment" so all apps tend to try doing anything inside evolving toward unmaintainable monsters no one can handle their codebase and at a certain point in time they became a kind of a framework where "features" became "ideas of someone" without a coherent vision or a target "for the application" like Eclipse from an IDE to a platform some use to code, some others to pay taxes (yes, Italian gov. have made Desktop Telematico witch is a custom Eclipse to fill taxes).
As the classic Greenspun's tenth rule we witness the same: we damn need an OS as a single user-programmable application like Emacs or doing ANYTHING is a nightmare and there is no long lasting solution.
It seems with some JavaScript generated from Java via Gwt. Regardless, I prefer it over VSCode UI.
We have no plans to stop development or maintenance on RStudio, and are committed to it for our users, both paid and community. While Positron and RStudio have some features in common, some R-focused features will remain exclusive to RStudio. If you're currently using RStudio and are happy with the experience, you can continue to enjoy RStudio. RStudio includes 10+ years of applied optimizations for R data analysis and package development.
Cross-posting the FAQ: https://github.com/posit-dev/positron/wiki/Frequently-Asked-...
I prefer coding in VSCode but prefer data exploration in RStudio.
One issue with this is the lack of copilot. Copilot can be installed on VSCodium [1] but it breaks often. The other is MS’s proprietary Remove Development extension that enables a lot of functionality in VSCode. There is an open equivalent but I haven’t tried it [2]
Otherwise, I’d be all in for this!
If I am doing ML most of the time I am using a server GPU with a large dataset. This is one of the use cases that made me use VSCode more than RStudio.
When I have really wanted to use RStudio I use the web version, but its less pleasant than a simple ssh remote.
I think if you are spamming a lot of code that looks like:
thing = params.get('thing')
thong = params.get('thong')
...
thunk = params.get('thonk')
then you probably aren't delivering a lot of value with your code.In fact that overly verbose code could be a liability because it has to be reviewed (possibly multiple times as people look over the code to see how values are getting assigned).
When you start using LLMs you realise how really is boilerplate. Error checking, unit tests, if statements, loops, so much code is boilerplate not just badly written code.
I've heard that one way to use it effectively is writing a detailed comment about what you want to do, then let it suggest the code. I personally don't like those style of comments, so I'd have to:
- enable copilot if I have it disabled (I have a keymap for this in vim)
- write the comment
- carefully review that the suggestion is correct and complete
- accept the suggestion, then go back up to delete the comment
kinda inconvenient, but if I was blanking on a bunch of stdlib functions maybe it would help? But accepting the copilot suggestions doesn't add imports, whereas accepting language server suggestions often does (e.g. with gopls).
this says more about your employer than CoPilot ?
Meanwhile, if said devs do that playing in other languages/frameworks, they’re chastised for not focusing on business value.
I do the same as you: write such comments, let Copilot draft the code, then I delete the comment.
Since the comment is detailed, I think of such use of Copilot as a “pseudocode to code compiler”.
Most of the time it is just autocomplete on steroids. Scanning the suggestion and hitting tab is almost always faster especially if you have to write a lot of repetitive code.
Generating useful documentation.
I work a lot with legacy codebases. Oftentimes I use copilot to explain what a chunk of code is doing, or give it a requirement and see what it generates based on context. At this point I think working with a legacy codebase as a new maintainer would be a lot harder if I didn't have copilot.
imo this is kinda underrated, I use it all the time and it's a real time saver for me. It seems to have a good recall for things you've been doing recently, so it'll make suggestions based on a method I just added or a variable I just declared.
Another fun one I had recently was implementing elo ranking. I'd done some reading so knew what I was doing, but based on the function name and comment it generated the whole thing for me, including all the correct argument names and object properties. Watching it pop up line by line was very cool!
Automatic not being useful was the least of the issue, it used crazy CPU, slowed down the IDE, and drained my battery too.
The other way that works OK is when drafting new code, I can write some comments about what I want to do and trigger Copilot via a keybinding to draft a first version of the code. It can be useful when working with new libraries since I often don’t know what functions to lookup in the documentation.
However it also made it impossible to build the kinds of experiences we wanted to with Positron. Positron has a bunch of top-level UI as well as some integrated/horizontal services that don't make sense as extensions. We built those into the core of the system which is why a fork was necessary.
It's a goal for Positron to be extensible, so it has its own API alongside VS Code's API, and both the R and Python language systems are implemented as extensions.
It's a development burden for sure -- but still an order of magnitude cheaper than trying to build a good workbench surface from scratch.
> You may not provide the software to third parties as a hosted or managed service, where the service provides users with access to any substantial set of the features or functionality of the software.
> You may not move, change, disable, or circumvent the license key functionality in the software, and you may not remove or obscure any functionality in the software that is protected by the license key.
When I open it with positron, it is treated like a text file, at least as far as I can tell by looking at the many icons and pulldown menus.
It is a weird choice, making a new application that cannot handle the key file type from its ancestor.
We do hope to add better GUI tooling for project-level actions; more info here: https://github.com/posit-dev/positron/issues/1486
They've added support in blink as well which is my favorite iOS purchase for productivity on my iPad https://blink.sh/
https://omz-software.com/pythonista/
Pyto is maybe less approachable but more up to date, with clang compiler and LLVM bitcode interpreter:
Juno is Python notebooks:
https://juno.sh/https://juno.sh/
In general I prefer Blink Code:
I used Carnets for the Advent of Code last year. I ran out of steam long before it did.
Because Microsoft does not allow third-party IDEs to access the official VS Code
Marketplace ...
Anyone know why?My wild guess is it means MS doesn't want third parties to build their own VS Code based IDEs (like this one)?
"Visual Studio Code is designed to fracture" - https://ghuntley.com/fracture/
This Microsoft "DevDiv" (see the link) sounds like the classic EEE dressed up as "open", "hip" and with all the right buzzwords.
personally, i see the value of rstudio (and in extension positron) while learning in a course, but i struggle to find its place beyond data exploration.
despite the licensing stuff, if they can provide some based defaults (removing microsoft telemetry "sauce"), it can be an ergonomic way to bring math-sided team members to share the same development platform.