R is a joy if you treat it like Awk
dwrodri.blog
dwrodri.blog
There are lots of valid criticisms of R, but this article doesn't touch on them. It's so off-base that it's the proverbial "not even wrong".
There is definitely more that R offers than what I discuss here. In retrospect, I will be more restrained on my opinions when I have little experience in my pocket. That being said, it was absolutely my intention to present CLI that deviate from R's intended use. There is already plenty out there on R's intended use.
What's that adage that goes something like, "if you want to get an answer on the Internet, don't pose a question..."?
I'm relatively new to technical writing and the discussion from all of these comments (yours included) has been really helpful for guiding how I write future posts.
HN is (in my experience) a lot more cynical and straightforward than other places[1]. It's something I've learned to appreciate and also take it with a grain of salt.
I agree with you where some here don't. I think there are often better tools for any single data cleaning task.
R's strength is being second best at an enormous range of tasks (and often being first to get new techniques) and packaging that with analysis and visualization.
Not a specific dig against you, but I find it useful to write (and say) anything with the assumption that the author I'm (hypothetically) addressing is a direct witness.
It really helps with online civility :)
I use Rscript all the time. I was taught to use it by one of R's developers. While not typical, it is absolutely an "intended" use case.
In fact, if you would like to learn more, Software Carpentry has an entire module on using R as a CLI: https://swcarpentry.github.io/r-novice-inflammation/05-cmdli...
Cleaning a broken csv or whatever: no, it is crap for that. You use awk/sed/tr and all that for such problems.
If you're the type who don't want to deal with R, I guess you can use it from the CLI. A couple of the R deploys I've done work like this.
The real problems with R are .... oh man .... so many. R inferno covers a lot of them as a language/environment. Weak database connectivity is another one. The thing which makes me batshit is the nodejsbro-ification of the package management system. Aka people chaining together things like node works; R's package manager isn't designed for this. But also the way code, packaged and otherwise simply rots between the many, many upgrades.
You could probably run and deploy scikit learn/pandas based code from 5 years ago without much problem. In R, you have to make a build with the salted package dependencies ... and for all I know stuff it in docker.
Anyway unlike python, it basically has every data transformation and statistical tool under the sun. I guess this is the price we pay.
EDIT: And yes, as citrate05 says, it's just an additional set of libraries. There's no changes to the language itself.
Edit: This of course goes hand-in-hand with the claim that it is easier/faster to write R scripts. If you're not familiar with it, the tidyr and dplyr packages in particular (part of the tidyverse) are fantastic in the verbs they provide for thinking about data cleaning.
R has inbuilt great parallel tools (check for example the doSnow and future frameworks);
the best packages for data manipulation are mostly written in C (for example data.table and a good part of the tidyverse);
and with frameworks like Drake you can easilly create a Dag out of it that can process complex iterations millions of times. Check the uses of the Rcpp package that makes interfacing C code to R a breeze.
But of course, if you were comparing R to a pure compiled language, you are out of luck.
managing python dependencies is no fun either, tbf
I understand what the author is stating, but I just feel like this is from inexperience with R and ignoring the vast amount of packages available within it that are specifically targeted at data science. There are some valid issues and criticisms of R, but I think this article only focuses on the application of R in a single context. A significant portion of data science is about cleaning the data and an entire suite of packages known as tidyverse solves these problems (for me anyways) while also being very simple and easy to understand. I mean tidyverse supports piping, which is exactly what this article is saying to use.
Obviously your mileage may vary, but this post irks me the wrong way.
Pretty much the argument of using apply vs for loops. Most of the time, apply is going to be significantly more computationally efficient. However, some people think of the problem in different context than others, and if it works for them then it works.
For those who come from the world of large enterprise statistical and data reporting tools such as SAS, awk shares an exceedingly strong resemblance to the SAS DATA step, whilst R effectively provides a host of analysis and graphics tools the correspond to numerous other SAS procedure and products.
The hacks for pipelining R are cool and useful. Thanks.
> I used to resent R, it was shoved upon me as a strange tool that promised to replace Python, but failed miserably.
well I'm glad the author found a way to make R work for their purposes, but this just reeks of inexperience with R..
What motivated me to write this post was the lack of discussion of R in this use case. I actually stumbled into using "Rscript -e" one liners while looking to do basic stats in a Linux CLI.
That being said, I still stand by my point. Taking data from "out in the wild" (log files, tarballs of images, unstructured text) and making use of it can be frustrating in R because cleaning up edge cases, removing unwanted data, and getting everything into the correct container/type often involved unintuitive chaining of function calls. This is coming from the perspective of someone who worked with Python/awk/sed prior to being exposed to R.
If you have good counterarguments, I'd be more than happy to hear them and addresss them in an end note in the post.
I've working on python for my current project and constantly longing for R's syntax specifically for cleaning data. It's so much better.
Just last night, I wanted to perform an anti-join to find discrepancies between two data sets for debugging
One line in R:
anti_join(a_tibble, another_tibble, by = c("id_col1", "id_col2"))
Witchcraft in Pandas: (from stack overflow): Method 1
# Identify what values are in TableB and not in TableA
key_diff = set(TableB.Key).difference(TableA.Key)
where_diff = TableB.Key.isin(key_diff)
# Slice TableB accordingly and append to TableA
TableA.append(TableB[where_diff], ignore_index=True)
Method 2:
rows = []
for i, row in TableB.iterrows():
if row.Key not in TableA.Key.values:
rows.append(row)
pd.concat([TableA.T] + rows, axis=1).T
I actually fired up R and re-imported the .csv data just for this. Took 15 secs while my colleague was still stuck debugging his own weird for loops.I have a hunch that there could be a really good follow-up post to this that takes these R hacks to the next level by extending it to work better with pre-structured text (where R really shines) and CSV files as arguments.
dataframe.join(
something_else,
how='left_anti'
)
pandas is such a shitshow. Every time i use it, im in a world of pain googling the finicky syntax for selecting columns, aggregating, filtering. I never touched R but pandas is so terrible for me. Nowadays it's either raw numpy arrays, plain sql or pyspark...May I ask what you use R for? I like learning languages for fun and I've been meaning to do some NBA analytics stuff. I'd love to have a REPL style interface to just do one-off math and analytics or short scripts. I haven't dug into the data science stuff yet but I'm disinterested in Python for some reason (maybe because I used to write Ruby for a living).
R is a thoroughbred at doing data analysis from your laptop. It's bad at living on a server and operating any sort of app.
For your use case (NBA analytics) I think it is a great tool and would recommend R. In fact, here are a couple packages/tutorials specifically for doing that: http://asbcllc.com/nbastatR/ https://blog.revolutionanalytics.com/2016/09/analyzing-nba-b...
np.setdiff1d(array_one, array_two)Sounds like a SQL use case to me:
select * from table_1 except select * from table_2 (one of a half dozen ways you could do this)
This is a fair response.
outer_join = df1.merge(df2, left_on='id_col1', right_on='id_col2', indicator=True, how='outer')
outer_join[outer_join['_merge'] == 'right_only']R is not made for cleaning up weird text files. Yes of course it can be done but that’s like the joke that everything is within walking distance if you have enough time. I recently had to use R to fix a 50gb csv where 10 of the columns were long json strings and needed to turn that into a data frame. That experience alone made me buy a book on awk and sed
first off I found it interesting to frame R as being presented as an improvement over python. should be the other way around. python for DS came second, and was supposed to improve over R
anyway I'm not sure i can even begin to discuss your use cases, not having that much relevant experience there. I use R (and python) for general analytics tasks and for building production models. in these more traditional DS environments I strongly believe R is far superior to python for data munging and visualization. When I say this I am comparing data.table (and to a lesser extent, tidyverse) to pandas. I don't even want to get started on everything I hate about pandas. So while we are both "cleaning data" you seem to be talking about a stage before someone like me would even be looking.
My language must have been ambiguous in the article. I intended to frame R as arriving second because that was my personal experience. My writing philosophy for blog posts is that my personal opinions and experiences should stand out from the technical detail. My reasoning is that even for those who disagree with the personal content will be able to discern for themselves the value of the post.
Disagree. R has an incredible ecosystem for parsing, cleaning, and manipulating data. Even ignoring the tidyverse, base R provides more than enough functionality to clean and analyze data--you just need to spend some time learning it. If you use the tidyverse, it's even easier. The only other ecosystem that comes close to R is Julia, which was designed taking many of the best parts of R into consideration.
The problem with logs is that lines have different number of columns depending on what that line is logging.
What a given line is logging can usually be determined by a regex (e.g. ends with a ip-address, starts with this words, etc., etc)
I'm genuinely curious, as this is a use-case where I have a lot of difficulty using python or R. I can see how grep + awk are strong as a preprocessor, as they can scan through a larger than memory file, and select columns from rows that match your criteria.
However, if your data fits in memory and you have a non-trivial analysis to perform, R is a great choice. I had a file with tabular data interspersed with metadata at random points and it was straightforward to parse and store the data in a custom data structure.
What is the Julia equivalent of the tidyverse
For example, in Julia you can work with the dataframes in a SQL/LINQ/dplyr fashion like this:
x_thread = @linq df |>
transform(y = 10 * :x) |>
where(:a .> 2) |>
by(:b, meanX = mean(:x), meanY = mean(:y)) |>
orderby(:meanX) |>
select(:meanX, :meanY, var = :b
(sorry not sure how to format code properly on HN)Theres also a version of ggplot, although I’m not sure if its a rewrite ot just a wrapping around the R version.
Edit: Actually, this I believe is a part of the DataFramesMeta.jl package... So the answer would be rather “most of the things are baked into Julia and when they arent, there’s a package for that”.
I haven't found a good way around very inconsistently formatted csv files in any language (a row only represents column 3,4 and 6 if it starts with a comma, all other rows have all columns, but are space separated and values may contain comma, etc, etc)
If you want to use R for data analysis I suggest the R package Tidyverse or use a product like OpenRefine.
If you just need to run summary statistics with an application then using Python is just fine.
I also get what the commenters are saying, because R is a useful interactive language too, and it's also pretty good for data cleaning. Although it's significantly slower than Python, which is why I do all cleaning that cuts down the data before loading it into R.
As an example, I generate some benchmarks with every release of Oil:
https://www.oilshell.org/release/0.7.pre5/benchmarks.wwz/osh...
and the tables are manipulated with R, but running the benchmarks is done with shell, and creating the HTML is done with Python:
https://github.com/oilshell/oil/blob/master/benchmarks/repor...
(and yes this page is meant to be motivation to speed up my principled but slow shell parser)
R and tidyverse are the best tools for manipulating tables by far. I wrote an intro here:
What Is a Data Frame? (In Python, R, and SQL) http://www.oilshell.org/blog/2018/11/30.html
There is some performance issues, but I am OK to trade that off for the convenience of using "try:" rather than R's "tryCatch". Having tryCatch as a function rather than built into the syntax is unacceptable to me. But there are some libraries in R that don't have elegant alternatives in Python or which you have multiple options or perhaps not the time to unlearn.
Like I do all my other data sources, as code blocks inside an emacs-org notebook. If you are doing data science, you quickly find that it's management and combination of the various particular projects that becomes the most daunting (imho), and your data science notebook becomes the most important part of that organization. In that arena for me it's pretty much either jupyter or emacs org-mode.
What does awk give you that these don't?
https://github.com/dkogan/feedgnuplot/ https://github.com/dkogan/vnlog/
Joke. I am a heavy R user.
Oh and there’s also Rio if you want to explore injecting R into your command line workflow.
awk '/Recovery time:/{print $2}' output.log | boxplotI'm quite familiar with asciinema[2] as well, which I've considered using for creating animated examples. In general, I like to err on the side of caution, and optimize my website for people with poor/metered connections.
Also, I didn't expect this to get this much attention! Mea culpa for not putting more work into the post itself.
1 = https://sw.kovidgoyal.net/kitty/index.html 2 = https://asciinema.org/
What I was trying to say is: how about extending your blog post with one line to show people how to make the boxplot appear in their terminal as soon as they run the command? (I think at the moment it's going to appear as a PDF file named Rplots.pdf in the current directory, right?)
UPDATE: fixed color palette issue, although it still looks quite bad.
That is a great quote!