Here's an example from the docs
flights[carrier == "AA",
lapply(.SD, mean),
by = .(origin, dest, month),
.SDcols = c("arr_delay", "dep_delay")]
that's clearly less clear than SQL SELECT
origin, dest, month,
MEAN(arr_delay), MEAN(dep_delay)
FROM flights
WHERE carrier == "AA"
GROUP BY arr_delay, dep_delay
or pandas flights[filghts.carrier == 'AA'].groupby(['arr_delay', 'dep_delay']).mean()
But once you get used to it data.table makes a lot of sense: every operation can be broken down to filtering/selecting, aggregating/transforming, and grouping/windowing. Taking the first two rows per group is a mess in SQL or pandas, but is super simple in data.table flights[, head(.SD, 2), by = month]
That data.table has significantly better performance than any other dataframe library in any language is a nice bonus! SELECT
origin, dest, month, AVG(arr_delay), AVG(dep_delay)
FROM flights
WHERE carrier == 'AA'
GROUP BY origin, dest, month
and flights[flights.carrier == 'AA'].groupby(['origin', 'dest', 'month'])[['arr_delay', 'dep_delay']].mean()flights.groupby("month").head(2)
Not only is does this have all the same keywords, but it is organized in a much clearer way to newcomers and labels things to look up in the API. Whereas your R code has a leading comma, .SD, and a mix of quotes and non-quotes for references to columns. You even admit the last was confusing to learn. This can all be crammed in your head, but not what I would call thoughtfully designed.
| Date | EventType |
and I want to find the count, and the first and last date of an event of a certain type happening in 2020: events[
year(Date) == 2020L,
.(first_date = first(Date), last_date = last(Date), count = .N),
EventType
]
Using first and last on ordered data will be very fast thanks to something called GForce.When exploring data, I wouldn't need or use any whitespace. How would your Pandas approach look like?
mask = events["Date"].year == 2020 events[mask].groupby("EventType").agg(first_date=("Date", min), last_date=("Date", max), count=("Date", len))
Anyway, I don't understand why terseness is even desirable. We're doing DS and ML, no project never comes down to keystrokes but ability to search the docs and debug does matter.
- How many events by type
- When did they happen
- Are there any breaks in the count, why?
- Some statistics on these events like average, min, max
and so on. Terseness helps me in doing this fast.
R is a language by and for statisticians. Python is a programming language that can do some statistics.
- Lots of high quality statistical libraries, for one thing.
- RStudio's RMarkown is great; I prefer it to Jupyter Notebook.
- I personally found the syntax more intuitive, easier to pick up. I don't usually find myself confused about the structure of the objects I'm looking at. For whatever reason, the "syntax" of pandas doesn't square well (in my opinion) with python generally. I certainly want to just use python. But, shrug.
- The tidyverse package, especially the pipe operator %>%, which afaik doesn't have an equivalent in Python. E.g.
with_six_visits <- task_df %>%
group_by(turker_id, visit) %>%
summarise(n_trials = n_distinct(trial_num)) %>%
mutate(completed_visit = n_trials>40) %>%
filter(completed_visit) %>%
summarise(n_visits = n_distinct(visit)) %>%
mutate(six_visits = n_visits >= 6) %>%
filter(six_visits) %>%
ungroup()
Here I'm filtering participants in an mturk study by those who have completed more than 40 trials at least six times across multiple sessions. It's not that I couldn't do the same transformation in pandas, but it feels very intuitive to me doing it this way.- ggplot2 for plotting; its really powerful data visualization package.
Truthfully, I often do my data text parsing in Python, and then switch over to R for analysis, E.g. python's JSON parsing works really well.
(The below is off topic, but I don't use R so I'd love to know whether I'm reading the code correctly)
"Here I'm filtering participants in an mturk study by those who have completed more than 40 trials at least six times across multiple sessions."
A user with this pattern of trials seems like they would fit the above definition:
Session 1: 82 trials Session 2: 82 trials Session 3: 82 trials
But the code seems to want 6 distinct sessions with >40 trials each. Have I misunderstood?
Also, is 'mutate' necessary before 'filter' or is that just to make the intent of the code clearer to your future self?
There were 50 trials in each session; so I counted a session completed if they did more than 40 in that session. They needed to have completed at least six sessions.
The mutate is unnecessary. I forget why I did that.
%>% is now simply >|
Hence my question:
What's the advantage of Python if you already know R?
I've heard they have similarities. Is there anything Python does better than R in terms of statistical analysis, charting, etc.?
AFAIK in statistical modelling Python is better only in neural networks, so if you do not need to do fancy things with images, text, etc. you do not need Python. R is still the king.
In terms of charting and dashboards, I would say that if you work high level R and Python are both pleasant. R has ggplot, but Python has Plotly Express. R has Shiny, but Python has Dash and Streamlit. You can do great with both.
In Python, there is statsmodels. Here, you'll find a lot of GLM stuff, which is sort of an older approach. Modern inferential statistics, if not just Bayesian, is usually in the flavor of semi-parametric models that rely on asymptotics.
As R is used by professional researchers, it is simply more on the edge of things. Python has most of the "statistics course" schoolbook methods, but not much beyond that.
For example, it has become very common to have dynamic panel data which require dynamic models. Now if you want to do a Blundell-Bond type model in PYthon you have to... code it yourself using GMM, if it exists even.
For statistics, that's pretty much like saying you have a Deep Learning package that maybe has GRU but no transformer modules at all. So yeah, you can code it yourself. Or you use the other one.