Give me a better query language and I will gladly drop Pandas and Data.Table.
More usefully, most databases allow you to write user-defined functions in other languages, including imperative SQL dialects.
Pyspark on the other hand just sticks in my brain, somehow. Chained pyspark method calls looks much neater.
Essentially, `inplace=True` rarely actually saves memory, and causes problems if you like chaining things together. The people who maintain the library/populate the discussion boards are generally pro-chaining, so `inplace` is slowly and quietly on its way out.
The documentation should clearly say if the inplace argument causes an internal copy, because the availability of it implies that it doesn't. I've used inplace many times with Pandas because I've had code that I know is working with large amounts of data and I've thought "well I probably shouldn't chain these and cause tons of unnecessary allocations just to have prettier code".
It's such a shame that python doesn't have a better DF library.
Base R is OK, but dplyr is magical.
For instance, integer indexing in base R is df[row,col] rather than the iloc pandas stuff.
plot, print and summary (and generic function OOP more generally is really underappreciated).
Python is a better programming language, but R is a better data analysis environment.
And dplyr is an incredibly fluent DSL for doing data analysis (not quite as good for modelling though).
Seriously, I read the original vignette for dplyr in late 2013/early 2014 and within two weeks I'd switched most of my new analytical code over to it. So very, very good. Less idea-impedance match than any other environment, in my experience.
1. Not so easy way to rename columns during aggregation
2. The group by generates its own grouped by data and hence you almost always need `reset_index`
3. Sometimes group by can convert a dataframe to series
4. Now `.loc` has provided bit consistent indexing/slicing, but earlier you had `.ix` `.iloc` and what not
These are something I can remember from top of my head. Of course all of these have solutions, but it makes pandas much more verbose. In R, these are just much more succinct.
https://www.rstudio.com/resources/rstudioglobal-2021/bringin...
Wait... I had no idea dplyr's database support was this good. https://db.rstudio.com/dplyr/
I learned Pandas first. I have no issue with indexing, different ways of referencing cells, modifying individual rows and columns, numerous ways of slicing and dicing. It gets a little sprawling but there's a method to the madness. I can come back to it months later and easily debug. With SQL, it's just madness and 10x more verbose.
That's fair, I was using the opportunity to complain about pandas and didn't point out all the problems SQL has for some tasks. What I really want is a dataframe library for Python that's designed in a more sensible way than pandas.
> long lines of().chained().['expressions'].like_this(0)
a _bad thing_?
IMHO these pandas chains are easy to read and communicate quite clearly what's being done. If anything, I've found that in my day-to-day while reading pandas I parse the meaning of those chains at least as efficiently as from comments of any level of specificity, or from what other languages (that I have had experience with) would've looked like.
...and then BOOM, pandas chain! One statement containing 29+ concepts and their implications.
Followed by more low density code. The rollercoaster leads to complaints because it feels harder. Not because of any actual change in difficulty.
/that's my current working theory, anyway