Pandas Illustrated: Visual Guide to Pandas
betterprogramming.pub
betterprogramming.pub
https://www.pearson.com/en-us/subject-catalog/p/pandas-for-e...
https://betterprogramming.pub/pandas-illustrated-the-definit...
One of my daughters is a panda bear fanatic and I thought this would be a resource I could share with her.
I mean, data analysis is useful and all, but not what the heart wanted at the moment.
If your code feels like it dealing with a matrix and not a table, it’s probably doing something funny.
Pretty much everything you need in pandas is as performant as you ought to need for doing tabular data manipulation in Python. Except dataframe.apply
df = pd.DataFrame({"foo": np.random.randn(100000)})
pandas map: df["foo"].map(lambda x: x * 2)
18.1 ms ± 109 µs per loop (mean ± std. dev. of 7 runs, 100 loops each)pandas apply:
df["foo"].apply(lambda x: x * 2)
17.9 ms ± 46.6 µs per loop (mean ± std. dev. of 7 runs, 100 loops each)Vectorised function, using underlying numpy operations:
df["foo"] * 2
267 µs ± 11.8 µs per loop (mean ± std. dev. of 7 runs, 1000 loops each)Series.map is not compiling your lambda's to C and running it. If there is a built-in method available it usually will be faster. Notable exception are pandas str methods which devolve into Python code but generally with more overhead than map/apply.
Vectorized, choice between lazy optimization and eager.
And if it’s bigger you probably ought to be using a different toolset than Python altogether.
In my experience it is exceedingly rare to find a situation where you have > 1M rows of data and need to do tabular data manipulation, and it’s not coming from a managed database sort of setup where the manipulation could be better done in the database.
Pandas built ins can cover everything with good performance except dataframe.apply imo
That and using dicts as maps.
> Pandas is an industry standard for analyzing data in Python.