Modern Polars: A comparison of the Polars and Pandas dataframe libraries
kevinheavey.github.io
kevinheavey.github.io
For example, from the link, here's how Polars and Pandas handles manipulating data in a subset of a dataframe:
f = pl.DataFrame({'a': [1,2,3,4,5], 'b':[10,20,30,40,50]})
# Polars
f.with_column(
pl.when(pl.col("a") <= 3)
.then(pl.col("b") // 10)
.otherwise(pl.col("b"))
)
# Pandas
f.loc[f['a'] <= 3, "b"] = f['b'] // 10
Its not clear in the Polars approach that the column "b" is being modified. An additional minor nitpick here is the use of when/then/otherwise for their conditional logic. Aren't these just if/else-if/else conditions? It's seems more in line with mathematical/python convention to use if/else... am I missing something?The Pandas equivalent, on the other hand, is much more concise, and more explicit. It also seems more mathematical to me. Polars mutates the dataframe, whereas in Pandas a function is applied to a dataframe indexed like a matrix. Pandas also benefits from it's reliance on symbolic notation, it makes everything visually clearer, whereas in Polars, the use of pl.col("b") and other similar methods contribute to multiple nested brackets and redundant naming calls contributing to less interpretability.
I know there's a lot of thought thats been put into Polars, so I assume I'm missing some of the advantages of the Polars approach, and would appreciate anyone who can shed some light on it.
I do understand, and partially agree, with the idea that indexing in Pandas leads to a lot of bugs. But in the example above, Pandas isn't really using indexing, it's using a boolean map to "index" the values from the same dataframe, so should be fairly robust. Is there a reason why Polars is trying to avoid this kind of filtering in the row/column indices?
> Aren't these just if/else-if/else conditions? It's seems more in line with mathematical/python convention to use if/else... am I missing something?
Yes, they are. But if you look at pandas `f['a'] <= 3` a boolean mask is created on eagerly, on the fly. Pandas has zero chance to do anything clever here.
And yes, `when.then.otherwise` is exactly `if else`, but if `if else` is already a keyword in python so we cannot use them. `when, then, otherwise` are close synonyms.
The benefit of using the `when().then().otherwise()` expression is that it is lazy. We don't do anything until we need to materialize the result. Then the optimizer has a chance to see the query a a whole and determine if the `mask` can be reused, is not needed, should be done somewhere else, etc.
> Polars mutates the dataframe,
Almost all polars methods are pure. There will be no dataframe mutated, but a new dataframe created.
> Is there a reason why Polars is trying to avoid this kind of filtering in the row/column indices.
Yes there is. Ambiguity. I want things to be explicit. So the method names should make clear that you are selecting rows:
`df.filter`
or selecting columns:
`df.select`
or slicing
`df.slice`
In pandas this can all be done with bracket notation. I often read code something like this
`df[foo] = bar` and wondered what kind of datatype was stored into `foo`.
Indexes has the same read complexity. I often read/saw queries that showed a different outcome after a `reset_index` call. I like things to be more explicit. This may cost some keystrokes, but future me/us can more easily understand what is going on.
Isn't this just an implementation detail? It seems like it wouldn't be tough to turn this into syntactic sugar rather than a forced eager evaluation. IE, `f['a'] <= 3` could just as easily evaluate into a computation graph rather than the evaluation of that graph. For example, I could imagine something like so:
```
from polars.dataframe import LazyDataFrame, DataFrame
def fn():
...
ldf = LazyDataFrame(df)
# this mutates the computation graph but doesn't evaluate
ldf.loc[f['a'] <= 3, "b"] = f['b']
df = DataFrame(ldf)
return df
```This is a toy example so I'm not sure if the part around evaluation makes complete sense, but it seems like how pandas eagerly evaluates the frame is a shortcoming of its implementation and model, rather than the syntactic sugar itself.
To be even more specific, this is the way SQLAlchemy does it. You could have something like this:
```
from models import Contact
def fn():
...
# doesn't evaluate; could trivially be done as Contact[Contact.name == 'John']
filtered_contact_exp = Contact.filter(Contact.name == 'John')
# actually evaluates
filtered_contacts = filtered_contact_exp.all()
return filtered_contacts
```And SQLAlchemy knows not to actually trigger the evaluation until you do something like `.all()`. Why not adopt this kind of pattern with Polars?
I’ve used pandas a lot, but I’ve come to the opposite conclusion.
In my experience, these pandas expressions end up being bracket soup, and become increasingly fragile to hold in your head while you try and figure out just which n rows and columns you’re looking at.
Couple that with pandas opaqueness around copy-vs-view and the blurring of lines between API’s for selection, vs API’s for mutation and you get an unpleasant experience.
This particular pandas example is simpler, but it doesn’t take much IME for pandas df’s to end up far more unreadable.
I’ll gladly take polars saner API if it means I don’t have to play “data frame lisp bracket-matching” games ever again.
Modin (https://github.com/modin-project/modin) seems more promising at this point, particularly since a migration path for standing Pandas code is highly desirable.
Additionally, Pandas seems an organically grown API. These days with more experience and more data frame implementations to learn from, it should be possible to do better, something I only partially see when looking at Polars.
Lately, I've used DuckDB to write SQL that manipulates pandas data frames.
Glad I’m not the only one who’s noticed this.
Coupled with this (which leads to: [I’m bad at coding so I won’t spend effort doing it even 1/2 way good]), and pandas having the most abstraction obfuscation of underlying data types, production can become a hot flaming mess that takes months to fix and scale up even linearly w/#of customers :sweat:
Second. I understand that because of the the places I work I encounter this more than 'standard' (say web dev), but it's painful to see how much time and money this attitude seems to cost. Anecdotal rant incming, typical example encountered multiple times: person is really good at math but subpar at programming, but just enough to make it through a PHD (though I'm like 99% sure it's impossible there were no mistakes in that code). Anyway: pretty much every meeting the "I'm bad at programming" and "I don't really know anything about language/framework/thing X" is mentioned and used as if it's a valid excuse for messing up. But the worst part is: instead of just acting on it and learning and trying to improve, there's hardly any progress and without strict guidance anything touched by said persons turns into a trainwreck in no time. Again anecdotal, but I see this much less often with engineers.
I’ve been using pandas for years and had no issues picking up the syntax. Can’t recommend giving it a try enough.
The problem is that it's a huge pile of hacks, exceptions, anti patterns, and regressions.
The API is inconsistent, loose, full of obscure options added as quickfixes.
I found RedFrames [1] recently which wraps Pandas dataframes with a more consistent interface, it's probably what I'd use if I had to write data transformations that had to be compatible with Pandas.
[1] https://dplyr.tidyverse.org/ [2] https://juliadata.github.io/DataFramesMeta.jl/stable/
`pl.from_numpy` and `series.to_numpy` are your friend here. For 1D columns, we often can be zero copy as well.
Besides that we support numpy ufuncs for `Series` and `Expressions`. As OP pointed out:
https://kevinheavey.github.io/modern-polars/performance.html...
Numpy can be used to speed up some functions by utilizing numpy ufuncs. Numpy drops the GIL and therefore they can still be executed in parallel.
1. Polars if data fits in ram
2. Vaex if data do not fit in ram
3. Spark with the dataframe api (koalas) if data do not fit in a computer
Polars is great and delivers as promised
Also dask is more flexible than spark, since it lets you deal with numpy arrays and arbitrary objects better than spark can.
Calling `collect(streaming=True)` on a `LazyFrame` will allow you to process datasets that don't fit into memory. This currently works for groupbys, joins, many functions, filter etc.
We will extend this to sorts and likely other operations as well.
I know its a boring use case, but the challenge with it is that it is a complete waste of money and carbon footprint to use Spark to process a 20 MB CSV or table with few thousand records, but tools like Pandas fall apart when you hit a 50 GB CSV or table with few billion records.
Something more efficient (say, in Rust and not Python or Java) and yet scalable (due to not fitting everything into memory) would be a great help here.
And we will extend functionality for out of core processing. A single node can do a lot!
dask.bag has generic parallel processing capabilities. Query a database, a rest api, something. Then merge into dataframes across dask workers.
1. Pandas if you stay in RAM, if the team and org already know this, but learn about reduced-ram types (eg float32 rather than float64, categorical for strings and dt if low cardinality, new Arrow strings in place of default Object str). Pandas 1.5 has an experimental copy-on-write option for more predictable (but probably still not "predictable") memory usage, try to use a subset of team-agreed functions (eg merge over join) due to varied defaults that'll confuse colleagues (eg inner Vs left and other differences). Buying more ram is normally a cheap (if inelegant) fix.
2. Dask as it is an easy transition from Pandas (and it scales numpy math, arbitrary python non-math functions and lots more), lots of cloud scaling options too. Stays within Python ecosystem for reduced cognitive load. It is probably less resource efficient than Vaex/Polars
3. Ignore Dask and stick with Spark if your team already uses it, as it'll scale to larger workloads and you've taken the cognitive and engineering hit (pragmatism over purity)
Vaex and Polars are definitely interesting (hi Ritchie!), and great if you're doing research and are comfortable with potentially changing APIs but you have no legacy systems to worry about. You might buy yourself a lot of future manoeuvring room. You'll find fewer clues to tricky problems in SO than for Pandas, and have a harder time hiring experienced help.
It depends on what let determine the order. Hiring experience and available content, I wholeheartedly agree with your list.
But if we order by performance/memory efficiency, A single threaded, (eager), library simply will be no comparison and should not top that list. In every TPCH query we ran, polars is orders of magnitudes faster than pandas.
https://www.pola.rs/benchmarks.html
Interopability with legacy systems should not be a concern. Polars is backed by arrow memory and arrow is becoming the default data transformation layer. Other than that, you can easily convert to pandas or numpy. That single copy is often no comparison with the time lost in a pandas join. Polars and pandas can work hand in hand, you don't have to fully replace one.
It is 2023, polars is used in production and is here to stay. IMO it should seriously be considered if performance and consistency is important to you.
I have no doubts that polars is faster than pandas. But the published TPCH results [0] are fairly outdated based on polars-0.13.51 while the current polars is 0.15.13. Are there any plans to refresh the benchmarks?
Right there is the disagreement. Like many (most?) people, all of my data munging is in small/medium data where 10 million+ rows is rare. A multiple of pandas performance will not be noticed for the majority of my operations.
Transitioning to a new api on performance alone is not enough to sway me. After all, I write in Python ;). If I were concerned about better throughput, my first alternative would be Dask - it should give better local performance, but could theoretically scale to enormous data without any code changes.
secondly, your data probably fits in RAM if you actually try. your can get a computer with 60TB of RAM which is an awful lot of data.
Between Polars and Spark Dataframe APIs (not Koalas) as well as the occasional dplyr, I will gladly abandon Pandas.
Merge df 1 and 2 on country, state and forecasted date, then create a new column of the diff between the 2 temp columns, then drop the 2 original temp columns.
In a format where your indexes are forecasted dates on the rows and multiindex of country, state on the columns, you just have to do: df1 - df2
The way I see pandas is a toolkit that lets you easily convert between these 2 representations of data. You could argue that polars is better than pandas for working with data in long format, and that a library like xarray is better than pandas for working with data in the dimensionally relevant structure, but there is a lot of value in having both paradigms in one library with a unified api/ecosystem.Why isn't that saying to assign the value of column b to these locations? Reading the code (and not being a Pandas user) I expected it to be
f.loc[f['a'] <= 3, "b"] = f['a']
Also the "// 10" comment is most confusing as looking at the result it matches 10, 20 & 30 in column b and replaces them with the matching values from column a
[1] https://quarto.org/docs/get-started/authoring/text-editor.ht...
May need to scope if it's worth updating our open-source connectors.