(to be clear, I'm a fan of "DuckDB via Ibis", I'm much less a fan of statements to the effect of "use Polars via Ibis instead of Polars directly", as I think they miss a lot :wink:)
DuckDB was chosen as the default because Polars was too unstable to be the default backend two years ago.
Pandas turns 10x programmers into 1x programmers.
The Modern Pandas blog series from Tom Augsburger is excellent, and the old Wes McKinney book (creator of pandas) gave a good glimpse into how to effectively use it.
I don't think many people learn it correctly so it gets used in a very inefficient and spaghetti script style.
It's definitely a tool that's showing it's age, but it's still a very effective swiss army knife for data processing.
I am glad we're getting better alternatives, Hadley Wickams tidyverse really showed what data processing can be.
I’ve tried to make mockups of similarly inspired API and I think they did a great job with clever tricks like the `_` reference (does that clash with IPython/Jupyter’s use?).
I do wonder how they manage compatibility with so many backends, it seems like many features will be directly tied to your backend (e.g. Trino using Java Pattern for regex vs BigQuery using Re2) in hard to explain ways. But maybe that’s not a big concern, because very often you’re only going to be using a couple of backends.
I’ll have to try Ibis out for myself, it looks like it can unify a lot of the work I have to do. I’ve moved away from pandas for all but the last mile of computation, so this might be a direct good SQL replacement.
However, if you import it from Ibis then it ceases to be used as "most recent result" and remains the ibis underscore object (unless of course you explicitly assign it to something else).
Regarding backend compatibility, there are definitely a few kinds of things that we don't currently abstract over. One is regular expression syntax and another is floating point math (e.g., various algebraic properties that are violated that result in slightly different outputs).
Hope you give it a go, and please report issues at https://github.com/ibis-project/ibis.
I think that’s the right policy to take, and I did notice the support matrix on your website which addresses my earlier question:
Look, I applaud your skill, but at some point even a master craftsman realizes that the swiss army knife may not be the best tool, and a leatherman offers certain advantages.
From my experience the biggest impediment to using R in production is many orgs don’t have a blessed way to run it.
R is my favourite language for data processing, the manual section Computing on the Language[1]is why R is such an ergonomic tool. I had hoped Julia would catch up, but Julia’s macros are not comparable in their depth.
I think pandas is probably the data equivalent of editing files using default vim or processing data with awk.
[1] https://rstudio.github.io/r-manuals/r-lang/Computing-on-the-...
It's faster than pandas in some cases and folks should put it into production immediately!
Once I adopted method chaining, a lot of the issues that I had with pandas in the past due to poor style (e.g. SettingWithCopyWarning) pretty much disappeared.
- Consistent Expression and SQL-like API.
- Lazy execution mode, where queries are compiled and optimized before running.
- Sane NaN and null handling.
- Much faster.
- Much more memory efficient.
I'd suggest less snark, you're not doing yourself any favors.
Why do people feel the need to jump in and police tone like this? Who are you? You're not doing yourself any favours, either.
>common literary technique to emphasize a point.
"Common" is a stetch, and who cares.
From the community guidelines(https://news.ycombinator.com/newsguidelines.html): "Be kind. Don't be snarky. Converse curiously; don't cross-examine. Edit out swipes."
> Who are you?
A member of the community.
Did you know that more information is communicated via tone than words?
The original purpose of my comment was to, in a light-hearted way, point out redundant information, not to incite a flame war about tone on HN.
This thread reminds me of why I don't come here very often.
I dont care so much about the memory and CPU stuff, I mostly leave the heavy lifting to an SQL engine.
Although the Null handling seems very compelling, I guess it comes at a cost of incompatibility with existing libraries, otherwise Pandas would have implemented it as well?
I am curious about the SQL api though.
If it were better, we'd use it internally in Ibis for the Polars backend implementation.
If you're going down the mixed SQL, DataFrame API route then Ibis is probably the best solution out there for that.
I work on Ibis, so take what I say with a grain of salt. There may yet be other libraries out that there that have similar functionality.
And by now I know that very well.
Like someone-screaming-in-my-ears-know.
I am starting to think that Polars is showing all the signs of a hype or a cult.
I am still not convinced, particularly since the community feels more like a marketing department than someone who wants to genuinely help.
I can do that thing you describe with SQL.
If you mean whether I run it distributedly a la Spark then no. If you mean whether I test it on various machines with different RAM sizes then yes.
> I dont care so much about the memory and CPU stuff, I mostly leave the heavy lifting to an SQL engine.
Well, I care. Both pandas and polars are, to my view, single-machine dataframe library, so the memory and CPU constraints are rather stringent.
My comparison is based solely on my experience: reading csv files that are 20% to 50% the size of RAM, pandas takes (or errors out after) 2 to 10 minutes, while polars finishes in 20 seconds. Queries in pandas are almost always slower than polars.
But reading your comment, it seems you and I have different use cases for dataframe libraries, which is fine. I mostly use them for exploratory analysis, so the SQL api is not that much of a plus to me, but the performance is.
Many cloud providers now offer serverless SQL and Spark capacities (serverless=no set up for you). This is the magnitude change for me.
With pandas you can maybe process 10 million rows, with polars maybe 50 million. But with a distributed service maybe 100 times more?