A bird's eye view of Polars
pola.rs
pola.rs
[1]: https://quarto.org/
But I do imagine someone out there has a data engineer hat they put on, and starts rewriting Python calls to polars into pure rust code, compiling and then deploying.
Here's how I work.
* Hack together a PoC with python, sed, awk, grep, cut, xsv etc.
* Clean that up, run it on larger sample sets (samples made with said sed/awk/cut etc)
* Attempt to run it on the full dataset.
* Rewrite it in rust.
Step 2 and 3 are hit-or-miss in python. I find it near impossible to do any refactoring without static types and/or tests. And quite often, I'm looking at a run for over an hour to have it crash on that one broken line. Whereas the same happens in the rust version in seconds: crucial for my trial-and-error style of building.
So: Python because I must, rust, as soon as it's clear what I'm going to do.
Polars biggest downfall is that pandas/matplotlib are so ubiquitous in data science and polars just plays so differently than pandas including using hvplot as its default plotting package, etc. It really is trying to do much of its ecosystem exactly how it wants to maximize productivity, speed, etc. This may slow down the adoption of it, but hopefully it will push others in the better direction.
huh? Polars doesn't have a default plotting package, it's always something you add. matplotlib supports polars out of the box through the dataframe exchange protocol, which is ubiquitous enough due to the proliferation of other dataframe libraries (dask, vaex, modin/ray) that you really get about every other tool in the ecosystem with polars for free.
1. load and process / aggregate in polars to get the smaller dataset that goes into your plot. 2. df.to_pandas() 3. apply your favourite vis library that works with pandas.
There's no use case i can think of where building a data viz interface more specific to polars than this is beneficial or necessary.
What issues are you running into with the Python ecosystem?
I ask because I'm in the middle of writing Effective Polars, and my experience is that many things like Xgboost, matplotlib, etc work fine with Polars. (Sadly/oddly/ironically some libraries now have issues with Pandas 2 pyarrow types but work with Polars.)
I think my only gripe is that the Rust API seems (to me at least) to be less well documented than the Python API. I guess Python is the de facto data science language, so maybe that explains it.
This is a huge difference in productivity, especially when running code and doing a lot of slicing in notebooks.
The main one in my team is ubiquity- i.e. lots of people know pandas, who might not be traditional "developers". I.e. data scientists, data analysts etc. Having a data scientist put together some code, it gets optimized by an engineer, and they can talk back and forth about the same code is a massive benefit.
Shifting to polars (and keeping that ability to collaborate) would require not just training the engineers to use a new framework, but all the analysts, data scientists etc that they are adjescant to. That's a huge business cost, and in a lot of cases it might be worth it. But I wouldn't describe it as "getting 95% speed increase for free".
I understand why you wouldn’t do this on an organizational level for production workflows, but for personal workflows in my opinion, it’s a no-brainer to incrementally learn and adopt it.
No matter the target, I test rust things using python. idk. Food for thought.
And to answer your question, we use the Rust API because all of our backend services are in Rust, and we like to stay in Rust whenever possible.
Guess that would help if I already knew what a dataframe is!
There are some areas where Hacker News might uniquely shine in answering questions, but "what are dataframes" isn't one of them. If someone sees "dataframe" and is confused, they have all the opportunity to search up "dataframes 101" and get up to speed.
“SQLite but only for Python and worse. Kind of.”
“That annoying transitional step you have no direct need for but that you have to do anyway, because all data-wrangling tools in Python assume a dataframe”
(I use this stuff daily)
[edit] “Imagine if you could read in from a database or CSV and then work with the returned rows directly and as if they were still a table/spreadsheet. Except by ‘directly’ I mean ‘after turning it into a dataframe and discarding/fucking-up all your data type information’. And then with another fucking-everything-up step if you want to write it out so you can do real stuff with it elsewhere.”
I have been programming for over 30 years on all sorts of systems and the Pandas DataFrame API is completely beyond me. Just trying to get a cell value seems way more difficult than it should be.
Same with xarray datasets.
I just loaded the same CSV into Pandas and Polars and Polars did a much better job of it.
Oh you meant multiindex... yeah, slicing multiindexes sucks :)
The whole thing though is just shockingly unhelpful and half-baked for something that taken over so completely in its niche. I guess my perspective is that of someone trying to build reliable automation, though—it’s probably really nice if you’re just noodlin’ in notebooks or the repl or whatever.
To be fair, Polars has the benefit of hindsight and designing their interfaces and syntax from scratch. The poor choices in Pandas were made long ago, and its adoption and evolution into the most popular dataframe library for python feels like mostly about timing the market than having the best software product.
After you have highlighted trends (the majority of the work), then you might go spelunking at individual examples to see why something is funny.
Dataframes are popular among people who are trying to use the data in a way that they don't intend to do anything more with the processed data after they've reported their findings.
I've actually personally found that DuckDB is tremendously slow against the cloud, though perhaps I'm going through the wrong API?
I'm using https://duckdb.org/docs/guides/import/s3_import.
My data is hive partitioned, when I monitor my network throughput, I only get a few MB/s with DuckDB but can achieve 1-2GB/s through polars.
Very possible it's a case of PEBKAC though.
Why isn’t there a decent file format for tabular data? https://news.ycombinator.com/item?id=31220841
> There is Parquet. It is very efficient with it’s columnar storage and compression. But it is binary, so can’t be viewed or edited with standard tools, which is a pain.
I can open parquet in excel
You can query/filter/sort and pick /transform data on multiple directions (row/column)
It’s used for data mining, ML etc.
It’s basically like working with a spreadsheet/sql table
Every row is a distinct entity. Every column is usually stored / treated as its own distinct array.
My first assumption was that it was somehow related to polar coordinates.
I have recently added a warning to Polars for this on import, could you confirm you get this warning if (before installing native Python) you update your Polars package?
If for whatever reason you really want to keep using the Rosetta version of Python you should install the polars-lts-cpu package instead.
IIRC pyarrow has some trouble like that, to pick one from the same ecosystem.
this to me seems like a good argument for only using ibis, but Im happy to be convinced otherwise
In your amazon example the data is probably optimized for fast lookups on a small number of records, like an in-memory cache.
Algorithms like joins, group bys, distinct, etc, are designed for out-of-core processing and can spill to disk if available RAM is low.
https://docs.pola.rs/py-polars/html/reference/dataframe/api/...
This was captured well in their company announcement blogpost [0]:
> A strict, consistent and composable API. Polars gives you the hangover up front and fails fast, making it very suitable for writing correct data pipelines.
https://docs.pola.rs/user-guide/migration/spark/
Look at the examples on this page of the Spark vs. Polars DataFrame APIs. (Disclaimer: I contributed this documentation. [1])
Having used SQL and Spark DataFrames heavily, but not Polars (or Pandas, for that matter), my impression is that Spark's DataFrame is analogous to SQL tables, whereas Polars's DataFrame is something a bit different, perhaps something closer to a matrix.
I'm not sure how else to explain these kinds of operations you can perform in Polars that just seem really weird coming from relational databases. I assume they are useful for something, but I'm not sure what. Perhaps machine learning?
Here's one of them:
# Polars
df.select(
pl.col("foo").sort().head(2),
pl.col("bar").sort(descending=True).head(2),
)
In SQL and Spark DataFrames, it doesn't make sense to sort columns of the same table independently like this and then just juxtapose them together. It's in fact very awkward to do something like this with either of those interfaces, which you can see in the equivalent Spark code on that page. SQL will be similarly awkward.But in Polars (and maybe in Pandas too) you can do this easily, and I'm not sure why. There is something qualitatively different about the Polars DataFrame that makes this possible.
Long story short, the memory model operates on columns of data as opposed to rows, so fields in a conceptual "row" aren't necessarily an atomic unit.
Sadly, AI is quite poor at Polars right now. (It is ok but not great with Pandas).
However, getting pandas data into Polars is easy. If you already have the code and it is not a bottleneck, I would just wrap it.
(Disclosure: currently writing a book on Polars and have a big chapter on porting hairy pandas code to Polars.)
How does polars sql context stack up against alternatives e.g. perhaps duckdb? If I'm in a notebook and I want to suck in and process a lot of data, which has the least boilerplate, the strongest support and the most efficiency (both RAM usage and speed)?
"Strongest support" is probably Pandas, in that it is very widely used and easy to get help with. DuckDB lets you write SQL and is very fast.
Since then Polars has improved downloading speeds 20x with shipping a proper async runtime in the engine.