Guess that would help if I already knew what a dataframe is!
Guess that would help if I already knew what a dataframe is!
Every row is a distinct entity. Every column is usually stored / treated as its own distinct array.
My first assumption was that it was somehow related to polar coordinates.
There are some areas where Hacker News might uniquely shine in answering questions, but "what are dataframes" isn't one of them. If someone sees "dataframe" and is confused, they have all the opportunity to search up "dataframes 101" and get up to speed.
“SQLite but only for Python and worse. Kind of.”
“That annoying transitional step you have no direct need for but that you have to do anyway, because all data-wrangling tools in Python assume a dataframe”
(I use this stuff daily)
[edit] “Imagine if you could read in from a database or CSV and then work with the returned rows directly and as if they were still a table/spreadsheet. Except by ‘directly’ I mean ‘after turning it into a dataframe and discarding/fucking-up all your data type information’. And then with another fucking-everything-up step if you want to write it out so you can do real stuff with it elsewhere.”
I have been programming for over 30 years on all sorts of systems and the Pandas DataFrame API is completely beyond me. Just trying to get a cell value seems way more difficult than it should be.
Same with xarray datasets.
I just loaded the same CSV into Pandas and Polars and Polars did a much better job of it.
Oh you meant multiindex... yeah, slicing multiindexes sucks :)
The whole thing though is just shockingly unhelpful and half-baked for something that taken over so completely in its niche. I guess my perspective is that of someone trying to build reliable automation, though—it’s probably really nice if you’re just noodlin’ in notebooks or the repl or whatever.
To be fair, Polars has the benefit of hindsight and designing their interfaces and syntax from scratch. The poor choices in Pandas were made long ago, and its adoption and evolution into the most popular dataframe library for python feels like mostly about timing the market than having the best software product.
After you have highlighted trends (the majority of the work), then you might go spelunking at individual examples to see why something is funny.
Dataframes are popular among people who are trying to use the data in a way that they don't intend to do anything more with the processed data after they've reported their findings.
I've actually personally found that DuckDB is tremendously slow against the cloud, though perhaps I'm going through the wrong API?
I'm using https://duckdb.org/docs/guides/import/s3_import.
My data is hive partitioned, when I monitor my network throughput, I only get a few MB/s with DuckDB but can achieve 1-2GB/s through polars.
Very possible it's a case of PEBKAC though.
Why isn’t there a decent file format for tabular data? https://news.ycombinator.com/item?id=31220841
> There is Parquet. It is very efficient with it’s columnar storage and compression. But it is binary, so can’t be viewed or edited with standard tools, which is a pain.
I can open parquet in excel
You can query/filter/sort and pick /transform data on multiple directions (row/column)
It’s used for data mining, ML etc.
It’s basically like working with a spreadsheet/sql table