https://bugs.debian.org/838338
These days Debian packaging has become a bit irrelevant, since you can just shove upstream releases into a container and go for it.
> apt-cache search hdf5 | wc -l
134
> apt-cache search netcdf | wc -l
70
> apt-cache search parquet | wc -l
0Are there dark secrets?
it's very modern and perhaps hasn't been around long enough to have debian maintainers feel it's vetted.
for instance, documentation for Python bindings is more advanced than for Rust bindings, but the package itself uses Rust at the low level.
FWIW i think i share your general aversion to _not_ using packages, just for the tidiness of installs and removals, though i'm on fedora and macos.
Python 3.10.12 (main, Nov 20 2023, 15:14:05) [GCC 11.4.0]
on linux
Type "help", "copyright", "credits" or "license" for more
information.
>>> import pandas
>>> pandas.read_parquet('sample3.parquet')
Traceback (most recent call last):
File "<stdin>", line 1, in <module>
File "/usr/lib/python3/dist-packages/pandas/io/parquet.py",
line 493, in read_parquet
impl = get_engine(engine)
File "/usr/lib/python3/dist-packages/pandas/io/parquet.py",
line 53, in get_engine
raise ImportError(
ImportError: Unable to find a usable engine; tried using:
'pyarrow', 'fastparquet'.
A suitable version of pyarrow or fastparquet is required for
parquet support.https://github.com/hangxie/parquet-tools/blob/main/USAGE.md#...
If you follow that link, you'll see polars and parquet are a large highly configurable collection of tools for format manipulations across many HPC formats. Debian maintainers possibly don't want to bundle the entirety, as it would be vast.
Might this help you, though?
https://cloudsmith.io/~opencpn/repos/polar-prod/packages/det...
https://pandas.pydata.org/docs/reference/api/pandas.read_par...
https://packages.debian.org/buster/python3-pandas
Perhaps more alarm is called for when this python+pandas and parquet does not work on Debian, but that is not the case today.
ps- data access in clouds often uses the S3:// endpoint . Contrast to a POSIX endpoint using _fread()_ or similar.. many parquet-aware clients prefer the cloudy, un-POSIX method to access data and that is another reason it is not a simple package in Debian today.
https://github.com/pandas-dev/pandas/blob/main/pandas/io/par...
says "fastparquet" engine must be available if no pyarrow
Polars and DuckDB are much better about memory management.