- unstable execution (random crashes)
- out-of-memory errors where I would've hoped for DuckDB to gracefully take the slow route to completion if no more memory is available (tried all the different conf settings)
- unstable execution (random crashes)
- out-of-memory errors where I would've hoped for DuckDB to gracefully take the slow route to completion if no more memory is available (tried all the different conf settings)
You mentioned OOMs, this has been a focus for a while and ha gotten steadily better over the past few releases. 0.9 added spill to disk to prevent most OOMs. And 0.10, released a couple of weeks ago, fixes a bunch more memory usage problems. The storage format, which another commenter brought up, is now fully backwards compatible.
I'd suggest giving it another try, especially once 1.0 comes out.
Example of a query that should never, ever, out-of-memory, but absolutely will in the latest DuckDB:
COPY
(
SELECT
rs.my_int,
rs.my_bigint
FROM
READ_PARQUET('s3://some/folder/my-large-files-*.parquet')
AS rs
)
TO
'/my/home/folder/my-large-file.parquet'
(
FORMAT PARQUET,
ROW_GROUP_SIZE 100000,
COMPRESSION 'ZSTD'
)
;
This query should simply read the two column series selected based on the parquet metadata and then stream the data to the disk.And yet it will try to load data in memory before crashing.
There were some recent fixes: https://github.com/duckdb/duckdb/issues/10737
Why not just make shims to migrate dbs for future compatibility? So you could read db 1.0 in v2.0 but only insofar as to migrate it to v2. The implication that you don't want to promise backwards read compatibility feels antithetical to a db driver.
For example, if I have an ancient mssql db that was started in 2001, I'm confident that I can grab the latest mssql driver and still use it. I don't have to track down mssql 2007 to migrate incrementally. Not sure about postgres or mysql but I assume it's the same there. Sqlite is definitely backwards read compatible.
> Major versions usually change the internal format of system tables and data files. These changes are often complex, so we do not maintain backward compatibility of all stored data.
Three separate occasions with different uses all leading to crashes in the first hour of using DuckDB is enough that I frankly see no point in trying it again; I don't expect it to ever magically become reliable.