I like the simplicity of DuckDB's proposal, but haven't seen much info about how fast to expect it to be in comparison with traditional RDBs, for smaller, mostly-read-only applications.
I like the simplicity of DuckDB's proposal, but haven't seen much info about how fast to expect it to be in comparison with traditional RDBs, for smaller, mostly-read-only applications.
For a dataset that size, I'd probably use SQLite to avoid having to manage a persistent MySQL process, especially when it's being used as an alternative to CSV files. That is, unless there's a MySQL/Postgres server already running I can just create a new database on.
DuckDB automatically creates indexes for all general-purpose columns. However, they're not persisted.
*I use the HoneySQL (clojure) library to programmatically build up queries and execute them via the JDBC driver.
[1] https://www.vantage.sh/blog/querying-aws-cost-data-duckdb