57 karma · joined September 1, 2020
I've just come up through the SQL side of analytics and I'm moving into Data Science and I feel like DuckDB is a superpower for people with my background.
Does that help? Happy to answer any other questions!
If so, DuckDB can process bulk queries about 20x faster than SQLite per CPU core because it is vectorized and column oriented. Then with multiple cores you can easily reach 100x SQLite speed. DuckDB has node bindings and is an in-process DB like SQLite.
If reading from disk is your bottleneck, I would recommend storing your data in compressed parquet files and reading them with DuckDB's parquet reader.
One drawback is that indexes are not persistent to the filesystem in DuckDB yet, but full table scans are much faster than SQLite since it is columnar.
What are the current challenges to scaling this more broadly?
Can kites be made larger for more power, or what is the limiting factor for larger individual units?