264 karma · joined October 29, 2022
Incidentally the most effectively way to stall a project according to the CIA is to have a huge guiding committee with clearly diverging interests.
Redis will win because it's focused on its users. It's competitors will lose. Like OpenSearch, like OpenCL etc.
I may sound like a rich apologist, but please believe me when I say it is harder to spend 10 hrs a day cranking on a risky startup if you know you can be clubbing with daddy's money
You don't have to use the system I am building, but it's worth thinking about that design.
I work on Quokka (https://github.com/marsupialtail/quokka). I support Iceberg reads. Recently we are adding SQL support from just parsing the DuckDB logical plan, though that is very challenging as well.
The Python world lacks a standard for a plug and play SQL query optimizer. Apache Calcite is good for the JVM world, but not great if you are trying to cut out the JVM.
I have users demanding Iceberg writes and Hudi reads/writes. I don't know what to tell them, since I don't have the resources to add a reader/writer myself for those projects.
Hopefully as DuckDB becomes more popular we will see Python bindings for these popular data lake formats this year.
I am not focused on complex data types though.
The key difference here is that these systems don't store data, so fault tolerance means recovering within a query instead of not losing data.
I agree with many of the points made in the blog by Andy. Writing a distributed database has become way easier due to open source components like Ray, Arrow, Velox, DuckDB, SQLGlot etc.
I personally believe we will see a switch from JVM based technologies to Rust/C based with Python wrapper
For push vs. pull, I'd recommend: https://news.ycombinator.com/item?id=27006476.
On single machine, you really should just use Polars. Quokka is faster than Pandas because it can take advantage of multiple cores, but so can Polars -- and it is likely to be faster.