- Is it possible to handle data larger than fits into RAM?
- Any benchmark? like: https://h2oai.github.io/db-benchmark/ ( see 50GB + Join -> "timeout" | "out of memory" )
- Is it possible to handle data larger than fits into RAM?
- Any benchmark? like: https://h2oai.github.io/db-benchmark/ ( see 50GB + Join -> "timeout" | "out of memory" )
https://github.com/h2oai/db-benchmark/pull/182
Also we do support running TPC-H benchmarks. For the queries we can run, those are already finishing faster than Spark. We are planning to do more benchmarking and optimizations in the future.
That's really promising, considering how fast Polars is. Both are written in Rust and use Apache Arrow, so they can even co-exist in the same context.
Not at the moment, but the community has plans to add support for disk spill.
> - Any benchmark? like: https://h2oai.github.io/db-benchmark/ ( see 50GB + Join -> "timeout" | "out of memory" )
One of the committer Daniel is working on a h2oai db benchmark PR for Datafusion :)
https://github.com/apache/arrow-datafusion/tree/master/balli...