HNHacker News
TopNewBestAskShowJobs

sammysidhu

14 karma · joined September 8, 2021

submissionscomments
sammysidhu··on Open-Sourcing SEC Edgar on Hugging Face
Amazing work leveraging Daft for this!
sammysidhu··on Cutting LLM Batch Inference Time in Half: Dynamic Prefix Bucketing at Scale
Part of the Daft team here! Happy to answer any questions
sammysidhu··on Distributed data engine for unstructured data in Rust
Hi, One of the authors of Daft here! Happy to answer any questions
sammysidhu··on Show HN: Sourcetable – AI Spreadsheet and Data Platform
Congrats on the launch! It's been great working with you from the Daft side
sammysidhu··on Lessons Learned from Scaling to Multi-Terabyte Datasets
https://github.com/Eventual-Inc/Daft Is also great at these types of workloads since it’s both distributed and vectorized!
sammysidhu··on Pg_lakehouse: Query Any Data Lake from Postgres
We're actually using pyiceberg to retrieve metadata! All our IO and decoding happens in the rust side once the data has been passthrough.

We expose something called a ScanOperator which allows integration into various catalogs through a thin layer that exposes ScanTasks.

Iceberg's impl: https://github.com/Eventual-Inc/Daft/blob/416009138359a9d410...

sammysidhu··on Ask HN: Who is hiring? (March 2024)
Eventual | Engineers; Rust | SF | https://www.ycombinator.com/companies/eventual/jobs

We're the people behind Daft (On the front page today!), a distributed dataframe built in Rust. We're VC backed and founded by folks from the self driving industry.

We're looking for developers who want to help build the next generation Spark!

sammysidhu··on Daft: Distributed DataFrame for Python
we love rust too!
sammysidhu··on Daft: Distributed DataFrame for Python
This would be more comparable to being a backend for ibis. We're working on adding the remaining operations (like regex on strings or trigonometry functions) that ibis requires!
sammysidhu··on Daft: Distributed DataFrame for Python
horizontal scaling however provides you with more aggregate network bandwidth. Most enterprises run workloads that downloads data from a data lake (usually S3), which is usually the bottleneck. Having horizontal scaling here allows the query leverage much higher network than just having a large single machine.
sammysidhu··on Daft: Distributed DataFrame for Python
One of the daft developers here - Thank you!
sammysidhu··on Transparency in Recoolit's carbon credits
Love the data driven approach to massive problem
sammysidhu··on Daft: A High-Performance Distributed Dataframe Library for Multimodal Data
(one of the Daft maintainers here) Great call out! I went ahead and make an issue for us to work on this: https://github.com/Eventual-Inc/Daft/issues/1016
sammysidhu··on Daft: A High-Performance Distributed Dataframe Library for Multimodal Data
Hi! (one of the Daft maintainers here), thanks for the feedback. Ultimately you're right that supporting the full Polars syntax in a distributed fashion is very difficult. There are libraries out there that do "Pandas but distributed" but from what I have seen is that they prioritized API coverage rather than performance or memory consumption. So you end up in a similar boat to the situation you mentioned.

We're trying to start with a simpler API that maps well to a distributed query query that we can execute well and then add the features that people request for.

I would love to know what you would want to see in Daft!

sammysidhu··on Daft: A High-Performance Distributed Dataframe Library for Multimodal Data
Hi (one of the maintainers here), that is a good suggestion! I wasn't aware of that project. I went ahead and made an issue to add `export DO_NOT_TRACK=1` as one of the variables we track! https://github.com/Eventual-Inc/Daft/issues/1015
sammysidhu··on Daft: A High-Performance Distributed Dataframe Library for Multimodal Data
Hi Sammy here, one of the Daft maintainers. Happy to answer any questions that you all might have!