HNHacker News
TopNewBestAskShowJobs

legg0myegg0

57 karma · joined September 1, 2020

submissionscomments
legg0myegg0··on ClickHouse, Inc.
Does anyone happen to know which country the new company is incorporated in? I'm still looking for a chance to use ClickHouse because it sounds so excellent!
legg0myegg0··on Fastest table sort in the West – Redesigning DuckDB's sort
I am not affiliated, just a happy user! I have spoken with one of the developers, but that's it!

I've just come up through the SQL side of analytics and I'm moving into Data Science and I feel like DuckDB is a superpower for people with my background.

Does that help? Happy to answer any other questions!

legg0myegg0··on Fastest table sort in the West – Redesigning DuckDB's sort
This is so fast!! If anybody is using Pandas to keep rows in order and has hesitated to use DuckDB for that reason, hesitate no more! Give it a shot!
legg0myegg0··on Apache Arrow Datafusion 5.0.0 release
How would you compare the goals, vision, and current status of DataFusion with DuckDB? (www.duckdb.org) Could DuckDB be an execution engine for Ballista?
legg0myegg0··on Searching for a better persistent cache
I'm not sure if it would help in your case, but could you process all categories at once with a larger SQL query?

If so, DuckDB can process bulk queries about 20x faster than SQLite per CPU core because it is vectorized and column oriented. Then with multiple cores you can easily reach 100x SQLite speed. DuckDB has node bindings and is an in-process DB like SQLite.

If reading from disk is your bottleneck, I would recommend storing your data in compressed parquet files and reading them with DuckDB's parquet reader.

One drawback is that indexes are not persistent to the filesystem in DuckDB yet, but full table scans are much faster than SQLite since it is columnar.

https://github.com/duckdb/duckdb/tree/master/tools/nodejs

https://duckdb.org/

legg0myegg0··on Show HN: Work with CSV files using SQL. For data scientists and engineers
Came here just to recommend DuckDB! :-) Huge fan. It's unreasonably fast for how easy it is to use.
legg0myegg0··on Show HN: Work with CSV files using SQL. For data scientists and engineers
Odd that DuckDB didn't work for you on Windows! I only use it on Windows and love it!
legg0myegg0··on Querying Parquet with Precision Using DuckDB
Here is a DuckDB FDW for Postgres! I have not used it, but it sounds like what you need! https://github.com/alitrack/duckdb_fdw
legg0myegg0··on Querying Parquet with Precision Using DuckDB
The best part of this is just how easy it is! Just a pip install and you're up and running using industry standard Postgres SQL!
legg0myegg0··on Query Engines: Push vs. Pull
The DuckDB folks are migrating from pull to push and put together this interesting documentation of their reasoning! They use a vectorized model instead of a compiled one, so it's another interesting comparison point. It seems like push will simplify how they handle parallelism.

https://github.com/duckdb/duckdb/issues/1583

legg0myegg0··on SQLite the only database you will ever need in most cases
I completely agree!! We see 20-100x performance from DuckDB over SQLite for OLAP style queries.
legg0myegg0··on Why hasn't Presto become industry standard?
Dremio also appears to take a similar approach, but with more advanced caching features / query pushdown. Plus it has Apache Arrow at its heart. I think that would be my choice of solution in this space
legg0myegg0··on SQLite is not a toy database
Check out DuckDB! It is designed for OLAP instead of OLTP, but it uses Postgres syntax and types! It's columnar and lightning fast for big queries.
legg0myegg0··on We Don’t Use Docker
Try DuckDB! I've been getting 20x SQLite performance on one thread, and it usually scales linearly with threads!
legg0myegg0··on Boston Dynamics not happy about killer-robot image
Well then maybe don't design killer robots...??
legg0myegg0··on Not-a-Boring Competition
Go Diggerdoos!! Go Tech Go!
legg0myegg0··on Ask HN: What are the best charting / graphing JavaScript tools
We use plotly.js! It is a layer built on top of D3 and has some great looking charts. It also has some statistical and 3D plots that come in handy for us.
legg0myegg0··on Kite Power for Mauritius
I absolutely love this concept! I remember the Popular Science article years ago that described this for use on ships!

What are the current challenges to scaling this more broadly?

Can kites be made larger for more power, or what is the limiting factor for larger individual units?

legg0myegg0··on Photon powered Engine (enhance Apache Spark 3.0’s performance by up to 20x)
DuxkDB's query engine is inspired by the same paper! You'll be surprised what you can process on a single node with DuckDB - it takes 33 Spark nodes to match performance! That still puts it ahead of Photon, and it is open source.
legg0myegg0··on DuckDB – An embeddable SQL database like SQLite, but supports Postgres features
It's also possible to return a result set as an Arrow table, so round trip SQL on Arrow queries is possible (Arrow to DuckDB to Arrow)! It's not 100% zero-copy for strings, but it should work pretty well!
legg0myegg0··on DuckDB – An embeddable SQL database like SQLite, but supports Postgres features
I work at a Fortune 100 company and we have this in production for our self-service analytics platform as a part of our data transformation web service. Each web request can do multiple pandas transformations, or spin up it's own DuckDB or SQLite db and execute arbitrary SQL transformations. It fits our use case like a glove and is super fast to/from Pandas.
legg0myegg0··on DuckDB – An embeddable SQL database like SQLite, but supports Postgres features
Before the latest optimization, and only using 1 core, vs. SQLite we were seeing 133x performance on a basic group by or join, and about 4x for a pretty complex query. It was roughly even to Pandas in performance, but it can scale to larger than memory data and now it can use multiple cores! As an example, I could build views from 2 Pandas DataFrames with 2 columns and 1 million rows each, join them, and return the 1 million row dataset back to Pandas in 2 seconds vs. 40 seconds with SQLite/SQLAlchemy... Pretty sweet. DuckDB is going to be even faster now I bet!
legg0myegg0··on The database I wish I had
I think DuckDB checks a number of these boxes! It is embedded and written in C so compilable to WASM. It is also 10x faster than SQLite and interoperable with Apache Arrow! It might be a good place to start anyway!