HNHacker News
TopNewBestAskShowJobs

houqp

687 karma · joined March 31, 2013

submissionscomments
houqp··on SQLite-based databases on the Postgres protocol
Very cool and well executed project. Love the sprinkle of Rust in all the other companion projects as well :)

The ROAPI(https://github.com/roapi/roapi) project I built also happened to support a similar feature set, i.e. to expose sqlite through a variety of remote query interfaces including pg wire protocols, rest apis and graphqls.

houqp··on SQLite-based databases on the Postgres protocol
bottomless looks really nice, thanks for sharing!
houqp··on Run SQL on CSV, Parquet, JSON, Arrow, Unix Pipes and Google Sheet
Haha, yeah, we should definitely put a little bit more efforts into SEO :) Everyone is so focused on the hard-core engineering at the moment. I think Matthew from the community is actually working on a new comprehensive benchmark for us at the moment, which I hope will be published soon.
houqp··on Run SQL on CSV, Parquet, JSON, Arrow, Unix Pipes and Google Sheet
Datafusion out performs spark by a large margin. It is on par with photon based on my experiences, see benchmarks at https://github.com/blaze-init/blaze.
houqp··on Run SQL on CSV, Parquet, JSON, Arrow, Unix Pipes and Google Sheet
Yes, I designed the code base so that the core of the IO and query logic are abstracted into a Rust library called columnq. My plan is to wrap it with pyo3 so the full API can be accessed as a Python package! If you are interested in helping with this, please feel free to submit a PR. The core library is located at https://github.com/roapi/roapi/tree/main/columnq
houqp··on Run SQL on CSV, Parquet, JSON, Arrow, Unix Pipes and Google Sheet
Most users already have pip installed, so they won't need to install a rust toolchain.
houqp··on Show HN: Query Google Sheet data using PostgreSQL clients
It's manual for now, ROAPI is designed for slowly moving datasets. You just hit the data update API to force a refresh.

I have plan to add automated streaming data update in the background, starting with delta lake tables. It should all be very straight forward to implement.

houqp··on Show HN: Query Google Sheet data using PostgreSQL clients
We cache all the data in memory in Arrow format so queries don't need to go through google api, it will only hit google api when a data refresh is needed.
houqp··on Ask HN: Is any company still hiring entry level engineers?
We are hiring new grads across the board if you have proofs for exceptional skills. Feel free to take a look at all our engineering postings at: https://neuralink.com/careers.
houqp··on Blaze: A Rust-based vectorized accelerator to speed up your Spark jobs
Up to 2x speed up with 20% of the resource consumption, pretty wild performance boost without changing a single line of code!
houqp··on Ask HN: What to do about ‘Good at programming Bad at Leetcode’
There are more companies noticing this now and have stopped asking these questions. For example, we at Neuralink[1] only give out practical programming challenges. If are you good at building practical systems, you should be able to ace our coding interviews without any preparation. No leetcode and no whiteboarding. In fact, I prefer to hire those who doesn't waste time practicing leetcode.

Over the past couple years, I have interviewed at a handful of other startups who also have similar coding interview philosophies.

[1]: https://neuralink.com/careers/

houqp··on Diagrams: Open-Source Alternative to Lucidchart
I love the product, not only does it have an easy to use web/offline app, I am also able to checkin the diagram source into git for version control, see: https://github.com/roapi/docs/blob/main/src/images/roapi.dra.... Then I can use automations to generate images based off that source file: https://github.com/roapi/docs/blob/main/src/images/roapi.svg.
houqp··on Dsq: Commandline tool for running SQL queries against JSON, CSV, Parquet, etc.
Neat! I have also built a similar project in Rust https://github.com/roapi/roapi/tree/main/columnq-cli :)
houqp··on Roapi: An API Server for Static Datasets
I think these two systems explore different design spaces, the biggest difference I would say is Roapi can apply more read optimizations by exploiting the fact that it doesn't need to support frequent online updates from the client. Most of the datasets it serves will be static. For data-sources that supports streaming updates like delta tables, the update frequency will be much lower than what clickhouse supports.
houqp··on Roapi: An API Server for Static Datasets
In its current form, the main use-case is to load data into memory first then serve them through query apis. Thomas has made some effort to support querying data directly from remote source without loading them into memory: https://github.com/roapi/roapi/pull/71. The underlying query engine, Apache Arrow Datafusion, supports running query on data stream on the granularity of partitions. This is not heavily used in roapi at the moment because I want to nail the in memory serving use-case first.
houqp··on Roapi: An API Server for Static Datasets
Thanks, nice work on qocache and qframe too :)
houqp··on Roapi: An API Server for Static Datasets
yeah, that's a good idea. thanks for the suggestion :)
houqp··on Roapi: An API Server for Static Datasets
Yes, I am aiming for production grade online serving + many more query frontends and data types.
houqp··on Roapi: An API Server for Static Datasets
I looked into Datasette before starting ROAPI. From a product/use-case point of view, to me Datasette focuses more on quick and easy ad-hoc data exploration type of work. ROAPI focuses more production ready online serving of static datasets. So I would expect users to use ROAPI to power micro-services in production with high QPS.

From a technical design point of view, ROAPI authors owns the full stack end to end from query parsing, data format parsing to query execution because I am also a maintainer of Apache arrow and it's sub-project datafusion. The whole project is built with Rust end to end from scratch. Datasette is mostly a wrapper around sqlite. It translates user actions into SQL queries, then execute them on sqlite. In ROAPI, we work at a lower level. We translate REST APIs, GraphQL and SQLs into datafusion logical plans and execute them. Datafusion is also a analytical compute engine optimized for columnar data, so it will be a lot faster for OLAP workload, while sqlite is optimized for OLTP. I also plan to add other type of query capabilities like nearest neighbor vector search for ML applications, etc.

houqp··on Roapi: An API Server for Static Datasets
This is true, the core of it is Apache arrow datafusion query engine, which is also a project I help maintain. I doubt you will be able to beat it with PHP though ;) The VM overhead alone will cause a big hit to your performance even if we can get JIT to work.
houqp··on Roapi: An API Server for Static Datasets
That's right, it's intended to be more lightweight since it's built with only Rust from the ground up. Apache Drill also only focuses on serving SQL as the user interface while ROAPI wants to provide a pluggable interface to support all use-cases. For example, we can plan graphql and rest api calls into query plan and efficiently execute them using Datafusion.
houqp··on Roapi: An API Server for Static Datasets
Author of the project here, thanks for writing about ropai! Happy to answer any question.
houqp··on Show HN: Columnq brings OLAP to Unix pipes
Thanks! It's using Datafusion as the query engine: https://github.com/apache/arrow-datafusion
houqp··on Apache Arrow Datafusion 5.0.0 release
I didn't dive into Vaex's implementation, but based on the example code, I would say they are similar in the sense that they all provide a Dataframe interface for end users to perform compute on relational data.

It looks like Vaex focuses more on end users like data scientists while Datafusion focuses more on being a composable embedded library for building analytical engines. For example, InfluxDB IOx, Ballista and ROAPI all uses Datafusion as the compute engine.

On top of that, Datafusion also comes with a builtin SQL planner so users can choose between Dataframe and SQL interfacts.

houqp··on Apache Arrow Datafusion 5.0.0 release
Datafusion, and Ballista by definition, also provides a Dataframe API that let's you construct queries programmatically. It also has preliminary support for UDFs.

We also have community members implementing Spark native executors using Datafusion, which showed significant speed improvements in the initial PoC.

houqp··on Apache Arrow Datafusion 5.0.0 release
ETL pipeline is a perfect fit for Datafusion and its distributed version Ballista. Personally, this is the main reason I am investing my time into Datafusion.
houqp··on Apache Arrow Datafusion 5.0.0 release
Indeed, big shout out to the InfluxDB team!
houqp··on Apache Arrow Datafusion 5.0.0 release
> - Is it possible to handle data larger than fits into RAM?

Not at the moment, but the community has plans to add support for disk spill.

> - Any benchmark? like: https://h2oai.github.io/db-benchmark/ ( see 50GB + Join -> "timeout" | "out of memory" )

One of the committer Daniel is working on a h2oai db benchmark PR for Datafusion :)

houqp··on Apache Arrow Datafusion 5.0.0 release
You beat me to it, was about to post the github link :) Readme is a good starting place to learn more about the project.
houqp··on Apache Arrow Datafusion 5.0.0 release
One of the Arrow Datafusion committers here. Happy to help answer any question.
Page 1 of 3Next →