HNHacker News
TopNewBestAskShowJobs

synsqlbythesea

8 karma · joined January 3, 2026

submissionscomments
synsqlbythesea··on Show HN: SQL service for tables with billions of rows or up to 1M columns
Hi Islambaraka.

Thank you for your question. You have hit on an important design choice. SynSQL relies on a strict fan out across all partitions.

- Wide tables: As you can see on the multi_omic live demo, queries include the hash key. And the hash key is always used in an implicit "order by" clause. Each chunk involved executes a query focused on the columns it contains.

- Deep tables: All data brokers execute the incoming query on their data set (partition).

Why fan out even for 'where pk = 1' ? The engine broadcasts to all partitions. It delegates expression evaluation downstream. It is designed for parallel scans. The overhead is perfectly acceptable for its target: analytical workload.

Best, Stéphane.

synsqlbythesea··on Ask HN: Distributed SQL engine for ultra-wide tables
That’s a fair question!

A concrete case where this comes up is multi-omics research. A single study routinely combines ~20k gene expression values, 100k–1M SNPs, thousands of proteins and metabolites, plus clinical metadata — all per patient.

Today, this data is almost never stored in relational tables. It lives in files and in-memory matrices, and a large part of the work is repeatedly rebuilding wide matrices just to explore subsets of features or cohorts.

In that context, a “wide table” isn’t about transactions or joins — it’s about having a persistent, queryable representation of a matrix that already exists conceptually. Integration becomes “load patients”, and exploration becomes SELECT statements.

I’m not claiming this fits every workload, but based on how much time is currently spent on data reshaping in multi-omics, I’m confident there is a real need for this kind of model.

synsqlbythesea··on Ask HN: Distributed SQL engine for ultra-wide tables
Thanks — both are great systems.

ClickHouse and Scuba are extremely good at what they’re designed for: fast OLAP over relatively narrow schemas (dozens to hundreds of columns) with heavy aggregation.

The issue I kept running into was extreme width: tens or hundreds of thousands of columns per row, where metadata handling, query planning, and even column enumeration start to dominate.

In those cases, I found that pushing width this far forces very different tradeoffs (e.g. giving up joins and transactions, distributing columns instead of rows, and making SELECT projection part of the contract).

If you’ve seen ClickHouse or Scuba used successfully at that kind of width, I’d genuinely be interested in the details.

synsqlbythesea··on Ask HN: Distributed SQL engine for ultra-wide tables
From what I understand, Exasol is a very fast analytical database for traditional data warehouses. My engine doesn't replace a data warehouse; it solves a type of table that data warehouses simply can't handle: tables with hundreds of thousands or millions of columns with an access model that guarantees interactive response times even in these extreme cases.
synsqlbythesea··on Ask HN: Distributed SQL engine for ultra-wide tables
In a few words: table data is stored on hundreds of MariaDB servers. Each table is user designed hash key columns(1->32) to manage automatic partitioning. Wide tables are split in chunks. 1 chunk = the hash key + columns = one MariaDB server. The data dictionary is stored on mirrored dedicated MariaDB servers. The engine in itself uses a massive fork policy. In my lab, the k1000 table is stored on 500 chunks. I used a small trick : where I say 1 MariaDB server you can use one database in a MariaDB server. So I have only 20 VmWare Linux servers with 25 database each containing 25 databases.