HNHacker News
TopNewBestAskShowJobs

nhourcard

82 karma · joined November 1, 2019

submissionscomments
nhourcard··on How we made WINDOW JOIN parallel and vectorized
TSBS (Time Series Benchmarking Suite) is as close as it gets
nhourcard··on Ask HN: Who is hiring? (October 2025)
QuestDB | C++/Rust Database Engineer | REMOTE| | Full-time | https://www.questdb.com

Work with world-class engineers - including two teammates who placed in the top 10 of the One Billion Row Challenge (1BRC).

As a specialized database, QuestDB stores, processes and analyzes time series data in real-time, with a focus on reliability, extreme performance and simplicity. It provides best-in-class hardware efficiency and robust features, saving costs and accelerating time-to-value.

Our open source repository has gathered 16k stars and QuestDB is one of the fastest growing database in the category. We are a product-first company with a large community of developers. As a team, we are globally distributed, remote-first and backed by leading venture capital firms and Y Combinator.

Teams have had success with QuestDB across a wide range of industries, such as Financial Services, Energy and Space Exploration. Category leading companies such as OKX, Mizuho, and Airbus rely on QuestDB for large-scale, data-intensive production systems. Emerging and disruptive startups also leverage QuestDB to gain a significant edge within traditional industries.

To apply: https://questdb.com/careers/core-database-engineer/

nhourcard··on The future of kdb+?
TLDR from the article;

Alternatives (which are open source) to KDB+ are split into two categories:

New Database Technologies (tick data store & ASOF JOIN): Clickhouse & QuestDB

Local Quant Analysis: Python – with DuckDB & Polars

Some personal thoughts:

Q is very expressive, and impressive performance can be extracted from kdb+, but the drawbacks are proprietary formats, vendor lock-in, costs, proprietary language and reliance on external consultants to make the system run adequately, which can increase operational costs.

I'm personally excited to see the open-source alternative stack emerging. Open Source time-series databases and tools like duckdb/polars for data science are a good combination. Storing everything in open formats like Parquet and leveraging high-performance frameworks like Arrow is probably where things are heading.

Seeing some disruption in this industry specifically is interesting; I think it will be beneficial, particularly for developers.

NB: disclosing that I'm from questdb to put thoughts in perspective

nhourcard··on Ask HN: Who is hiring? (August 2024)
QuestDB | Remote | Full-time | questdb.io

QuestDB is one of the fastest-growing time series databases and is growing its engineering team. QuestDB is open source, and the code is available under Apache 2.0 at https://github.com/orgs/questdb/

Our team is obsessed with performance; it is a blend of low latency zero-GC Java, C++ and Rust (for the entire distributed architecture). Check out our blog post for heavily technical content about the project: https://questdb.io/blog/tags/engineering/

VC funded with plenty of runway, investors include YCombinator, and founders of Github, Docker, Supabase, Posthog, NGINX, Mesosphere, Citus Data and more.

Job application here: https://questdb.io/careers/core-database-engineer/

nhourcard··on pg_timeseries: Open-source time-series extension for PostgreSQL
Interesting release, it feels that the time-series database landscape is evolving toward:

a) columnar store & built from scratch, with convergence toward open formats such as parquet & arrow: influxdb 3.0, questdb

b) Adding time-series capabilities on top of Postgres: timescale, pg_timeseries

c) platforms focused on observability around the Prometheus ecosystem: grafana, victoria metrics, chronosphere

nhourcard··on Thoughts on low latency trading if exchanges went full cloud
About the market data & monitoring element, beyond the usual suspects Clickhouse and kdb+ - worth noting that Aquis mentioned in this article uses QuestDB ( https://questdb.io/case-study/aquis/)
nhourcard··on All you need is Wide Events, not "Metrics, Logs and Traces"
In QuestDB, only SYMBOL columns can be indexed. However, sometimes, queries can run faster without indexes. This is because, under the hood, QuestDB runs very close to the hardware and only lifts relevant time partitions and columns for a given query. Therefore table scans between given timestamps are then very efficient. This can be faster than using indexes when the scan is performed with SIMD and other hardware-friendly optimizations.

When cardinality is very high, indexes make more sense.

nhourcard··on Ask HN: Who is hiring? (January 2024)
QuestDB | Core Database Engineer (Java, C++, Rust) | full time | Remote | https://questdb.io/

About QuestDB We are building an open source time series database focused on performance and simplicity.

QuestDB is used to store, process and analyze time series data in real-time across a wide range of industries, such as Financial Services, Energy, Manufacturing, Web3 and Space Exploration. Fortune 500 companies such as Airbus and Yahoo deploy QuestDB for large-scale, data-intensive production systems, some of which serve close to a billion users.

Our open source repository has gathered 13k+ stars in three years, and is the fastest-growing in our category. We are a product-first company with a large community of developers. We are a globally distributed remote-first team backed by leading venture capital firms and Y Combinator.

The role As a Core Database Engineer, you will bring your experience in design, development, and testing to improve our open source time series SQL database. You will continuously improve the system's performance, ensuring that QuestDB remains scalable and easy to use as we roll out new features built with C++ and Java (zero-GC). You will have the opportunity to interact with and gather feedback from QuestDB's growing community of users and contributors. You'll have the chance to work in an open and collaborative environment to improve user experience and the system's consistency.

Page: https://questdb.io/careers/core-database-engineer/

nhourcard··on Building a faster hash table for high performance SQL joins
indeed! thank you for that :)
nhourcard··on Solving duplicate data with performant deduplication
Not for CSV import via SQL COPY sadly
nhourcard··on Solving duplicate data with performant deduplication
Sorry to hear about your painful past experience and thanks for sharing this feedback. QuestDB improved a lot in the past two years and it's much more robust and production-ready now.

A couple of things the engineering team did to address this as it grew from the early days at Ycombinator:

QuestDB introduced fuzz tests for all major database components [1] which helped find and fix lots of bugs. Besides that, the SQLancer team [2] was kind enough to add QuestDB support and find bugs around tricky SQL edge cases. We do our best to squash all critical bugs and constantly grow our integration tests suite.

On the SQL side specifically, we brought many SQL-related fixes over the last few years, introduced EXPLAIN and query plan testing, and finally improved memory management of cached queries. We're prioritising new bugs should they arise on GitHub or elsewhere.

Large companies such as Airtel [3], Mizuho Bank, Airbus, NetApp, Yahoo [4], Central Group [5] (largest retailer in Asia), Cloudera use QuestDB in production today.

There is still a lot of work to be done to bring better functionality. If you're keen to try QuestDB for one of your projects, we'd love to hear your feedback and see how it works for you.

[1] https://questdb.io/blog/fuzz-testing-questdb/ [2] https://github.com/sqlancer/sqlancer/ [3] https://questdb.io/case-study/airtel-xstream-play/ [4] https://questdb.io/case-study/yahoo/ [5] https://questdb.io/case-study/central-group/

nhourcard··on Solving duplicate data with performant deduplication
I'd be curious to hear how RDS is starting to fall over with time series data, is it a bottleneck on ingestion, queries, or both?
nhourcard··on Ask HN: Does anyone use InfluxDB? Or should we switch?
Awesome, don't hesitate to reach out to the team at https://slack.questdb.io/ if you have any questions!
nhourcard··on Ask HN: Does anyone use InfluxDB? Or should we switch?
I agree that the difference in versions is confusing!

If InfluxDB is not the right fit, QuestDB can be a good drop-in replacement as it uses a high-performance implementation of the InfluxDB Line Protocol for ingesting data and SQL for queries. There are several SQL extensions for time-series data to simplify queries, such as SAMPLE BY (downsampling), WHERE ... IN (time intervals), LATEST ON (latest records) and ASOF JOIN (time-series joins).

An intro about QuestDB can be found here: https://questdb.io/docs/ And a live demo with three datasets (one of them has got data being streamed live): https://demo.questdb.io/

Another well-known time-series database includes TimescaleDB, an extension built on top of Postgres.

I should disclose that I'm a co-founder of QuestDB.

nhourcard··on WeWork Skips $95M in Interest Payments
our wework in waterloo is quite full, especially on tuesday/thursday. Astonishingly, there are still two baristas serving coffees from state of the art La Marzocco coffee machines. All this while the share price collapsed and the market cap is $150M.
nhourcard··on Every database will become a vector database sooner or later
Different needs dictate different design choices for optimality.

Could not agree more. Even for time series, which could be seen as a subset of OLAP, trade-offs and design choices inherent to time-series data are necessary. As an example of a TSDB that I know well, QuestDB: Data is always ordered by time once it lands on the disk, the data is partitioned by time, and the ingestion protocol is conceived to stream large volumes of data, which can be either continuous or in bursts.

nhourcard··on Influxdb made the switch from Go to Rust
have you given QuestDB a try? it includes its implementation of the InfluxDB Line Protocol, adds SQL for queries and can sustain a higher ingestion rate, without high cardinality limitations
nhourcard··on Ask HN: What does the db business model look like?
Each database will differentiate itself and shine for specific workloads. A case in point is QuestDB, which deals very well with heavy ingestion (high cardinality / out-of-order data). This lends itself well to use cases around financial tick data and IoT data.
nhourcard··on Leveraging Rust in our Java database
Folks with a background in electronic trading (FX, hedge funds, trading firms etc) are familiar with zero-gc Java. London / NY / HK are a good pool of talent in that respect
nhourcard··on Leveraging Rust in our Java database
thanks for the kind words! We want to stick with SQL - having done a few extensions to make it easier to work with time series data such as SAMPLE BY, LATEST ON, etc. Window functions that the product has been lacking for some time are next to bridge the gap vs other more mature platforms while offering something very new and unique on the performance side, especially ingestion related.
nhourcard··on Leveraging Rust in our Java database
Also this looks to be more of an analytical, column oriented, database. So I can imagine they're optimizing more for throughput than transactional latency.

Yes that is the case

nhourcard··on Announcement regarding possible offer
Their market cap is $30m, which is more or less what a YC company with good prospects will be valued at nowadays for a seed round. Crazy
nhourcard··on DuckDB's AsOf Joins: Fuzzy Temporal Lookups
QuestDB: https://questdb.io/docs/reference/sql/join/#asof-join
nhourcard··on Uses and abuses of cloud data warehouses
QuestDB, kdb+ and others mentioned are more geared toward time-series workloads, while Clickhouse is more toward OLAP. There are also exciting solutions on the streaming side of things with RisingWave etc.
nhourcard··on Do we really need a specialized vector database?
Time-series features can be implemented in a general purpose DB engine, but the architecture of the database will be limiting for lots of use cases requiring performance, especially for high cardinality datasets.

A columnar database engineered from the ground up for time-series should have better foundations to allows fast ingest (to the tune of million of rows/sec per server) and also be very efficient for time-based queries that can be done via languages such as SQL.

Using QuestDB as an example: Data is stored in chronological order and is optimized for sequential ingestion with a timestamp component (re-ordering data on the fly if it comes out-of-order). The data is also partitioned by time. The InfluxDB Line Protocol is better suited for streaming type of ingest versus transactional inserts via Postgres.

nhourcard··on Daft: A High-Performance Distributed Dataframe Library for Multimodal Data
one of the best bands in the world has Daft in their name!
nhourcard··on Launch HN: Clearspace (YC W23) – Cut back on screen time
love this. note that i struggled to find UK's country code in your list.
nhourcard··on Investigating Linux phantom disk reads
MongoDB is one of the most successful open-source databases of all time. The parent company is a listed company and worth $15BN, 3x more than Elastic to put some perspective.

This reflection [1] came from the founders of RethinkDB, a competitor of MongoDB at the time:

"It turned out that correctness, simplicity of the interface, and consistency are the wrong metrics of goodness for most users. The majority of users wanted these three trade-offs instead:

- A use case. We set out to build a good database system, but users wanted a good way to do X (e.g. a good way to store JSON documents from hapi, a good way to store and analyze logs, a good way to create reports, etc.).

- Timely arrival. They wanted the product to actually exist when they needed it, not three years later.

- Palpable speed [...]. MongoDB mastered these workloads brilliantly, while we fought the losing battle of educating the market."

MongoDB narrowed things down for a specific use case, and became the best for that use case. This comes with trade-offs. MongoDB was probably not the best database for healthcare back in the days, but that is OK. It did the job very well for other use cases and industries. And over time, they fixed the issue around losing data and became more stable. Essentially, they made developers feel like superheroes, and over time improved their product, and eventually grabbed a massive market share.

[1] https://www.defmacro.org/2017/01/18/why-rethinkdb-failed.htm...

nhourcard··on Ask HN: Who is hiring? (April 2023)
QuestDB | Backend Engineer (Python) & Front-end UI & Growth Engineer| Remote, full time

QuestDB is an open-source time-series database focused on performance and simplicity. 11k GitHub stars, 2k slack community, large users and customers such as Yahoo, Airbus and Central Group. Top 10 db-engines for time-series databases.

Job description & application:

- https://questdb.io/careers/senior-backend-engineer-python/

- https://questdb.io/careers/growth-engineer-open-source/

- Front-end UI: contact us at careers@questdb.io

nhourcard··on Show HN: QuestDB with Python, Pandas and SQL in a Jupyter notebook – no install
[One edit, adding one additional paragraph at the end]

Note that I'm one of the co-founder of QuestDB, but let me try to be as objective and un-biased as possible. Under the hood, InfluxDB and QuestDB are built differently. Both storage engines are column-oriented. InfluxDB's storage engine uses a Time-Structured Merge Tree (TSM), while QuestDB uses a linear data structure (arrays). A linear data structure makes it easier to leverage modern hardware with native support for CPU's SIMD instructions [1]. Running close to the hardware is one of the key differentiators of QuestDB from an architectural standpoint.

Both have a Write-Ahead Log (WAL) that makes the data durable in case of an unexpected failure. Both use the InfluxDB Line Protocol to ingest data efficiently. Hats off to InfluxDB's team, we found the ILP implementation very neat. However, QuestDB's implementation of ILP is over TCP rather than HTTP for performance reasons. QuestDB is Postgres Wire compatible, meaning that you could also ingest via Postgres, although for market data it would not be the recommended way.

One characteristic of QuestDB is that data is always ordered by time on disk, and out-of-order data is dealt with before touching the disk [2]. The data is partitioned by time. For queries spanning time intervals, the relevant time partitions & columns are lifted to memory, while others are left untouched. This makes such queries (downsampling, interval search etc) particularly fast and efficient.

From a developer experience standpoint, one material difference is the language: InfluxDB has got its own native language, Flux [3], while QuestDB uses SQL, with a bunch of native SQL extensions to manipulate time-series data efficiently: SAMPLE BY, LATEST ON, etc [4]. QuestDB also includes SQL Joins and time-series join (ASOF Join) popular for market data. Since QuestDB speaks the postgresql protocol, developers can use their standard Postgres libraries to query from any language.

From a performance perspective, InfluxDB is known to struggle with ingestion and queries alongside high-cardinality datasets [5]. QuestDB deals with such high cardinality datasets better and is particularly good at ingesting data from concurrent sources, with a max throughput can now reach nearly 5M rows/sec on a single machine. Benchmarks on TSBS [6] with the latest version will follow soon.

InfluxDB is a platform, meaning that they provide an exhaustive offering around the database, while QuestDB is less mature. QuestDB is not yet fully compatible with several tools (say a dashboard like metabase for example), as some popular ones have been prioritised instead (Grafana, Kafka, Telegraf, Pandas dataframes). The charting capabilities of InfluxDB's console are excellent, while QuestDB users would mostly rely on Grafana instead.

[Adding this via post edit #1] One area Influx currently has edge is storage overhead. QuestDB does not support compression yet. Time-series data can often be compressed well [7]. Chances are QuestDB will use more disk space to store the same amount of data.

Hope this helps!

[1] https://news.ycombinator.com/item?id=22803504 [2] https://questdb.io/blog/2021/05/10/questdb-release-6-0-tsbs-... [3] https://docs.influxdata.com/influxdb/cloud/query-data/get-st... [4] https://questdb.io/blog/2022/11/23/sql-extensions-time-serie... [5] https://docs.influxdata.com/influxdb/cloud/write-data/best-p... [6] https://github.com/timescale/tsbs [7] https://www.vldb.org/pvldb/vol8/p1816-teller.pdf

Page 1 of 2Next →