82 karma · joined November 1, 2019
Work with world-class engineers - including two teammates who placed in the top 10 of the One Billion Row Challenge (1BRC).
As a specialized database, QuestDB stores, processes and analyzes time series data in real-time, with a focus on reliability, extreme performance and simplicity. It provides best-in-class hardware efficiency and robust features, saving costs and accelerating time-to-value.
Our open source repository has gathered 16k stars and QuestDB is one of the fastest growing database in the category. We are a product-first company with a large community of developers. As a team, we are globally distributed, remote-first and backed by leading venture capital firms and Y Combinator.
Teams have had success with QuestDB across a wide range of industries, such as Financial Services, Energy and Space Exploration. Category leading companies such as OKX, Mizuho, and Airbus rely on QuestDB for large-scale, data-intensive production systems. Emerging and disruptive startups also leverage QuestDB to gain a significant edge within traditional industries.
To apply: https://questdb.com/careers/core-database-engineer/
Alternatives (which are open source) to KDB+ are split into two categories:
New Database Technologies (tick data store & ASOF JOIN): Clickhouse & QuestDB
Local Quant Analysis: Python – with DuckDB & Polars
Some personal thoughts:
Q is very expressive, and impressive performance can be extracted from kdb+, but the drawbacks are proprietary formats, vendor lock-in, costs, proprietary language and reliance on external consultants to make the system run adequately, which can increase operational costs.
I'm personally excited to see the open-source alternative stack emerging. Open Source time-series databases and tools like duckdb/polars for data science are a good combination. Storing everything in open formats like Parquet and leveraging high-performance frameworks like Arrow is probably where things are heading.
Seeing some disruption in this industry specifically is interesting; I think it will be beneficial, particularly for developers.
NB: disclosing that I'm from questdb to put thoughts in perspective
QuestDB is one of the fastest-growing time series databases and is growing its engineering team. QuestDB is open source, and the code is available under Apache 2.0 at https://github.com/orgs/questdb/
Our team is obsessed with performance; it is a blend of low latency zero-GC Java, C++ and Rust (for the entire distributed architecture). Check out our blog post for heavily technical content about the project: https://questdb.io/blog/tags/engineering/
VC funded with plenty of runway, investors include YCombinator, and founders of Github, Docker, Supabase, Posthog, NGINX, Mesosphere, Citus Data and more.
Job application here: https://questdb.io/careers/core-database-engineer/
a) columnar store & built from scratch, with convergence toward open formats such as parquet & arrow: influxdb 3.0, questdb
b) Adding time-series capabilities on top of Postgres: timescale, pg_timeseries
c) platforms focused on observability around the Prometheus ecosystem: grafana, victoria metrics, chronosphere
When cardinality is very high, indexes make more sense.
About QuestDB We are building an open source time series database focused on performance and simplicity.
QuestDB is used to store, process and analyze time series data in real-time across a wide range of industries, such as Financial Services, Energy, Manufacturing, Web3 and Space Exploration. Fortune 500 companies such as Airbus and Yahoo deploy QuestDB for large-scale, data-intensive production systems, some of which serve close to a billion users.
Our open source repository has gathered 13k+ stars in three years, and is the fastest-growing in our category. We are a product-first company with a large community of developers. We are a globally distributed remote-first team backed by leading venture capital firms and Y Combinator.
The role As a Core Database Engineer, you will bring your experience in design, development, and testing to improve our open source time series SQL database. You will continuously improve the system's performance, ensuring that QuestDB remains scalable and easy to use as we roll out new features built with C++ and Java (zero-GC). You will have the opportunity to interact with and gather feedback from QuestDB's growing community of users and contributors. You'll have the chance to work in an open and collaborative environment to improve user experience and the system's consistency.
A couple of things the engineering team did to address this as it grew from the early days at Ycombinator:
QuestDB introduced fuzz tests for all major database components [1] which helped find and fix lots of bugs. Besides that, the SQLancer team [2] was kind enough to add QuestDB support and find bugs around tricky SQL edge cases. We do our best to squash all critical bugs and constantly grow our integration tests suite.
On the SQL side specifically, we brought many SQL-related fixes over the last few years, introduced EXPLAIN and query plan testing, and finally improved memory management of cached queries. We're prioritising new bugs should they arise on GitHub or elsewhere.
Large companies such as Airtel [3], Mizuho Bank, Airbus, NetApp, Yahoo [4], Central Group [5] (largest retailer in Asia), Cloudera use QuestDB in production today.
There is still a lot of work to be done to bring better functionality. If you're keen to try QuestDB for one of your projects, we'd love to hear your feedback and see how it works for you.
[1] https://questdb.io/blog/fuzz-testing-questdb/ [2] https://github.com/sqlancer/sqlancer/ [3] https://questdb.io/case-study/airtel-xstream-play/ [4] https://questdb.io/case-study/yahoo/ [5] https://questdb.io/case-study/central-group/
If InfluxDB is not the right fit, QuestDB can be a good drop-in replacement as it uses a high-performance implementation of the InfluxDB Line Protocol for ingesting data and SQL for queries. There are several SQL extensions for time-series data to simplify queries, such as SAMPLE BY (downsampling), WHERE ... IN (time intervals), LATEST ON (latest records) and ASOF JOIN (time-series joins).
An intro about QuestDB can be found here: https://questdb.io/docs/ And a live demo with three datasets (one of them has got data being streamed live): https://demo.questdb.io/
Another well-known time-series database includes TimescaleDB, an extension built on top of Postgres.
I should disclose that I'm a co-founder of QuestDB.
Could not agree more. Even for time series, which could be seen as a subset of OLAP, trade-offs and design choices inherent to time-series data are necessary. As an example of a TSDB that I know well, QuestDB: Data is always ordered by time once it lands on the disk, the data is partitioned by time, and the ingestion protocol is conceived to stream large volumes of data, which can be either continuous or in bursts.
Yes that is the case
A columnar database engineered from the ground up for time-series should have better foundations to allows fast ingest (to the tune of million of rows/sec per server) and also be very efficient for time-based queries that can be done via languages such as SQL.
Using QuestDB as an example: Data is stored in chronological order and is optimized for sequential ingestion with a timestamp component (re-ordering data on the fly if it comes out-of-order). The data is also partitioned by time. The InfluxDB Line Protocol is better suited for streaming type of ingest versus transactional inserts via Postgres.
This reflection [1] came from the founders of RethinkDB, a competitor of MongoDB at the time:
"It turned out that correctness, simplicity of the interface, and consistency are the wrong metrics of goodness for most users. The majority of users wanted these three trade-offs instead:
- A use case. We set out to build a good database system, but users wanted a good way to do X (e.g. a good way to store JSON documents from hapi, a good way to store and analyze logs, a good way to create reports, etc.).
- Timely arrival. They wanted the product to actually exist when they needed it, not three years later.
- Palpable speed [...]. MongoDB mastered these workloads brilliantly, while we fought the losing battle of educating the market."
MongoDB narrowed things down for a specific use case, and became the best for that use case. This comes with trade-offs. MongoDB was probably not the best database for healthcare back in the days, but that is OK. It did the job very well for other use cases and industries. And over time, they fixed the issue around losing data and became more stable. Essentially, they made developers feel like superheroes, and over time improved their product, and eventually grabbed a massive market share.
[1] https://www.defmacro.org/2017/01/18/why-rethinkdb-failed.htm...
QuestDB is an open-source time-series database focused on performance and simplicity. 11k GitHub stars, 2k slack community, large users and customers such as Yahoo, Airbus and Central Group. Top 10 db-engines for time-series databases.
Job description & application:
- https://questdb.io/careers/senior-backend-engineer-python/
- https://questdb.io/careers/growth-engineer-open-source/
- Front-end UI: contact us at careers@questdb.io
Note that I'm one of the co-founder of QuestDB, but let me try to be as objective and un-biased as possible. Under the hood, InfluxDB and QuestDB are built differently. Both storage engines are column-oriented. InfluxDB's storage engine uses a Time-Structured Merge Tree (TSM), while QuestDB uses a linear data structure (arrays). A linear data structure makes it easier to leverage modern hardware with native support for CPU's SIMD instructions [1]. Running close to the hardware is one of the key differentiators of QuestDB from an architectural standpoint.
Both have a Write-Ahead Log (WAL) that makes the data durable in case of an unexpected failure. Both use the InfluxDB Line Protocol to ingest data efficiently. Hats off to InfluxDB's team, we found the ILP implementation very neat. However, QuestDB's implementation of ILP is over TCP rather than HTTP for performance reasons. QuestDB is Postgres Wire compatible, meaning that you could also ingest via Postgres, although for market data it would not be the recommended way.
One characteristic of QuestDB is that data is always ordered by time on disk, and out-of-order data is dealt with before touching the disk [2]. The data is partitioned by time. For queries spanning time intervals, the relevant time partitions & columns are lifted to memory, while others are left untouched. This makes such queries (downsampling, interval search etc) particularly fast and efficient.
From a developer experience standpoint, one material difference is the language: InfluxDB has got its own native language, Flux [3], while QuestDB uses SQL, with a bunch of native SQL extensions to manipulate time-series data efficiently: SAMPLE BY, LATEST ON, etc [4]. QuestDB also includes SQL Joins and time-series join (ASOF Join) popular for market data. Since QuestDB speaks the postgresql protocol, developers can use their standard Postgres libraries to query from any language.
From a performance perspective, InfluxDB is known to struggle with ingestion and queries alongside high-cardinality datasets [5]. QuestDB deals with such high cardinality datasets better and is particularly good at ingesting data from concurrent sources, with a max throughput can now reach nearly 5M rows/sec on a single machine. Benchmarks on TSBS [6] with the latest version will follow soon.
InfluxDB is a platform, meaning that they provide an exhaustive offering around the database, while QuestDB is less mature. QuestDB is not yet fully compatible with several tools (say a dashboard like metabase for example), as some popular ones have been prioritised instead (Grafana, Kafka, Telegraf, Pandas dataframes). The charting capabilities of InfluxDB's console are excellent, while QuestDB users would mostly rely on Grafana instead.
[Adding this via post edit #1] One area Influx currently has edge is storage overhead. QuestDB does not support compression yet. Time-series data can often be compressed well [7]. Chances are QuestDB will use more disk space to store the same amount of data.
Hope this helps!
[1] https://news.ycombinator.com/item?id=22803504 [2] https://questdb.io/blog/2021/05/10/questdb-release-6-0-tsbs-... [3] https://docs.influxdata.com/influxdb/cloud/query-data/get-st... [4] https://questdb.io/blog/2022/11/23/sql-extensions-time-serie... [5] https://docs.influxdata.com/influxdb/cloud/write-data/best-p... [6] https://github.com/timescale/tsbs [7] https://www.vldb.org/pvldb/vol8/p1816-teller.pdf