HNHacker News
TopNewBestAskShowJobs

bluestreak

1,042 karma · joined May 13, 2016

www.questdb.io
submissionscomments
bluestreak··on Building a faster hash table for high performance SQL joins
In certain situations, crossing the JNI boundary can be advantageous. When data resides in "native" memory, outside the Java heap, the coordination of concurrent execution logic can be handled in Java, while the actual execution occurs in C++ or Rust. In this context, the negligible penalty for crossing the JNI boundary once per 10 million rows pales in comparison to the substantial benefits achievable through C++ optimization or CPU-specific assembly.
bluestreak··on Building a faster hash table for high performance SQL joins
this is THE number one question being asked about QuestDB :) There is a thread that might help https://news.ycombinator.com/item?id=37557880
bluestreak··on Leveraging Rust in our Java database
One of our distribution channels is Maven Central where we ship Java 11 compatible library. Embedded users preclude us from leveraging latest Java features.
bluestreak··on Leveraging Rust in our Java database
Java tooling was excellent back when QuestDB was started and still is excellent today compared to C++.
bluestreak··on Investigating Linux phantom disk reads
great, we are on the same page! `fallocate()` (or posix one) is called on large swades of file. 16MB default. Not too often to hurt performance. I wonder if zeroing the file with `fallocate()` will result in actual disk writes or is it ephemeral?
bluestreak··on Investigating Linux phantom disk reads
> On x86, and I think every architecture, when you write to a memory mapping that is not already backed by a writable page, the kernel is notified that user code is trying to write. And the kernel needs to fill in the contents of the page, which requires a read if the page isn’t already loaded.

This is very true. Perhaps there wasn't enough context to what the article is describing. The read problem started to occur on database that is subject to constant write workload. Data is flowing in all the time at variable rate. Typically blocks are "hot" and being filled in fully within seconds if not millis.

Zeroing the file is an option to try. QuestDB allocates disk with `posix_fallocate()`, which doesn't have the required flags. We would need to explore `fallocate()`. Thanks.

bluestreak··on Investigating Linux phantom disk reads
It is not always read-write-modify. There is no evidence of this pattern in Ubuntu when there is no memory pressure. Merge occurs when block is partially updated after kenel had lost state of the block, which can happen under memory pressure.
bluestreak··on Show HN: rust-maven-plugin: Compile Rust JNI libraries in Java Maven projects
I am excited for what Memory API brings. The biggest pain with native memory management for us are "struct" or lack of those. Extremely bug prone memory arithmetic I personally am very motivated to see the back of.

FFI - we need to wait and see how this actually evolves.

We are not looking at any of this just yet. While it is undoubtedly fun, there are few other things we're busy with.

bluestreak··on AVX 512 will be the future
I couldn't agree more. I don't think compiler vectorization is that useful even for columnar (!) database we're building. The specialized JIT doesn't even use AVX512 because too much effort for little to no gain.
bluestreak··on Importing 3m rows/SEC with io_uring
Thank you for sharing the post! I’m vlad, cto of questdb - we went all the way trying to optimize CSV ingestion with io_uring, and ended up benchmarking against some OLAP peers in the process, Any feedback is greatly appreciated !
bluestreak··on No, QuestDB is not Faster than ClickHouse
The query Clickhouse picked on does not actually leverage time order. Perhaps clickhouse vendors on this thread can comment on relevance of the date partitioning for this query. My best guess is that it might help the execution logic to create data chunks for parallel scan.

QuestDB does also use partitions for this purpose but we also calculate chunks dynamically based on available CPU to distribute load across cores more evenly

bluestreak··on No, QuestDB is not Faster than ClickHouse
"fair" means that we comparing apples to apples. Ad-hoc, unindexed predicate, compiled by QuestDB into AVX2 assembly (using AsmJIT) vs same predicate complied by Clickhouse (I'm assuming by LLVM). One can perhaps view this as comparing SIMD-based scans from both databases. Perhaps we generate better assembly, which incidentally offers better IO.

We all understand that creating very specific index might improve specific query performance. Great, Clickhouse geared the entire table storage model to be ultra specific for latitude search. What if you search by longitude, or other column? Back to the beginning.

JIT-compiled predicates offer arbitrary query optimisation with zero impact on ingestion. This is sometimes useful.

What would you offer assuming that we reached out, other than creating an index?

Clickhouse does better than we do in other areas. It JITs more complicated expressions, such as some date functions. It optimises count() queries specifically. For example we collect "found" rowed_ids in an array. Clickhouse does not specifically for count(). We still have work to do. On other hand we ingested this very dataset about 5x quicker than clickhouse, which we left out because article is not about "QuestDB is faster than Clickhouse"

bluestreak··on No, QuestDB is not Faster than ClickHouse
Full disclosure: I am CTO of QuestDB and I took part in JIT implementation. The quote above is not mine, it was written by Clickhouse staff. "utilizes its full indexing strategy" statement is false and is news to me.
bluestreak··on No, QuestDB is not Faster than ClickHouse
I am in fact very proud of my team, who worked very hard on both implementation and the article. It is disappointing to read unfounded insults where we made every effort to be fair.
bluestreak··on No, QuestDB is not Faster than ClickHouse
Our article in question can be found here: https://questdb.io/blog/2022/05/26/query-benchmark-questdb-v...

The intent of the article was to showcase JIT-optimised WHERE clause and we did not use any indexes on QuestDB.

bluestreak··on DeWitt Clause, or can you benchmark %database% and get away with it
It does result in cease and desist threats quite often. We have been on the receiving end of one.
bluestreak··on 4Bn rows/sec query benchmark: ClickHouse vs. QuestDB vs. Timescale
Aggregation is also optimised quite a bit via SIMD and map-reduce. They are as fast as the “where” predicates. Multiple field keyed aggregation is not as optimal yet. I would also suggest our demo site (free and fully open) to see how queries that you use work.

Cleanup is semi manual for now. Time partitions can be removed or detached via SQL. We’re working on automating that.

bluestreak··on 4Bn rows/sec query benchmark: ClickHouse vs. QuestDB vs. Timescale
Last year we released QuestDB 6.0 and achieved an ingestion rate of 1.4 million rows per second (per server). We compared those results to popular open source databases [1] and explained how we dealt with out of order ingestion under the hood while keeping the underlying storage model read-friendly. Since then, we focused our efforts on making queries faster, in particular filter queries with WHERE clauses. To do so, we once again decided to make things from scratch and built a JIT (Just-in-Time) compiler for SQL filters, with tons of low-level optimisations such as SIMD. We then parallelized the query execution to improve the execution time even further. In this blog post, we first look at some benchmarks against Clickhouse and TimescaleDB, before digging deeper in how this all works within QuestDB's storage model. Once again, we use the Time Series Benchmark Suite (TSBS) [2], developed by TimescaleDB,: it is an open source and reproducible benchmark.

We'd love to get your feedback!

[1]:https://news.ycombinator.com/item?id=27411307

[2]:https://github.com/timescale/tsbs

bluestreak··on Are you sure you want to use MMAP in your database management system? [pdf]
Questdb's author here. I do share Ayende's sentiment. There are things that the OP paper doesn't mention, which can help mitigate some of the disadvantages:

- single-threaded calls to 'fallocate' will help avoiding sparse files and SIGBUS during memory write - over-allocating, caching memory addresses and minimizing OS calls - transactional safety can be implemented via shared memory model - hugetlb can minimize TLB shootdowns

I personally do not have any regrets using mmap because of all the benefits they provide

bluestreak··on Building inter-thread messaging from scratch
There is an option to park threads, publisher can signal the consumer if needed.

This model is non-blocking to allow applications to have fewer thread and those threads to process more than one queue or more generally do more than one task. The rationale here is that waking threads up and parking them dramatically increases queue latency.

QuestDB is indeed using this model, but it does not have dedicated threads for each queue. Each thread is able to back-off filling up an already full queue and instead take role of consumer and drain the queue instead. Having non-blocking queues helps to promote work stealing and generally doing other tasks in the same thread.

bluestreak··on Show HN: Demo geospatial and timeseries queries on 250k unique devices
Hello HN,

We previously shared benchmark results showing write speeds of 1.4 million rows per second [1] and how we built this ingestion system. Over the last few months, we've developed functionality that supports geospatial data in our time series database. We decided to add geohashes to our type system along with language features to support handling this type.

We put a lot of thought into optimizing these features from a storage, performance, and usability perspective and we've put together a blog post [2] that gives a tour of the features including implementation details. We've also updated our live demo [3] which now includes an example data set of 250k objects and sample queries to run against this data set.

Given that it's the first time we support geodata in our database, we learned a lot about usage patterns and characteristics of spatial data. We would be happy to hear your thoughts on how we went about adding support for this, where we could improve, and other types of geospatial data we could including in future.

Thanks, Vlad

[1] https://news.ycombinator.com/item?id=27411307

[2] https://questdb.io/blog/2021/10/04/geospatial-timeseries-dem...

[3] https://demo.questdb.io/

bluestreak··on How we achieved write speeds of 1.4M rows per second
I have not had any experience with these trees, but I will try to read and make sense of this paper, thanks! There is C++ implementation here is someone else is interested: https://github.com/rahulyesantharao/b-epsilon-tree
bluestreak··on How we achieved write speeds of 1.4M rows per second
Reading data via index would lead to scattered memory reads on actual data, which is a dead-end from an optimisation standpoint unfortunately.

Using index would have been much simpler but also slower.

bluestreak··on Building a new vector based storage model
Query performance would be affected in so far as ingest jobs share the same thread pool as query jobs. As I am writing this I am also realising that perhaps we should have an option to separate these jobs... If we ignore resource usage and commit() latency, query performance would remain unaffected. Reader remains lockless largely unchanged code-wise. This was one of our major objectives to maintain data model as seen by the readers. I hope I'm making sense here?
bluestreak··on Building a new vector based storage model
Thank you! We are not yet distributed. That’s coming right up along with Jensen style tests. We are really serious about testing!
bluestreak··on Building a new vector based storage model
We launched QuestDB last summer [1, 2]. Our storage model is vector-based and append-only. This meant that all incoming data had to arrive in the correct time order. This worked well for some use cases but we increasingly saw real-world cases where data doesn't always land at the database in chronological order. We saw plenty of developers and users come and go specifically because of this technical limitation. So it became a priority to deal with out-of-order data.

The big decision was which direction to take to tackle the problem. LSM trees seemed an obvious choice, but we chose an alternative route so we wouldn't lose the performance we spent years building. Our latest release supports out-of-order ingestion by re-ordering data on the fly. That's what this article is about.

Also, we had many people asking about the differences between QuestDB and other open-source databases and why users should consider giving it a try instead of other systems. When we launched on HN, readers showed a lot of interest in side-by-side comparisons to other databases on the market. One suggestion [3] that we thought would be great to try out was to benchmark ingestion and query speeds using the Time Series Benchmark Suite (TSBS) [4] developed by TimescaleDB. We're super excited to share the results in the article.

[1] https://news.ycombinator.com/item?id=23975807

[2] https://news.ycombinator.com/item?id=23616878

[3] https://news.ycombinator.com/item?id=23977183

[4] https://github.com/timescale/tsbs

bluestreak··on Redpanda – A Kafka-compatible streaming platform for mission-critical workloads
this would persist kafka traffic to database to make it available long term and at high speed :)
bluestreak··on Redpanda – A Kafka-compatible streaming platform for mission-critical workloads
are there plans to support Kafak connector API, specifically PostgreSQL?
bluestreak··on NYC taxi meter and options pricing
This formula breaks down when you have a kids that go to school and worse still, two separate schools. Not everyone can hop on a bike.
bluestreak··on We chose Java for our high-frequency trading application
It is a good question. I feel there is a happy medium where boilerplate is in Java and more intricate data processing routines are in C++. Makes both worlds simpler. I like C++ better. Those things you mentioned - Java truly sucks at indeed, but you don't always need them. When it comes to IDE, testing, compilation speed, cross platform code and finding talent - Java is way easier than C++.
Page 1 of 5Next →