1,042 karma · joined May 13, 2016
This is very true. Perhaps there wasn't enough context to what the article is describing. The read problem started to occur on database that is subject to constant write workload. Data is flowing in all the time at variable rate. Typically blocks are "hot" and being filled in fully within seconds if not millis.
Zeroing the file is an option to try. QuestDB allocates disk with `posix_fallocate()`, which doesn't have the required flags. We would need to explore `fallocate()`. Thanks.
FFI - we need to wait and see how this actually evolves.
We are not looking at any of this just yet. While it is undoubtedly fun, there are few other things we're busy with.
QuestDB does also use partitions for this purpose but we also calculate chunks dynamically based on available CPU to distribute load across cores more evenly
We all understand that creating very specific index might improve specific query performance. Great, Clickhouse geared the entire table storage model to be ultra specific for latitude search. What if you search by longitude, or other column? Back to the beginning.
JIT-compiled predicates offer arbitrary query optimisation with zero impact on ingestion. This is sometimes useful.
What would you offer assuming that we reached out, other than creating an index?
Clickhouse does better than we do in other areas. It JITs more complicated expressions, such as some date functions. It optimises count() queries specifically. For example we collect "found" rowed_ids in an array. Clickhouse does not specifically for count(). We still have work to do. On other hand we ingested this very dataset about 5x quicker than clickhouse, which we left out because article is not about "QuestDB is faster than Clickhouse"
The intent of the article was to showcase JIT-optimised WHERE clause and we did not use any indexes on QuestDB.
Cleanup is semi manual for now. Time partitions can be removed or detached via SQL. We’re working on automating that.
We'd love to get your feedback!
- single-threaded calls to 'fallocate' will help avoiding sparse files and SIGBUS during memory write - over-allocating, caching memory addresses and minimizing OS calls - transactional safety can be implemented via shared memory model - hugetlb can minimize TLB shootdowns
I personally do not have any regrets using mmap because of all the benefits they provide
This model is non-blocking to allow applications to have fewer thread and those threads to process more than one queue or more generally do more than one task. The rationale here is that waking threads up and parking them dramatically increases queue latency.
QuestDB is indeed using this model, but it does not have dedicated threads for each queue. Each thread is able to back-off filling up an already full queue and instead take role of consumer and drain the queue instead. Having non-blocking queues helps to promote work stealing and generally doing other tasks in the same thread.
We previously shared benchmark results showing write speeds of 1.4 million rows per second [1] and how we built this ingestion system. Over the last few months, we've developed functionality that supports geospatial data in our time series database. We decided to add geohashes to our type system along with language features to support handling this type.
We put a lot of thought into optimizing these features from a storage, performance, and usability perspective and we've put together a blog post [2] that gives a tour of the features including implementation details. We've also updated our live demo [3] which now includes an example data set of 250k objects and sample queries to run against this data set.
Given that it's the first time we support geodata in our database, we learned a lot about usage patterns and characteristics of spatial data. We would be happy to hear your thoughts on how we went about adding support for this, where we could improve, and other types of geospatial data we could including in future.
Thanks, Vlad
[1] https://news.ycombinator.com/item?id=27411307
[2] https://questdb.io/blog/2021/10/04/geospatial-timeseries-dem...
Using index would have been much simpler but also slower.
The big decision was which direction to take to tackle the problem. LSM trees seemed an obvious choice, but we chose an alternative route so we wouldn't lose the performance we spent years building. Our latest release supports out-of-order ingestion by re-ordering data on the fly. That's what this article is about.
Also, we had many people asking about the differences between QuestDB and other open-source databases and why users should consider giving it a try instead of other systems. When we launched on HN, readers showed a lot of interest in side-by-side comparisons to other databases on the market. One suggestion [3] that we thought would be great to try out was to benchmark ingestion and query speeds using the Time Series Benchmark Suite (TSBS) [4] developed by TimescaleDB. We're super excited to share the results in the article.
[1] https://news.ycombinator.com/item?id=23975807
[2] https://news.ycombinator.com/item?id=23616878