831 karma · joined April 25, 2021
https://quickwit.io https://github.com/quickwit-oss/quickwit
HN has been interwoven with Quickwit's journey from the very beginning. Looking back, it's striking to see how our progress is literally chronicled in our HN front-page posts:
- Searching the web for under $1000/month [0]
- A Rust optimization story [1]
- Decentralized cluster membership in Rust [2]
- Filtering a vector with SIMD instructions (AVX-2 and AVX-512) [3]
- Efficient indexing with Quickwit Rust actor framework [4]
- A compressed indexable bitset [5]
- Show HN: Quickwit – OSS Alternative to Elasticsearch, Splunk, Datadog [6]
- Quickwit 0.8: Indexing and Search at Petabyte Scale [7]
- Tantivy – full-text search engine library inspired by Apache Lucene [8]
- Binance built a 100PB log service with Quickwit [9]
- Datadog acquires Quickwit [10]
Each of these front-page appearances was a milestone for us. We put our hearts into writing those engineering articles, hoping to contribute something valuable to our community.
I'm convinced HN played a key role in Quickwit's success by providing visibility, positive feedback, critical comments, and leads that contacted us directly after a front-page post. This community's authenticity and passion for technology are unparalleled. And we're incredibly grateful for this.
Thank you all :)
[0] https://news.ycombinator.com/item?id=27074481
[1] https://news.ycombinator.com/item?id=28955461
[2] https://news.ycombinator.com/item?id=31190586
[3] https://news.ycombinator.com/item?id=32674040
[4] https://news.ycombinator.com/item?id=35785421
[5] https://news.ycombinator.com/item?id=36519467
[6] https://news.ycombinator.com/item?id=38902042
[7] https://news.ycombinator.com/item?id=39756367
[8] https://news.ycombinator.com/item?id=40492834
I also wrote a benchmark on Loki vs. Quickwit: https://quickwit.io/blog/benchmarking-quickwit-loki
To further reduce the storage costs, you can use S3 Storage Classes or cheaper object storage like Alibaba for longer retention. Quickwit does not handle that, so you need to handle this yourself, though.
Let's estimate the costs of compute.
For indexing, they need 2800 vCPUs[1], and they are using c6g instances; on-demand hourly price is $0.034/h per vCPU. So indexing will cost them around $70k/month.
For search, they need 1200 vCPUs, it will cost them around $30k/month.
For storage, it will cost them $23/TB * 20000 = $460k/month.
Storage costs are an issue. Of course, they pay less than $23/TB but it's still expensive. They are optimizing this either by using different storage classes or by moving data to cheaper cloud providers for long term storage (less requests mean you need less performant storage and usually you can get a very good price on those object storages).
On quickwit side, we will also improve the compression ratio to reduce the storage footprint.
[1]: I fixed the num vCPUs number of indexing, it was written 4000 when I published the post, but it corresponded to the total number of vCPUs for search and indexing.
Repository: https://github.com/quickwit-oss/quickwit Latest release: https://quickwit.io/blog/quickwit-0.8
We will follow Tracecat closely. I'm convinced this will impact our roadmap, and I'm happy to receive any feedback so you can get the most out of Quickwit for Tracecat.
Good luck, guys!!!
PS: it’s tantivy!!!
Do you know if the dataset is public? We should just offer a cheap alternative and ditch BigQuery.
From the README, it works only on SSD.
All those projects serve different purposes, and several are not actively maintained.
- Meilisearch: It provides a search-as-you-type experience and comes with many features; I don't know it very well, but I think it targets first e-commerce/application search.
- Quickwit: it's a distributed search engine for append-only data and works well on S3, a good fit for observability/security/financial/... data.
- Sonic: it looks like it targets search-as-you-type use cases and does not provide many features (which can be a very good feature in itself as it remains very light).
- Tantivy: It's a library; you need to build your server on top of it if you want an HTTP API. toshi, lnx did. It's used by a lot of search projects like tabbyML, Milvus, bloop, paradedb, airmail...
I'm curious about the comparison between those two.
I like the method. I found the app is still rough on the edges, and now I want to code a small one dedicated to science fields for him :)
There are a lot of OSS vector search databases out there, we could probably list the main ones:
- Qdrant: https://github.com/qdrant/qdrant
- Weaviate: https://github.com/weaviate/weaviate
- Milvus: https://github.com/milvus-io/milvus
What else?
I would love to see some benchmarks there.
For Quickwit, our users proved to us it was scaling up to petabytes. So we consider this scale factor in the "alternative". But... we don't have a dedicated metrics storage engine yet, so if you want to store metrics in Quickwit, it won't be efficient in the current version. It will come later this year.
Postgresql is not mandatory; it's also possible to use Quickwit with a metastore on S3. For large use cases, Postgresql is the way to go. I've seen users using Quickwit with metastore on S3, RDS, and Aurora.
On the UI side, we have several users who have their own UI. Jaeger is used just for the UI part so it's quite simple to have it in HA, I don't thing it's hard to have HA for Grafana but I'm not sure on this point.
Which docker compose did you look at?
That being said, we also hear users complaining about OpenDashboard/Kibana, looking for an alternative different from Kibana/Grafana explore view (the view used for log and tracing search). You will also find users satisfied by the Grafana Explore view.
Personally, I don't find the Grafana Explore view great for log searches. I saw that Grafana recently made some improvements, and I need to dig into that to adapt the Quickwit Grafana plugin. I don't have a clear opinion on Kibana, one of my dreams is to build a better UI for log/traces search anyway, not yet on the roadmap though :)
It's possible to add updates/deletes to Quickwit, but this is a lot of work, and for now, we have not prioritized this development.
Do you mind sharing your use case?
I like the example of Grafana with all their AGPL projects (Grafana, Loki, Tempo, ...). There are a LOT of companies using Grafana with the AGPL version.
The inverted index and columnar storage are part of tantivy [0], which is the fastest OSS search library out there (except for the academic project pisa) [1]. We maintain it, and we decided to build the distributed engine on top of it.
[0] tantivy github repo: https://github.com/quickwit-oss/tantivy
[1] tantivy bench https://tantivy-search.github.io/bench/