HNHacker News
TopNewBestAskShowJobs

francoismassot

831 karma · joined April 25, 2021

Co-founder at Quickwit

https://quickwit.io https://github.com/quickwit-oss/quickwit

submissionscomments
francoismassot··on How we replaced Elasticsearch and MongoDB with Rust and RocksDB
it's tantivy :)
francoismassot··on Datadog acquires Quickwit
Co-founder of Quickwit here. Seeing our acquisition by Datadog on the HN front page feels like a truly full-circle moment.

HN has been interwoven with Quickwit's journey from the very beginning. Looking back, it's striking to see how our progress is literally chronicled in our HN front-page posts:

- Searching the web for under $1000/month [0]

- A Rust optimization story [1]

- Decentralized cluster membership in Rust [2]

- Filtering a vector with SIMD instructions (AVX-2 and AVX-512) [3]

- Efficient indexing with Quickwit Rust actor framework [4]

- A compressed indexable bitset [5]

- Show HN: Quickwit – OSS Alternative to Elasticsearch, Splunk, Datadog [6]

- Quickwit 0.8: Indexing and Search at Petabyte Scale [7]

- Tantivy – full-text search engine library inspired by Apache Lucene [8]

- Binance built a 100PB log service with Quickwit [9]

- Datadog acquires Quickwit [10]

Each of these front-page appearances was a milestone for us. We put our hearts into writing those engineering articles, hoping to contribute something valuable to our community.

I'm convinced HN played a key role in Quickwit's success by providing visibility, positive feedback, critical comments, and leads that contacted us directly after a front-page post. This community's authenticity and passion for technology are unparalleled. And we're incredibly grateful for this.

Thank you all :)

[0] https://news.ycombinator.com/item?id=27074481

[1] https://news.ycombinator.com/item?id=28955461

[2] https://news.ycombinator.com/item?id=31190586

[3] https://news.ycombinator.com/item?id=32674040

[4] https://news.ycombinator.com/item?id=35785421

[5] https://news.ycombinator.com/item?id=36519467

[6] https://news.ycombinator.com/item?id=38902042

[7] https://news.ycombinator.com/item?id=39756367

[8] https://news.ycombinator.com/item?id=40492834

[9] https://news.ycombinator.com/item?id=40935701

[10] https://news.ycombinator.com/item?id=42648043

francoismassot··on Ask HN: What's your preferred logging stack in Kubernetes
Latest HN thread on quickwit (Binance built a 100PB log service with Quickwit): https://news.ycombinator.com/item?id=40935701

I also wrote a benchmark on Loki vs. Quickwit: https://quickwit.io/blog/benchmarking-quickwit-loki

francoismassot··on Binance built a 100PB log service with Quickwit
Indeed. They benefit from a discount, but we don't know the discount figure.

To further reduce the storage costs, you can use S3 Storage Classes or cheaper object storage like Alibaba for longer retention. Quickwit does not handle that, so you need to handle this yourself, though.

francoismassot··on Binance built a 100PB log service with Quickwit
They have 181 trillion logs
francoismassot··on Binance built a 100PB log service with Quickwit
Good question.

Let's estimate the costs of compute.

For indexing, they need 2800 vCPUs[1], and they are using c6g instances; on-demand hourly price is $0.034/h per vCPU. So indexing will cost them around $70k/month.

For search, they need 1200 vCPUs, it will cost them around $30k/month.

For storage, it will cost them $23/TB * 20000 = $460k/month.

Storage costs are an issue. Of course, they pay less than $23/TB but it's still expensive. They are optimizing this either by using different storage classes or by moving data to cheaper cloud providers for long term storage (less requests mean you need less performant storage and usually you can get a very good price on those object storages).

On quickwit side, we will also improve the compression ratio to reduce the storage footprint.

[1]: I fixed the num vCPUs number of indexing, it was written 4000 when I published the post, but it corresponded to the total number of vCPUs for search and indexing.

francoismassot··on Turbopuffer: Fast search on object storage
But you don’t have fast search on those files stored on object storage.
francoismassot··on Turbopuffer: Fast search on object storage
If you don't need vector search and have very large Elasticsearch deployment, you can have a look at Quickwit, it's a search engine on object storage, it's OSS and works for append-only datasets (like logs, traces, ...)

Repo: https://github.com/quickwit-oss/quickwit

francoismassot··on Tantivy – full-text search engine library inspired by Apache Lucene
One workaround is to use the JSON field, see doc https://github.com/quickwit-oss/tantivy/blob/main/doc/src/js...
francoismassot··on Pg_lakehouse: Query Any Data Lake from Postgres
Well, MongoDB was under AGPL v3.0 :)
francoismassot··on Show HN: OneUptime – open-source Datadog Alternative
Quickwit is an alternative with a strong focus on scalability (max we have seen is 40PB) with a decoupled compute and storage architecture. But we do only logs and traces for now.

Repository: https://github.com/quickwit-oss/quickwit Latest release: https://quickwit.io/blog/quickwit-0.8

francoismassot··on Show HN: Tracecat – Open-source security alert automation / SOAR alternative
This is awesome; we need this kind of alternative to overpriced software like Splunk. We built and open-sourced Quickwit to see this kind of tool built on top of it.

We will follow Tracecat closely. I'm convinced this will impact our roadmap, and I'm happy to receive any feedback so you can get the most out of Quickwit for Tracecat.

Good luck, guys!!!

francoismassot··on Quickwit 0.8: Indexing and Search at Petabyte Scale
tantivy, not tantivity!!!!!
francoismassot··on Quickwit 0.8: Indexing and Search at Petabyte Scale
Thanks! Quickwit is the distributed engine built on top of tantivy, we basically separated compute and storage for search, I wrote this blog post to introduce the architecture: https://quickwit.io/blog/quickwit-101

PS: it’s tantivy!!!

francoismassot··on Quickwit 0.8: Indexing and Search at Petabyte Scale
Some companies are using it with AWS Lambda to scale to 0.
francoismassot··on Quickwit 0.8: Indexing and Search at Petabyte Scale
Building the inverted index is quite CPU-intensive, and we are also merging index files called "splits".
francoismassot··on Warning: $14k BigQuery charge in 2 hours
BigQuery is just too costly...

Do you know if the dataset is public? We should just offer a cheap alternative and ditch BigQuery.

francoismassot··on Show HN: Host a planet-scale geocoder for $10/mo
Oh I forget to add stract is using tantivy too, I really hope this project will take off.

https://stract.com/

https://github.com/StractOrg/stract

https://news.ycombinator.com/item?id=39254172

francoismassot··on Show HN: Host a planet-scale geocoder for $10/mo
> "Do store the Sonic database on SSD-backed file systems only."

From the README, it works only on SSD.

All those projects serve different purposes, and several are not actively maintained.

- Meilisearch: It provides a search-as-you-type experience and comes with many features; I don't know it very well, but I think it targets first e-commerce/application search.

- Quickwit: it's a distributed search engine for append-only data and works well on S3, a good fit for observability/security/financial/... data.

- Sonic: it looks like it targets search-as-you-type use cases and does not provide many features (which can be a very good feature in itself as it remains very light).

- Tantivy: It's a library; you need to build your server on top of it if you want an HTTP API. toshi, lnx did. It's used by a lot of search projects like tabbyML, Milvus, bloop, paradedb, airmail...

francoismassot··on Show HN: Host a planet-scale geocoder for $10/mo
The geocoder is built on top of tantivy which is fast and uses low resources too (https://github.com/quickwit-oss/tantivy).

I'm curious about the comparison between those two.

francoismassot··on Anki – Powerful, intelligent flash cards
I used it with my 10 years old boy for spelling.

I like the method. I found the app is still rough on the edges, and now I want to code a small one dedicated to science fields for him :)

francoismassot··on Qdrant, the Vector Search Database, raised $28M in a Series A round
Congrats to Qdrant's team, $28M for a Series is really nice.

There are a lot of OSS vector search databases out there, we could probably list the main ones:

- Qdrant: https://github.com/qdrant/qdrant

- Weaviate: https://github.com/weaviate/weaviate

- Milvus: https://github.com/milvus-io/milvus

What else?

francoismassot··on Ceph: A Journey to 1 TiB/s
Does someone knows how Ceph compares to other object storage engine like MinIO/Garage/...?

I would love to see some benchmarks there.

francoismassot··on Show HN: Quickwit – OSS Alternative to Elasticsearch, Splunk, Datadog
Quickwit is under AGPLv3. Are you saying that AGPLv3 is not FOSS?
francoismassot··on Show HN: Quickwit – OSS Alternative to Elasticsearch, Splunk, Datadog
I must admit that 'alternative' is always a tricky word... Datadog, Elasticsearch, and Splunk are giant beasts, and the alternative makes sense only on a subset of features (and hopefully, we will successfully execute our 2024 roadmap to reduce the difference)

For Quickwit, our users proved to us it was scaling up to petabytes. So we consider this scale factor in the "alternative". But... we don't have a dedicated metrics storage engine yet, so if you want to store metrics in Quickwit, it won't be efficient in the current version. It will come later this year.

francoismassot··on Show HN: Quickwit – OSS Alternative to Elasticsearch, Splunk, Datadog
Quickwit supports different data sources: Kafka, Pulsar, Google pubsub, ... and we have our own ingest API (not HA right now, but it will be the case in the next release in 1 month or so).

Postgresql is not mandatory; it's also possible to use Quickwit with a metastore on S3. For large use cases, Postgresql is the way to go. I've seen users using Quickwit with metastore on S3, RDS, and Aurora.

On the UI side, we have several users who have their own UI. Jaeger is used just for the UI part so it's quite simple to have it in HA, I don't thing it's hard to have HA for Grafana but I'm not sure on this point.

Which docker compose did you look at?

francoismassot··on Show HN: Quickwit – OSS Alternative to Elasticsearch, Splunk, Datadog
Good point. Several users are asking us the OpenDashboard/Kibana compatibility, and this is on the 2024 roadmap.

That being said, we also hear users complaining about OpenDashboard/Kibana, looking for an alternative different from Kibana/Grafana explore view (the view used for log and tracing search). You will also find users satisfied by the Grafana Explore view.

Personally, I don't find the Grafana Explore view great for log searches. I saw that Grafana recently made some improvements, and I need to dig into that to adapt the Quickwit Grafana plugin. I don't have a clear opinion on Kibana, one of my dreams is to build a better UI for log/traces search anyway, not yet on the roadmap though :)

francoismassot··on Show HN: Quickwit – OSS Alternative to Elasticsearch, Splunk, Datadog
So yes, currently we only support infrequent deletes for GDPR reasons mainly.

It's possible to add updates/deletes to Quickwit, but this is a lot of work, and for now, we have not prioritized this development.

Do you mind sharing your use case?

francoismassot··on Show HN: Quickwit – OSS Alternative to Elasticsearch, Splunk, Datadog
Well, I would say it depends. We have many companies using the AGPL version without buying a license. We also know that some companies have strict policies and will forbid using AGPL software unless taking a commercial license. We're happy with both users.

I like the example of Grafana with all their AGPL projects (Grafana, Loki, Tempo, ...). There are a LOT of companies using Grafana with the AGPL version.

francoismassot··on Show HN: Quickwit – OSS Alternative to Elasticsearch, Splunk, Datadog
So Quickwit is primarily a search engine and thus relies on an inverted index. We also implemented our schemaless columnar storage optimized for object storage.

The inverted index and columnar storage are part of tantivy [0], which is the fastest OSS search library out there (except for the academic project pisa) [1]. We maintain it, and we decided to build the distributed engine on top of it.

[0] tantivy github repo: https://github.com/quickwit-oss/tantivy

[1] tantivy bench https://tantivy-search.github.io/bench/

Page 1 of 4Next →