Let's estimate the costs of compute.
For indexing, they need 2800 vCPUs[1], and they are using c6g instances; on-demand hourly price is $0.034/h per vCPU. So indexing will cost them around $70k/month.
For search, they need 1200 vCPUs, it will cost them around $30k/month.
For storage, it will cost them $23/TB * 20000 = $460k/month.
Storage costs are an issue. Of course, they pay less than $23/TB but it's still expensive. They are optimizing this either by using different storage classes or by moving data to cheaper cloud providers for long term storage (less requests mean you need less performant storage and usually you can get a very good price on those object storages).
On quickwit side, we will also improve the compression ratio to reduce the storage footprint.
[1]: I fixed the num vCPUs number of indexing, it was written 4000 when I published the post, but it corresponded to the total number of vCPUs for search and indexing.
1PB with triple redundancy costs around ~$20k just in hard drive costs per year. That's ~$2.5M per year just in disks.
I'd be impressed if they're doing this for less than $1.5M per month (including SWE costs).
Obviously, if they can, saving $1.5M a month vs BigQuery seems like maybe a decent reason to DIY.
The money motivation to self host on bare metal at this scale is huge.
The cost per year is much higher - that's using a 5-year amortization.
You can get a spinning disk of 18TB (not need for SSD if you can parallel write) for 224€. Let's round that to $300 for easy calculations.
To store 100 petabytes of data by purchasing disks yourself, you would need approximately 5556 18TB hard drives totaling $1,666,800.
Of course, you'll pay more than the disks.
Let's add the cost of 93 enclosures at $3,000 each ($279,000), and accounting for controllers, network equipment ($100,000), and power and cooling infrastructure ($50,000, although it's probably already cool where they will host the thing), that would be a about $2.1 M.
That's total, and that's for the uncompressed data.
You would need 3 times that for redundancy, but it would still be 40% cheaper over 5 years, not to mention I used retail price. With their purchasing power they can get a big discount.
Now, you do have the cost of having a team to maintain the whole thing but they likely have their own data center anyway if they go that route.
Do note that you can put, like, at most?, 1TB of hot/warm data on this 18TB drive.
Imagine you do a query, and 100GB of the data to be searched are on 1 HDD. You will wait 500s-1000s just for this hard drive. Imagine a bit higher concurrency with searching on this HDD, like 3 or 5 queries.
You can't fill these drives full with hot or warm data.
> To store 100 petabytes of data by purchasing disks yourself, you would need approximately 5556 18TB hard drives totaling $1,666,800.
You want to have 1000x more drives and only fill 1/1000 of them. Now you can do a parallel read!
> You would need 3 times that for redundancy
With erasure coding you need less, like 1.4x-2x.
Backplace storage pods are an initial investment of 5 Million, thats probably the best bet you could do and on that savings level, having 1-3 good people dedicated to this is probably still cheaper.
But you could / should start talking to the big cloud providers to see if they are flexible enough going lower on the price.
I have seen enough companies, including big ones, being absolut shitty in optimizing these types of things. At this level of data, i would optimize everyting including encoding, date format etc.
But i said it in my other comment: the interesting questions are not answered :D
To further reduce the storage costs, you can use S3 Storage Classes or cheaper object storage like Alibaba for longer retention. Quickwit does not handle that, so you need to handle this yourself, though.
So the underlying storage is still Object storage, so base that around your calculations depending if you are using S3, GCP Object Storage, self hosted Ceph, MinIO, Garage or SeaweedFS.
> Size on S3 (compressed): 20 PB
There are also charts about vCPUs and RAM for the indexing and searching clusters.
Maybe Quickwit is that indexing layer in this case? I haven't dug too much into the general state of cloud dw indexing.
There are no equivalent technology, apart maybe:
- Chaossearch but it is hard to tell because they are not opensource and do not share their internals. (if someone from chaossearch wants to comment?)
- Elasticsearch makes it possible to search into an index archived on S3. This is still a super useful feature as a way to search punctually into your archived data, but it would be too slow and too expensive (it generates a lot of GET requests) to use as your everyday "main" log search index.