Let's just consider IO, as it is the main effect here. With a columnar, you have to read all of the fields targeted by your query.
If this is log, let's assume
- 200B per line of logs
- 100TB of logs = 500 billions lines of logs
- 30 days of retention.
- a body text field taking 70% of your data.
- highly compressible (10x compression ratio). Note this will come with a higher cpu cost, but let's focus on an IO lower bound.
If you have 100TB of data, regardless of the query, you will have to read (and decompress, but let's not talk about cpu) 7TB worth of data.
Now with an inverted index? You will have to read the posting lists only. The posting lists are delta-encoded and bitpacked.
Assuming the probability of presence of a given term in a log line is p, the worst thing that can happen is having a token that is in 99% of the documents.
In that case, you will have to read 2.06 bits per documents. That's 128GB for the worst posting list.
If you are looking for a single keyword, in the worst possible case, you will have to read 54 times less data than with the columnar solution.
In practice, users search for several keywords, but also considerably less pathological than the example I just gave. Overall you will typically end up reading 20 to 100 times less data than with the grep solution.
I left CPU aside, but actual search engines are also much more CPU efficient.
But I said there was a trade-off... where is it? Well you had to pay a much higher cost at indexing.
Some search engine implementation makes it seem like indexing is more expensive than it should be. Quickwit/tantivy are especially efficient there. With a 4vCPUs VM, you can expect to index at 2TB/day. So in the example above, you will have to dedicate 8vCPUs for indexing. This is perfectly reasonable.
BUT if your retention is much much shorter (few days), indexing might not be worth it.
If your volume of data is small too, you probably do not need to care at all about efficiency.
The inverted index and columnar storage are part of tantivy [0], which is the fastest OSS search library out there (except for the academic project pisa) [1]. We maintain it, and we decided to build the distributed engine on top of it.
[0] tantivy github repo: https://github.com/quickwit-oss/tantivy
[1] tantivy bench https://tantivy-search.github.io/bench/