In short, we needed to approximate the count.
We had some kind of hard limit for matching documents during a regular keyword search (80k maybe? It's been years since I worked there). The ranking was done after this, as well as the aggregation of data for the facets. So if your query of "cute bunnies" (faceted on file type) filled up all 80k results by the time the query processor made it through 25% of the data, and those 80k results contained 5k gifs, then we'd display '20k' next to the 'gifs' check box.
A related issue is that counts tend to treat all results as equal. If you retrieva a lot of results but most of them are not relevant -- as can happen with full-text search -- then the counts can be misleading. You may have the converse problem if your retrieval excludes a lot of relevant results. So, if you are implementing a faceted search application where you use and show counts, you should keep in mind that it will only work if your retrieval does a reasonable job of balancing precision and recall.
Finally, remember that supply != demand. The distribution of a facet in your index may be different from the distribution of that facet in searcher intent. A bit more on that here: https://dtunkelang.medium.com/search-intent-not-inventory-28...
generally what you have is lightweight searches that you can query and get just the count of documents you would receive, in the case of facets you can generally get a list of facets or even a list of facets relating to a particular search and a count for each of these facets.
My application Datasette runs a separate SQL query for each one, which works fine if you are using SQLite and only have a few hundred thousand rows of data.
Keyword -> (num_docs, doc1,doc3,doc4,…)
So you can quite quickly look up the number of matches.
With faceted search you usually need to do things like "the user searched for 'x' and filtered for 'price less than $Y', now show counts for each of the different categories" - where pre-calculated counts won't help you.
Search indexes still help here, because they are really good at fast set intersections - so you take the set of document IDs matching your filters so far, then intersect them against the set of IDs that are listed for each of those categories and count the size of that intersection.