How not to break a search engine
about.sourcegraph.com
about.sourcegraph.com
Something tells me your results are a little more personalized or you are in a country specific search that isn't the US.
Shitstorms these days blow over quickly. Someone is always ready for the next big outrage.
off the top of my head, there's two meaning - performance benchmarking (how fast the search results comes back), and accuracy/fit-for-purposeness benchmarking (how good it is at finding something the user intends).
Performance is easy. It's the accuracy/fit-for-purposeness that would be an interesting benchmark.
I wonder if you have to use an empirical measurement for accuracy - that is, give a random sample of people a target piece of code (or file) to find, and see how long or how many queries it takes to find it.
However, you can do things like generate search terms from your top N documents through some method, and then do the queries and confirm the document you generated the term from shows up in the top M results.
This can be circular though if you're not careful; the top N documents may not include important documents that nobody could find.
The whole book is superb, and while some of the content is a bit long in the tooth at this point, the evaluation chapter has aged extremely well.
Oh, and if you’re interested in how to evaluate search engine user interfaces, the equivalent chapter in Hearst’s book on Search UI has you covered: https://searchuserinterfaces.com/book/sui_ch2_evaluation.htm...