https://cs.uwaterloo.ca/~jimmylin/publications/Nogueira_Lin_...
Do you have any thoughts on this or similar approaches in production?
Results for R@1000 look pretty impressive, and I'll check out the project code. Given the high recall and low MRR, using this for the initial recall step with a rerank is definitely worth looking at. if the high recall carries over to your own data and you can rerank those top 1k to increase precision, then you've got something good.
Of course there are many caveats here, and this is by no means a solved problem, but works well for queries that are more than a couple terms.
I'd treat it as a base signal like BM25 for a field - which is frequently tuned and used as one of many features for relevance.
for reference: https://ai.googleblog.com/2019/06/introducing-tensornetwork-...
From my perspective of usually dealing with much smaller corpuses, SBERT helps with in a couple ways here practically: it reduces size requirements by an order of magnitude, and also represents sentence context easier than blending individual tokens - which IMO more closely matches the goal of information needs with longer queries.