Some gotchas I experienced (but I might be using the wrong embedding/vector DB: spaCy/FAISS):
- Short user questions might result a low signal query vector, e. g. user : "Who is Keanu Reeves?" -> false positives on Wikipedia articles which only contain "Who is"
- Typos and formatting affects the vectorization, a small difference might lead to a miss, e.g. "Who is Keanu Reeves?" -> match, "Who is keanu Reeves?" -> no match, no match with any other capitalization.
If there's only a single document, a simple keyword search might lead to better results.
In my experience, false positives (retrieving an irrelevant text and generating completely wrong answer) are a bigger problem than negatives (not retrieving text, possibly can't answer question).
Has somebody experience with Apache Lucene / Solr or Elasticsearch?