Another idea I've had is to "overfit" a generative model like GPT on a dataset but pay more attention to how url and the like are tokenised
Another idea I've had is to "overfit" a generative model like GPT on a dataset but pay more attention to how url and the like are tokenised
Here you go https://twitter.com/theseamouse/status/1614453236349693953
You have late-interaction models, which replace the dot product with a few transformer layers and are able to learn complex semantics.
Of course this would adversely affect latency and embedding size, so you might want to compress and cache the answers, hence (shameless plug):
https://huggingface.co/sentence-transformers/multi-qa-MiniLM...
[0] https://python.langchain.com/en/latest/modules/chains/index_...
That's what hypothetical embeddings solve: https://summarity.com/hyde
There are also encoding schemes for question-answer retrieval (e.g. ColBERT)
If the embeddings are worth their salt, then they should not be influenced by paraphrasing with different words. Try the OpenAI embeddings or sbert.net embedding models.
Also would you just return a list of likely candidates and loop over the result set to see if any info is relevant to the question and then have the the final pass try to answer the question.