Just had a look at the code. It’s a cool project that’s clearly had a lot of thought put into it.
If the devs are still around, I’d love to hear about your experiences with embeddings.
If the devs are still around, I’d love to hear about your experiences with embeddings.
2. We don't use any vector datastores (yet). You can do a lot in memory, it's faster and it does exact matches (no KNN, approx matching)
Feel free to ask if you were looking for something more specific?
1. content / question vector mismatch
2. what types of embedding you experimented with storing per-chunk (text only? Hypothetical question? Metadata?)
3. choice of embeddings model (eg OpenAI vs instructorEmbeddings or an alternative from the MTEB leaderboard)
It’s a great project, going to have a deeper dig today.