Using vectors for retrieval augmentation seems really powerful at first, but in practice we've found it very finicky. How long should each chunk be? How should you split? Are vector embeddings actually a good space/distance for your domain?
One improvement is to use refinement (ie iterate over many relevant chunks, refining your answer). This is more expensive and slower but less lossy than vectors. LlamaIndex does refinement well.
Would be curious if people have found better patterns. We've been playing around with using one LLM chain to do the retrieval, then passing the retrieved chunks into the main LLM call. Seems to work better for certain contexts.