Or anything where the question answer isn't 'close' to the words used in the question?
How well does this work vs giving it the whole thing as a prompt?
I assume worse but I'm not sure how this approach compares to giving it the full thing in the prompt or splitting it into N sections and running on each and then summarizing.
The problem is that humans have continuous information retrieval and storage where the current crop of embedding systems are static and mostly one shot.
This weird leaky memory has advantages and disadvantages. Forgetting is useful, it removes garbage.
Machine models could vary the balance of temporal types, drop out Etc. We may get some weird behavior.
I would guess we will see many innovations in how memory is stored in systems like these.
Background: https://summarity.com/hyde
Demo: https://youtu.be/elNrRU12xRc?t=1550 (or try it on findsight.ai and compare results of the "answer" vs the "state" filter)
For even deeper retrieval consider late interaction models such as ColBERT
Does the embedding structure somehow expose the themes? And if so, is it more the embeddings that are answering the question by how it groups things?