The article is referring to the problem of having a limited context length in LLMs. That is you can only pass X tokens in the prompt.
For example, let’s say you have a prompt that lets you answer questions about a book. If the book is long enough, you won’t be able to include it as is in the prompt, so you have to figure out what are the most relevant passages you must include to answer a given question. What you usually do is find the passages that are the most semantically similar to your question.
Chunk vectors are the vectorized passages of the book (i.e., a numerical vector that represents a passage), and the prompt vector is usually the vectorized question.
To find the most similar vectors you need a distance measure, cosine similarity being the most popular.
The output of finding the most similar vectors is the vectors + it’s metadata (chunk, page, chapter, etc)