I think the most common design pattern nowadays goes like this:
1. Chunk all your data (e.g. per paragraph of content)
2. Generate an embedding for each chunk
3. Index embeddings in a vector database
4. When a query comes in, find chunks relevant to the query (based on embeddings similarity) and ONLY send the relevant chunks + query to a LLM to formulate the answer
Quickly glancing through the repository from this post, I can see that it also follows this pattern. It uses OpenAI's embedding API for 2. and Pinecone DB for 3.