Natural language search and answer generation _are_ completely separate.
Search is often (but not exclusively) performed using cosine similarity over semantic vectors. Such vectors are produced using embedding models, which represent the meaning of the document via an arbitrary length vector called an embedding, and 768 is a common vector size for this.
You calculate the embedding for all documents in your database ahead of time (during insertion), and you calculate the embedding for the user's query, then search for documents closest to the query using a similarity metric of choice such as cosine.
Nothing prevents you from serving documents found this way directly, instead of using them to generate answers. Part of the Google search pipeline involves something like this. Many full-text search products also do this (Algolia is such an example https://www.algolia.com/blog/ai/what-is-vector-search/)
LLMs and generation are used to synthesize the answer and tie it back into the question using a layer of soft judgment based on the LLMs prior knowledge. This does work out great in some contexts, less so in others, as you pointed out. But these components aren't coupled in any way.