To my understanding, RAG is perfect for this.
Models are more for completion. Think of it like autocomplete. If you wanted a model to be good at storytelling, you'd train a model for that. Or say, writing Assembly code. It's like you write "Go to" and the completion model figures out the next word, which may be "jail", "Mexico" or "END".
Fine tuning is a way to bias the completion towards something. In general, it's better to fine tune a general model like Llama or GPT-4 than train it from scratch.
Embeddings models are there to decide which words are related to another. So you might say cat and dog are near each other. Or cat and gato. But cat and "go to" are far from each other. Where encoding turns letters and numbers to bits, embeddings turn words, phrases, images, sounds into vectors.
Since vectors are a little different to bits, they're stored in vector DBs. Vector DBs are often a pain in the ass to deal with. And embedding is super cheap. So, often RAG embeds the entire book each time in tutorials. This is not good practice.
RAG is really a fancy term meaning query, then generate based on that query. So tutorials would embed a million words, toss that in memory, query the memory, then throw it out. That's... wasteful. But not as wasteful as training a model. You should store it in a vector DB, then query that DB.
If it's for a few uses, then just use langchain for RAG. If it's for over 100, then you want to convert your text into embeddings and put that on a vector DB. If it's small and static, LanceDB is fine. Or pgvector (Supabase supports this).
If you want scale, there's plenty of others, but the price goes up fast. Zillis and qdrant seem be good at higher levels, especially if the text is updated continuously.