RAG can cost a lot of money if not done thoughtfully. Most embedding and chat completion model providers charge by the token (think number of words in the request). You'll pay to have the data in the database transformed into embeddings, that is mostly a one-time fixed cost. Then every time there is a search query in RAG, that question needs to be transformed. The chat completion model (like ChatGPT 4), will charge for the number of tokens in the request + number of tokens in the response.
Self-hosting can be a big advantage for cost control, but it can be complicated too. Tembo.io's managed service provides privately hosted embedding models, but does not have hosted chat completion models yet.