They support uploading documents to it for that via that code interpreter, and they're adding connectors to applications where the documents live, not sure what more you're expecting.
They support uploading documents to it for that via that code interpreter, and they're adding connectors to applications where the documents live, not sure what more you're expecting.
Edit: spelling
That being said you can use fine tuning to improve retrieval, which indirectly improves recall. You can do things like fine tune the model you're getting embeddings from, fine tune the LLM to craft queries that better match a domain specific format, etc.
It won't replace the expensive on-the-fly retrieval but it will let you be more accurate in your replies.
Also retrieval can be infinitely faster than inference depending on the domain. In well defined domains you can run old school full text search and leverage the LLMs skill at crafting well thought out queries. In that case that runs at the speed of your I/O.