> as every time you make a call, the system loads a large chunk or many chunks and sends them to the model along with your prompt,
This is how RAG works.
While you can come up with work-arounds like using lesser LLMs as a pre-filtering step the fact is that if you need GPT to read the doc you need GPT to read the doc.