At the time I was working with either GPT 3.5 or 4, and – looking at my old code - I was limiting myself to around 14k tokens.
> Were you happy with the results?
It was somewhat ok. I was testing the system from a (copyrighted) psychiatry textbook book and getting feedback on the output from a psychotherapist. The idea was to provide a tool to help therapists prep for a session, rather than help patients directly.
As usual it was somewhat helpful but a little too vague sometimes, or missed important specific information for situation.
It is possible that it could be improved with a larger context window, having more data to select, or different prompting. But the frequent response was along the lines of, "this is good advice, but it just doesn't drill down enough".
Ultimately we found that GPT3.5/4 could produce responses that matched or exceeded our RAG-based solution. This was surprising as it is quite domain specific, but also it seemed pretty clear that GPT must have be trained on very data very similar the (copyrighted) content we were using.
Further steps would be:
1. Use other LLM models. Is it just GPT3.5/4 that is reluctant to drill down?
2. Used specifically-trained LLMs(/LORA) based on the expected response style
I'd be careful of entering this kind of arms race. It seems to be a fight against mediocre results, and at any moment OpenAI at al may release a new model that eats your lunch.