I think one thing you can try is to figure out what lies where, so chunking arbitrarily will not work as well as chunking with headings for e.g
For e.g, if the question is: "What should the opposition leader have mentioned in their response that the minister would have found difficult to answer."
An embedding based search will find it fairly difficult to match this against a text. Based on my experience, you have to figure out what is a nonanswer first (i don't think that's easy but gpt4 is very good at a lot of language based stuff.)
you can try Q1: question A1: Answer then prompt GPT, do you think A1 answers the question and then save it.
And then,
Q1: question, A1: Answer, Q2: Follow up based on the following questions, do you think A1 answers Q1 and then save it to a db.
You can then augment it in the code with our own knowledge of how politicians lie, using certain words etc :) to improve what gpt4 might miss...