Great data point, and I think it's the same failure mode: retrieval is good at "what's relevent", bad at "what's still in force". Validity is a property you have to assign, not something similarity can recover.
Honest anser on where I am: the rule side is the stronger half. Each rule can carry a citation (the exact policy sentence it encodes), and the full_audit response records every rule evalution, the facts it tested, and the chunks retrieved. So the verdict and the clause behind it are captured at decision time, independent of any later retrieval.
What I don't hae yet is first-class versioning. Saving rules replaces the set, and a decision doesn't record which version of the rules it ran against. Each decision gets an ID and the audit response is self-contained, but today there's no endpoint to fetch it back later, so the caller has to store it. For "why was this rejected six months ago", the answer has to come from that stored, not a re-run.
The RAG side has exactly the gap you describe. Chunks carry a source but no effective date or supersession, so nothing stops it from retrieving an outdated passage to narrate a new decision.
Your "latest defination wins" result pushes me towards the obvious fix: make validity deterministic too, rather than hoping ranking sorts it out. The plan:
1. Immutable rule-set versions, with a hash stamped on every decision.
2. Documents with effective_from/supperseded_by, filtered by the decisions built from that record and the citations of that version, never from live re-retrieval.
Curious whether your 38/40 with full history failed on the same cases each time, or randomly? That would say a lot about whethe it's a context-length problem or a "which one wins" problem.