That's why I ended up gating at the tool-call level with session state instead of at the network level. Same curl gets a different verdict depending on whether the session already read something secret-shaped. Deterministic rules only, and the approval prompt times out to deny, logged as a timeout rather than a denial so I can tell the two apart later.
On the policy-model approach: I'd be careful about what a model's truthy value is allowed to do. A classifier crossing a threshold is a guess with a confidence attached, and a guess that hard-blocks real work gets the whole guardrail disabled by the end of the day. The split I settled on is that a probability can ask (escalate to a human) but only a predicate can refuse. Curious whether you're letting the model produce the deny directly, or routing its output through a human when it's uncertain.
For what it's worth this is what I've been building: https://github.com/DobermanCore/Doberman-Core. Apache-2.0, 100% open source