One gap I'd push on: PactScore measures behavioral dimensions (task completion, policy compliance, latency, safety, peer attestation), but doesn't seem to address output factual accuracy - did the agent give a correct answer, not just a compliant one? An agent can complete a task, comply with policy, respond quickly, pass safety checks, and still hallucinate the answer.
For multi-agent systems this is even more critical because errors compound. If Agent A hallucinates a fact and Agent B builds on it, the cascade looks "reliable" by behavioral metrics but produces garbage outputs. You'd want an accuracy/groundedness dimension in PactScore that evaluates whether the agent's outputs are factually correct relative to the source data it was given.
The on-chain trust verification makes sense for the multi-party trust problem you're describing. Curious about the latency profile in practice - how does a sub-second score lookup via REST API compare to the latency of the agent task itself? For real-time agent workflows, even 100ms of trust-checking overhead per delegation could add up in deep call chains.