Signal Trust Is Not a Score: Runtime Assurance Contracts for High-Risk AI Agents
Summary
This paper argues that existing benchmarks, audits, and agent protocols describe an AI agent's performance, permissions, and repair procedures, but fail to specify how observed evidence during a consequential task should change the agent's authority in real time. The authors call this gap the 'assurance-transition gap.' They propose Runtime Assurance Contracts (RAC), a policy-level formal schema binding autonomy boundaries, evidence state, and non-compensatory gates, under which a failed or unknown mandatory gate forces retry, switch, escalation, deferral, or stop regardless of aggregate performance score. In a deterministic failure-injection study of 280 constructed agentic-coding cases, a score-only rule using published example weights admitted 80 of 100 injections that should have been blocked while passing all 40 injections that required review. In a further prospective synthetic holdout of 24 episodes covering 72 action attempts, two blinded LLM judges assigned identical labels to every attempt, and both RAC and a separately implemented stateful baseline matched those labels. The authors caution that these results test the mechanism only on synthetic cases and establish neither deployed safety nor cross-domain effectiveness.
Classification
Evidence 1
- Trust Is Not a Score: Runtime Assurance Contracts for High-Risk AI Agents arXiv (cs.AI / cs.CY) 2026-09-30 accessed 2026-10-01T23:00:11+00:00
Part of trends 0
No objects.
Directly linked issues 0
No objects.
Public id: fm-9d7fe2c4e3cc
