Research demo

RunTime vs baseline LLM: a recorded retail execution benchmark.

The same set of retail service tasks is run two ways: an open-ended frontier-model agent, and a Proverify RunTime compiled with Razor policy guardrails. Step through the canonical exchange, or browse all 18 scenario variants.

Loading recorded benchmark data…
Scenario Multi-item delivered exchange

Metrics at this step

Claim boundary. This is a recorded tau2-bench-derived retail demo, not an official tau benchmark score and not any third party’s commercial agent. The point is architectural: a pure LLM agent can propose and execute policy-violating consequential actions, while ProVerify can block those actions before the tool layer because authorization is compiled into deterministic runtime control.