Research demo
RunTime vs baseline LLM: a recorded retail execution benchmark.
The same set of retail service tasks is run two ways: an open-ended frontier-model agent, and a Proverify RunTime compiled with Razor policy guardrails. Step through the canonical exchange, or browse all 18 scenario variants.
Loading recorded benchmark data…
Metrics at this step
Customer message
Perturbation
Case metrics
Aggregate across all 18 recorded combinations
Claim boundary. This is a recorded tau2-bench-derived retail demo, not an official tau benchmark score and not any third party’s commercial agent. The point is architectural: a pure LLM agent can propose and execute policy-violating consequential actions, while ProVerify can block those actions before the tool layer because authorization is compiled into deterministic runtime control.