A recent McKinsey article, “The cost of intelligence: How CIOs can manage AI demand at scale,” moves the AI-cost discussion beyond prompt prices. It defines tokenomics broadly — model selection, routing, orchestration, agent behavior, workflow design, infrastructure utilization, waste — which makes this a control-system problem, not a procurement spreadsheet.

The numbers explain the urgency

The underlying survey is small enough to treat directionally — 120 participants, 75 qualified respondents across five industries — but the pattern is hard to ignore:

~4×increase in AI spend as organizations move from isolated use cases to enterprise-wide adoption
93%of qualified respondents exceeded their AI budgets
20–30%of AI spend often unaccounted for because purchases are fragmented
30×possible variation in token use for the same task (Stanford Digital Economy Lab)

Tokens aren't simply expensive — demand is hard to predict, purchases are distributed, agents multiply calls and retries from one request, and the same task can follow wildly different computational paths.

The core operating principle Govern the completed business outcome, not the token cost. That reframes the question from “How many tokens did we buy?” to “What did the workflow accomplish, what did it cost, what authority did it exercise, and can we prove the result?”

Two control planes, not one

The "AI control plane" spans the management layer between users, applications, agents, models and platforms, with three functions: visibility and attribution, policy and governance, and routing and optimization. Proverify.ai splits that into two distinct accountabilities, because they answer different enterprise questions.

Requirement What has to be controlled Where Proverify addresses it
Visibility & attribution Request-level usage across vendors, models, teams and workflows; allocation to business units, products, use cases, owners and cost centers; unit economics such as cost per task or case. TokenOps controls the ledger: usage and spend visibility, cost and outcome attribution, investment and budget governance, and a financial/operational record of where AI capital is deployed.
Policy & governance Approved access, permissions, budgets, spend thresholds, operating limits, exception paths and enforcement before unmanaged cost or risk occurs. RunTime controls the operation: permissions, limits, escalation paths and operating rules across the models, tools, systems and data a workflow can use, with actions recorded for review.
Routing & optimization Choose the right model/provider for cost, quality, latency and risk; reduce unnecessary context, retries, tool calls and agent loops; avoid paying frontier-model prices when a simpler path works. TokenOps + RunTime connect economics to execution. TokenOps shows where capital is consumed and what outcome it funds; RunTime bounds how the workflow executes. For routine work it has learned, Proverify's current RunTime positioning is software-like execution rather than repeated open-ended AI inference.
Proof of control Architecture diagrams and policy statements are insufficient if the workflow cannot demonstrate correct behavior, exceptions and evidence under real operating conditions. VERIFY is Proverify's production-grade runtime rubric: the test for whether a specialist runtime is actually governed, verifiable and trustworthy enough for a real business process.

TokenOps: the accounting problem

AI spend is spread across cloud providers, model vendors, AI-enabled software, experimentation environments and business-unit purchases. 20–30 percent of it can go unaccounted for in fragmented environments, and overruns typically surface only after the fact.

TokenOps is built to make that legible — not another vendor-cost dashboard, but a ledger that connects consumption to the business activity that caused it: model, workflow, owner, unit, use case and, critically, outcome. That's how an organization moves from invoice reconciliation to numbers management can act on: cost per resolved case, per completed review, per approved transaction, per revenue-generating workflow.

That same ledger is the prerequisite for forecasting, showback or chargeback, sourcing decisions, and a permanent AI FinOps capability. Forecasting without attribution is guessing; optimization without an outcome denominator can just push cost down while quality, latency, or risk deteriorates.

RunTime: the operating problem

The highest-impact savings levers aren't finance actions — they're execution controls: route to the cheapest model that meets the quality bar; limit context; cap outputs; control retries, tool calls, and agent loops; define escalation paths; cache repeated work; redesign poor workflows instead of bolting on more agents.

Those controls only become durable once embedded in the operating environment. RunTime is the execution side: it defines what a workflow may do, which systems and data it may touch, where authority stops, and when a human takes over. For repetitive work, Proverify's RunTime goes further — it learns bounded workflows and executes them like software, escalating anything unfamiliar or risky. That shifts the economics from continuously paying a probabilistic model to reconsider routine work, toward paying AI only where judgment is actually needed.

The largest savings come from optimizing the systems that consume tokens, not from shopping for cheaper ones. A controlled runtime is where model choice, retries, escalation, and workflow shape become enforceable behavior — not a slide in a governance deck.

VERIFY: the evidence problem

Guardrails and policy engines are necessary, but enterprises still need a test for whether the claimed controls actually work on the named workflow. The question isn't whether a vendor can point to a governance feature — it's whether the system can demonstrate production-grade behavior under routine requests, exceptions, uncertain inputs, and actions that exceed its authority.

The VERIFY Runtime Rubric is Proverify's answer: a way to evaluate whether a specialist runtime is production-grade before the enterprise confuses "agent deployed" with "operation controlled." It isn't a third control plane — it's the discipline for validating that the financial and operational controls are real enough to trust.

Sourcing is a continuous decision now

The old buy-versus-build question isn't sufficient anymore. Enterprises now choose continuously among buying, building, hosting, routing, and switching across proprietary, open-weight, and local models — a multimodel, dynamic-demand world, not traditional single-vendor procurement.

That's consistent with Proverify's broader thesis in Autonomous RunTimes (ART): the durable enterprise asset is the specialist operating capability — the workflow, authority structure, integrations, evidence, and operating history — not dependence on any one frontier model. Providers can change underneath the operation; the business process and its controls should stay owned.

Where to apply caution

  • The survey stats are directional, not universal benchmarks: 120 enterprise participants, only 75 qualified respondents across five industries.
  • The 20–30 percent savings range is anecdotal and survey-based, not a guaranteed result for every enterprise or workload.
  • Prompt-caching savings of up to ~90 percent apply mainly to workloads with large, stable, repeated context — don't generalize them to all AI spend.
  • A control plane doesn't create business value by itself. It creates the visibility and enforcement needed to decide whether a workflow is worth operating.
  • Any Proverify claim should be proven on a named workflow with an economic baseline, an operating envelope, and evidence of actual results.

A better buying test

This becomes actionable once procurement shifts from platform claims to a controlled workflow test. Before scaling, a buyer should be able to run:

  1. Choose one repetitive, measurable workflow. Define the completed business outcome and the current human/system baseline before discussing token savings.
  2. Reconcile the economics. Attribute model, API, infrastructure and workflow costs to that outcome. Expose retries, long context, agent loops and other hidden consumption.
  3. Run under bounded authority. Specify which systems, data and actions are permitted; which conditions require escalation; and what happens when the workflow is uncertain.
  4. Verify behavior before autonomy. Run in observation or shadow mode, test routine and exception cases, retain evidence, and require a clear rollback or human handoff path.
  5. Expand only on measured improvement. Scale when cost per valid outcome improves without unacceptable degradation in quality, latency, control or auditability.
Bottom line Enterprise AI economics is a demand-management problem, not token penny-pinching. Proverify's addition is to split the job cleanly: TokenOps governs where AI capital goes, RunTime governs how the operation executes, and VERIFY tests whether the result is production-grade. Control the money, control the workflow, then prove both.

Related Proverify thinking

Source and interpretation. McKinsey & Company, “The cost of intelligence: How CIOs can manage AI demand at scale,” July 20, 2026, including its Enterprise AI FinOps survey (May 2026; 120 enterprise participants, 75 qualified respondents across five industries). This article is Proverify.ai's interpretation of that published framework; McKinsey does not endorse Proverify.ai, TokenOps, RunTime or VERIFY.