/
/
The Metrics that Matter in Agentic Commerce

The Metrics that Matter in Agentic Commerce

Share this article:
Oliver Tan,
MD Rezolve Ai APAC
Table of Contents
Oliver Tan,
MD Rezolve Ai APAC

When AI starts to shape product discovery and transactions, retail leaders need a scorecard that follows the decision all the way through. 

The first measurement mistake in enterprise AI was mistaking activity for value. 

We counted conversations, session length, response rates and engagement because they were easy to count. More interaction looked like success, even when customers were taking longer to get an answer or the economics underneath had not improved. 

Agentic commerce could repeat that mistake with more sophisticated-looking metrics. 

A brand can now measure whether it appears in an AI answer, how prominently it appears, whether it is recommended and what the AI says about it. Those are useful advances. They are not the same as commercial value, and neither is the same as trustworthiness. 

That matters because AI is no longer simply another traffic source. Search and answer surfaces such as ChatGPT are becoming places where products are discovered, compared and recommended. In some cases, the transaction starts there too. 

Salesforce’s latest State of Commerce research found that traffic referred from AI chats grew between 150% and 428% year over year in every quarter measured, based partly on behavioural data from more than 1.5 billion global shoppers. The use of agentic search as the first step in the shopping journey grew 200% year over year. Salesforce, State of Commerce, 30 September 2026 

The old funnel is changing. AI is becoming a decisioning surface.  That calls for a different measurement architecture. 

The activity trap 

There is already evidence of how easily one metric can mislead. 

Adobe found that, in March 2025, AI-referred traffic to US retail sites converted 38% worse than non-AI traffic. Twelve months later, it converted 42% better. In the later period, AI-referred visitors were also spending 48% longer on site and viewing 13% more pages per visit. Adobe Digital Insights, 2026 

Engagement alone could have told the wrong story.  

The same goes for AI visibility. A brand can be frequently mentioned, appear ahead of competitors and receive positive recommendations, yet still produce poor commercial outcomes. The product may be unavailable. The price may be wrong. Checkout may fail. The recommendation may not fit the customer. The transaction may come back as a return. 

The reverse is also possible. A retailer can have excellent product data and reliable commerce infrastructure, yet remain invisible to the AI systems mediating discovery. 

These are different failures. They should not be compressed into one AI score. 

At Rezolve, we think about agentic-commerce measurement across four planes: 

Measurement plane The management question 
Visibility & Influence Does AI consider us, and how does it represent us? 
Commercial Execution Does AI-mediated intent become a correct and economically valuable transaction? 
Assurance & Control Can the system be trusted to make and execute that decision appropriately? 
Measurement Discipline Is the evidence reliable enough to make a decision from it? 

Each plane answers a different question. None can stand in for another. 

1. Visibility and influence: are you in the decision? 

The IAB’s 2026 Measuring Visibility in the AI Era framework organises AI visibility around four ideas: 

  • Presence: Does the brand appear? 
  • Prominence: Where and how prominently? 
  • Portrayal: In what context, and with what accuracy? 
  • Persuasion: Does visibility drive action? 

Its underlying measures include mention rate, citation rate, share of voice, position, framing, factual accuracy and recommendation strength. 

The distinctions matter. 

Being mentioned is not the same as being cited. Being cited is not the same as being recommended. And being recommended is not the same as being selected or purchased. 

A useful scorecard should ask whether products appear for relevant customer intents, where they appear relative to alternatives, how the brand and product are framed, whether the information is accurate, and how strong the recommendation is. 

Visibility is a leading indicator. A product that never enters the AI’s consideration set has little chance of capturing downstream demand. 

But visibility has a boundary. The IAB framework focuses on organic AI visibility measurement. It does not prescribe GEO or AEO optimisation, and it does not solve downstream commerce attribution. Visibility tells you whether you entered the decision. It does not tell you whether the decision created value. 

2. Commercial execution: did the decision create value? 

Once an AI system recommends a product or initiates a purchase, the question changes. 

It is no longer only: Did the AI choose us? 

It becomes: Was it a good transaction? 

Conversion, revenue, checkout completion and transaction share all still matter. But an autonomous or semi-autonomous system can optimise the wrong outcome efficiently. 

Imagine an agent raises conversion by 10%. That sounds successful. Now imagine it achieved that increase by over-discounting, recommending poorly fitting products or applying promotions too generously. Returns increase. Gross margin falls. Customer-service contacts rise. 

Conversion improved. The business did not. The more useful unit of measurement is decision-adjusted commercial value. Retailers should examine: 

  • Agent-attributed conversion: Did AI-mediated intent produce an order? 
  • Agent-mediated revenue: How much revenue was materially influenced or executed by an agent? 
  • Margin quality: What contribution remains after discounts, incentives and channel economics? 
  • Transaction completion: Did the intent survive inventory, identity, payment, shipping and fulfilment? 
  • Decision quality: Was the product appropriate to the customer’s stated requirements? 
  • Post-purchase failure: What followed in returns, exchanges, cancellations, disputes and service recovery? 

Deloitte’s 2026 work on agentic-commerce economics identifies revenue, service cost, return rates, dispute reserves and net margin among the measures that matter to this emerging channel. Deloitte, Agentic commerce: Redefining retail economics, 2026 

The commercial decision does not end at the Buy button. A transaction that converts today and comes back tomorrow is not equivalent to one that satisfies the customer’s intent. Nor should a transaction that converts only because margin was unnecessarily surrendered be valued the same way. 

The scorecard has to follow the economics far enough downstream to know the difference. 

3. Assurance and control: can the system be trusted to act? 

Trust is not another stage after checkout. It surrounds the entire transaction. 

An AI system can fail during discovery by inventing a product feature. It can fail during recommendation by using stale availability. It can fail during checkout by applying the wrong eligibility rule. It can complete a transaction successfully but still fail what we call a recursive dissonance test: a final check that the outcome aligns with the customer’s original intent and constraints. 

It can also fail after purchase by taking an action the customer never authorised. 

Assurance has to operate as a control plane across visibility and execution. It includes factual accuracy, catalogue freshness, price and inventory consistency, policy adherence, authorization boundaries, tool and API reliability, appropriate escalation, decision traceability, exception recovery and customer recourse. 

Even the difference between hallucination and factual inaccuracy matters. A hallucination is an unsupported claim generated by the model. Factual inaccuracy can arise when the AI reflects information that exists but is wrong or out of date. 

They can look identical to the customer. Operationally, they require different remedies. One needs better grounding or model behaviour. The other needs the source of truth fixed. 

That is why machine-readable product data, live inventory, structured policy information and reliable commerce APIs matter. They are not visibility metrics. They are the operational infrastructure that makes accurate visibility and reliable execution possible. 

Governance belongs here too. NRF and PwC’s 2026 work on agentic AI in retail treats governance and security as foundational as agents move from assisting customers toward browsing, comparing and purchasing on their behalf. NRF/PwC, Managing and Governing Agentic AI in Retail, March 2026 

The more authority an agent receives, the less defensible it becomes to measure success without measuring control. 

4. Measurement discipline: is the evidence strong enough to act on? 

Even the right metrics can produce the wrong conclusions when the measurement itself is weak. 

Generative AI outputs are probabilistic and often non-deterministic in practice. Ask the same question twice and you may not receive the same brands, sources or recommendations. 

The IAB makes an important distinction between directional and decision-grade measurement. Directional measurement can tell us that something appears to be changing. Decision-grade measurement demands more: sufficient sampling, representative queries, reproducibility, methodology disclosure, validation, appropriate platform coverage and an understanding of expected variation. 

A single AI response is not measurement. Visibility is a distribution, not a fixed value. Observed movements need to be distinguished from normal variation in the system. 

Before acting on an AI metric, executives should know: 

  • What exactly are we measuring? 
  • Against what population or baseline? 
  • How stable is the result? 
  • Could a platform or model change have caused the movement? 
  • Can we reproduce it? 
  • Is attribution sufficiently credible for the decision we are about to make? 

A number does not become decision-grade because it appears on a dashboard. Precision should not be confused with certainty. 

Use a harder test 

Before celebrating any AI metric, classify it first. 

Is it an outcome, a diagnostic or a control? 

Then ask: If this outcome improves, are the customer and merchant actually better off? 

If the answer is uncertain, it probably is not an outcome metric. 

Machine readability can improve without sales improving. Conversation volume can rise without customer value increasing. Escalations can fall because the AI became better, or because it became dangerously overconfident. Even visibility can rise without producing profitable demand. 

Those measures are useful. They simply answer different questions. The mistake is not measuring them. The mistake is allowing one class of metric to impersonate another. 

Measure the chain, not the metric 

Agentic commerce is compressing discovery, evaluation and transaction into the same interface. Measurement has to evolve with it. 

Visibility and Influence tells you whether you enter the AI-mediated decision and how you are represented. Commercial Execution tells you whether that decision becomes economically valuable business. Assurance and Control tells you whether the information, decisions and actions can be trusted. Measurement Discipline tells you whether the evidence behind all three is strong enough to justify action. 

None substitutes for the others.  A retailer can dominate AI visibility and lose money on the resulting transactions. It can convert brilliantly while accumulating returns and disputes. It can build a highly controlled system that customers never encounter. 

The objective is not to optimise a metric. It is to preserve customer intent and economic value across the entire chain, from discovery to decision to execution. 

That is the scorecard agentic commerce now needs.

Share this article
Author
Oliver Tan

MD Rezolve Ai APAC

Oliver leads APAC and drives agentic commerce initiatives for Rezolve Ai (Nasdaq: RZLV) to deliver the next-generation of AI-driven experiences with partners.

Related blogs

September 11, 2026
Why “Agentic” Doesn’t Mean What Most Vendors Think It Does
July 1, 2026
Prestige by Design. Building AI That Knows the Difference Between La Mer and MAC
January 29, 2026
Agentic Commerce Is Growing Up: Why Emerging Protocols Matter – and How Rezolve Ai Is Built for What Comes Next
Our new site is on the way,
and it's built for conversation.

Get a sneak peek while we finish the final touches.