There is a difference between AI that sounds impressive and AI that sells. You see it in the numbers a business chooses to track.
Most AI commerce programmes are judged on engagement: session length, conversation quality and chat satisfaction scores. Those measures can show that a system is pleasant to use. They say much less about whether it helped a shopper move from intent to a completed, profitable order.
Engagement is also easy for an AI system to inflate. A longer session can mean better discovery. It can also mean the customer is struggling to get a trustworthy answer. More interaction does not automatically mean more value. Sometimes it simply makes the cost of failure visible.
Conversion rate, average order value and cart abandonment were built for a world in which people made every decision themselves. They still matter in AI-mediated commerce, but they no longer tell the whole story. The measures below follow the customer through three stages of the journey and show where value is really being created.
Why Engagement Metrics Are No Longer Enough
For two decades, commerce measurement followed a static funnel: traffic in, conversion out, with average order value and abandonment in between. It assumed the shopper would search, compare, and buy. That model starts to break down when an AI system takes on part of that work.
Engagement on its own can mislead. Adobe found that, in March 2025, shoppers who reached US retail sites from AI assistants spent longer on site and browsed more pages than shoppers from other channels. Yet they converted 38% worse than traffic from non-AI sources. Twelve months later, the result had reversed. By March 2026, AI-referred traffic converted 42% better than non-AI traffic, with revenue per visit 37% higher.
Engagement barely changed during that reversal. Conversion did. AI became more effective, and consumers trusted it more. A team measuring engagement alone could have called 2025 a success when the channel was not selling, then missed the point at which it started to.
The riskiest metric is the one that improves while the economics do not. Conversation volumes, session length and lower escalation rates can all rise while conversion quality, margin, return rates or customer confidence get worse. Activity is not value.
Apply the Brutally Simple Test
Use one question for every AI metric: if this metric doubled tomorrow, would the customer be better served and would the merchant be economically better off?
If the honest answer is no, it is not an outcome metric. It may be an operating signal. It may be a vanity metric used to justify infrastructure. Session length, conversation volume, and even a lower escalation rate can help diagnose a problem. They do not prove commercial value unless the system is also helping customers choose correctly, protecting margins and cutting the cost of failure.
Engagement still has a place. It should sit below the commercial outcome, not take its place.
The Case for Changing the Scorecard
AI-mediated commerce is harder to measure in owned channels because, when customers shop through a third-party agent, much of the decision can happen before they reach your store.
Bain estimates the US agentic commerce market could $300 to $500 billion by 2030, and Coresight Research puts total US retail sales mediated by agentic experiences at $943 billion by the same year. This is already material, even if agreement on how to measure it is still taking shape.
The AI-mediated journey can be measured across three stages: aspiration, checkout and trust. Each stage can fail in a different way, so each needs its own set of measures.
| Stage | The question it answers | The metrics that matter |
| Aspiration | Are you visible and relevant before the purchase decision is made? | AI Citation Rate; Generative Engine Optimisation (GEO) Score; Machine Legibility Score; Win-Rate in Agent Comparisons |
| Checkout | Does the decision become a completed, profitable order? | Agent-Driven Conversion Rate; Agent-Mediated Revenue Per Visit (RPV); Channel Transaction Share; Checkout Completion; Margin Protection; Decision Quality and Recovery Cost |
| Trust | Can shoppers rely on what the AI tells them and how it executes? | Product Data Accuracy; Hallucination Rate; Escalation Quality; API Response Quality; Catalog Sync Latency |
The distinction between these measures matters. Visibility and machine legibility are leading indicators. They show whether an agent can find and assess your offer. Engagement and escalation are operating signals that help explain what is happening in the journey. The commercial outcome is whether the system helps customers choose correctly, keeps their intent intact, follows the right policy and payment conditions, and reduces failure and recovery costs.
Aspiration: Are You Visible Before the Decision Is Made?
A commerce AI first has to connect what a shopper wants with what they should buy. In an AI-mediated journey, that work begins before most retailers start measuring. If an agent never surfaces your product, on-site optimisation cannot fix it.
AI engine optimisation (AEO) and citation visibility across consumer-facing LLMs are becoming part of product discovery. Retailers that do not structure product data or optimise for AI engines risk disappearing from customer journeys, whatever the strength of their brand. That makes AI visibility a leading indicator for every revenue measure that follows.
AI Visibility Is the New Top of Funnel
The channel is large enough to track. Traffic from AI sources to US retail sites grew 393% year over year in the first quarter of 2026. Thirty-nine percent of consumers say they have used AI for online shopping, and 85% of that group say it improved the experience.
Most retail sites are not built for this yet. Adobe scored US retail pages on how much of their content LLMs can read, and the gap is wide. Retail homepages average 75% readability, which leaves about a quarter of homepage content invisible to AI. Product pages perform worse, at 66%.
| Page type | Average AI visibility score |
| Returns and exchanges | 82% |
| Contact | 81% |
| Customer service and help | 79% |
| Homepage | 75% |
| Category pages | 74% |
| Product pages | 66% |
That product pages score lowest should concern commerce leaders. This is where the buying decision happens and where an agent needs the most detail.
The Four Visibility Metrics Worth Tracking Now
Visibility comes down to four measurable components. There is no published industry benchmark for them yet, but retailers can track all four now and see whether they improve over time. They are leading indicators, not proof of value by themselves. A stronger score matters only if it drives more relevant consideration and better commercial decisions later in the journey.
| Metric | What it measures | How to start measuring it |
| AI Citation Rate | How often an AI assistant retrieves, references or recommends your products during discovery | Build a fixed set of 50 to 100 buying queries relevant to your categories. Run them across the major assistants on a set schedule, then log whether your brand and products appear and where they rank. |
| Generative Engine Optimisation (GEO) Score | How discoverable and machine-readable your catalog and brand context are to AI models | Audit product pages for structured schema, conversational descriptions and Q&A content. Score them against a consistent internal rubric and retest quarterly. |
| Machine Legibility Score | Whether attributes, pricing rules, stock levels and policies are exposed in structured formats rather than buried in prose | Check that each product has consistent identifiers, one authoritative price and eligibility source, and return and warranty terms expressed as conditions rather than buried in PDF text. |
| Win-Rate in Agent Comparisons | How often your products win when an agent weighs them against competitors | This is the hardest measure to capture today. Track a fixed query set, run repeated sessions and note the final selections. |

Machine legibility is often underestimated. EY puts the issue plainly: agents assess structured signals, not storytelling or imagery. If your differentiation is not machine-legible, it is effectively invisible. Commercial information scattered across catalogs, contract tools, PDFs and local spreadsheets may be manageable for a person. To an agent, it looks unreliable, and unreliable offers get bypassed.
Feed freshness matters too. Agentic platforms can refresh product feeds every 15 minutes, while stale data can lower rankings or remove products from listings altogether. A feed that is accurate once a day is not accurate enough for an agentic journey.
Checkout: Does the Decision Become a Completed Order?
This stage asks whether intent survives the trip to checkout. It remains the best-documented failure point in commerce.
Checkout Completion Remains Relevant to Agentic Commerce
Across  50 separate studies, Baymard Institute puts average online cart abandonment at about 70%. For every ten shoppers who add an item to a cart, roughly seven leave without buying. The figure has changed little over a decade, despite sustained investment in payment technology and checkout design.
That makes checkout the clearest test for agentic commerce. Conversational checkout aims to turn a long, form-heavy process into a single flow, right where the 70% drop-off occurs. The question is whether AI closes that gap or simply moves it. A checkout that feels quick but is disconnected from live inventory and shipping rules creates a different broken promise, not a completed sale.
The Metrics That Show Agents Are Actually Selling
Cart abandonment tells you how the site performs. It does not tell you how agents perform. The following measures separate agent-driven revenue from everything else. First, segment agent traffic in your analytics.
| Metric | What it measures | How to start measuring it |
| Agent-Driven Conversion Rate | How often an agent interaction results in a completed transaction, bypassing browsing stages entirely | Set attribution rules using available referrer, integration and session data. Report conversion for that segment separately from human traffic rather than blending it into a site average. |
| Agent-Mediated Revenue Per Visit | Revenue generated per agent interaction, which can differ from human RPV when the journey carries more context and checkout becomes simpler | Track revenue against agent sessions and compare the pattern with human RPV, particularly on lower-consideration repeat purchases. |
| Channel Transaction Share | The split between traditional human web and app traffic and agent-completed transactions | Report agent-attributed orders as a percentage of total orders each month, so you spot the trend early rather than after the shift is under way. |
| Margin Protection | Whether automated decisions defend your economics or erode them to close | Compare realised margin on agent-completed orders with your standard margin. Check that discounting rules are enforced at the point of decision, not after it. |
| Decision Quality and Recovery Cost | Whether customers choose the right item under the right policy and payment conditions, and what failures cost after the sale | Compare return, exchange, cancellation, dispute and support-recovery rates and costs for agent-assisted orders with an appropriate control group. Review whether the agent preserved the shopper’s stated intent. |
This final measure stops a surface-level conversion gain becoming an expensive failure. An agent that steers a shopper to the wrong product, applies an ineligible promotion or loses context between discovery and payment can lift a front-end number while creating returns, service contacts and lost confidence later. The commercial decision does not end at order confirmation.
Industry benchmarks for these measures are only starting to emerge. Your own trendline is the benchmark that matters, and it only exists if you start capturing it now..
Trust: Can You Rely on What the AI Says?
Trust has no direct equivalent in the old funnel, and it may be the most important stage in agentic commerce. When an AI answers a shopper directly, it needs to be accurate. Otherwise, it makes a broken promise in your brand’s voice.
Trust in AI is becoming easier to measure commercially. Adobe found that 66% of consumers believe AI tools provide accurate results and identified rising confidence as a factor in the conversion reversal described earlier.

How to Measure Accuracy, Accurately
General-purpose AI models are still some way from being fully reliable. Independent testing published via Statista shows leading models hallucinating at rates below 2% when they are grounded in verified source data. Broader benchmarks place hallucination rates across many models at 15% or higher on open-ended tasks.
Accuracy in agentic commerce depends on how closely the AI is tied to real, structured product data. Three emerging measures help you track it.
| Metric | What to ask | How to start measuring it |
| Product Data Accuracy | How often are price, availability or specification details wrong? | Sample a fixed set of AI responses each month and check each product claim against the live catalog. |
| Hallucination Rate | Is the system answering from a live catalog or generating plausible text? | Log responses that cannot be traced to a source record, then track that share over time. |
| Escalation Quality | Does the system hand off uncertainty when it should, and does the recovery protect the customer? | Measure handoffs as a share of sessions, then sample them for necessity and outcome. Treat a rate near zero as a warning rather than a win, especially if returns, complaints or recovery costs rise. |
Zero escalations are not the goal. The goal is correct decisions and safe recovery when the system should not decide alone. A lower escalation rate is only useful when conversion quality improves, returns stay stable or fall, and customer confidence is protected.
Two Technical Metrics Your Commerce Team Should See
| Metric | What it measures | Why it belongs on a commercial dashboard |
| API Response Quality | The speed, completeness and structural accuracy of data returned to an inquiring agent | If a response is slow or incomplete, the agent needs follow-up queries to make a choice and may resolve the choice elsewhere. |
| Catalog Sync Latency | The delay between an inventory or price change in your systems and that change being readable by external agents | Every minute of lag creates a window in which an agent can promise a shopper something you cannot honour. |
Catalog sync latency deserves particular attention because it turns a technical gap into a customer-facing failure. When an agent sells a product you no longer hold, the customer does not blame the agent.
Start Measuring What Sells
AI is taking on more of the buying journey, and the scorecard built for static funnels wasn’t designed for that role. Conversion, AOV and engagement still matter. On their own, they are not enough.
The right measures follow a shopper’s intent from start to finish. Citation rate and machine legibility show whether your products are considered at all. Agent-driven conversion, channel transaction share, margin protection and decision quality show whether the channel is selling well. Accuracy, appropriate escalation and catalog sync latency show whether it can be trusted. A strong number at one stage counts for little if intent leaks at the next.
Start by segmenting agent traffic in your analytics. You cannot recover most agentic measures retrospectively, and your own trendline will be the only useful benchmark for a while. Set a baseline before you deploy anything. Without documented pre-deployment performance, no uplift claim, including a vendor’s, can be checked later.
Then return to the brutally simple test. If a metric doubled tomorrow, would the customer be better served and would the merchant be economically better off? If not, it is a signal to investigate, not an outcome to celebrate.
As AI moves from assisting the journey to executing it, ask whether the deployment helps customers choose correctly, keeps their intent intact, follows the right policy and payment conditions, sells profitably, and lowers the cost of recovery when something goes wrong. Everything else is measuring the conversation.
Book a conversation today to map your agentic commerce journey against the metrics that matter.