Hype vs reality in agentic commerce: demand is real, spend isn't (yet)

Hype vs reality in agentic commerce: demand is real, spend isn't (yet)
Photo by Shutter Speed / Unsplash

I'm building in agentic commerce, which gives me a vested interest in the optimists being right.

They have good material. Morgan Stanley, Bain, and McKinsey all draw the same line - 20% of commerce goes agentic by 2030. Seven protocols shipped in a single year, from Visa, Google, Shopify, Mastercard. On paper, the wave is already crashing. But in all of my conversations with merchants and execs, there's still headscratching.

Cue actual agentic trips - when in doubt, just run agents. We've run thousands of trips - agents turned loose in real stores, told to go buy - and they tell a quieter story. The agent does the "work": gathers, compares, builds the case, adds three items to the cart. Then it stops and asks the human which ones to take out. Fewer than one in ten ever attempt checkout. Coinbase's agent wallets moved $17,000 in two months - not a rounding error short of the forecast, a number well below even a rounding error.

So, like a lot of people deeply watching the space, I'm holding two numbers that don't agree. Demand is real: 20 to 40% of shopping searches now start on an agentic surface. Spend isn't. The question I sit with isn't which number is lying - neither is. It's "where's the value, today, in the distance between search and spend?"

Two stories wearing one coat

The forecast bundles two stories that move at very different speeds. People want agents to help them shop - research, compare, narrow, surface the thing they'd have found on page nine. And people want agents to transact for them - take the card, click buy, own the decision.

The first is happening now, fast. The second is happening-ish. In the most advanced setups, among the most advanced developers, the crypto and shadow-card transactions are real. But fewer than 5% of agentic users live in that advanced class, and only a sliver of those are actually letting agents spend. The headline treats the two stories as one curve. On the ground they're a year or more apart, and the distance isn't closing on the timeline the deck assumes.

What the trips actually show

Watch a few thousand sessions instead of imagining them, and the behavior sorts itself. The overwhelming majority of what an agent does in a store is gather, reason, and compare. Around 60% of agentic searches are broad - "recommend me something, across retailers" - and an agent handed "find me the best deal" doesn't come back with a completed order. It comes back with a case.

A fair challenge: maybe that's just early tooling, and checkout is a quarter from being solved. Here's what's underneath it. The blocker isn't model quality. It's trust, liability, and authentication - and those don't improve on a model-release schedule. A merchant loses little by letting an agent read its catalog. It loses a lot by letting an unverified agent move money on a customer's behalf - and so does the customer, and so does the model provider. Everyone in the chain has a reason to keep a human's thumb on the final click.

From the trips

Even when we tell the agent to finish the job, the instinct wins. It fills the cart, then hands the decision back to the person.

So the gathering layer races ahead, because nobody gets hurt when it's wrong. The spending layer crawls, because everybody does.

We're all counting the 2010s number

Here's where I think the measurement goes sideways - me included, for a while. Conversion was the currency of 2010s e-commerce, so conversion is what everyone reaches for now: agent-driven GMV, and it rounds to zero. Stare at that number and you land in one of two camps - it's all hype, or it's a launchpad and any day the curve goes vertical.

Both camps are watching the countable thing. The thing that's actually moving is harder to count - call it influence: how often an agent shaped a purchase a human ultimately finished. It never shows up in GMV. It shows up as a person who bought the jacket the agent surfaced, on the site the agent shortlisted, framed by the comparison the agent built - and then typed in their own card, so the sale logs as 100% human.

I catch it in my own behavior. A real share of my purchases now start in a chat, narrow to two or three options there, and finish on the merchant's site the way they always did. My statement says nothing changed. Everything changed. The agent was already the most important actor in the funnel, and it left no fingerprints on the receipt.

Where the value is today

Back to the question I sit with: where's the value, today, in the distance between search and spend? Not in racing the checkout. Trust and liability will keep that gated for years, however good the model gets. The value today is in the gathering layer - the part that's already real, already high-volume, and mostly unmeasured.

It's a complicated layer. About 60% of GPT commerce searches never leave the chat at all; the model's memory carries the recommendation. And when agents do leave to browse, they leave on cheap web-surfing - fewer than one in ten of the trips we've watched use the advanced browser or MCP capabilities that would make them powerful. The merchant-side tech is further along than the behavior - Shopify's MCP setup is genuinely good. The gap isn't the rails. It's that user-to-agent behavior hasn't caught up to them yet.

That points the work somewhere specific. For merchants, the job isn't "get ready for agent checkout." It's "understand how agents read you, because they already are, and most are guessing." For builders, the durable opportunity is in that gathering layer - measurement, legibility, trust - not in a transaction rail the whole chain is incentivized to keep slow.

The spend will come; I'm betting my days on it. But it's trailing the demand, not riding it, and the distance between the two is where the next few years of real work actually live.

If you take one thing: stop counting only what agents buy, and start counting what they say. The first number is small and loud. The second is large, quiet, and already in the funnel.