Skip to content
Professional

17 min read

AI can think. It still can’t shop.

A post-mortem · agentic commerce

Last September, OpenAI promised a million merchants inside ChatGPT. By February, thirty had turned up. By March it was switched off. The models were never the problem — and the thing that is the problem has no engineering fix.

Exhibit A

The short, strange life of Instant Checkout — the product that was supposed to make ChatGPT a shopping mall. Every line on it is drawn from public reporting or the companies’ own statements. The last line is the one the industry keeps not printing.


Start with the number.

On 29 September 2025, OpenAI shipped Instant Checkout. You could buy things without leaving ChatGPT. Etsy sellers first, and then — the announcement said — over a million Shopify merchants. Glossier. SKIMS. Spanx. Vuori. Shopify’s president called it the new frontier for online retail. Seven hundred million weekly users were about to become seven hundred million shoppers.

By February 2026, Forrester counted the Shopify merchants who had actually gone live. Roughly thirty. Shopify’s own president put it closer to a dozen, and said the holdup was on the AI side.

On 4 March, Walmart’s head of AI acceleration described the whole thing as “a very temporary moment in time.” Three weeks later, OpenAI switched it off. Discovery stays in the chat, it said. Checkout goes back to your website.

One hundred and seventy-six days. That was the working life of the flagship product of the most-funded consumer AI use case of the decade. And it wasn’t a quiet little experiment either — Walmart had loaded two hundred thousand products into it.

Checkout inside an AI assistant will account for roughly a tenth of one percent of US retail e-commerce this year. eMarketer expects it to stay under two percent even in 2029.

For scale: US retail media — sponsored listings and banners on retailer websites, the least glamorous line item in all of marketing — will take about $71 billion this year. Hold that thought. It comes back.

Now let me ruin that argument — whiplash ahead

If this were simply a bubble popping, it would be a boring essay and you’d have read it four times already.

The same eighteen months produced the best traffic story in retail. Adobe, working off more than a trillion visits to US retail sites, has AI-referred traffic up 393% year on year in the first quarter of 2026, and up 1,324% since it started counting in October 2024. December peaked at 1,151%.

And then the bit that made people sit up straight. In March 2025, shoppers arriving from an AI assistant converted 38% worse than everyone else. In March 2026, they converted 42% better. Same channel, same shops, twelve months apart. Revenue per visit up 37%. Time on site up 48%.

The worst-performing traffic source in US retail became the best-performing one inside a year.

And in the retailers’ own houses, where nobody was making protocol announcements, the assistants were quietly working:

So which is it? A five-month flameout, or the fastest-growing channel in retail?

Both. At the same time. And the reason both are true is the whole point of this piece.

The split — Discovery won · transaction lost

AI won product discovery outright. It lost the transaction badly. Those are two different businesses, and only one of them was ever a modelling problem.

Let me be fair to the models, because the lazy version of this argument blames them and the lazy version is wrong. On OSWorld — the benchmark for AI doing real work on a real computer — success went from 12% to 66.3% in a single year. That is an astonishing curve. It also means the best agents still fail one task in three. Fine for drafting an itinerary. Not fine for a non-refundable flight, or a ₹58,000 sofa in the wrong grey.

But a perfect model would not have saved Instant Checkout, and here is how you know. Go and look at what it actually couldn’t do at launch. No multi-item carts. No promo codes. No shipping promises. According to The Information, OpenAI hadn’t built the machinery to remit US state sales tax. And the product data came from scraping retailer websites, which meant price and stock were regularly, confidently wrong.

None of that is a reasoning failure. That is a shop with no shelf, no till and no address.


On that benchmark: OSWorld drops an agent into a real desktop and asks it to finish real tasks across real apps. Going 12% → 66.3% in a year is one of the steepest capability curves anyone has published. It is also why “the models aren’t good enough” has quietly stopped being a defensible explanation for any of this.


The kirana we threw away — four missing rails

I want to put an old thing next to a new thing, because it explains more than the protocol diagrams do.

If you grew up middle-class in India in the nineties, you have already lived through agentic commerce. We called it sending the kid downstairs with a chit of paper.

The chit said Surf, bada wala and nothing else, and it worked. Bhaiya knew which Surf. He knew you took Amul and not the other one, that the second week of the month was tight, and that it could go on the khaata. If he was out of your brand he substituted something you’d accept, because you were coming back on Thursday and he would hear about it if he got it wrong. And when the packet turned out to be past date, you walked it back and he took it. No receipt. No ticket number. No five-to-seven business days.

Four things were holding that up.

He knew who you were. He knew what was on his shelf that minute. He carried the risk when he got it wrong. And he was paid out of the margin on what he sold you, not out of showing you an advertisement.

The internet dropped all four and handed you a search box. Which was a perfectly fine trade, because a human was doing the clicking, and the human quietly supplied the missing four. You were your own identity. You judged the shelf by squinting at photographs. You absorbed your own mistakes. And you looked at the ads, which is how the whole arrangement got paid for.

Take the human out of the loop and you find out exactly how much of the loop was the human.


For readers outside India

Kirana — the independent neighborhood grocery, still roughly 80% of Indian retail. Bhaiya — literally “brother”, what you call the man behind the counter. Khaata — the running credit ledger he keeps in a cloth-bound book, settled monthly, secured by nothing but the fact that you live around the corner.


01 A name — status: building

To a retailer’s servers, a shopping agent and a scraper look identical.

Amazon sued Perplexity in November 2025 over Comet, its agentic browser. The complaint is worth reading because it is so unusually direct: Amazon says it warned Perplexity at least five times from November 2024, put up a technical block in August 2025, and watched Perplexity route around it inside twenty-four hours. It also alleges Comet dressed itself up as an ordinary Chrome session to avoid detection.

A district judge granted an injunction in March. Amazon had, the ruling noted, spent more than five thousand dollars responding — a wonderfully deadpan sentence to find in a filing about the future of commerce. On 4 August 2026 the Ninth Circuit vacated it. When you tell your assistant to do something on Amazon, the court held, it is you accessing Amazon, with help. Narrow ruling. The case grinds on.

But look at what Amazon argued in its own papers: agent traffic forces it to filter fraudulent impressions before it can charge advertisers. Perplexity said the quiet part louder — agents skip the ads, and that is what this is really about. Both statements are true. It’s a business-model fight wearing a security lanyard.

The actual fix is being built somewhere far less dramatic. Web Bot Auth — Cloudflare’s proposal, now in IETF standardization — lets an agent cryptographically sign every request so a merchant can verify who is knocking. Visa’s Trusted Agent Protocol and Mastercard’s Agent Pay both sit on top of it. Amex is in. AWS, Shopify, Akamai and Vercel have implemented. Cloudflare already handles more than ten billion AI bot requests a week and expects bot traffic to overtake human traffic next year.

robots.txt was a polite request. This is proof of identity. It is the least interesting thing in this essay and probably the most important.

02 A shelf — status: unsolved

Roughly six in ten e-commerce catalogues carry missing product IDs, inconsistent attribute naming, or stale stock states. Around 78% of agent recommendation errors trace straight back to incomplete product schema.

Read that again in plain language. The agent is not hallucinating. The agent is reading your spreadsheet, and your spreadsheet is wrong.

Anyone who has spent time in analytics feels this one in their spine. The model is never the hard part. The hard part is that color is a free-text field, three teams populate it, and one types “Navy Blue”, one types “navy”, and one types NB-04 because that’s what the vendor emailed. You cannot reason your way out of a bad join. You can only go and fix the join, which takes eighteen months and gets nobody promoted.

Walmart put 200,000 products into ChatGPT and the onboarding was still cumbersome and the inventory and shipping data still frequently wrong. Walmart. If the company with the best supply-chain data on the planet can’t hand a clean shelf to an agent, the long tail is nowhere close.

There is a real prize sitting here, to be fair. Merchants with attribute fill rates above 95% get surfaced; below 80% they get skipped entirely. Near-complete catalogues report three to four times the visibility in AI recommendations. It is the least glamorous competitive advantage ever invented and it is lying unclaimed in your PIM.

03 A khaata — status: unsolved

Americans returned $849.9 billion of merchandise last year. That is 15.8% of all retail sales, and 19.3% of everything bought online. Apparel is worse.

Now put an agent in the middle — one that fails a complex task about a third of the time — and ask the only question that matters. Who eats it?

If the agent orders the wrong size, is that customer error or a defect? If it buys from a merchant who doesn’t ship to your pin code, whose refund is that? If it gets talked into an upsell by a sponsored prompt, is that fraud, advertising, or Tuesday? The card networks are drafting rules for this right now, in public, which is a reliable signal that nothing is settled.

Until it is, no CFO signs off on autonomous purchasing at volume. Honestly, no CFO should. This is the rail people find least exciting and it is the one actually gating the money.

04 A till — status: unresolved · the big one

If I could keep only one section of this essay, it would be this one.

The web paid for product discovery with advertising. That is the deal we all signed. You get free search and free comparison; retailers and brands buy their way to the top of it. US retail media will take about $71 billion this year, growing 18%, and roughly 60% of that is search advertising on retailer sites. Globally, retail media crosses $196.7 billion in 2026 — bigger than linear and connected TV combined, for the first time.

An agent does not look at advertisements. It has no eyes to catch. One agency executive called agents an existential threat to retail media, and for once the phrase is doing honest work.

So the funding mechanism for discovery evaporates at exactly the moment discovery becomes the thing AI is best at. And nobody has agreed what replaces it. Three candidates are on the table. All three are being tried. None is winning.

Take a cut. OpenAI charged merchants about 4% on Instant Checkout; add Shopify’s processing and you land near 9.2% all in. Roughly a third of what Amazon’s marketplace costs, which sounds reasonable right up until you notice that Google and Microsoft charge nothing at all for chatbot purchases. Hard to levy a 9% tax when the shop next door is free.

Put the ads back in the chat. Walmart’s Sponsored Prompts. Amazon’s ads inside Rufus, live since November. Google’s Direct Offers inside AI Mode, which surfaces a discount when the model reckons you’re wavering. Similarweb estimates about 26% of ChatGPT responses now carry an ad, concentrated in the first answer of a session. This works. It is almost certainly the answer. It also quietly rebuilds the exact thing agents were sold to us as an escape from.

Charge the shopper. Nobody has seriously tried. Which tells you something.

And now the asymmetry that I think explains most of what is on this page. The companies technically capable of building agentic commerce — Google, Amazon, Meta — are precisely the companies earning most from the funnel an agent would collapse. The companies with every incentive to collapse it — Perplexity, the startups, the protocol shops — have neither the distribution nor the merchant relationships.

Capability and incentive are sitting in different buildings. That is a structural problem, not an engineering one. No amount of model progress touches it. It gets resolved when somebody decides who pays the shopkeeper, and that decision has not been made.


The number to argue with

Amazon’s marketplace costs a seller roughly 15% before advertising. OpenAI wanted about 9.2% all in. Cheaper — and still dead on arrival, because Google and Microsoft were charging nothing. A toll road only works when there is no free bridge, and in agentic commerce there are currently three free bridges.


The surveys are a Rorschach test

This month you will read that 69% of Americans are happy to let AI buy things without approving each transaction. Croud ran it. It is a real survey of two thousand people.

This month you will also read Gartner, from January: willingness to let AI actually make the purchase decision tops out at 11%, and that ceiling is for household supplies. Not handbags. Toilet cleaner.

And Accenture, 25,590 people across sixteen countries: 74% will hand over the boring stuff, 32% will let an agent choose if they get to approve the payment, 9% will let it buy outright.

These are not contradictions. They are a ladder, and every rung sheds people.


Why the spread is this wide

Every one of these is a stated preference. Nobody in any of these panels had money on the table. The revealed number — people who actually completed a purchase inside an assistant — is the 0.1% at the top of this page. When stated and revealed preference disagree by three orders of magnitude, believe the till.


Look at where the drop is. It isn’t where a technologist would put it. It isn’t at can the model do this. It is at whose fault is it if this goes wrong.

Bain found 25% of US consumers trust retailers to run the whole shopping experience end to end. For AI platforms like ChatGPT, it is 7%. That single gap explains most of the last six months of corporate strategy, including why OpenAI’s retreat routed shoppers back to the merchant’s own site, and why Walmart was so happy to help.

Travel tells the identical story, which is the tell that this is structural rather than category-specific. Expedia’s 2026 research: travellers use AI for inspiration constantly, and roughly 70% still finish the booking with a brand they already trust. Two-thirds wouldn’t let an assistant book at all. Only 8% currently lean on AI agents when planning. Booking.com, meanwhile, finds 89% want AI in future trip planning. Both numbers are real. Planning is free. Booking is a non-refundable mistake with your family watching.

Five things that are working (not like the demo)

1 · The assistant that lives in the shop. Rufus, Sparky, Vaani, Flipkart’s SLAP. These win by standing still and solving three rails at once. The retailer knows who you are, because you’re logged in. It knows what’s on the shelf, because it’s their shelf. And the returns desk is down the corridor. Trust follows: 25% versus 7%.

2 · Agents pointed at the seller, not the shopper. This is the pivot almost nobody wrote up. In India, Amazon, Flipkart and Meesho have quietly moved AI investment away from shopper-facing assistants and towards seller-side stacks — listing quality, catalogue enrichment, pricing, demand forecasting. Expedia is doing the same for hotels through Partner Central agents. Akeneo shipped agentic PIM in July.

It is dull. It compounds. And it matches what MIT found looking across industries: budgets pile into sales and marketing, returns show up in the back office. Their headline — 95% of GenAI pilots producing no measurable P&L impact — got quoted to death last year, mostly by people who skipped the part explaining why.

3 · Repeat purchase, not discovery. B2B tail spend. MRO. Consumable replenishment. Groceries. Here the supplier, the price and the terms are already agreed, and the agent just executes inside rules somebody already signed. Forrester expects a third of B2B payment workflows to involve agents by the end of this year.

Notice what that actually is. It’s the khaata. Identity, shelf and liability all pre-solved by a contract, leaving the agent only the easy part.

4 · The protocol layer, having a very strange year. Google launched UCP at NRF in January with Shopify, Etsy, Target and Wayfair. On 24 April, Amazon, Meta, Microsoft, Salesforce and Stripe joined its Tech Council.

Sit with the Amazon one for a second. August 2025: blocks AI crawlers. November: sues Perplexity. February: rewrites its seller agreement to force agents to identify themselves. April: joins the governance body of the open standard. Eight months, complete reversal. Companies do that when they’ve worked out the thing is happening with or without them. And under UCP the retailer stays merchant of record — the industry admitting, in schema, that the transaction isn’t going anywhere.

5 · The browser agent, which sidesteps the problem entirely. Comet, Atlas, Copilot Mode, Gemini in Chrome. No merchant integration, no protocol, no onboarding queue — the agent just uses the website like a person, because it is being a person’s hands. August’s ruling made that path considerably safer. It is the least elegant answer available and it may well arrive first, in the way screen-scraping quietly ran Indian fintech for a decade while everyone waited politely for clean APIs.

Everyone is right, because nobody agrees what counts

Before you take any forecast in this category seriously, notice that the entire disagreement is definitional. Narrow, broad and consultant estimates sit roughly 243 times apart. Almost none of that gap is disagreement about growth rates. All of it is disagreement about what counts as an agentic transaction. Toggle it above and watch a $1.6 billion reality become a $5 trillion market without a single new fact entering the room.

That isn’t analysts being dishonest. It is a category that hasn’t yet decided what it is.

Now the hopeful part

Every commerce shift I can find looked exactly like this before it worked.

E-commerce in 1998 didn’t need better websites. It needed SSL so you’d type your card number into a stranger’s form, PayPal so strangers could pay strangers, and a courier network that could actually find your house. Boring, invisible, unfundable-looking plumbing, laid by people nobody profiled. Then one morning it was all there and Amazon looked inevitable in hindsight.

2026 is a plumbing year, and the plumbing is unusually good.

Web Bot Auth is in IETF standardization with the card networks behind it. UCP has ten of the largest names in commerce in one room and an actual specification. AP2 has sixty-plus partners. Universal Cart went live in the US in May and is heading for Canada, Australia, the UK, YouTube, hotel booking and food delivery. Visa runs a directory of agent keys. The sales tax thing will get solved, because sales tax always gets solved.

And the demand was never fake. It was pointed somewhere other than where the money went. AI-referred shoppers are now the best-converting traffic many retailers have. Meesho put a voice agent into tier-2 and tier-3 India in four languages and got a million and a half users in the first month, with better conversion and fewer returns — a more interesting result than anything that shipped out of San Francisco last year, and one almost nobody wrote up.

My honest read: agentic commerce doesn’t arrive as a moment. It arrives as an accumulation of extremely boring decisions about identity, feeds and liability, and one ordinary Thursday you will reorder the dog food by talking to your phone and never think about it again. Nobody will call it a revolution, because by then it will just be Thursday.

And the question I can’t put down—

The kirana bhaiya worked for you. That is what made the whole arrangement safe. He carried your preferences, your credit and your risk, because you were coming back on Thursday and your mother knew his mother.

An agent works for whoever pays for the shelf.

If the answer to who funds discovery turns out to be sponsored prompts — and I think it will, because of the three candidates it is the only one that has ever worked at scale — then the thing standing between you and every product on earth is an intermediary with a quarterly revenue target, speaking in the register of a helpful friend, and vastly better than a banner ad at not looking like one.

We spent twenty years learning to see through advertising. We got quite good at it. Banner blindness is practically a sense organ by now.

Nobody has any practice at seeing through a recommendation.

So the question was never whether the agent learns to shop. It will; the plumbing says so. The question is the one that got answered by default in 1998 and is being answered by default again right now, in specifications almost none of us will ever read.


When the agent finally learns to shop —
who is it shopping for?


I would very much like to be wrong about the answer. Ask yours. See what it says.


Notes on the numbers

Vendor data, honestly labelled. Adobe’s traffic and conversion figures come from its own analytics customers — real, large, and not independently audited. Similarweb and Grips are panel estimates. Treat the direction as solid and the decimals as decorative.

The merchant count. Forrester’s Emily Pfeiffer put live Instant Checkout Shopify merchants at around 30 as of February 2026. Shopify’s own president said roughly a dozen. Other reporting says fewer than 15. Nobody says anything close to a million.

Survey rungs aren’t one series. The delegation cliff stacks five different panels asking differently worded questions. It is a shape, not a time series. The shape has been remarkably stable across every study I could find.

Things move. Written 19 August 2026, two weeks after the Ninth Circuit vacated the Amazon injunction. The underlying lawsuit is still live, and half of this could be stale by Diwali.

First published on Substack.