OpenAI pulled its in-chat checkout in March after six months. Varun Raj Duvalla, a senior data scientist who builds financial forecasting systems at PayPal, argues that the retreat bought the industry time without touching the problem underneath it: the machines doing the shopping leave behind almost none of the evidence the credit system was built to read.
The number that ended it was thirty. When OpenAI launched Instant Checkout in September 2025, letting people buy from Etsy sellers without leaving a ChatGPT conversation, the promise attached to it was over a million Shopify merchants to follow. By February, roughly thirty were actually live. Walmart, which had put around 200,000 products into the system, found that shoppers who checked out inside the chat converted at about a third the rate of the ones sent to walmart.com to finish.
In March, OpenAI folded the feature, saying it would let merchants run their own checkout and focus instead on helping people find things. Six months, one retreat. Plenty of people read that as the end of a hype cycle.
They are reading the wrong signal. The checkout button was never the interesting part.
What survived the retreat
The standard underneath it is still there. The Agentic Commerce Protocol, open-sourced with Stripe, outlived the product it was built for, and PayPal adopted it in October 2025. Amazon Web Services shipped a payments layer for autonomous agents in May. Visa followed within a day with cards designed for agents rather than people. The plumbing kept getting built while the storefront experiment was being dismantled.
Which leaves a strange situation. Software is increasingly doing the browsing, the comparing and the deciding, even where a human still clicks the final button on a merchant’s own site. The transaction looks ordinary when it lands. Everything that happened before it does not.
Varun Raj Duvalla builds systems that depend on exactly that missing part. A senior data scientist at PayPal Holdings, Inc., he designs forecasting models for consumer credit products across the United States and United Kingdom, the sort of work that determines how much capital a business sets aside and how confident it should be about the customers it is lending to. He spoke in a personal capacity.
His view is that the March pullback changed the timeline and nothing else.
“A checkout feature failing tells you something about that feature, not about the direction,” Duvalla said. “The behavior underneath has already shifted. People are letting software do the searching and the shortlisting. Whether the final click happens in a chat window or on a retailer’s own page is close to irrelevant to the systems I work on.”
Hesitation was carrying information
The part he considers underappreciated is what gets lost when software handles the middle of a purchase.
Online retail is engineered against the human nervous system. Two left in stock. A countdown in the corner. Free shipping if you add eleven more dollars, which is why an eleven-dollar item exists. An email the next morning about the thing sitting in an abandoned cart. None of that is aimed at anybody’s judgment, and it works often enough that a whole profession tunes it.
“Hesitation was never noise in these systems, it was one of the most useful things we had,” Duvalla said. “How long somebody lingers, how often they come back to a page, whether they buy at nine in the morning or at midnight. Those patterns carry real predictive weight. An agent produces none of them. The purchase arrives clean, and clean is far less informative than it sounds.”
Put differently, the industry spent two decades building models that read human indecision, and indecision is being quietly edited out of the record.
The person behind the proxy
A second effect follows, and Duvalla treats it as the harder one. When software mediates a purchase, what reaches the lender is the software’s output, not the person’s process.
“You are no longer observing a customer, you are observing a representative,” he said. “Two people in completely different financial situations can produce histories that look nearly identical, because both delegated to the same tool with similar instructions. The variation you relied on to tell them apart has been flattened by whatever is sitting in between.”
For credit, that flattening is not academic. Behavioral signals are one input among several in assessing risk, and they are the input most exposed here. Regulators have been tightening from a different direction: the European Union’s AI Act treats creditworthiness assessment as a high-risk use, and the Consumer Financial Protection Bureau moved in 2024 to bring buy now, pay later products under Truth in Lending obligations, a position since contested.
“If a model loses the ability to distinguish between two customers, it does not announce that,” Duvalla said. “It keeps producing an answer. Somebody gets approved who should have been looked at harder, or somebody gets declined for reasons nobody can reconstruct later. That is not a technical footnote. It is a decision about a person made on evidence that no longer means what it used to.”
Nobody is writing it down
Asked what he would fix first, Duvalla names a missing column rather than a modeling technique.
“Almost nowhere is anyone reliably recording whether a transaction was initiated by a person or by an agent,” he said. “Without that flag you cannot measure the effect, you cannot separate the two populations, and you cannot tell whether a model is degrading or the customer base is changing underneath it. Everything downstream depends on a piece of provenance most systems are not capturing today.”
It is an unglamorous request, and consistent with the wider record on artificial intelligence programs that underdeliver. RAND has reported that more than 80 percent of such projects fail, roughly double the rate for conventional technology work, with weak data foundations rather than weak algorithms cited repeatedly as the cause.
His second suggestion is stranger and, he concedes, unpopular internally. Deliberately hold back a slice of human-only transactions as a reference group, even as agent-assisted volume grows, because without something to compare against there is no way to know what the shift has actually done.
The quiet question
The industry will spend the next two years arguing about standards, and about which intermediary collects a fee when software buys a pair of shoes. Those arguments come with enormous sums attached and they will get the coverage.
Duvalla’s closing point is that the argument is happening one layer above the problem.
“Whoever wins the protocol fight, every lender still has to answer the same question,” he said. “How do you assess somebody you can only see through a piece of software? Nobody has a settled answer. The decisions are being made anyway, weekly, at scale.”
Which leaves an odd conclusion about a story most people filed under failed product launch. The consequential thing about machines doing the shopping may have very little to do with shopping. It is that a system built to read people has started receiving its information secondhand, and has not yet been told.
