From Chat to Errands: What Agent Commerce Looks Like
Sitting in a moving car, you say "order me a hot latte." The infotainment system recognizes the intent, places the order, matches the nearest store, pre-pays, and by the time you arrive the drink is ready. No app opened, no thumb on a screen.
That was the live demo at Alipay's AI ecosystem partner conference on August 17. Behind the one sentence is a full industry chain: the car maker's infotainment system on the user side, Alipay's "Abao" agent as the orchestration layer in the middle (intent recognition, personal preferences, merchant capability and location, AI payments), and a physical merchant like Luckin Coffee completing fulfillment on the ground. Every link crosses hardware, an AI platform, and a physical business. If any one of them lacks standardized coordination, the whole task falls apart.
This is the clearest sign yet that AI's center of gravity has moved out of the lab and into real-world service delivery.
Two Bottlenecks That Kept Agents Stuck in Demos
Why haven't agents already been running everyone's errands? Language understanding and speech recognition stopped being the hard part years ago. The bottlenecks were on the supply and distribution sides.
On supply, most real-world services simply cannot be called by AI. A merchant's inventory, ordering system, and redemption flow live in separate digital silos. An agent can understand "buy a coffee" or "ship a package," but it cannot open an order, complete a payment, and trigger fulfillment when the underlying systems were never built to be called.
On distribution, user entry points are fragmented across phones, cars, smart home devices, and government mini-apps. A merchant that plugs into one platform does not get stable traffic; a merchant that wants all channels must re-adapt to multiple technical standards. The two bottlenecks together keep most agent projects at the demo stage, confined to small closed scenarios.
Alipay's Three-Part Answer: A Base, a Protocol, and Incentives
Alipay's response has three layers.
First, a tiered agent foundation for merchants, upgraded in three stages: stage one is managed AI (standardized toolkits for low-barrier entry), stage two is autonomous intelligence (agent building, distribution, intent insight, and operating assistance), and stage three is multi-agent collaboration, where agents form a new commercial network.
Second, the AHA (Agent Hub Access) protocol, a shared industry-wide standard that lets phones, cars, hardware, and merchant agents talk to each other without repeated re-integration. It is effectively a common language between agents, and the piece that makes cross-device distribution possible.
Third, an incentive policy that lowers the entry barrier for small merchants and government agencies to plug into the AI ecosystem.
The numbers so far: Abao has AI-enabled more than 10,000 services, spans 5 phone brands and 16 major car makers across 8 core scenarios. KFC users complete full voice ordering and payment. Fengchao's parcel lockers serve 368 million users daily through voice-command pickup, home services, and storage. Huangshan scenic area has voice ticketing and tour guidance. Jiangxi province's digital government assistant handles social security inquiries by voice.
Why This Is the Internet's E-Commerce Moment
The shape of this transition is familiar. The early internet delivered browsing and email — genuinely new, yet not life-changing. It only completed its commercial loop when real services (shopping, bill payment) moved online. AI is replaying that arc: phase one proved machines can understand language; phase two asks whether they can get things done.
Venture money has already voted. Sequoia's Roelof Botha has said the firm's money is not for paying sky-high training costs but for companies that use models well rather than build models. Palantir made the same point from the enterprise side: customers do not care which foundation model is inside, only whether AI delivers measurable returns. As open-source models commoditize raw capability, the marginal return from stacking ever-bigger frontier models keeps falling — the industry narrative has shifted from benchmarks to ROI.
The Moat Has Moved: From Models to Connections
This is the structural claim underneath the event: compute and parameters are no longer the moat in AI. The scarce resources now are the physical services an agent can reliably call, the terminal distribution network that covers every scenario, and the fulfillment infrastructure — identity verification, payment settlement, risk control, and accountability — that turns a conversation into a trusted transaction.
That is why A2A (agent-to-agent) interoperability stops being a research concept and becomes a commercial necessity. No single company owns every interaction point; users commute in cars, live on phones, and sit in front of smart speakers at home. Open, cross-brand coordination is the only path to scale. Ant Group CEO Han Xinyi predicted at the conference that agent commerce will explode within 6 to 12 months, with agents becoming the new carrier between hundreds of millions of users and tens of millions of merchants.
Enterprise adoption is following a parallel path — we covered how Oracle is turning the database into an agent platform — but consumer agent commerce raises the stakes: the winning layer is not the model but the protocol that connects users, merchants, and fulfillment.
China also brings a specific advantage to this race: a mobile internet substrate with mature payments and service networks already in place. It does not need to rebuild supply-side infrastructure; standardized agent connectors can sit directly on top of it.
What to Do Now
If you run a business: make your services AI-callable. Standardize ordering, inventory, and redemption APIs so an agent can complete a transaction end to end. The companies that get called first will be the ones that win agent traffic.
If you build agents: design for cross-platform distribution from day one, and assume A2A protocols are coming — build against them before they are forced on you.
If you watch the market: stop treating benchmark tables as the primary signal. The meaningful metric in the second half of AI is how many real-world tasks a system can complete, and who owns the connection layer that makes completion possible.