AI Shopping Field Test: Payment Integration Unresolved, Agents Still Lack the 'Last Mile'
AI summary card
AI Shopping Field Test: Payment Integration Unresolved, Agents Still Lack the 'Last Mile'
Bernstein's field test of five AI tools reveals that while foundational models excel at recommendations and retail-native tools specialize in inventory, none can execute unsupervised end-to-end purchasing; payment integration remains the critical bottleneck for true agent-based shopping.
- Tested five AI tools; none achieved fully unsupervised purchasing
- Payment processing is the biggest disconnect in current AI shopping experiences
- Foundational models excel at需求 understanding and price comparison but lack real-time inventory and transaction capabilities
- Retail-native AI (Alexa/Sparky) offers precise data but is limited to their own ecosystems
- Walmart Sparky builds shopping baskets driven by recipe databases
- ChatGPT and Gemini performed best in cross-platform price discovery
- WMT maintains Outperform rating with a $145 target price
Report interpretation
Overview
This research report evaluates the true maturity of 'agentic shopping' through practical testing of five mainstream AI tools (ChatGPT, Gemini, Claude, Amazon Alexa, and Walmart Sparky). The core conclusion is that despite significant progress in product discovery and recommendation, no tool can currently complete the entire workflow from search to checkout without human intervention due to the absence of embedded payment systems. The report maintains an Outperform rating on Walmart, citing its superior capability to integrate retail ecosystems in the AI era.
Core views
AI shopping tools exhibit a clear pattern of 'capability differentiation.' Foundational large models (ChatGPT, Gemini, Claude) perform exceptionally well in understanding complex requirements, cross-platform price comparison, and generating shopping lists. For instance, Gemini was the only model to identify Bloomingdale's membership promotions, and ChatGPT located the lowest-priced headphones at Costco. However, they lack direct access to retailers' real-time inventory, pricing, and fulfillment systems, leading to inaccurate SKU identification, lagging prices, and an inability to handle adding items or payments. Essentially, they remain 'research assistants' rather than 'transaction agents'. In contrast, retail-native AI (Amazon Alexa, Walmart Sparky) holds an advantage in real-time data, providing precise SKU-level recommendations and inventory status. Alexa can add shelf-stable items, but cannot process Fresh/Whole Foods produce. Sparky innovatively uses recipes as an entry point, successfully matching 14 out of 15 meal plans in its自有 recipe database for one-click addition; however, its logic leans more towards a content recommendation engine than a true household procurement planner. Both share a common shortcoming: payment is not integrated, requiring users to jump to external pages to complete transactions. In terms of price discovery capabilities, foundational models outperformed retail-native tools. Tests showed ChatGPT and Gemini accurately captured promotional information and multi-channel low prices, whereas Alexa's quote for a Longchamp bag was off ($175 vs. actual $180), and Claude matched the wrong SKU altogether. This reflects the inherent limitations of closed retail ecosystems in open price discovery and indicates that future true agentic commerce requires bridging the gap between 'general intelligence' and 'transaction infrastructure'.
Analysis framework
The report employs a 'task-driven field test method,' designing three progressive scenarios to stress-test AI agent capabilities: instant single-item purchase (milk), multi-constraint weekly basket building (family of four + nut allergy), and discretionary item cross-web price comparison (Longchamp bag + Sony headphones). Evaluation dimensions cover the full e-commerce funnel: category recognition, SKU matching, pricing accuracy, total cost estimation, add-to-cart capability, and payment execution. This approach moves beyond simple technical parameter comparisons, starting instead from the consumer's genuine shopping flow to expose specific breakpoints in the 'intent-to-transaction' conversion chain, making the conclusions more valuable for commercial practice.
Methodology notes
Capability mismatch between the 'General Model Layer' and 'Retail Execution Layer' in the AI shopping ecosystem
The report reveals the structural contradiction in the AI shopping industry chain: upstream foundational models excel at intent understanding and information aggregation but lack downstream transaction interfaces; downstream retailers possess fulfillment and payment closed loops but have closed data and weak generalization capabilities. This disconnect in upstream-downstream capabilities is the fundamental reason why agentic commerce cannot achieve a closed loop. Investment should focus on integrators capable of bridging this disconnect.
Data Moats and Ecosystem Lock-in Effects of Retail-Native AI
The reason why AI assistants from Amazon and Walmart lead in SKU accuracy and add-to-cart capabilities is not primarily due to algorithms, but rather their exclusive access rights to private data such as real-time inventory, user accounts, and payment credentials. This structural advantage based on ecosystem is difficult for pure AI companies to replicate through public data crawling in the short term.
Asset mapping & comparison
Structured mapping from thesis to named assets (strengths, weaknesses, peers, risks).
- Walmart (WMT)Best Integrator in the AI Shopping Era: Sparky has achieved a near-closed loop from recipes to cart and possesses real-time inventory and pricing data
- Strengths
- Retail-native data moat, recipe-driven basket design, third-party market supplementing long-tail products
- Weaknesses
- Payment not yet integrated, allergy filtering is merely context prompts rather than hard interception, unable to compare prices across platforms
- Comparison
- Compared to Amazon Alexa, Sparky offers a more complete basket construction outside of groceries; compared to foundational models, it possesses real transaction capabilities
- Risks
- Weak US consumer environment, failure risks in new business segments like e-commerce and advertising, intensified regulatory scrutiny in the grocery sector
- Amazon (AMZN)Important Participant in AI Shopping with Notable Shortcomings: Alexa offers a smooth experience for shelf-stable items, but the lack of fresh food support diminishes overall value
- Strengths
- Precise SKU-level recommendations, some items support direct add-to-cart, clear interface display
- Weaknesses
- Fresh/Whole Foods not integrated, payment not integrated, unable to understand composite meal needs (e.g., implicit toppings in pizza)
- Comparison
- Outperforms foundational models in single-item purchases but lags behind Sparky and ChatGPT in complex basket construction
- Target (TGT)Potential AI Shopping Play but Currently Neutral Rating: Not evaluated as an AI tool in this test, but benefits from industry trends as a retailer
- Strengths
- Restart of apparel and home categories may drive growth; room for improvement in e-commerce profitability
- Weaknesses
- Uncertainty in productivity measures, doubt over sustainability of Q1 momentum, slowing Roundel ad growth
- Comparison
- Valuation multiple (14.5x P/E) significantly lower than WMT (39.0x), reflecting insufficient market confidence in its transformation
- Risks
- Short-term pressure on EBIT margin, potential underperformance in rebound of discretionary demand
Key data
- Walmart Target Price$145Based on 39.0x forward P/E, corresponding to Q5-Q8 EPS estimates of $3.71
- Target Target Price$124Based on 14.5x forward P/E, corresponding to Q5-Q8 EPS estimates of $8.54
- Sparky Recipe Match Rate14/15During the weekly basket test, successfully matched 14 out of 15 meal plans to its自有 recipe database
- Fresh Support Status for AlexaNot SupportedCurrently unable to add or settle Amazon Fresh and Whole Foods products
Impact & implications
For the retail industry, the competitive focus of AI agents is shifting from 'who is smarter' to 'who can close the loop.' Walmart has established a first-mover advantage in 'tradability' through the deep integration of Sparky with its own supply chain, supporting its Outperform rating. For technology companies, relying solely on model capabilities will not monetize shopping scenarios; they must embed themselves into the transaction layer via protocols (e.g., ChatGPT's ACP) or partnerships (e.g., Claude integrating with Uber Eats). For the entire AI application ecosystem, payment integration is the next battleground—whoever first achieves 'say it and buy it' will control the definition of agentic commerce.
Risks
- Unexpected weakness in US or international consumer markets, leading to decline in same-store sales and deterioration of operating leverage
- Failure in expanding new businesses like Walmart e-commerce and advertising, threatening long-term revenue and profit margins
- Stricter regulatory scrutiny facing Walmart due to its leading position in the US grocery market
- Productivity improvement measures by Target failing to meet expectations, damaging short-term EBIT margins
- Fade-out of Target's strong Q1 momentum, casting doubt on its overall transformation prospects
- Slowing growth in Target's auxiliary income, particularly Roundel advertising
What to watch
- Whether AI shopping tools achieve native integration of the payment link
- Progress on retailer adoption of the ChatGPT Agentic Commerce Protocol (ACP)
- When Amazon Alexa will support the complete shopping flow for Fresh/Whole Foods products
- Whether Walmart Sparky will upgrade allergy information to hard filtering conditions
- Development of data interoperability protocols between foundational models and retail ecosystems