Researchers at the Wharton School tested how AI shopping agents recommend products when the search process changes. Even minor shifts in context led to significant variations in purchase decisions. The team tested six current models, both mini variants and frontier-level, tasking each one to act as a personal shopping assistant picking a fitness watch from a fixed product grid. They used the ACES simulator (Agentic e-Commerce Simulator), which shows the AI agent a screenshot of a product page. The agent analyzes the image, optionally pulls in recommendation sources, and then picks a product. This is how the models received product information in the study.
A single source was enough to flip the recommendation. Even without external sources, the models showed different baseline preferences. But when the agent saw just one external source before the product page, recommendations shifted dramatically in some cases. The researchers tested three sources: a Reddit thread recommending the Garmin Forerunner 55, a Wirecutter review for the Fitbit Inspire 3, and a Strategist article about the WHOOP 5.0. Wirecutter had the strongest pull. The probability of picking the Fitbit Inspire 3 jumped by 90 percentage points for Claude Opus 4.8 compared to the control condition, and by 99 percentage points for Gemini 3.5 Flash. A single external source like Wirecutter shifts product picks drastically for most models. Each model also shows a different baseline preference in the control condition.
In a second experiment, agents saw combinations of two or three sources. Multiple sources didn't balance out the recommendations, though. Wirecutter tended to dominate for most models whenever it was part of the mix, though the strength of the effect varied. More sources actually led to more variability, according to the study. Combining multiple sources doesn't balance out recommendations. Wirecutter dominates whenever it's included in the mix.
Source: thedecoder