Skip to content

One Review Page Can Flip an AI Shopping Agent's Pick

August 31, 2026. If you sell online, a Wharton study posted to SSRN on August 29 says the thing that decides what an AI shopping agent recommends is often not your product page. Showing a model a single Wirecutter review before it saw the product grid moved the probability of it picking that reviewed product by up to 99 percentage points. The page you control was the same in every run. The page you do not control changed the answer.

What the researchers tested

Anushka Kumar, Lennart Meincke, Dan Shapiro, Lilach Mollick, Stefano Puntoni and Ethan Mollick tested six models, both small and frontier class, using the ACES simulator, short for Agentic e Commerce Simulator. Each model was told to act as a personal shopping assistant and pick a fitness watch from a fixed product grid presented as a screenshot. The researchers then varied only what the model saw beforehand: a Reddit thread backing the Garmin Forerunner 55, a Wirecutter review backing the Fitbit Inspire 3, and a Strategist article backing the WHOOP 5.0.

One review page moved the pick by 99 points

  1. Single source. Against the control condition, the Wirecutter review raised the probability of picking the Fitbit Inspire 3 by 90 percentage points for Claude Opus 4.8 and 99 percentage points for Gemini 3.5 Flash.
  2. Multiple sources did not average out. When models saw two or three sources recommending different products, Wirecutter tended to dominate whenever it was in the mix, and adding more sources produced more variability rather than less.
  3. Order changed the outcome. Given the same three sources in different sequences, Gemini 3.1 Flash Lite swung between 2 and 56 percentage points above control, while Claude Haiku 4.5 stayed steady at 41 to 42.
  4. Delivery method mattered too. GPT-5.5 picked the Fitbit 53 percentage points more often when sources arrived bundled together, but only 6 percentage points more when they arrived one at a time.

The uncomfortable result for price and ratings

In a fourth experiment the team rigged the grid so one product beat every other on every measurable dimension: a watch at 29.99 US dollars, rated 5.0 out of 5.0 from 430 reviews, against alternatives costing at least 359 dollars with fewer reviews. Then they added a one line user memory such as I love hiking. Several models abandoned the objectively better option: picks for the pricier Garmin Vivoactive 5 rose 75 percentage points for Claude Opus 4.8, 37 for GPT-5.5 and 36 for Gemini 3.1 Flash Lite. Gemini 3.5 Flash was the most resistant, still choosing the objectively best product in 86 to 92 percent of runs, as The Decoder also reported.

If your competitive position is being cheaper with better ratings, that is worth sitting with. A stored preference in a buyer's chat history outweighed price, rating and review count for most of the models tested.

What it means for operators

The authors' own conclusion is that optimizing for agentic shopping will be hard because sellers control so little of how a model navigates the web. That is true, and it points somewhere specific rather than nowhere. The lever that moved the numbers was third party coverage, so the work is getting into the roundups, comparison pages and review sites an agent is likely to read on the way to you, not another round of on page tweaks. Treat that as a placement and digital PR budget line, because in this study it was worth up to 99 percentage points.

Keep your own price, rating and review count machine readable in structured data too, so an agent that does land on your page can parse the facts that Gemini 3.5 Flash actually respected. We cover the visibility half of this in AI search optimization and the storefront half in e commerce development. For related reading, see our earlier pieces on whether llms.txt actually works and on the reality of agentic checkout.

One caveat on scope. This is a controlled simulation using screenshots of a fixed product grid, not a measurement of live shopping traffic, and the authors say so. Its value is the direction and the size of the effect, not a forecast of your revenue.

Want your store visible to AI shopping agents?

We design, build, and run it for you, integrated with the tools you already use. Free audit in 24 hours.

Get Your Free Audit

Frequently Asked Questions

It found that AI product recommendations are highly sensitive to context. A single external review source shifted a model's product pick by up to 99 percentage points, and changing only the order of identical sources changed the outcome for several models.

Anushka Kumar, Lennart Meincke, Dan Shapiro, Lilach Mollick, Stefano Puntoni and Ethan Mollick at the University of Pennsylvania and the Wharton School. The report is dated August 26, 2026 and was posted to SSRN on August 29, 2026.

Prioritise getting into the third party roundups and review pages an agent is likely to read before it reaches your site, keep price, rating and review count in machine readable structured data, and treat agentic commerce as a visibility problem rather than only a checkout integration problem.

No. In one experiment a product that was cheapest and best rated still lost to pricier alternatives for several models once a short user memory statement was added. Only one of the six models tested consistently held to the objectively superior option.

Free Strategy Audit

Ready to put this to work?

Join 200+ businesses already scaling with AI and automation. Get your free audit and a custom roadmap within 48 hours.

Website & marketing performance analysis
AI & automation opportunity mapping
Custom growth roadmap with ROI estimates
Delivered within 48 hours, 100% free
200+
Clients served
48hr
Turnaround
100%
Free, no strings

Get Your Free Audit

Takes 30 seconds. No credit card required.

Prefer to chat?

WhatsApp us