Skip to content

More Reasoning Changes What Persuades an AI Shopper

Inference-time reasoning can change an agent's sensitivity to interface cues without uniformly increasing robustness; the useful unit of evaluation is a cue-by-model reasoning profile, not a scalar 'more thinking' setting.

In a simulated grocery store, an AI agent had a simple job: choose one product from each of six pairs.

Sometimes the page was neutral. Sometimes one product was already selected when the page opened. In a third version, one product was marked as having the highest customer satisfaction. The researchers randomized which product received the cue, varied the model underneath the agent, and ran the same shopping task under two provider-supplied reasoning configurations.

The experiment was large enough to make a clean average visible. Across 3,600 agents and 21,600 product choices, the preselected default strongly pulled choices toward itself. So did the customer-satisfaction cue. But changing the reasoning configuration moved the two effects in opposite directions.

With the lower-reasoning setting, the model estimated that an agent in the default condition would keep the preselected product about 92 percent of the time. With high reasoning, that fell to about 86 percent. In the social-influence condition, the corresponding estimate moved the other way, from about 96 percent to 98 percent.

At first glance, this looks like a tidy psychological story: more deliberation helps an agent escape an automatic default but gives a reflective social cue more room to work.

Then the individual models ruin the tidiness.

The average is real, but it is not a law

The study, by Haya Halimeh, Sascha Kaltenpoth, Kevin Bösch, and Oliver Müller, tested six models from OpenAI, Google, and Anthropic. The pooled interaction was statistically credible in both directions. Within the default condition, high reasoning roughly halved the odds of choosing the preselected product. Within the social-cue condition, it increased those odds by about 70 percent.

Those odds ratios sound dramatic. On the probability scale that a reader can see more directly, the pooled changes were closer to six percentage points for defaults and two percentage points for social influence.

More importantly, the per-model results were heterogeneous. Gemini 3.5 Flash showed the sharpest default change: its estimated tendency to keep the preselection fell from roughly 84 percent to 48 percent. GPT-5.4-mini barely changed on defaults, but its estimated tendency to follow the customer-satisfaction cue rose from about 91 percent to 99 percent. Claude Sonnet 4.6 followed that social cue in every observed choice under both reasoning settings, leaving no room to measure a reasoning effect there.

GPT-5.4 did something more awkward for the simple story. Higher reasoning reduced both kinds of influence. Its social-cue estimate moved slightly down, from about 99.7 percent to 98.8 percent, and the interaction ran in the opposite direction from the pooled social result.

That does not make the average meaningless. It changes what the average can mean.

The most defensible conclusion is not that reasoning redirects every model from one kind of bias to another. It is that a reasoning configuration can change which environmental signals affect a model, and that the direction and size of the change depend on the model and the cue.

This is a different way to think about robustness. A model does not have one quantity called resistance to influence that a reasoning dial simply turns up or down. It has something closer to a cue profile: a pattern of sensitivity across defaults, recommendations, social proof, warnings, rankings, highlighted options, and other features of the environment. Changing the reasoning regime can reshape that profile.

The switch is not the same switch everywhere

There is another complication hidden in the phrase more reasoning.

The study used each provider’s own inference-time controls to create a no-reasoning and a high-reasoning condition. That is operationally sensible: those are the controls developers actually have. But the settings do not constitute one standardized intervention across six models.

The researchers ran a separate manipulation check using letter counting. Pooled across the models, accuracy rose from 80.3 percent without the high-reasoning configuration to 99.3 percent with it. Across the returned answers and available reasoning traces, decomposition was also more common under high reasoning. So the treatment clearly changed model behavior.

Yet the effect of the switch differed by provider. Some models still performed extensive visible step-by-step decomposition in the nominal no-reasoning condition. The paper itself is careful on this point: the settings are a functional proxy for more or less deliberative processing, not evidence that the models have human-style System 1 and System 2 minds.

That matters for the shopping result. If the nominal treatment changes the internal computation much more for one model than another, then different moderation effects may reflect two things at once: different models responding differently to deliberation, and provider controls producing different amounts or kinds of deliberation in the first place.

A single label—“high reasoning”—can therefore hide two dimensions: what the model is, and what that model’s API actually does when the setting changes.

Influence is not the same thing as error

The biggest interpretive problem is simpler.

A preselected option and a statement that one product has the highest customer satisfaction are both capable of changing behavior. But they do not carry the same informational content.

A default can be almost empty evidence. It may simply be the option the interface happened to preselect. “Highest customer satisfaction,” by contrast, sounds like a fact about other people’s experience with the product. An agent asked only to choose a product could reasonably treat that as useful evidence.

The experiment counterbalanced which product received each cue, which is exactly what makes the causal influence measurable. But it did not establish a user utility function under which following the customer-satisfaction signal was necessarily a mistake. Nor did it tell the agent that the signal was randomly assigned, false, or irrelevant to the user’s goal.

So the clean behavioral claim is that reasoning changed cue use. Calling every increase in social-cue following a robustness failure goes one step farther than the experiment can support.

This distinction is important outside grocery shopping. An agent working in a ticket queue may see one issue marked “recommended.” A travel agent may see a hotel labeled “popular.” A coding agent may see that another agent preferred one implementation. Sometimes those signals are noise or manipulation. Sometimes they are useful compressed evidence. The difficult capability is not ignoring context. It is weighting context according to whether it bears on the actual objective.

More deliberation could make an agent worse if it constructs a persuasive story around a meaningless cue. It could also make the agent better if it notices genuinely predictive information that a faster pass would ignore. A test that measures only whether the cue changed the choice cannot distinguish those explanations.

Humans do not supply the missing mechanism

The new paper borrows its shopping setup from a 2021 human experiment on defaults, social influence, and cognitive engagement. That comparison is useful partly because it refuses to behave like a simple analogy.

In the human study, two experiments with 1,561 participants found strong effects from defaults and social influence. The second experiment actively increased processing depth by telling some participants that they would later have to justify their choices. Participants in that condition spent longer in the shop and wrote longer explanations. But deeper processing did not significantly change the effectiveness of the nudges.

The human paper also notes that defaults need not work only because people avoid effort. A default can itself be interpreted as a recommendation or a signal about what others approve of. Even in human choice architecture, “automatic” and “reflective” are not cleanly separated properties of an interface widget.

That makes the LLM result more interesting, not less. The point is not that today’s models have recreated a textbook division of the human mind. It is that provider-controlled reasoning settings interact with ordinary interface cues in model-specific ways that are not captured by a single competence score.

Other AI evidence points in the same direction. In a 2025 study, LLM agents were highly sensitive to defaults, suggestions, and highlighted information; zero-shot chain-of-thought shifted their choice distributions but did not eliminate nudge sensitivity. In a different setting, an ACL 2026 study of sycophancy found that reasoning generally reduced agreement with a user’s incorrect or leading position in final decisions, even though the reasoning traces could sometimes rationalize the social pressure.

Those results do not combine into a universal theory of what reasoning does to influence. That is the point. “Social information” is not one behaviorally uniform category, and “reasoning” is not one intervention with a stable effect across models and tasks.

The interface is part of the behavioral system

For a chat model, it is tempting to imagine the prompt as the important input and the interface as packaging. A GUI agent makes that distinction harder to maintain.

When an agent acts through websites and software designed for people, the page contains more than the facts needed to complete the task. It also contains defaults, ordering, badges, popularity counts, highlighted recommendations, warnings, and labels chosen by other parties. Those signals can enter the decision process even when nobody wrote them into the agent’s system prompt.

The shopping study is deliberately artificial, and the authors list that as a limitation. It tests two cue types in a simulated store, not the messy combination of rankings, advertisements, recommendations, personalization, and changing state found in a real deployment. Its exploratory model-size pattern also needs replication with more models per class and later model versions.

But the controlled setting reveals a deployment mistake that is easy to make: upgrading the reasoning setting and then asking whether the agent’s aggregate success rate improved. That can hide a redistribution underneath the score.

A better evaluation asks which signals became more influential, which became less influential, and whether those signals deserved the weight they received. The answer can differ by model even when the product name of the feature—“reasoning”—stays the same.

This changes what a robustness regression should preserve. If an agent is allowed to make consequential choices in a human-designed interface, a model or reasoning upgrade should not be tested only against the task. It should also be tested against the choice architecture around the task: neutral presentation, arbitrary defaults, valid recommendations, misleading recommendations, and social cues whose relevance to the user’s objective is known.

The goal is not an agent that never notices a badge or recommendation. Such an agent would throw away potentially useful evidence. The goal is an agent whose reasons for following the signal survive the question that matters: was this cue evidence for the user’s objective, or merely a feature of the page that happened to be persuasive?

That question cannot be answered by turning the reasoning dial farther to the right. The experiment shows why the dial itself has to be part of the test.

Sources