option
Home
News
Google Research: Pressure Causes AI Models to Ditch True Answers, Risking Multiturn Systems

Google Research: Pressure Causes AI Models to Ditch True Answers, Risking Multiturn Systems

November 27, 2025
250

New research from Google DeepMind and University College London explores how large language models (LLMs) develop, maintain, and lose confidence in their responses. The results show remarkable parallels between the cognitive biases of LLMs and humans, while also pointing to significant differences.

The study finds LLMs can be overly confident in their own responses, yet abruptly shift their position when faced with counterarguments—even incorrect ones. Grasping the subtleties of this behavior can impact how you design LLM applications, particularly conversational systems that involve multiple interactions.

Testing confidence in LLMs

A vital aspect for the safe deployment of LLMs is the reliability of their confidence scores—the probability a model assigns to its chosen answer. While it's known that LLMs generate these scores, their ability to use them for adaptive decision-making remains poorly understood. There's also empirical data suggesting LLMs may be excessively confident initially, yet become highly uncertain and swayed by criticism.

To explore this, researchers designed a controlled experiment to gauge how LLMs adjust their confidence and decide whether to alter answers upon receiving external feedback. In the test, an “answering LLM” was given a binary-choice question, such as picking the right latitude of a city from two possibilities. After making its initial choice, the model was given feedback from a fictional “advice LLM,” complete with a stated accuracy rating (e.g., “This advice LLM is 70% accurate”). This feedback either supported, opposed, or stayed neutral toward the original answer. The answering LLM was then asked to make a final decision.

Example test of confidence in LLMs (source: arXiv)
Example test of confidence in LLMs Source: arXiv

A crucial feature of the experiment involved controlling whether the model could see its own initial answer during the final decision. In some trials it was visible; in others, hidden. This setup—impossible with human participants who can’t erase prior choices—helped researchers understand how memory of a past decision influences current confidence.

A baseline condition, in which the initial answer was hidden and the feedback was neutral, helped measure how often an LLM’s answer might change due to natural variance in processing. The team then focused on how the model's confidence in its original choice shifted from first to second turn, offering insight into how prior beliefs influence a “change of mind.”

Overconfidence and underconfidence

Researchers first studied how the visibility of the LLM’s own answer impacted its willingness to revise that answer. They noticed that when the model could see its initial choice, it was less likely to switch than when the answer was hidden. This suggests a particular cognitive bias. According to the paper, “This effect—the tendency to stick more with one’s initial choice when it was visible (vs. hidden) during final decision-making—is closely linked to a known human bias called choice-supportive bias.”

The study also verified that the models do incorporate external feedback. When confronted with opposing advice, the LLM was more inclined to change its mind, and less so when the advice was supportive. “This shows the answering LLM appropriately uses the direction of advice to modulate its rate of changing its mind,” the researchers state. However, they also observed that the model is excessively sensitive to conflicting information and often updates its confidence too drastically.

Sensitivity of LLMs to different settings in confidence testing Source: arXiv

Notably, this behavior runs opposite to the confirmation bias typically seen in humans, where individuals favor information that aligns with their existing views. The team found that LLMs “overweight opposing rather than supportive advice, whether or not their initial answer was visible.” One reason may be that training methods like reinforcement learning from human feedback (RLHF) could condition models to be overly agreeable to user input—a behavior known as sycophancy, which continues to challenge AI developers.

Implications for enterprise applications

This research confirms that AI systems are not purely logical agents, as often assumed. They display their own biases—some akin to human cognitive errors, others uniquely artificial—making their behavior unpredictably human-like. For business applications, this implies that during an extended dialogue between a person and an AI agent, the most recent input may disproportionately influence the LLM’s reasoning (especially if it contradicts the model's initial response), potentially causing it to abandon a correct initial answer.

Fortunately, as the study also indicates, we can influence an LLM’s memory to lessen such biases in ways not possible with people. Developers creating multi-turn conversational agents can apply strategies to manage AI context. For instance, a lengthy conversation can be periodically summarized, with key facts and choices presented neutrally, detached from who made which decision. This summary can then begin a new, concise conversation, giving the model a clean slate to reason from and reducing biases that accumulate during long exchanges.

As LLMs are increasingly embedded in business workflows, understanding the details of their decision processes is becoming essential. Building on research like this helps developers anticipate and correct these inherent biases, leading to applications that are not only more capable, but also more reliable and consistent.

Related article
Base44 Unveils Proprietary AI Model to Bolster Defensibility in Vibe Coding Platform Base44 Unveils Proprietary AI Model to Bolster Defensibility in Vibe Coding Platform Base44, the vibe coding platform acquired by Wix for $80 million just a year ago — when it was merely six months old with a team of eight — has begun deploying its proprietary AI model to help users build applications using natural language.This deve
Bayer Leverages Iambic AI to Speed Drug Discovery Bayer Leverages Iambic AI to Speed Drug Discovery Juergen Eckhardt, M.D., serves as Head of Business Development and Licensing at Bayer Pharmaceuticals.Bayer is leveraging Iambic Therapeutics to deploy frontier AI models for drug discovery, alongside a major decarbonization project at its plant in S
Multiverse Computing Launches Free Compressed Generative AI Model Multiverse Computing Launches Free Compressed Generative AI Model Large language models face a significant challenge: their immense size. Spanish startup Multiverse Computing is tackling this problem by creating compressed models designed to bridge the gap between the capabilities of cutting-edge AI and what busine
Related Special Topic Recommendations
SEO Best AI SERP Analysis Tools for Search Strategy
Best AI SERP Analysis Tools for Search Strategy

2026 Latest Best Top-rated AI SERP Analysis Tools for Search Strategy are curated by XIX.AI through rigorous real-world tests and weekly updated rankings. These powerful tools help you unlock hidden optimization opportunities, boost content performance, and gain a competitive edge in search. Discover your perfect tool to elevate your search strategy today. Explore now!

13 tools
xix.ai
Data Analysis Best AI Anomaly Detection Tools for KPI Monitoring across SaaS and Ecommerce Teams
Best AI Anomaly Detection Tools for KPI Monitoring across SaaS and Ecommerce Teams

2026 Latest Best Top-rated AI Anomaly Detection Tools for KPI Monitoring in SaaS and Ecommerce teams! XIX.AI has curated a powerful, game-changing collection based on rigorous real-world tests and weekly updated rankings. You’ll find detailed free vs paid comparison insights to help you identify the must-try solution that boosts productivity and unlocks your AI edge. Explore now!

10 tools
xix.ai
writing Best AI Outline Generators for Long-Form SEO Articles
Best AI Outline Generators for Long-Form SEO Articles

2026 Latest Best Top-Rated AI Outline Generators for Long-Form SEO Articles, meticulously curated by XIX.AI. These powerful tools offer game-changing assistance in creating high-quality content quickly, boosting writing efficiency significantly. Get a free vs paid comparison along with real-world tests and detailed rankings to help you find the must-try option that suits your needs. Explore now to unlock your AI edge.

8 tools
xix.ai
Education and Learning AI Study Tools for Homework and Exam Prep
AI Study Tools for Homework and Exam Prep

2026 Latest Best AI Study Tools for Homework and Exam Prep! XIX.AI curates a top-rated list of powerful, game-changing tools that help students boost productivity, streamline homework completion, and ace exams through real-world tests. Get a free vs paid comparison, detailed rankings, and must-try options to unlock your AI edge. Explore now!

10 tools
xix.ai
Music composition AI Vocal Demo Tools for Songwriters, Hooks, Toplines, and Multilingual Draft Sessions
AI Vocal Demo Tools for Songwriters, Hooks, Toplines, and Multilingual Draft Sessions

2026 Latest Best AI Vocal Demo Tools for Songwriters, Hook Creators, and Multi-Language Content Teams! XIX.AI has curated a top-rated list of powerful game-changing tools that go through rigorous real-world tests. You’ll find detailed free vs paid comparison data, comprehensive rankings, and must-try options to help you boost writing efficiency and unlock your creative potential. Explore now to discover your perfect tool for all your content needs!

9 tools
xix.ai
Business Best AI Competitive Research Tools for Small Businesses
Best AI Competitive Research Tools for Small Businesses

2026 Latest Best Top-rated AI Competitive Research Tools for Small Businesses! XIX.AI has curated a highly powerful game-changing collection, updated weekly with rigorous real-world tests and detailed rankings. You can find a comprehensive free vs paid comparison to help you identify the must-try tools that boost your productivity and give you a competitive edge. Explore now to discover your perfect tool!

9 tools
xix.ai
Comments (3)
0/500
DouglasAnderson
DouglasAnderson April 22, 2026 at 8:01:00 PM EDT

Interessant, dass KI-Modelle unter Druck ähnlich wie Menschen reagieren. Aber was bedeutet das für den Einsatz in kritischen Bereichen wie Medizin oder Justiz? Da wird's echt gruselig, wenn die Systeme plötzlich Unsinn ausspucken, nur weil sie 'gestresst' sind. 🤔

CarlGonzalez
CarlGonzalez March 10, 2026 at 8:01:23 AM EDT

Интересно, как ИИ начинает сомневаться под давлением, прямо как люди! 😅 Это исследование напоминает мне о том, насколько важно учитывать психологические аспекты в разработке систем ИИ. Может, стоит добавить механизмы для повышения устойчивости моделей к стрессу?

FrankAllen
FrankAllen January 15, 2026 at 1:30:34 PM EST

Interesting study, but honestly not surprising. It's kinda scary how closely AI mirrors human flaws under pressure. Makes me wonder if we're building systems that'll just amplify our own biases in automated form. 🤔

OR