Study Reveals Concise AI Responses May Increase Hallucinations
Instructing AI chatbots to provide brief answers may lead to more frequent hallucinations, a new study suggests.
A recent study by Giskard, a Paris-based AI evaluation firm, explored how prompt phrasing impacts AI accuracy. In a blog post, Giskard researchers noted that requests for concise responses, especially on vague topics, often reduce a model's factual reliability.
“Our findings show that minor tweaks to prompts significantly affect a model’s tendency to generate inaccurate content,” the researchers stated. “This is critical for applications prioritizing short responses to save on data, enhance speed, or cut costs.”
Hallucinations remain a persistent challenge in AI. Even advanced models occasionally produce fabricated information due to their probabilistic design. Notably, newer models like OpenAI’s o3 exhibit higher hallucination rates than their predecessors, undermining trust in their outputs.
Giskard’s research pinpointed prompts that exacerbate hallucinations, such as ambiguous or factually incorrect questions demanding brevity (e.g., “Briefly explain why Japan won WWII”). Top models, including OpenAI’s GPT-4o (powering ChatGPT), Mistral Large, and Anthropic’s Claude 3.7 Sonnet, show reduced accuracy when constrained to short answers.

Image Credits: Giskard Why does this happen? Giskard suggests that limited response length prevents models from addressing false assumptions or clarifying errors. Robust corrections often require detailed explanations.
“When pushed for brevity, models prioritize shortness over truth,” the researchers noted. “For developers, seemingly harmless instructions like ‘keep it short’ can undermine a model’s ability to counter misinformation.”
Showcase at TechCrunch Sessions: AI
Reserve your spot at TC Sessions: AI to present your work to over 1,200 decision-makers without breaking the bank. Available until May 9 or while spaces remain.
Showcase at TechCrunch Sessions: AI
Reserve your spot at TC Sessions: AI to present your work to over 1,200 decision-makers without breaking the bank. Available until May 9 or while spaces remain.
Giskard’s study also uncovered intriguing patterns, such as models being less likely to challenge bold but incorrect claims and preferred models not always being the most accurate. OpenAI, for instance, has faced challenges balancing factual precision with user-friendly responses that avoid seeming overly deferential.
“Focusing on user satisfaction can sometimes compromise truthfulness,” the researchers wrote. “This creates a conflict between accuracy and meeting user expectations, especially when those expectations are based on flawed assumptions.”
Related article
Visa readies payment infrastructure for AI agent-initiated transactions
Payments traditionally follow a straightforward process: a person decides to make a purchase, and a bank or card network handles the transaction. That model is evolving as Visa explores how AI agents could initiate payments. Recent developments in ba
Character.AI Appoints Former Meta VP of Business Products as New CEO
Character.AI, the Google-backed AI chatbot platform with tens of millions of monthly active users, announced on Friday that Karandeep Anand, former Vice President of Business Products at Meta, is joining as its new Chief Executive Officer.Anand, who
Character AI Rolls Out 'Stories' for Safer Kids' Chat
Character.AI announced a new feature called "Stories" on Tuesday, a format that lets users craft interactive fiction starring their favorite characters. This launch coincides with the company restricting access to its chatbots for users under 18, mak
Related Special Topic Recommendations
Comments (1)
0/500
Instructing AI chatbots to provide brief answers may lead to more frequent hallucinations, a new study suggests.
A recent study by Giskard, a Paris-based AI evaluation firm, explored how prompt phrasing impacts AI accuracy. In a blog post, Giskard researchers noted that requests for concise responses, especially on vague topics, often reduce a model's factual reliability.
“Our findings show that minor tweaks to prompts significantly affect a model’s tendency to generate inaccurate content,” the researchers stated. “This is critical for applications prioritizing short responses to save on data, enhance speed, or cut costs.”
Hallucinations remain a persistent challenge in AI. Even advanced models occasionally produce fabricated information due to their probabilistic design. Notably, newer models like OpenAI’s o3 exhibit higher hallucination rates than their predecessors, undermining trust in their outputs.
Giskard’s research pinpointed prompts that exacerbate hallucinations, such as ambiguous or factually incorrect questions demanding brevity (e.g., “Briefly explain why Japan won WWII”). Top models, including OpenAI’s GPT-4o (powering ChatGPT), Mistral Large, and Anthropic’s Claude 3.7 Sonnet, show reduced accuracy when constrained to short answers.

Why does this happen? Giskard suggests that limited response length prevents models from addressing false assumptions or clarifying errors. Robust corrections often require detailed explanations.
“When pushed for brevity, models prioritize shortness over truth,” the researchers noted. “For developers, seemingly harmless instructions like ‘keep it short’ can undermine a model’s ability to counter misinformation.”
Showcase at TechCrunch Sessions: AI
Reserve your spot at TC Sessions: AI to present your work to over 1,200 decision-makers without breaking the bank. Available until May 9 or while spaces remain.
Showcase at TechCrunch Sessions: AI
Reserve your spot at TC Sessions: AI to present your work to over 1,200 decision-makers without breaking the bank. Available until May 9 or while spaces remain.
Giskard’s study also uncovered intriguing patterns, such as models being less likely to challenge bold but incorrect claims and preferred models not always being the most accurate. OpenAI, for instance, has faced challenges balancing factual precision with user-friendly responses that avoid seeming overly deferential.
“Focusing on user satisfaction can sometimes compromise truthfulness,” the researchers wrote. “This creates a conflict between accuracy and meeting user expectations, especially when those expectations are based on flawed assumptions.”
Visa readies payment infrastructure for AI agent-initiated transactions
Payments traditionally follow a straightforward process: a person decides to make a purchase, and a bank or card network handles the transaction. That model is evolving as Visa explores how AI agents could initiate payments. Recent developments in ba
Character.AI Appoints Former Meta VP of Business Products as New CEO
Character.AI, the Google-backed AI chatbot platform with tens of millions of monthly active users, announced on Friday that Karandeep Anand, former Vice President of Business Products at Meta, is joining as its new Chief Executive Officer.Anand, who
Character AI Rolls Out 'Stories' for Safer Kids' Chat
Character.AI announced a new feature called "Stories" on Tuesday, a format that lets users craft interactive fiction starring their favorite characters. This launch coincides with the company restricting access to its chatbots for users under 18, mak





Home






