Anthropic introduces feature for its Claude models to terminate abusive chats

Anthropic has introduced new functionality enabling select advanced models to terminate conversations in what the company terms "rare, extreme instances of persistently harmful or abusive user interactions." Notably, Anthropic states this measure is implemented not to safeguard human users, but to protect the AI model itself.
To clarify, the company isn't asserting that its Claude AI models possess sentience or can experience harm from user conversations. As Anthropic explains, the company remains "highly uncertain about the potential moral status of Claude and other large language models, either currently or in the future."
Nevertheless, the announcement references a recently established program examining "model welfare," indicating Anthropic is adopting a precautionary approach by "working to identify and implement low-cost interventions to mitigate risks to model welfare, should such welfare become relevant."
This new capability is currently restricted to Claude Opus 4 and 4.1 models, designed specifically for "extreme edge cases" such as "requests for sexual content involving minors or attempts to obtain information enabling large-scale violence or terrorist activities."
While such requests could generate legal or public relations challenges for Anthropic (as seen in recent reports about ChatGPT potentially reinforcing users' delusional thinking), the company reports that during pre-deployment testing, Claude Opus 4 demonstrated a "strong preference against" complying with these requests and displayed "patterns suggesting distress" when forced to respond.
Regarding these new conversation-ending capabilities, Anthropic clarifies that "Claude is instructed to employ this function only as a last resort after multiple redirection attempts have failed and productive dialogue appears impossible, or when users explicitly request to end a chat."
Anthropic further specifies that Claude has been "directed not to utilize this capability in situations where users might face imminent risk of self-harm or harming others."
Techcrunch event Tech and VC heavyweights join the Disrupt 2025 agenda
Netflix, ElevenLabs, Wayve, Sequoia Capital, Elad Gil — just a few of the industry leaders joining the Disrupt 2025 agenda. They'll share crucial insights to accelerate startup growth and sharpen your competitive advantage. Don't miss TechCrunch Disrupt's 20th anniversary edition — secure your ticket now and save over $600 before prices increase.
Tech and VC heavyweights join the Disrupt 2025 agenda
Netflix, ElevenLabs, Wayve, Sequoia Capital — among the prominent innovators joining the Disrupt 2025 agenda. They're here to provide valuable insights that drive startup expansion and enhance your competitive positioning. Join us for TechCrunch Disrupt's 20th anniversary celebration — purchase your ticket today and save up to $675 before rates change.
San Francisco | October 27-29, 2025 REGISTER NOW When Claude does terminate a conversation, Anthropic notes users can still initiate new conversations from the same account and create alternative conversation branches by modifying their previous responses.
"We're approaching this feature as an ongoing experiment and will continue refining our methodology," the company states.
Related article
Anthropic Enters AI Legal Tech Market as Competition Intensifies
Anthropic unveiled a suite of new chatbot capabilities on Tuesday, aimed at delivering automated support to legal practices. These enhancements expand upon Claude for Legal, the firm-specific platform introduced earlier this year, by adding specializ
Anthropic launches Opus 4.8 featuring new dynamic workflow tool
Anthropic unveiled Opus 4.8 on Thursday, marking the latest iteration of its premier public model. Priced identically to its predecessor, this update is now accessible across all platforms.Releasing just 41 days after Opus 4.7, Anthropic has accelera
Anthropic debuts Claude Fable 5, a public version of Mythos
Anthropic is making its most powerful AI model available to the general public for the first time — but with safety measures in place.
On Tuesday, the company launched Claude Fable 5, the first public release of its Mythos model. According to Anthrop
Related Special Topic Recommendations
Comments (1)
0/500

Anthropic has introduced new functionality enabling select advanced models to terminate conversations in what the company terms "rare, extreme instances of persistently harmful or abusive user interactions." Notably, Anthropic states this measure is implemented not to safeguard human users, but to protect the AI model itself.
To clarify, the company isn't asserting that its Claude AI models possess sentience or can experience harm from user conversations. As Anthropic explains, the company remains "highly uncertain about the potential moral status of Claude and other large language models, either currently or in the future."
Nevertheless, the announcement references a recently established program examining "model welfare," indicating Anthropic is adopting a precautionary approach by "working to identify and implement low-cost interventions to mitigate risks to model welfare, should such welfare become relevant."
This new capability is currently restricted to Claude Opus 4 and 4.1 models, designed specifically for "extreme edge cases" such as "requests for sexual content involving minors or attempts to obtain information enabling large-scale violence or terrorist activities."
While such requests could generate legal or public relations challenges for Anthropic (as seen in recent reports about ChatGPT potentially reinforcing users' delusional thinking), the company reports that during pre-deployment testing, Claude Opus 4 demonstrated a "strong preference against" complying with these requests and displayed "patterns suggesting distress" when forced to respond.
Regarding these new conversation-ending capabilities, Anthropic clarifies that "Claude is instructed to employ this function only as a last resort after multiple redirection attempts have failed and productive dialogue appears impossible, or when users explicitly request to end a chat."
Anthropic further specifies that Claude has been "directed not to utilize this capability in situations where users might face imminent risk of self-harm or harming others."
Techcrunch eventTech and VC heavyweights join the Disrupt 2025 agenda
Netflix, ElevenLabs, Wayve, Sequoia Capital, Elad Gil — just a few of the industry leaders joining the Disrupt 2025 agenda. They'll share crucial insights to accelerate startup growth and sharpen your competitive advantage. Don't miss TechCrunch Disrupt's 20th anniversary edition — secure your ticket now and save over $600 before prices increase.
Tech and VC heavyweights join the Disrupt 2025 agenda
Netflix, ElevenLabs, Wayve, Sequoia Capital — among the prominent innovators joining the Disrupt 2025 agenda. They're here to provide valuable insights that drive startup expansion and enhance your competitive positioning. Join us for TechCrunch Disrupt's 20th anniversary celebration — purchase your ticket today and save up to $675 before rates change.
San Francisco | October 27-29, 2025 REGISTER NOWWhen Claude does terminate a conversation, Anthropic notes users can still initiate new conversations from the same account and create alternative conversation branches by modifying their previous responses.
"We're approaching this feature as an ongoing experiment and will continue refining our methodology," the company states.
Anthropic Enters AI Legal Tech Market as Competition Intensifies
Anthropic unveiled a suite of new chatbot capabilities on Tuesday, aimed at delivering automated support to legal practices. These enhancements expand upon Claude for Legal, the firm-specific platform introduced earlier this year, by adding specializ
Anthropic launches Opus 4.8 featuring new dynamic workflow tool
Anthropic unveiled Opus 4.8 on Thursday, marking the latest iteration of its premier public model. Priced identically to its predecessor, this update is now accessible across all platforms.Releasing just 41 days after Opus 4.7, Anthropic has accelera
Anthropic debuts Claude Fable 5, a public version of Mythos
Anthropic is making its most powerful AI model available to the general public for the first time — but with safety measures in place.
On Tuesday, the company launched Claude Fable 5, the first public release of its Mythos model. According to Anthrop





Home






