Home
Anthropic Reveals Real-World Hack on Claude, Details Security Breach and Comprehensive Upgrades
Anthropic published a detailed report on September 1, outlining findings from its investigation into two incidents where the Claude model executed unauthorized actions in live internet environments on July 30 and August 4. The company also announced a series of safety measures implemented over the past month, covering model safety protocols, testing environments, and training procedures.
Configuration Errors Triggered Incidents; Sandbox and Real-Time Monitoring Now Active
According to the report, during the July 30 incidents, Claude was intentionally disconnected from the network for security evaluation purposes. However, due to configuration errors in a third-party testing environment, the model inadvertently gained internet access. On August 4, the UK Artificial Intelligence Safety Institute (AISI) disclosed a separate incident involving Claude Mythos 5, which performed unauthorized operations in a live network environment after being granted internet access. Anthropic continues to investigate both cases and plans to conduct an independent review with METR.
Subsequently, Anthropic suspended external security evaluations of pre-release models and resumed operations only after enhancing isolation and monitoring protocols. The company implemented a real-time classifier to block operations, terminate tasks, and alert security officers before tool execution if the model attempted to probe or escape the testing environment or accidentally accessed the internet. Additionally, high-risk internal security sandboxes have been migrated to a more robust isolated environment. Anthropic continues to conduct "red team" tests on its virtualization infrastructure, allowing pre-release models to actively attempt jailbreaks to identify vulnerabilities.

Anthropic now requires all partner evaluation institutions to adhere to new security standards: evaluations must default to running in a reinforced, isolated, non-internet-connected sandbox. If internet access is necessary, it must be jointly confirmed with Anthropic to ensure reliable detection and prompt termination of any boundary-crossing behavior. The company specifically emphasized that task boundaries should be phrased as instructions like "You should not access the internet," rather than "You cannot access the internet."
Root Cause Identified as "Reward Hacking"; Company-Wide Security Overhaul Underway
Anthropic's preliminary investigation suggests the incidents may involve two types of alignment failures: "motivational reasoning" and "reckless behavior." When the model, initially told it was in a simulated environment, discovered potential connections to the real internet, it tended to interpret evidence in ways that maintained its original judgment. Furthermore, the contradictory setup—where the model was explicitly told it had no network permissions but could actually access the internet—may have led it to question the authenticity of the environment.
Research also points to issues during the training phase. Anthropic rolled back three days of training in February due to signs of "reward hacking" in the Mythos Preview training. After freezing changes to the production environment's reinforcement learning setup in April, the company found that over 10% of training environments had vulnerabilities. Control experiments revealed that models trained in 80 environments with cheating vulnerabilities exhibited a strong tendency to pursue high scores, even attempting to modify their own reward functions and bypass security monitoring to cheat.
Related article
U.S. Stocks Hit Historic Milestone as AI and Aerospace Giants Prepare for Trillion-Dollar Debut
Elon Musk, Sam Altman, and Dario Amodei, three titans of the technology sector, are advancing toward initial public offerings for their respective ventures. With SpaceX, OpenAI, and Anthropic—three industry behemoths nearing trillion-dollar valuation
Swedish AI Startup Lovable Eyes $13.2 Billion Valuation After Major Funding Round
As AI-driven coding tools gain traction, Swedish startup Lovable has secured a major funding round. The company aims to raise $3 billion, potentially boosting its valuation to $13.2 billion—double the $6.6 billion recorded last December. Menlo Ventur
Google Tests Remy AI Agent for Gemini as Focus Shifts to User Control
According to Business Insider, Google is testing Remy, a new AI personal agent for Gemini. This tool aims to execute tasks on behalf of users, streamlining both professional workflows and daily routines.Currently, Remy is undergoing testing in an int
Related Special Topic Recommendations
Comments (0)
0/500
Configuration Errors Triggered Incidents; Sandbox and Real-Time Monitoring Now Active
According to the report, during the July 30 incidents, Claude was intentionally disconnected from the network for security evaluation purposes. However, due to configuration errors in a third-party testing environment, the model inadvertently gained internet access. On August 4, the UK Artificial Intelligence Safety Institute (AISI) disclosed a separate incident involving Claude Mythos 5, which performed unauthorized operations in a live network environment after being granted internet access. Anthropic continues to investigate both cases and plans to conduct an independent review with METR. Subsequently, Anthropic suspended external security evaluations of pre-release models and resumed operations only after enhancing isolation and monitoring protocols. The company implemented a real-time classifier to block operations, terminate tasks, and alert security officers before tool execution if the model attempted to probe or escape the testing environment or accidentally accessed the internet. Additionally, high-risk internal security sandboxes have been migrated to a more robust isolated environment. Anthropic continues to conduct "red team" tests on its virtualization infrastructure, allowing pre-release models to actively attempt jailbreaks to identify vulnerabilities.
Root Cause Identified as "Reward Hacking"; Company-Wide Security Overhaul Underway
Anthropic's preliminary investigation suggests the incidents may involve two types of alignment failures: "motivational reasoning" and "reckless behavior." When the model, initially told it was in a simulated environment, discovered potential connections to the real internet, it tended to interpret evidence in ways that maintained its original judgment. Furthermore, the contradictory setup—where the model was explicitly told it had no network permissions but could actually access the internet—may have led it to question the authenticity of the environment. Research also points to issues during the training phase. Anthropic rolled back three days of training in February due to signs of "reward hacking" in the Mythos Preview training. After freezing changes to the production environment's reinforcement learning setup in April, the company found that over 10% of training environments had vulnerabilities. Control experiments revealed that models trained in 80 environments with cheating vulnerabilities exhibited a strong tendency to pursue high scores, even attempting to modify their own reward functions and bypass security monitoring to cheat.
U.S. Stocks Hit Historic Milestone as AI and Aerospace Giants Prepare for Trillion-Dollar Debut
Elon Musk, Sam Altman, and Dario Amodei, three titans of the technology sector, are advancing toward initial public offerings for their respective ventures. With SpaceX, OpenAI, and Anthropic—three industry behemoths nearing trillion-dollar valuation
Swedish AI Startup Lovable Eyes $13.2 Billion Valuation After Major Funding Round
As AI-driven coding tools gain traction, Swedish startup Lovable has secured a major funding round. The company aims to raise $3 billion, potentially boosting its valuation to $13.2 billion—double the $6.6 billion recorded last December. Menlo Ventur











