option
Home
News
AI Guardrails Stifle Offensive Cybersecurity Researchers

AI Guardrails Stifle Offensive Cybersecurity Researchers

September 6, 2026
36

AI Guardrails Stifle Offensive Cybersecurity Researchers

For months, AI leaders have implemented strict vetting protocols and safety guardrails to prevent malicious actors from exploiting their models. However, these restrictions are now obstructing legitimate network defenders and offensive cybersecurity researchers.

In June, the U.S. government imposed export controls on Anthropic’s highly anticipated AI models, Mythos and Fable. This decision was partly driven by reports suggesting that the models’ safeguards against facilitating cyberattacks could be bypassed.

Whether or not the incident was truly triggered by jailbreak concerns, Anthropic has consistently positioned Mythos as a powerful cyber tool accessible only to rigorously vetted users under strict constraints. (Note: Export controls on Fable 5 and Mythos 5 have since been lifted. Fable 5 returned to general availability on July 1, while Mythos 5 has been reintroduced exclusively to vetted U.S. organizations as part of the government’s review process.)

This gatekeeping approach is not unique to Mythos. Both Anthropic and OpenAI offer cybersecurity researchers programs to apply for vetted access, which, if approved, grants them models with fewer cybersecurity restrictions: OpenAI’s Trusted Access for Cyber and Anthropic’s Cyber Verification Program.

These guardrails have faced widespread criticism, particularly from researchers tasked with identifying unknown system vulnerabilities and developing exploits before criminals can.

During a recent cybersecurity podcast appearance, renowned security researcher Mark Dowd stated, “It’s not really comfortable to me that these random large companies are making arbitrary decisions about what is safe in security and what’s not.”

Dowd has spent decades discovering and selling “zero days” — previously unknown software flaws and their exploits — to Western governments rather than reporting them to software vendors for patching. Governments pay a premium for these vulnerabilities because they remain undisclosed, which is valuable for intelligence operations.

Dowd acknowledged his work might introduce bias, but he is not alone. Several professionals in offensive cybersecurity, who proactively probe systems for weaknesses, described to TechCrunch how they utilize AI tools and navigate their guardrails.

Chris Anley, chief scientist at security consulting giant NCC Group, explained that asking an AI model to attempt exploiting a bug is a crucial step in confirming it is a genuine vulnerability worth fixing. However, if a guardrail causes the model to refuse the request entirely, it hinders defenders, he noted.

“This is where the whole offensive versus defensive and guardrails part comes in, because ‘fix this code’ as a prompt is both an essential mechanism for defense but also a roadmap for finding critical vulnerabilities in the code base,” said Anley. “So at the same time, the same tool is both an offensive tool and a defensive tool, and the two can’t really be unpicked.”

It’s “like a hammer,” he continued. “You can’t build a house without a hammer. It’s definitely a tool but it’s also irreducibly a weapon as well.”

When encountering such roadblocks, he and his colleagues sometimes resort to open-source AI models that lack any guardrails.

Paolo Stagno, chief technology officer at CrowdFense, a prominent company that develops, acquires, and sells unknown vulnerabilities to government agencies, agreed with Dowd, stating that AI companies “essentially treat customers like children who need babysitting” through their vetted programs and guardrails.

Stagno mentioned that he and his team use frontier models for reverse engineering but avoid using AI to find vulnerabilities or build exploits. Feeding such work into cloud-based models risks leaking sensitive vulnerability data or having it absorbed into future training runs. For these tasks, he uses locally run open-source models to avoid sharing data externally.

Giuseppe Cali, a security researcher who discovers zero-days and develops exploits, stated that guardrails are not hindering his work. This is because he does not use AI for offensive purposes; instead, he uses it for initial reverse engineering to understand the code under analysis and to build supporting tools. He noted that AI tools accelerate this process, allowing him to focus on discovering vulnerabilities.

“I still want to own the actual bug discovery and weaponization myself and that wouldn’t change if all guardrails were lifted tomorrow,” said Cali. “I am jealous of my bugs, and I like this game too much to let models play it for me.”

One researcher at a smartphone-component manufacturer, who spoke on the condition of anonymity as he is not authorized to talk to the press, stated that his employer is not part of Anthropic’s CVP program. Consequently, its tools are barely useful for finding vulnerabilities due to overly strict guardrails.

“If it catches wind we’re doing anything security related, it just stops and isn’t usable,” the person said.

Chris Thompson, chief executive of cybersecurity firm RemoteThreat and founder of Offensive AI Con, an event focused on offensive security and AI, noted that in his experience with frontier AI models, guardrails are inconsistent and behave differently each day. This holds true even within the more lenient boundaries of Anthropic and OpenAI’s vetted programs.

“I think the practical impact is you spend a lot of time negotiating with the model instead of working on the core security program,” said Thompson. “Instead of analyzing a vulnerability and reasoning through the exploitability, you’re trying to find why you’re getting inconsistent results or why are models over-sanitizing the output.”

As a result, researchers are increasingly relying on or being pushed toward Chinese open-source models like GLM — freely downloadable models that can be run locally without vetting or usage restrictions, according to Thompson.

“You have these responsible researchers that are being pushed away from U.S.-governed systems to foreign-owned systems,” he said. “I think it’s more harmful than good to have these guardrails in place.”

Rather than tightening restrictions further, Thompson urged AI frontier labs to open their programs, provide responsible access, and hold abusers of their tools accountable. Otherwise, he argued, defenders will lose the AI race.

“There’s this big storm coming. There’s this big wave of attacks that are going to happen at speed and scale like never before,” said Thompson. “But the same security consulting firms and legit researchers that are trying to make a difference are being stifled right now.”

Related article
OpenAI Partners with Yubico to Bolster GPT-5.6 Security via Hardware Passkeys OpenAI Partners with Yubico to Bolster GPT-5.6 Security via Hardware Passkeys Jerrod Chong, CEO of Yubico | Image Credit: YubicoOpenAI’s upcoming GPT-5.6 will mandate hardware-backed passkeys starting in September, a move that Yubico CEO Jerrod Chong says confirms their product as the premier defense against account compromise
OpenAI Leads 100-Signatory Push for Cyber Defence OpenAI Leads 100-Signatory Push for Cyber Defence OpenAI’s letter begins stating "we have a limited window to strengthen cyber defences". Credit: Getty ImagesMajor tech firms – including AWS, Google and Microsoft – urge governments to fund defensive tools, while critics call the manifesto self-servi
Anthropic Unleashes AI Agents on Shared Task, Sparking Internal Rivalry Anthropic Unleashes AI Agents on Shared Task, Sparking Internal Rivalry What occurs when AI agents are pitted against one another? Anthropic’s recent tests reveal that the results can quickly become chaotic.On Thursday, Anthropic’s Frontier Red Team released new research analyzing how groups of AI agents interact when th
Related Special Topic Recommendations
writing Best AI Outline Generators for Long-Form SEO Articles
Best AI Outline Generators for Long-Form SEO Articles

2026 Latest Best Top-Rated AI Outline Generators for Long-Form SEO Articles, meticulously curated by XIX.AI. These powerful tools offer game-changing assistance in creating high-quality content quickly, boosting writing efficiency significantly. Get a free vs paid comparison along with real-world tests and detailed rankings to help you find the must-try option that suits your needs. Explore now to unlock your AI edge.

8 tools
xix.ai
Education and Learning AI Study Tools for Homework and Exam Prep
AI Study Tools for Homework and Exam Prep

2026 Latest Best AI Study Tools for Homework and Exam Prep! XIX.AI curates a top-rated list of powerful, game-changing tools that help students boost productivity, streamline homework completion, and ace exams through real-world tests. Get a free vs paid comparison, detailed rankings, and must-try options to unlock your AI edge. Explore now!

10 tools
xix.ai
Music composition AI Vocal Demo Tools for Songwriters, Hooks, Toplines, and Multilingual Draft Sessions
AI Vocal Demo Tools for Songwriters, Hooks, Toplines, and Multilingual Draft Sessions

2026 Latest Best AI Vocal Demo Tools for Songwriters, Hook Creators, and Multi-Language Content Teams! XIX.AI has curated a top-rated list of powerful game-changing tools that go through rigorous real-world tests. You’ll find detailed free vs paid comparison data, comprehensive rankings, and must-try options to help you boost writing efficiency and unlock your creative potential. Explore now to discover your perfect tool for all your content needs!

9 tools
xix.ai
Business Best AI Competitive Research Tools for Small Businesses
Best AI Competitive Research Tools for Small Businesses

2026 Latest Best Top-rated AI Competitive Research Tools for Small Businesses! XIX.AI has curated a highly powerful game-changing collection, updated weekly with rigorous real-world tests and detailed rankings. You can find a comprehensive free vs paid comparison to help you identify the must-try tools that boost your productivity and give you a competitive edge. Explore now to discover your perfect tool!

9 tools
xix.ai
Image editing Photoshop AI Retouch Tools for Ecommerce Apparel, Skin Cleanup, and Color Consistency
Photoshop AI Retouch Tools for Ecommerce Apparel, Skin Cleanup, and Color Consistency

2026 Latest Best Photoshop AI retouch tools for ecommerce apparel, skin cleanup, and color consistency! This top-rated curated list features powerful game-changing solutions that help you boost writing efficiency, streamline content creation, and achieve perfect visual results effortlessly. Each tool has undergone real-world tests through weekly updated rankings, complete with free vs paid comparison details. Backed by XIX.AI, it’s the must-try guide for anyone aiming to unlock your AI edge. Explore now!

10 tools
xix.ai
Prompt Best AI Prompt Libraries for ChatGPT Workflows
Best AI Prompt Libraries for ChatGPT Workflows

2026 Latest Best Top-Rated AI Prompt Libraries for optimizing all types of ChatGPT workflows. XIX.AI has curated a powerful, game-changing collection that goes through rigorous real-world tests to ensure top performance. You can find detailed free vs paid comparisons and expert rankings to help you choose the must-try tools that boost your productivity and unlock your AI edge. Explore now!

11 tools
xix.ai
Comments (0)
0/500
OR