AI Guardrails Hinder Legitimate Cybersecurity Research, Experts Say
AI News

AI Guardrails Hinder Legitimate Cybersecurity Research, Experts Say

5 min
7/24/2026
aisecuritycybersecurityAI security

TL;DR

Strict safety guardrails on frontier AI models from OpenAI and Anthropic are creating significant barriers for legitimate offensive cybersecurity researchers. While designed to prevent malicious use, these restrictions are blocking critical tasks like vulnerability analysis and exploit development. The U.S. government's recent export controls on Anthropic's models, combined with inconsistent enforcement of safety policies, are forcing researchers to find workarounds or abandon the tools entirely. Experts argue this approach weakens rather than strengthens overall security by hindering those who need to think like attackers to defend systems.

The Guardrail Paradox

For months, AI giants like OpenAI and Anthropic have implemented strict guardrails and vetted programs to prevent malicious actors from weaponizing their models. However, these same barriers are now impeding the work of legitimate network defenders and offensive security researchers who need to think like hackers to find and patch vulnerabilities. The situation has become so acute that in June 2026, the U.S. government imposed export control restrictions on Anthropic's much-hyped Mythos and Fable models, citing concerns about their potential use in cyberattacks.

The core problem lies in the nature of cybersecurity research itself. As Chris Anley, Chief Scientist at NCC Group, explains, a command to "fix" software code can serve as a guide for both defense and attack. "The same tool acts as both an offensive and defensive weapon, and they cannot be separated. If an AI model refuses to analyze code, it makes the work of defenders more difficult," Anley told TechCrunch. This fundamental tension between safety and utility is at the heart of the current debate.

Inconsistent Enforcement and Vetted Programs

Chris Thompson, CEO of RemoteThreat and founder of Offensive AI Con, reports that even within the looser boundaries of Anthropic and OpenAI's vetted programs, guardrails can be inconsistent and work differently every day. This unpredictability makes it difficult for researchers to rely on these tools for their work. One researcher at a smartphone-component manufacturer, who spoke on condition of anonymity, said his employer isn't part of Anthropic's CVP program and as a result, its tools are barely useful for finding vulnerabilities because the guardrails are too strict.

"If it catches wind we're doing anything security related, it just stops and isn't usable," the person said. This subjective enforcement is particularly problematic for organizations that work with thousands of different code repositories spread across the internet, as noted by security researcher Kinsbruner. The current approach, he argues, cannot scale to become an enterprise-grade cybersecurity solution.

continue reading below...

Language Gaps in AI Security

Another dimension of the problem is the uneven protection offered by AI guardrails across different languages. As DeepKeep highlighted in a recent blog post, the AI security layer on top of many products doesn't protect equally against jailbreaking and unsafe actions in every language. Security intent can shift during translation, with jailbreaks becoming less explicit or prompt injections losing their command structure in non-English languages.

This multilingual reality poses a particular risk for European organizations. While a model may speak and understand German or Spanish well, it doesn't mean security guardrails will perform equally well. Mixed-language inputs can preserve malicious instructions in one language while surrounding them with benign context in another, creating blind spots that attackers could exploit.

Export Controls and Government Response

The U.S. government's export control restrictions on Anthropic's models were prompted at least in part by a report claiming it was possible to bypass the models' guardrails. While these restrictions were later partially eased, cybersecurity professionals argue this approach is weakening defense systems. The Trump administration may be playing catch-up on threats that have been building for years, as it has more fully realized the national security implications of the technology.

Andrew Lohn and Jessica Ji of Georgetown University's Center for Security and Emerging Technology argue that blocking AI models won't stop the cyber threats they create. The real problem is that defensive efforts haven't kept up with the pace of AI progress. The federal government cut resources to key agencies like CISA and redistributed their authorities, creating a gap that AI companies have filled by taking on responsibilities that should be government-led.

The Researcher's Perspective

Not all researchers are equally affected. Giuseppe Cali, a security researcher who finds zero-days and develops exploits, said guardrails are not impeding his work because he doesn't use AI for offensive work. Instead, he uses it for initial reverse engineering, to understand code, and to build supporting tools. "I still want to own the actual bug discovery and weaponization myself and that wouldn't change if all guardrails were lifted tomorrow," Cali said.

However, prominent security researcher Mark Dowd believes that private companies independently deciding what is safe and what is dangerous negatively impacts the development of the field. He notes that major tech giants are imposing arbitrary restrictions without considering the complex aspects of cybersecurity. The debate highlights the need for a more nuanced approach that balances safety with the legitimate needs of security researchers.

The Path Forward

As 2026 turns out to be the year when predictions about AI-powered cyberattacks seem to be coming true, the real question isn't whether cybersecurity still matters, but rather how the risks that AI introduces will be managed before they outpace defenses. The transitional period we're in requires a fundamental rethinking of how AI safety and cybersecurity intersect.

AI companies like Anthropic and OpenAI have launched initiatives like Project Glasswing and Patch the Planet to shore up critical infrastructure and open-source software libraries. But these efforts, while valuable, are limited by corporate incentives that prioritize limiting liability over comprehensive security. The solution likely requires a collaborative approach between government, industry, and the security research community to develop guardrails that protect against actual threats without crippling the very people tasked with defending our digital infrastructure.