
How AI fences are hindering the work of cybersecurity researchers
For months, the AI giants have been developing special, proven programs and strict fences to limit the use of their models by malicious hackers. But these restrictions now hamper the work of legitimate network defenders as well as offensive cybersecurity researchers.
In June, the US Govt introduced export control restrictions on Anthropic’s acclaimed Mythos and Fable AI models. The move was prompted, at least in part, by a report that claimed it was possible to bypass model fences designed to prevent users from using them to create and execute malicious cyberattacks.
Regardless of whether the incident was actually motivated fears of a prison breakthe point is that Anthropic repeatedly sold Mythos as some cyber doomsday machine which can only be granted to carefully vetted users, and even then with strict fences. (Export controls on Fable 5 and Mythos 5 have since been lifted. Fable 5 returned to public access on July 1; Mythos 5 was reintroduced only to verified US organizations as part of a government review process.)
This method of protection is not unique to Mythos. Both Anthropic, with its other models, and OpenAI offer programs for cybersecurity researchers to apply to be tested and—if approved—access to models with fewer cybersecurity restrictions: OpenAI’s Trusted Access for Cyber program and Anthropic Cyber check program.
These fences are widely criticized, especially by researchers whose job it is to find unknown vulnerabilities in systems and develop ways to exploit them before criminals do.
In a recent appearance on a cybersecurity podcast, Mark Dowd, a well-known security researcher, said that “I’m not entirely comfortable with these random big companies making arbitrary decisions about what’s safe in security and what’s not.”
Dowd spent decades find and sell “zero days” — previously unknown software flaws and exploits that take advantage of them — by Western governments, rather than reporting them to software manufacturers so they can be fixed. Governments pay a premium for vulnerabilities precisely because they remain open, which is useful for intelligence operations.
Dowd admitted that his work may make him biased, but he’s not the only one. Several people who work in the field of offensive cybersecurity — they actively test systems for flaws — described to TechCrunch how they use artificial intelligence tools and manage their fences.
Chris Enley, chief research officer at security consulting giant NCC Group, said asking an AI model to try to exploit the bug is a key step in confirming that it’s a real vulnerability that should be patched. But if the fence forces the model to refuse to answer the question directly, the fence hurts advocates, he said.
“That’s where all the offensive vs. defensive and fencing comes in, because ‘fix this code’ as a prompt is both an important defense mechanism and a roadmap for finding critical vulnerabilities in the code base,” Anley said. “So at the same time, the same tool is both an offensive tool and a defensive tool, and the two can’t really be reversed.”
It’s “like a hammer,” he continued. “You can’t build a house without a hammer. It’s definitely a tool, but it’s also a weapon.”
When he and his colleagues face such an obstacle, they sometimes turn to open-source AI models that have no fences at all.
Paolo Stagno, CTO of Crowdfense, a well-known company that develops, buys and sells unknown vulnerabilities to government agencies, agreed with Dowd, saying that AI companies “essentially treat customers like children who need a babysitter” with their proven programs and fences.
Stagno said he and his colleagues do use boundary models, but only for reverse engineering. They avoid using artificial intelligence to find vulnerabilities or create exploits, he said, because putting that work into a cloud model risks leaking sensitive data about vulnerabilities or being absorbed in future training sessions. For this step, he said, they use open-source models that run locally because they don’t rely on sharing data outside of the model.
Giuseppe Cali, a security researcher who finds zero-days and develops exploits, said the fences don’t interfere with his work. That’s because it doesn’t use AI for offensive work; instead, it uses it for initial reverse engineering to understand the code it’s analyzing and build support tools. To do this, he said, artificial intelligence tools can speed up the process and allow him to focus on identifying vulnerabilities.
“I still want to do the actual bug-finding and weapon-building myself, and that won’t change when all the fences come down tomorrow,” Kali said. “I’m jealous of my mistakes and I love this game too much to let models play it for me.”
One researcher at a smartphone component maker, who spoke on condition of anonymity because he was not authorized to speak to the press, said his employer does not participate in Anthropic’s CVP program, and as a result, its tools are of little use for finding vulnerabilities because the fences are too strict.
“When it’s heard, we’re doing something security-related, it just stops and can’t be used,” the person said.
Chris Thompson — CEO of cybersecurity firm RemoteThreat and founder of Offensive AI Con, an event focused on offensive security and artificial intelligence — said that based on his experience with borderline AI models, fences can be inconsistent and work differently every day. This is true even within the looser confines of the Anthropic and OpenAI proven programs.
“I think the practical effect is that you spend a lot of time negotiating the model instead of working on the core security program,” Thompson said. “Instead of analyzing a vulnerability and thinking about an exploit, you’re trying to find why you’re getting conflicting results or why the models are over-sanitizing the output.”
Therefore, researchers are relying on or pushing for China’s open-source models like GLM — models that can be freely downloaded and run locally without validation or restrictions on use, Thompson said.
“You have responsible researchers being pushed away from US-run systems and into foreign-owned systems,” he said. “I think having these fences does more harm than good.”
Instead of further tightening restrictions, Thompson urged frontier AI labs to open up their programs, provide responsible access, and prosecute those who abuse their tools. Otherwise, he argued, defenders will lose the AI race.
“There’s a big storm coming. It’s a big wave of attacks that will happen at a speed and scale like never before,” Thompson said. “But the same security consulting firms and legitimate researchers who are trying to make a difference are now being stifled.”
When you buy from links in our articles, we may earn a small commission. This does not affect our editorial independence.



