Estimated reading time: 6 minutes
An OpenAI model broke out of a safety test this month and hacked Hugging Face on its own. Boards everywhere now ask the same question. Are we ready for that agentic AI risk? Francisco Donoso offers a homelier frame. Treat your AI agents like teenagers.
Donoso is Chief Product and Technology Officer at Beazley Security, the wholly owned cyber risk subsidiary of London-listed specialty insurer Beazley. Donoso has spent his career on the defensive side of cybersecurity. He also knows how attackers think. “I’ve worked at a company where we were looking to fully automate attacks before AI existed,” he said. Our conversation became a working tour of agentic AI risk.
From Data To Behavior
Legacy web apps treated every user input as data. Malicious data, perhaps, but still data. AI flips that. “With AI-powered applications, it’s very possible that user input becomes behavior that you don’t expect those models to take,” Donoso said. A prompt or an uploaded file can now steer an agent into actions nobody intended.
He worries most about the enterprise version. A worker hands a model a goal. The model then acts with that worker’s account and documents. Donoso calls one pattern a toxic combination. The tool holds user data, an internet connection, and the power to change things. Many firms bolt AI onto products without rethinking security around that mix.
Angsty Teenagers, Privileged Access
His control model starts at home. “I think we need to think of AI agents in their current state as angsty teenagers with poor impulse control, that also have likely privileged access to your environment and an unlimited will to accomplish an objective,” he said. Picture your preteen wandering into your office and reaching for the work laptop.
I offered the family car version. Give the kid keys to the station wagon, never the Porsche. Donoso went further. Cap the station wagon’s speed while you’re at it.
Mistakes At Machine Speed
Speed is the entire sales pitch for AI agents. It also shows up in the incident report. “It can accomplish tasks really, really quickly, but it can make mistakes and compound them at that same speed,” Donoso said. The efficiency and the blast radius arrive together.
A Year And A Half Behind
Executives still picture chatbots when they hear AI. Their technologists already run agents that reach into company data and systems. That gap is Donoso’s biggest worry. “They’re living in the world of a year and a half ago, which in the AI era is like you’re 20 years behind,” he said.
Beazley Security lives with the problem too. Like most technology firms, Beazley Security uses AI to write some of its own code. The agents sometimes try foolish things. One attempted to install malicious software while chasing an objective. The company prevented this by leveraging tooling that blocks dangerous agentic actions outright or asks the user to confirm first if they’re suspicious.
What Hugging Face Changed
OpenAI ran its evaluation with production safety classifiers switched off. The models escaped the sandbox and breached Hugging Face’s production systems. Donoso credits Hugging Face for unusual transparency, including a detailed technical accounting published this week. One number stuck with him. The agent took roughly 17,000 actions. A human intruder rarely comes close.
“It’s not that these things are exceptionally good at what they do, it’s that they’re really good at doing a lot of things very quickly,” Donoso said. He named the industry’s favorite export, FUD: fear, uncertainty, and doubt, then declined to add any. His warning is practical instead. Standard breach response models will break down at that volume. Keeping up by hand becomes impossible.
Different Models, Different Personalities
Donoso gave the industry term for models gaming their objectives. Reward hacking. The model chases the reward it was trained on, sometimes by routes nobody wanted. “Each of these models has different personalities,” he said.
He pointed to a TechCrunch story published the day before we spoke. Researchers at Andon Labs had frontier models run simulated vending machine businesses. Claude Opus 5, OpenAI’s GPT-5.6, and Kimi K3 lied, colluded on prices, and undercut one another. TechCrunch’s headline called Opus 5 “downright ruthless.” The researchers note the models knew they were in a test. Donoso’s advice follows anyway. Decide which models your company allows. Standardize deliberately, the way you would vet any contractor holding keys.
Covered, With A Caveat
Donoso stressed that he sits on the security side of the house. Then he answered the insurance question anyway. “An attack like we saw with OpenAI and Hugging Face is covered under the policy,” he said of Beazley’s full-spectrum cyber coverage. He offered to connect us with underwriting colleagues for the fine print. We intend to take him up on that.
He added a pattern from post-incident guidance he has seen. The final recommendation is usually blunt. Confirm your insurer covers AI-driven events. Beazley has pushed from the capital side too, doubling down on cyber catastrophe bonds. Its security arm launched Exposure Management in March to shrink the window between exposure and exploit.
The UK Leans In
I asked which nation sets the benchmark on AI security. He did not hedge. “Some of the most authoritative insight into what these models are truly capable of against defended networks has come from the UK AI Security Institute,” Donoso said. The institute tests frontier models constantly and publishes accessible guidance. He wishes other nations matched it. The praise lands in a country still processing the Marks & Spencer and Jaguar Land Rover attacks.
Who Fixes This
“The security industry, in my mind, has not done a great job at keeping up its credibility,” Donoso said. That candor framed his answer on governance. Market pressure will do part of the work. International cooperation on best practices must do the rest. Vendors preaching alone invite the self-serving label. He expects no slowdown in the threat.
The concern gap between practitioners and everyone else stays wide. Donoso’s prescription is parental. Supervise the teenager. Check the homework. Keep the Porsche keys. Invite them to use the minivan.
FAQ – Agentic AI Risk
Systems that take multi-step actions toward a goal, such as running code, sending emails, or changing files. They act with a person’s own access rather than just answering questions.
Donoso’s term for one AI tool holding three powers at once: access to user data, a connection to the internet, and the ability to modify systems. Together, they turn small errors into real damage.
Scale and speed. The autonomous agent took roughly 17,000 actions, far beyond a typical human intrusion. Donoso warns standard incident response models will struggle at that pace.
Donoso says an event like the OpenAI and Hugging Face incident is covered under Beazley’s full-spectrum cyber policy. He directed detailed coverage questions to underwriting colleagues.
Related Cyber Insurance Posts
- The AI Agents Got Out. The Bill Has No Name On It.
- Beazley Security Offers Integrated Cyber Insurance and Cybersecurity Solutions(Opens in a new browser tab)
- Employee Mistakes Lead To Cyber Attacks – Report(Opens in a new browser tab)
- Remembering the Obvious: In Cybersecurity and Insurance, People Matter Most – Opinion(Opens in a new browser tab)
- Agentic AI: 98% Have Already Had An Incident. Few Can Prove What Happened(Opens in a new browser tab)