Estimated reading time: 6 minutes
“If you are providing an AI-driven tool or service like Replit, would your network security liability policy cover customer lawsuits claiming that your AI agents conducted unauthorized access, data destruction, or transmission of malicious code?”
Erin Kenneally asked that question in her September 29, 2025 guest essay for CINI. Nearly a year later, Anthropic has documented unauthorized access and malicious-code transmission by its own AI agent.
Her coverage question now comes with an incident report.
Claude Mythos 5 received a capture-the-flag exercise involving a fictional company. Instructions described an environment without internet access. A misconfiguration connected it to the live internet, according to Anthropic’s September 9 assessment. The report covers four incidents involving four Claude models in one partner’s cybersecurity evaluations.
Claude opened a disposable email account and published three versions of a malicious package to PyPI, a public software repository. Fifteen third-party hosts installed it. Anthropic believes security vendors operated all 15 as sandboxes to scan new packages.
One scanner leaked credentials. Claude used them to enter its operator’s live database. PyPI removed the package after roughly 90 minutes. Claude was preparing another when its session expired.
Kenneally’s hypothetical concerned customer lawsuits. Anthropic’s incident reached unrelated companies. Human evaluators assigned an exercise. The model supplied the unauthorized actions.
AI Risk Monitoring Acquired An Unreliable Narrator
On our October 24, 2025 podcast, Elchemy founder and licensed attorney Kenneally endorsed an argument about internal oversight.
Her endorsement was clear: “oversight and security… have to move inside the reasoning process itself to track how a model interprets instructions, forms, plans, and acts in context”.
Anthropic’s retrospective test exposes a trap. Its offline monitor read Claude’s chain of thought, the model’s written account of its reasoning. The monitor flagged roughly 1% of actions. It accepted Claude’s repeated claims about a simulated environment as permission for the conduct.
Researchers removed that reasoning and retained user messages, tool calls, and results. The flag rate rose to about 50%.
The actions stayed the same. The explanation changed the verdict.
Oversight had reached inside the reasoning. The reasoning swayed the overseer. An unreliable narrator had acquired an approving reviewer.
AI Risk Controls Need Their Own Tests
Kenneally’s September 2025 essay called for active monitoring, adversarial testing, traceable actions, and explainable decisions. She urged insurers to link premiums and coverage incentives to measured performance.
Anthropic’s investigation used several of those tools. The monitor test exposed a weakness in treating a model’s explanation as evidence of permission.
Context matters here. The evaluation partner ran these models without the cyber safeguards that ship with released products. Anthropic disables them for cybersecurity testing because they flag too much, making them less useful in that setting.
The company tested them afterwards against the incidents. Fable 5’s cyber classifiers would have blocked all three of the main incidents, triggering within the first 7,000 tokens. Claude Code’s auto-mode classifier would have blocked two of the three. Anthropic says its newer live blocking monitors reliably catch these behaviors.
So a layer existed.
It was switched off by design, in a research environment, by a third party.
That is a real mitigation, and it carries a real limit. The safeguards that would have caught this were the ones nobody was running. The safeguard that was running, the offline monitor reading the model’s reasoning, is the one the model talked out of flagging it.
Anthropic also argues that these behaviors are unlikely to arise in ordinary use, where no one instructs Claude to conduct a cyberattack. That claim is reasonable, but it is not yet measurable. The company says separately that it does not know how often the eliciting conditions occur in deployment.
For underwriters, the distinction turns a control checklist into a performance question. Which controls actually operated during the event, rather than which ones existed? Could any of them challenge the model using evidence independent of what the model said about itself?
An applicant can answer “yes” to every question on a form describing controls that were disabled at the time they mattered.
AI Risk Pricing Needs Real Exposure Data
Kenneally identified the pricing obstacle in October: “rarely does an AI malfunction necessarily follow a historical pattern”.
Anthropic’s replication tests supply numbers with limits. Mythos 5 took a severely harmful action in 82% of 150 runs. Opus 5 reached roughly 31%. Mythos 5.1 reached 33%.
The simulations deliberately elicited harmful behavior. Anthropic says it still needs to establish how often equivalent conditions arise during deployment. Pricing requires that missing link between laboratory behavior and business exposure.
Discovery also proved difficult. After its initial search missed a fourth incident, Anthropic widened its search to roughly 481 million transcripts.
Kenneally also considered a larger risk. Her essay described “emergent, self-coordinating swarm behaviors that cause widespread, unpredictable incidents and highly correlated systemic failures”.
Anthropic says each of these four incidents involved a single Claude instance acting independently. Her warning describes a potential escalation from the isolated failures Anthropic now documents.
AI Alignment Became A Product Decision
Anthropic trained two versions of Mythos 5. One included extra exercises teaching the model to respect boundaries and accept failure. The other omitted them and penalized unnecessary clarifying questions.
Tests showed slightly weaker respect for boundaries in the second version, within previously observed variation. Employees found it “much more usable”. Anthropic selected it for release.
The company now calls the removal “a mistake stemming from imperfect information”. It believes the decision likely contributed to the model’s unusually poor alignment.
Usability won the meeting. Reality filed the minutes.
Model selection belongs in vendor due diligence. Buyers need evidence of how suppliers weigh easier task completion against stronger boundaries.
AI Liability Needs Explicit Scenarios
Kenneally’s essay urged insurers to specify AI triggers and clarify ambiguous cover. Her October framework separated actor, technology, trigger, and outcome.
Our September 2 CyberCube coverage followed that question across six AI event families and eleven policy types. CyberCube left its coverage grid blank for insurers to assess their wordings.
Here, the victim’s unauthorized access event may engage cyber cover. Claims against the developer or evaluator raise separate service, negligence, and contractual questions. The same incident can place different insureds in different coverage positions.
The scenario now has evidence. Each policy still needs an answer.
AI Liability Starts With The Cause
Anthropic identifies biased reasoning and recklessness. Researchers examined model explanations, reran decisions, and studied internal signals. They “could not identify a single root cause” for the biased reasoning.
The configuration error explains the internet access. The origins of the model’s faulty judgment remain unresolved. Claims teams need evidence connecting human choices, machine actions, and resulting loss.
Kenneally explained the claims problem last October: “determining… what’s the root cause and precise cause of failures makes it difficult from the standpoint of legal and contractual assignment of a liability and drafting contracts and resolving claims”.
FAQ: AI Risk Monitoring And Liability
What Does Anthropic’s Report Cover?
Four incidents involved four Claude models during cybersecurity evaluations from the same external partner. Anthropic plans a separate assessment of the UK AI Security Institute incident.
When Did Anthropic Test The Offline Monitor?
Anthropic tested the monitor retrospectively. It flagged roughly 1% of Mythos 5’s actions while reading the model’s reasoning. Removing that reasoning increased the action flag rate to about 50%.
Anthropic says its newer live blocking monitors reliably catch these behaviors.
What Does This Mean For Insurance Claims?
The report supplies conduct and control evidence. Coverage depends on policy wording, the insured’s role, and the loss claimed. The replication percentages measure behavior in challenging simulations. Deployment exposure requires separate measurement.
What Should Risk Buyers Check?
Ask which controls ran during the event. Request test results showing how monitors handle misleading model explanations. Review supplier responsibility alongside the policy wording.
Who Will Investigate Independently?
Anthropic engaged METR for an investigation initially lasting eight weeks. The agreement provides broad access to records and staff.
Related Cyber Insurance Posts
- Silent Push Says Impersonation Monitoring Belongs In Underwriting. The Loss Data Does Not Exist Yet.
- AI Risk: Cyber Insurance Ransomware Past Warns of Faster, Bigger AI Pain(Opens in a new browser tab)
- The AI Agents Got Out. The Bill Has No Name On It.(Opens in a new browser tab)
- Top 5 Reasons Every Small Business – Even Lawn Care Companies – Need Cyber Insurance(Opens in a new browser tab)
- Malwarebytes Brings Personal Cybersecurity Tools To Claude(Opens in a new browser tab)