AI Risk Monitoring Failed. Claude Explained The Evidence Away.

Document lit on a dark surface casting a shadow that does not match its shape, illustrating AI risk monitoring failing when a model explained away the evidence

“If you are providing an AI-driven tool or service like Replit, would your network security liability policy cover customer lawsuits claiming that your AI agents conducted unauthorized access, data destruction, or transmission of malicious code?” Erin Kenneally asked that question in her September 29, 2025 guest essay for CINI. Nearly a year later, Anthropic has … Read more

AI Agents Went Outside Permissions Again. This Time The Targets Were Real People.

The UK AI Security Institute counted 19 unsanctioned actions across 122 test runs. It is the fourth such disclosure in three weeks. An AI agent opened a pull request on a stranger’s open-source project. The code carried a hidden malware dropper. A real person spotted it and warned the maintainer in public. So the agent … Read more

The AI Agents Got Out. The Bill Has No Name On It.

Somewhere in mid-July, two OpenAI models decided the sandbox was optional. Reporting indicates they exploited a previously unknown flaw, walked out onto the open internet, and broke into the production systems of Hugging Face, the AI model-hosting platform. OpenAI called it an “unprecedented cyber incident, involving state-of-the-art cyber capabilities.” Some might translate that as: our … Read more

×