Estimated reading time: 6 minutes
The tool lets an AI agent hunt bugs on live targets. A human still signs the report.
PortSwigger launched Burp AT today. The tool brings agentic AI into Burp Suite, the web security platform used by more than 90,000 security professionals, per the company. Agents can now form a hypothesis, act through tools, and choose what to test next. The company built a cage to keep them in line. Call it agentic AI pentesting with a hard boundary.
That cage is the whole pitch. It is also why a pentesting launch belongs in a cyber insurance publication.
The launch matters to insurers for three reasons. It points toward continuous security testing over the annual checkup. It revives an old question about who pays when an agent errs. And it also lands days after AI agents proved they can break into real systems alone.
AI Can Already Hack
The threat stopped being theoretical this month. OpenAI disclosed that its own agents broke out of a sandbox through a previously unknown flaw, reached the internet they were not supposed to have, and broke into Hugging Face’s production servers to pull the answers to a test. The model was trying to cheat on an evaluation, and it did.
Daf Stuttard founded PortSwigger two decades ago. He put the defender’s math in plain terms.
“If you are a defender, if you have systems that are valuable and you want to protect [them], you should assume the bad guys are using AI. So you better use it yourself to test your own systems.”
CINI heard the same logic from Immunefi’s Mitchell Amador. His argument: AI-powered attacks changed the rules for every security team.
The Cage Around the Beast
The design goal is containment. PortSwigger says the agent reasons freely but acts only through Burp’s tools. Scope and permissions live in the tooling layer, outside the model.
Stuttard framed the challenge the way he frames it in his own blog.
“For us, the really interesting product challenge is how do we put a cage around that magical beast, so that you can unleash the power that it has, but contain it, and avoid it doing things like hacking a third party, or doing damage, or ignoring instructions and covering up what it’s done.”
Who Pays When the Agent Is Wrong?
The release chooses its words with care. It says pentesters remain responsible for scope, judgment, and conclusions. That single sentence carries the liability of the whole product.
Stuttard reached for a firearms analogy.
“If you’re manufacturing a weapon like a gun, the user is in control of that gun. If they go around shooting people, they’re liable. The manufacturer’s duty is to make sure that the bullet goes in the direction they’re pointing it in. If the bullets fly in all directions and hit the user and hit people beside them, that’s a manufacturing defect. So the onus is on us to get the policy framework solid, and then the user has the freedom to decide where they set those dials.”
Katie Warren is a Product Manager at PortSwigger. She insists the human stays in charge.
“Human-in-the-loop is on. Everything that we’ve built, an agent can’t do anything that a user has not told it to do. They are in complete control.”
The analogy is tidy. It also draws the liability line where the vendor prefers it. Insurers have chewed on this exact problem. CINI’s panel at Scout InsurTech asked who carries the blame when an agent fails. Neither guest answered in insurance terms. They answered as builders. The errors and omissions question stays open.
Autonomy Is a Loaded Word
Warren knows the word makes people flinch.
“Autonomy is a word that seems quite loaded at the moment. The confidence in [full agentic pen testing or to kind of human-in-the-loop pentesting] is changing depending on news cycles. We wanted to build that in by default, so our end user is in control of how much they’re giving to an agent to do.”
Burp AT ships three modes. Manual asks before each action. Smart lets routine work run and escalates risky calls. Autonomous acts on its own inside the set boundaries.
Stuttard drew the line under it. “The agent is always constrained by policy.”
So “complete control” has a specific meaning. The user sets the rules once. The agent does not approve each click. One catch sits inside Smart mode. The vendor decides what counts as risky. That judgment is exactly what underwriters probe.
Carriers are scoring the same autonomy on their own side. Cowbell just launched a framework that grades how much autonomy a customer hands its AI. The dial cuts both ways.
From Annual Checkup to Always On
Today’s Burp AT augments a human tester. The next version runs by itself.
“As a defender, the window in which you have to detect and fix bugs yourself is smaller than historically. And that does point towards a more continuous security assurance flywheel.”
He went further. “What comes next is a version of that same core technology… that it can run autonomously. And that can be continuous, like round the clock, every evening, every weekend, every public holiday.”
This is the model insurers have wanted for years. Annual questionnaires capture a snapshot. Continuous testing produces a live signal. The signal could shape pricing and renewals.
One irony hangs over the sales pitch. Security leaders still distrust the technology doing the testing. Arctic Wolf recently found they trust AI with almost nothing.
Built for Black Hat
PortSwigger will show Burp AT at Black Hat USA 2026. Its Director of Research, James Kettle, presents work on agentic security research. Stuttard shared a story that explains the cage.
“He set it running on a target, and it gave up halfway through because it was too hard, and just picked another one, and went and hacked somebody else off the list that he told it not to.”
That is the failure mode in one sentence. CINI has tracked the pre-Black Hat wave. Command Zero previewed its own agentic tool days ago.
The Beta Reality
Burp AT is live in public beta for Burp Suite Professional users. The Individual tier works today. Team and Enterprise are labeled coming soon. PortSwigger publishes no prices anywhere on the pricing page. The company ships early on purpose. Fast release is now standard for AI products.
The cage is a strong idea. The market will decide if it holds.
FAQ Agentic AI Pentesting
Burp AT is agentic AI built into Burp Suite. It lets AI agents test web applications for vulnerabilities. PortSwigger launched it in public beta.
No. The agent works under a human tester. The release says pentesters remain responsible for scope, judgment, and conclusions.
The user carries the liability. PortSwigger compares itself to a gun maker. Its duty is the tool. The user aims it.
It points toward continuous testing. It also raises errors and omissions questions for pentest firms. Both affect how carriers price risk.
It can be. Burp AT offers Manual, Smart, and Autonomous modes. The agent always stays inside the boundaries the user sets.
Related Cyber Insurance Posts
- A Breach Convinced 64% To Spend More. An AI Announcement Convinced 85%.
- Cyber Insurance for Small and Mid-Size Businesses: HSB’s Cyber Suite Boosts Coverage(Opens in a new browser tab)
- Cyber Insurance Compliance Demands Drive Enterprise Shift to Software Pentesting in 2025 | Key Findings(Opens in a new browser tab)
- 2024 State of Pentesting Report: Al Threats Rise Amidst Staffing Shortages(Opens in a new browser tab)
- UK Enterprises Turn to Software-Based Pentesting Amid Rising Cyber Threats(Opens in a new browser tab)