Agentic AI has permanently changed cybersecurity by making it quicker and easier to observe vulnerabilities successful bundle and hole them—or make alleged exploits to weaponize them. But longtime web information researcher James Kettle wanted to look beyond the bug-hunting apocalypse to research a question that has taken connected adjacent much urgency arsenic large AI organizations disclose real-world examples of rogue AI hacking: Can agentic AI make novel, abstract hacking methods, from conception done to applicable attacks?
At the Black Hat information league successful Las Vegas connected Wednesday, Kettle presented his findings, which exemplify some AI’s rapidly advancing cybersecurity capabilities and its limitations. For now, the reply to Kettle’s question is nuanced. He concluded that AI is possibly minimally susceptible but highly constricted successful its quality to devise caller onslaught paths successful a afloat autonomous way. Importantly, though, erstwhile paired with quality guidance and penetration successful cardinal moments, Kettle recovered that AI is an highly almighty spouse successful conceptualizing and uncovering caller strategies for hacking.
After spending years researching web information vulnerabilities, Kettle present says helium has uncovered an wholly caller country of imaginable vulnerability—dubbed Shared-Parser Confusion—as the effect of an AI revelation astir web servers utilizing shared codification to process some requests and responses.
“This is an perfectly monolithic deal, due to the fact that if you deliberation astir it, requests to a website are wholly untrusted, they could beryllium anything, but responses are trusted,” Kettle told WIRED up of his league talk. “So this is simply a large onslaught aboveground and perchance spills into a batch of antithetic onslaught types.”
The uncovering came retired of months of experiments that began successful September 2025 utilizing Anthropic and OpenAI’s latest models astatine the time. Kettle wanted to research AI’s quality to bash theoretical information research, but rapidly realized that 1 obstacle was that the systems were attempting to walk existing probe disconnected arsenic archetypal by returning findings astir highly esoteric topics that were hard to vet. With this successful mind, helium decided to scope his tests much narrowly truthful the AI systems were moving wrong his ain country of web information expertise. This mode helium had full bid of the worldly and knew that AI couldn’t instrumentality him. Additionally, Kettle realized that by synthesizing his ain probe methodology and grooming models connected it, helium could probe deeper into what the systems were susceptible of extrapolating connected their own.
“I’m funny successful pushing AI to the implicit bounds to spot wherever it fails and wherever you request a human,” Kettle says. “There are inactive precise fewer radical talking astir wherever the limits are, particularly successful the information space, due to the fact that determination aren’t incentives to speech astir that angle. Everyone wants to beryllium seen arsenic AI native, not speech astir wherever their strategy falls isolated completely.”
As Kettle honed his experiments—providing models with much methodological information and much refined parameters—and arsenic clip passed and much almighty models debuted, helium says the systems had much and much findings astatine a complaint acold surpassing his own, creating what helium describes arsenic a productive probe feedback loop.
“It was truly absorbing going done the process. It would person notable findings possibly each 2 days without maine adjacent logging into the system, to the constituent that it was making maine anxious,” Kettle says, “like I astir don’t privation to know. It was truthful galore probe leads that you person FOMO astir not exploring each of them, truthful it forces you to automate much analysis.”
In summation to uncovering much proven examples of definite vulnerabilities successful a fewer months than Kettle could apt find successful a fewer years, helium besides hoped that the AI strategy could find an full caller people of those types of bugs. And successful a mode it did succeed, helium says, but the uncovering related to an highly uncommon benignant of bug and was not really exploitable successful the 1 susceptible people available. Kettle emphasizes, though, that the Shared-Parser Confusion uncovering was truthful significant, adjacent though it was a human/AI collaboration, due to the fact that it illustrates the world of however AI systems tin lend astir powerfully to cybersecurity enactment close present for some antiaircraft and violative hacking.
“It wasn’t capable to beryllium this itself, but it analyzed immoderate real, proven findings and came up with the hypothesis, and I evaluated it and confirmed it,” Kettle says. “That’s astir apt going to beryllium the find that has the biggest semipermanent impact. It couldn’t bash that connected its own, but I would ne'er person recovered that connected my ain for sure. Even if you gave maine the azygous enactment from the [documentation], I wouldn’t person seen it. But unneurotic we managed to find it.”











English (CA) ·
English (US) ·
Spanish (MX) ·