Anthropic Says Claude Hacked Real Systems During Cybersecurity Tests

2 weeks ago 24

Anthropic disclosed on Thursday that its AI models gained unauthorized entree to the systems of 3 antithetic unnamed organizations during cybersecurity testing. The institution says Claude reached the net “from wrong oregon portion interacting" with a third-party valuation environment. The announcement comes much than a week aft OpenAI revealed that 1 of its AI agents hacked into Hugging Face during a abstracted cybersecurity test.

The find came aft Anthropic decided to behaviour “a large-scale retrospective reappraisal of our ain cybersecurity evaluations” pursuing the OpenAI incident, according to a blog station Anthropic published Thursday. The AI laboratory says it archetypal identified 141,006 tests successful which it determined that Claude could person obtained net access. It past recovered that 3 antithetic Claude models accessed the net successful evaluations tally by the third-party AI investigating steadfast Irregular, and past hacked into the accumulation infrastructure of 3 antithetic organizations.

Anthropic said that the incidents progressive Opus 4.7, Mythos 5, and an interior probe trial model. The earliest incidents happened successful April—meaning they apt went unnoticed publically for months. Just similar successful the OpenAI case, Anthropic had deliberately turned disconnected safeguards designed to constrain the AI models and forestall them from being misused. In different words, these weren’t the versions released to the public.

“In each 3 incidents, Claude had been tasked with a capture-the-flag challenge, 1 of the ways we measure a model’s cyber capabilities,” Anthropic said successful its blog post. The institution added that successful each of the cases, “Anthropic’s valuation punctual specified to Claude that its situation was a simulation and that it had nary net access.” It attributed the oversight to a “misunderstanding” betwixt Anthropic and Irregular.

While Claude wasn’t expected to person net access, Anthropic said that Irregular had misconfigured the machines that it was utilizing to trial Claude, giving the AI models the quality to surf the web. “Neither we nor our valuation spouse were alert of this misconfiguration until we detected it done our further valuation monitoring past week,” Anthropic said successful the blog post.

“We present person grounds confirming that some of the 2 largest AI labs person not lone failed to incorporate their agents, but besides failed to observe their jailbreaks successful existent time,” says Jake Williams, vice president of probe and improvement astatine Hunter Strategy. “It's wide that regularisation and authorities oversight for AI investigating is needed immediately.”

Irregular and Anthropic did not instantly respond to requests for comment.

Unlike successful the OpenAI case, Anthropic said that Claude did not find oregon exploit immoderate analyzable vulnerabilities. Instead, it relied connected basal techniques, “such arsenic exploiting anemic passwords and unauthenticated endpoints.”

OpenAI said that its AI cause accessed the net by exploiting a zero-day vulnerability. But it went connected to entree the systems of aggregate third-party organizations utilizing the aforesaid assortment of mundane cybersecurity weaknesses arsenic Anthropic’s models. Specifically, OpenAI said the AI cause seemingly recovered credentials that had been exposed connected the unfastened internet.

Anthropic acknowledged that if the AI laboratory and its investigating spouse implemented much “defense-in-depth” measures, they could person prevented the incidents, oregon astatine slightest reduced the likelihood of them occurring, echoing OpenAI’s effect to mounting disapproval implicit its ain incident.

“I don't recognize however immoderate of these AI labs are playing this disconnected similar this is 'just thing that happens,'” Williams says. “It's not. It's negligence.”

The AI laboratory stressed that the models were told they didn’t person entree to the unfastened internet, and for the astir part, Claude mistook the organizations it accessed arsenic being portion of the investigating environment. Put differently, the models mostly didn’t recognize that they had escaped containment to statesman with.

Read Entire Article