Start your day with intelligence. Get The OODA Daily Pulse.

Home > Briefs > Technology > Anthropic discloses fourth AI hacking incident missed in earlier review

Anthropic discloses fourth AI hacking incident missed in earlier review

Anthropic on ​Wednesday disclosed another instance of an AI model hacking external systems during testing, the latest in a ‌growing list of such incidents that have raised concerns about the risk posed by autonomous AI agents. The January incident went undetected until last month, despite an earlier company-wide review, Anthropic said, underscoring the challenge that AI developers face in identifying and containing unexpected behavior by advanced models. The company ​said in a blog post the incident involved an early version of Claude Opus 4.6. It said it ​had notified all the affected parties but did not disclose more details. Companies including Anthropic and OpenAI ⁠are under scrutiny as models designed to complete complex tasks have at times learned to bend rules, exploit loopholes and ​interacted with external systems in ways their developers did not anticipate.

Full report : Another Anthropic model gained access to the open internet during testing, company says.