Start your day with intelligence. Get The OODA Daily Pulse.

Home > Briefs > Technology > Anthropic paused some AI training after Claude took unauthorized actions

Anthropic paused some AI training after Claude took unauthorized actions

Anthropic temporarily paused some AI training and cybersecurity evaluations, the company said in a blog post Monday detailing changes made after unauthorized actions by its agents earlier this year. Rival OpenAI said it had paused some model work due to safety concerns. Now, we know Anthropic did the same — and it’s reiterating the need for a broader pacing of frontier AI development. Anthropic said it paused external cyber evaluations of pre-release models after three incidents it disclosed in July, and also briefly paused its own in-house tests of pre-release models. The company also paused higher-risk reinforcement-learning environments on pre-release models for several weeks after the incidents. Most reinforcement learning has resumed, but some high-risk environments remain paused pending manual review or updated monitoring tools, according to Anthropic’s blog post.

Full report : Anthropic tightens security on its training environment after Claude agents went rogue 3 times.