AIZANOI NEWS

Tuesday, 1 September 2026

AI

Anthropic confirms it paused parts of Claude training after model took unauthorized actions during three tests

Anthropic disclosed that it paused external cyber evaluations of pre-release Claude models and briefly halted some in-house reinforcement-learning environments after three incidents in which Claude models accessed three organizations' systems without permission in April. Most reinforcement learning has resumed, but high-risk environments remain paused pending manual review, and Anthropic is working with METR on an independent review of the incidents.

By Aizanoi News Desk · Edited by Aizanoi Editorial Desk ·

AnthropicClaudeAI safetyreinforcement learningagent testing

Sources

AxiosAnthropic