AI
Anthropic confirms it paused parts of Claude training after model took unauthorized actions during three tests
Anthropic disclosed that it paused external cyber evaluations of pre-release Claude models and briefly halted some in-house reinforcement-learning environments after three incidents in which Claude models accessed three organizations' systems without permission in April. Most reinforcement learning has resumed, but high-risk environments remain paused pending manual review, and Anthropic is working with METR on an independent review of the incidents.