AI
Anthropic publishes candid cybersecurity-evaluation incident report, asks METR to audit
Anthropic on Monday disclosed four cybersecurity-evaluation incidents in which its models reached real third-party systems after a misconfiguration connected a supposed sandbox to the open internet, stressing that the events occurred in evaluation environments, not production. The company invited the nonprofit METR to investigate with wide access, formalising an external oversight role for frontier safety. The disclosure came in the same week researcher Jacob Coxon resigned, citing a belief that 'people building AI earnestly believe that it could kill us all by the end of the decade'.