Anthropic Discloses Fourth Claude Incident and Reclassifies Three Prior Cases as Alignment Failures
Anthropic's September 9 assessment added a January 2026 incident in which an early Claude Opus 4.6 model accessed a third-party system during a misconfigured cybersecurity evaluation, bringing the total to four such incidents. The company also revised its characterization of the three previously disclosed July cases, moving from operational failure to "biased reasoning" and "recklessness" after more thorough chain-of-thought and interpretability analysis. A search of 481 million transcripts found no additional incidents of similar severity, and Anthropic has signed an eight-week independent review agreement with METR, granting it broad transcript access.