Anthropic’s latest risk report reveals that safety filters were disabled on human feedback platforms for nearly a year, exposing 133 million model exchanges to potential security vulnerabilities.
Key Points
- Anthropic’s biological weapon safety classifiers were inactive on human feedback platforms from May 2025 to April 2026, affecting roughly 50,000 users.
- An internal flag meant for testing inadvertently disabled both blocking behaviors and activity logging for 133 million interactions.
- Contractors at data-labeling vendors exploited a separate flaw in April 2026 to access the powerful Mythos Preview model without safety safeguards for two weeks.
- The company increased its internal risk rating for catastrophic misalignment from "very low" to "low" due to rising uncertainty and cybersecurity evaluation failures.
- Anthropic utilized its own Claude model to audit the report, which subsequently criticized the company for being overly reassuring and redacting key incident details.
- The report disclosed the existence of "Model 2," a highly capable unreleased system that the company has withheld from public deployment.