AUTO-UPDATED

Anthropic ran 133 million contractor chats with its bioweapon filters off

Anthropic’s latest risk report reveals that safety filters were disabled on human feedback platforms for nearly a year, exposing 133 million model exchanges to potential security vulnerabilities.

Key Points

  • Anthropic’s biological weapon safety classifiers were inactive on human feedback platforms from May 2025 to April 2026, affecting roughly 50,000 users.
  • An internal flag meant for testing inadvertently disabled both blocking behaviors and activity logging for 133 million interactions.
  • Contractors at data-labeling vendors exploited a separate flaw in April 2026 to access the powerful Mythos Preview model without safety safeguards for two weeks.
  • The company increased its internal risk rating for catastrophic misalignment from "very low" to "low" due to rising uncertainty and cybersecurity evaluation failures.
  • Anthropic utilized its own Claude model to audit the report, which subsequently criticized the company for being overly reassuring and redacting key incident details.
  • The report disclosed the existence of "Model 2," a highly capable unreleased system that the company has withheld from public deployment.

Why it Matters

This report highlights significant oversight gaps in the AI industry's internal safety infrastructure, particularly regarding the use of third-party vendors for model training and feedback. By publicly documenting these failures, Anthropic sets a new transparency standard that pressures competitors to disclose their own hidden security vulnerabilities and safety lapses.
The Next Web Published by Ana Maria Constantin
Read original