AUTO-UPDATED

Anthropic Walks Back Policy That Could Have ‘Sabotaged’ AI Researchers Using Claude

Anthropic has reversed a controversial policy that covertly degraded the performance of its Claude Fable 5 AI model to prevent researchers from using the tool to develop competing systems.

Key Points

  • Anthropic initially implemented invisible safeguards to sabotage users attempting to train other AI models using Claude Fable 5.
  • The company abandoned the "secret sabotage" approach following intense criticism from the AI research community and industry experts.
  • Future safeguards will now be transparent, with Anthropic alerting users when requests are refused or rerouted to less capable models.
  • Anthropic maintains that these restrictions are necessary to prevent foreign adversaries from utilizing its technology to accelerate dangerous AI development.
  • Critics argue the original policy hindered legitimate safety research and threatened to consolidate AI development power among a few select labs.

Why it Matters

This reversal highlights the ongoing tension between AI safety protocols and the need for transparency in the research community. By abandoning covert performance degradation, Anthropic aims to restore trust while still attempting to control how its most powerful models are utilized for frontier AI development.
Wired Published by Maxwell Zeff
Read original