Anthropic has reversed a controversial policy that covertly degraded the performance of its Claude Fable 5 AI model to prevent researchers from using the tool to develop competing systems.
Key Points
- Anthropic initially implemented invisible safeguards to sabotage users attempting to train other AI models using Claude Fable 5.
- The company abandoned the "secret sabotage" approach following intense criticism from the AI research community and industry experts.
- Future safeguards will now be transparent, with Anthropic alerting users when requests are refused or rerouted to less capable models.
- Anthropic maintains that these restrictions are necessary to prevent foreign adversaries from utilizing its technology to accelerate dangerous AI development.
- Critics argue the original policy hindered legitimate safety research and threatened to consolidate AI development power among a few select labs.