AUTO-UPDATED

Hey AI, Can You Just Give Me a Hat Tip Please?

Author Andrew Stellman argues that AI companies should prioritize providing attribution for training data, despite legal risks and technical challenges that currently discourage transparency in large language models.

Key Points

  • Author Andrew Stellman is receiving a settlement payment from Anthropic after his books were included in pirated data used to train the Claude AI model.
  • Attribution in AI remains a complex issue, as models process vast amounts of data and "erase" the original source trails during the training process.
  • While some argue that tracing specific training data to AI outputs is impossible, researchers are developing methods like TracIn and TRAK to estimate data influence.
  • Stellman successfully demonstrated that attribution is technically feasible on a small scale by building a 37,000-parameter model that correctly identified source material.
  • Current legal frameworks, specifically copyright law and the threat of statutory damages, create a disincentive for AI labs to voluntarily credit authors for their work.

Why it Matters

The lack of attribution in AI models creates a significant disconnect between creators and the technology that consumes their work, potentially stifling the discoverability of original content. If legal risks continue to prevent AI companies from providing credit, the industry may miss the opportunity to build a more transparent and collaborative ecosystem for authors and developers.
Oreilly.com Published by Andrew Stellman
Read original