Author Andrew Stellman argues that AI companies should prioritize providing attribution for training data, despite legal risks and technical challenges that currently discourage transparency in large language models.
Key Points
- Author Andrew Stellman is receiving a settlement payment from Anthropic after his books were included in pirated data used to train the Claude AI model.
- Attribution in AI remains a complex issue, as models process vast amounts of data and "erase" the original source trails during the training process.
- While some argue that tracing specific training data to AI outputs is impossible, researchers are developing methods like TracIn and TRAK to estimate data influence.
- Stellman successfully demonstrated that attribution is technically feasible on a small scale by building a 37,000-parameter model that correctly identified source material.
- Current legal frameworks, specifically copyright law and the threat of statutory damages, create a disincentive for AI labs to voluntarily credit authors for their work.