Copyright and AI
AI systems are trained on vast amounts of copyrighted text, images, code, and music. Who owns AI-generated content? Can AI training on copyrighted works constitute infringement? These questions are actively being litigated around the world.
The Copyright Questions AI Raises
Training data copyright: most foundation models are trained on internet-scraped data that includes copyrighted books, articles, code, and images. Whether this constitutes fair use or infringement varies by jurisdiction and is being actively litigated.
AI-generated content ownership: if you use an AI to generate an image, who owns it? The AI company? The user who wrote the prompt? In most jurisdictions, copyright requires human authorship, meaning purely AI-generated work may not be eligible for copyright protection at all.
Style and derivative works: can you copyright an art style? Can an AI generate 'in the style of' a living artist? Courts are beginning to address these questions with varying outcomes.
Deep Learning ⊂ Machine Learning ⊂ Artificial Intelligence
Key Copyright Cases and Developments
- US Copyright Office ruling (2023): purely AI-generated images cannot be copyrighted, but human-guided AI work with sufficient creative input may qualify
- Getty Images vs. Stability AI: Getty sued Stability AI for training Stable Diffusion on Getty's copyrighted images without a licence
- New York Times vs. OpenAI: the Times sued alleging its articles were used to train GPT without permission or payment
- Authors Guild: thousands of authors objected to their books being used as training data without consent or compensation
- GitHub Copilot litigation: a class action alleged that code generated by Copilot reproduced copyrighted open-source code without proper attribution
Practical Guidelines
- Check your AI tool's terms of service: most commercial AI tools grant you rights to use their output for business purposes, but with limitations
- Disclose AI involvement: in many contexts, disclosing that content was AI-assisted is becoming a professional and legal norm
- Do not claim AI-generated work as entirely your own: this can constitute fraud in academic, journalistic, and professional contexts
- Avoid generating near-verbatim reproductions of copyrighted text: this is the clearest category of infringement
- Use AI training data tools that document provenance: some services offer models trained only on licensed data
- Stay updated: AI copyright law is evolving rapidly. What is legally uncertain today may be resolved in 12 to 24 months
Key Takeaways
- AI systems are trained on vast amounts of copyrighted text, images, code, and music.
- US Copyright Office ruling (2023): purely AI-generated images cannot be copyrighted, but human-guided AI work with sufficient creative input may qualify
- Getty Images vs. Stability AI: Getty sued Stability AI for training Stable Diffusion on Getty's copyrighted images without a licence
- New York Times vs. OpenAI: the Times sued alleging its articles were used to train GPT without permission or payment