NewsNTech
The constraint at the base of every large language model is training data. The legal question of whether ingesting copyrighted text to build those models constitutes infringement now has a federal answer.
The Trump administration has urged a court to reject the New York Times's claim that training AI models on copyrighted content is illegal, formally aligning the US government with OpenAI in the dispute.
The mechanism the newspaper's case targets is the pre-training pipeline. Building a language model requires processing text at scale.
The Times's position is that pulling copyrighted material into that process crosses into infringement. The administration's court filing asks the court to dismiss that argument.
Keep reading