The constraint at the base of every large language model is training data. The legal question of whether ingesting copyrighted text to build those models constitutes infringement now has a federal answer. The Trump administration has urged a court to reject the New York Times's claim that training AI models on copyrighted content is illegal, formally aligning the US government with OpenAI in the dispute.

The mechanism the newspaper's case targets is the pre-training pipeline. Building a language model requires processing text at scale. The Times's position is that pulling copyrighted material into that process crosses into infringement. The administration's court filing asks the court to dismiss that argument.

Where this intervention lands is at a consequential point in the case. The training data layer is where foundation models are constructed. A ruling that accepts the Times's theory would challenge the legality of how those systems are assembled. The administration's brief makes the US government a named institutional supporter of OpenAI's defense on that specific question.

Related reading