NewsNTech
The core constraint here is the training data pipeline. Large language models require massive corpora to function, and the specific unit that drives the economics is the cost of acquiring or scraping that data.
The New York Times alleges that this engineering reality was deliberately exploited by OpenAI executives to bypass licensing costs.
Lawyers representing the newspaper claim that OpenAI co-founder Greg Brockman was motivated by the "gazillions" he hoped to gain from models trained on copyrighted content.
The mechanism of the allegation The New York Times asserts that OpenAI staff were aware of the existential threat posed to publishers by their technology. This is not a claim about accidental data ingestion.
Keep reading