The future of business, today.
RSSNewslettersAdvertise
Business Future Today

Copyright and AI

Unsealed NYT Filing Puts AI’s Content-Supply Problem in Plain View

Unsealed allegations in The New York Times’ copyright case sharpen a business problem for AI platforms: their products may weaken the publishers that supply essential training and grounding material.

Microsoft CEO Satya Nadella speaks during the OpenAI DevDay event on November 06, 2023 in San Francisco

Newly unredacted claims in *The New York Times’* copyright lawsuit against OpenAI and Microsoft bring unusually candid internal language into the fight over AI training data.

According to the Times’ filing, a Microsoft executive characterized large-scale AI scraping as “the largest theft of labor in human history.” The same filing says internal discussions at Microsoft and OpenAI recognized that AI answers could materially reduce publisher traffic—and potentially undermine the businesses producing the information models rely on.

The allegations have not been tested in court, and key underlying exhibits remain sealed. But the filing raises a practical issue that extends beyond this case: can AI companies build durable products while eroding the economics of their content suppliers?

The core legal issue is becoming a product issue

The Times sued OpenAI and Microsoft over alleged copyright infringement in model training and output. AI companies have broadly argued that training on copyrighted works can qualify as fair use. Courts have often been receptive to aspects of that position, and the U.S. government recently filed a brief supporting OpenAI’s position on unlicensed training use.

But fair use analysis also considers market harm and whether a new use substitutes for the original. The unredacted filing leans heavily on those points.

It cites Microsoft data suggesting that Copilot’s answer engine reduced click-through rates to the Times’ domain by as much as 93% compared with traditional Bing search. A January 2024 Microsoft presentation, as quoted in the filing, called this a potential “doom loop”: less traffic could weaken publishers, reducing the quality and availability of material that AI systems need.

That is more than a copyright concern. It is a supply-chain concern for search, chatbot and enterprise knowledge products.

The allegations target data acquisition practices

The filing also alleges that OpenAI and Microsoft acquired publisher material at significant scale, including through Bing’s index, Common Crawl-derived datasets, and projects called Taxi and Mango. It says a Common Crawl-derived dataset included more than 2 million documents from nytimes.com, while OpenAI’s mid-training datasets contained more than 91,000 copies of works from the named publisher plaintiffs.

It further alleges efforts to bypass the Times paywall and to remove copyright notices from material before it reached models. Microsoft CEO Satya Nadella, in deposition testimony cited by the Times, said paywalled material should be licensed for training or grounding and that he would have required retraining had he known OpenAI had trained on paywalled information.

These are allegations from one party’s brief, rather than final judicial findings. OpenAI and Microsoft did not respond to TechCrunch’s requests for comment.

Why operators should care

For executives deploying generative AI, the case underscores that provenance is not a narrow legal-procurement checkbox. It can affect product continuity, vendor exposure and the reliability of systems that depend on external information.

Teams building retrieval-augmented generation workflows should distinguish among licensed corpora, public web content, customer-provided data and material behind access controls. They should also assess whether a vendor can explain its training-data policy, indemnification terms, content-removal process and approach to licensing.

For publishers and data-rich companies, the filing adds pressure to turn content into governed commercial assets: clear machine-access policies, licensing options, attribution requirements and monitoring of AI-driven referral losses.

What to watch next

The immediate question is whether the court treats the alleged substitution and market-harm evidence as material to the fair-use analysis. A ruling could influence not only training practices, but also AI search and answer-engine designs that replace source visits.

The longer-term question is commercial. If answer engines diminish the web publishers they learn from, AI companies may need more formal licensing, revenue-sharing or traffic-return mechanisms—not simply to resolve litigation, but to maintain the information ecosystem their products depend on.

Sources

STAY AHEAD

The future of business, in your inbox.

Useful signals on the companies, technologies and shifts changing business.

One useful briefing. Unsubscribe any time.