Microsoft is trying to turn a large cache of Copilot chat logs into a central argument in its copyright defense: even when users ask questions that touch publishers’ material, the chatbot rarely produces output that could replace the original work.
In new filings connected to lawsuits brought by news publishers, including *The New York Times*, and by book authors, Microsoft says it supplied 8.2 million Copilot conversations for analysis. The company says those logs were selected because they contained keywords associated with the news plaintiffs’ websites, making them among the conversations most likely to reveal use of the plaintiffs’ content.
Microsoft’s conclusion is that meaningful overlap was scarce. It says 59,545 of the 8.2 million conversations contained at least 16 words in common with news content used to ground the system. An expert for the Center for Investigative Reporting identified 51 cases of “substantial overlap” with CIR work in the dataset, according to Microsoft. In the related authors’ case, Microsoft says an expert found just 24 Copilot responses with at least 30 matching words across the 8.2 million conversations reviewed, and that only 10 of 212 books evaluated had any matches.
The key distinction: training versus output
The evidence is aimed at a core question in the AI copyright fight: whether a model’s occasional ability to reproduce text is enough to make its underlying use of copyrighted material unlawful.

Microsoft argues that its use of material to train large language models is transformative and that Copilot serves a substantially different purpose from a newspaper article or book. Under that view, low levels of verbatim or near-verbatim output support a fair-use defense: the system may have learned patterns from protected works, but it is not routinely functioning as a substitute for them.
Publishers and authors make the opposite case. Their suits contend that Microsoft and OpenAI used their works to build products that now compete for readers and answers—and can, at times, regurgitate protected content. The practical dispute is not limited to long copied passages. For publishers, answer engines that summarize or satisfy an information need without a referral may still weaken traffic, subscriptions, licensing leverage, and the economic value of original reporting.
The figures in Microsoft’s filing are therefore significant, but not necessarily decisive. They reflect a dataset selected and characterized by Microsoft, and the precise methodology, prompts, definitions of “substantial overlap,” and role of retrieval or grounding will matter. A low count of exact matches does not by itself settle arguments about market harm, training inputs, or whether a system can generate infringing results under other conditions.
Why operators should pay attention
For AI product leaders, the case illustrates the growing importance of measurement and auditability. Companies deploying models will need more than broad claims that safeguards work. They may need evidence showing how often protected content appears in outputs, what prompts trigger it, how their systems respond to requests for full text, and whether retrieval features cite or link to sources.
That creates a product and governance agenda:
- Maintain testing and monitoring for memorization and high-overlap outputs.
- Separate policies for model training, web-grounded answers, and user-provided content.
- Build attribution, linking, and rights-management capabilities where products depend on timely publisher material.
- Preserve logs and evaluation methods that can withstand scrutiny in litigation or licensing negotiations.
What happens next
Microsoft filed the arguments while seeking summary judgment, a ruling that could end the dispute before trial. The publisher and author claims have been consolidated before one judge, while the Trump administration has also filed a statement of interest supporting OpenAI in the *Times* case.
If the court accepts Microsoft’s framing, AI developers will gain useful support for the proposition that rare output overlap reinforces a fair-use defense. If it does not, developers may face stronger pressure to license training and retrieval content, redesign answer experiences, or accept more explicit limits on how their models use protected works.
Either way, the durable lesson for builders is clear: output-copying rates are becoming a business-critical metric, not just a model-quality concern.



