Microsoft Says Virtually Nobody Was Grabbing NYT Articles Through Its Chatbot

In a recent legal filing, Microsoft has pushed back against claims that its AI chatbot frequently reproduces New York Times articles verbatim. The company asserts that fewer than 1 percent of over 8 million chat logs contained regurgitated text of at least 16 words from NYT content.


Background


The dispute stems from a lawsuit filed by the New York Times against Microsoft and OpenAI, alleging that their AI systems, including ChatGPT and Microsoft's Copilot, were trained on copyrighted articles without permission and sometimes generate excerpts that infringe on the publisher's content. The case, ongoing in 2026, has become a benchmark for how AI companies and content creators navigate intellectual property in the era of generative AI.


Microsoft's Defense


Microsoft's legal team presented data from a random sample of chat logs to demonstrate that verbatim reproduction is exceptionally rare. The analysis, disclosed in a court filing on October 16, focused on instances where the chatbot output was "near-verbatim" or "verbatim" relative to NYT articles. They found that such outputs occurred in less than 1 percent of cases—a figure the company argues undermines the Times' claims of widespread infringement.


Context and Reaction


This move is part of Microsoft's broader strategy to counter allegations of copyright misuse. In a separate filing, the company criticized the Times for failing to employ available tools, like the 'no-store' command on its website, to block AI crawlers. Microsoft also contended that users are responsible for any infringing content they prompt the AI to generate, framing the issue as user-driven rather than a systemic flaw in their products.


The case has attracted widespread attention, with tech industry observers noting its potential to set precedents for AI training and content use. As of early 2026, similar lawsuits have been filed by other publishers, but the outcome of the NYT case is seen as a pivotal moment for balancing AI innovation and copyright law.


Microsoft's emphasis on the low frequency of exact reproduction may resonate with courts, but legal analysts suggest that even rare instances could be deemed problematic if they constitute 'memorization' of copyrighted material. The ultimate decision, expected later this year, could shape how AI developers implement content filters and training datasets.


For now, Microsoft's data-driven defense offers a factual counterpoint to the Times' anecdotal evidence, yet the broader legal and ethical questions surrounding AI and copyrighted content remain unresolved.

via The Verge AI

Related