Home / Viral & Trending / Local newspapers sue OpenAI and Microsoft over use of paywalled articles

Local newspapers sue OpenAI and Microsoft over use of paywalled articles

The Seattle Times and Newsday filed a federal lawsuit on Thursday against OpenAI and Microsoft, alleging the technology companies misappropriated their copyrighted journalistic content to train generative artificial intelligence models. The complaint, filed in the U.S. District Court for the Southern District of New York, claims the tech giants systematically scraped thousands of paywalled articles to develop and refine their AI products, including ChatGPT, Microsoft Copilot, and the AI-enhanced Bing search engine.

This legal action marks a significant escalation in the ongoing conflict between the traditional news industry and the rapidly expanding artificial intelligence sector. The publishers contend that by utilizing their reporting without permission or compensation, the defendants are effectively "freeloading" on the multi-million dollar investments required to produce high-quality local journalism. The lawsuit seeks unspecified damages and a permanent injunction that would require the tech firms to destroy any training datasets or AI models containing the publishers’ intellectual property.

The Core Allegations: Bypassing Paywalls for AI Training

According to the legal filings, OpenAI and Microsoft utilized automated web crawlers to bypass digital paywalls and harvest vast quantities of proprietary data. The plaintiffs argue that this content was then fed into large language models (LLMs) to teach the software how to mimic human language, summarize complex events, and provide direct answers to user queries. The publishers allege that these AI systems frequently output near-verbatim reproductions of their articles, which directly undermines the subscription-based business models that sustain local newsrooms.

The Seattle Times and Newsday assert that their articles are not merely "data points" but are the result of rigorous investigative work, fact-checking, and editorial oversight. By incorporating this work into commercial AI products, the lawsuit claims, OpenAI and Microsoft are creating a competing product that provides the same information to users without the users ever needing to visit the original news websites. This dynamic threatens to siphon off both advertising revenue and digital subscription fees, which are the lifeblood of modern local news organizations.

Legal representatives for the newspapers highlighted specific instances where AI chatbots provided detailed summaries or direct quotes from paywalled investigative reports. They argue that because these summaries satisfy the user’s information needs, the "click-through" rate to the original source drops precipitously. This phenomenon, often referred to as "zero-click" searches, has become a central point of contention for digital publishers worldwide.

Protecting the Financial Future of Local Journalism

The financial stakes for local news organizations are particularly high. Alan Fisco, president and CEO of The Seattle Times, emphasized the economic reality of the situation in a memo to staff. Fisco noted that the organization spends millions of dollars annually to maintain its newsroom and produce original reporting. He argued that allowing tech companies to utilize this content for free represents a fundamental threat to the sustainability of the fourth estate.

"We feel strongly that we must defend our content from being used without our consent or compensation," Fisco stated. The sentiment reflects a broader industry-wide anxiety that the "AI revolution" could finish the job that the initial shift to digital advertising began: the hollowization of local newsrooms across the United States. Unlike national outlets with massive global reach, local newspapers rely on a specific, geographically targeted audience that is increasingly being diverted to AI-generated summaries.

The lawsuit points out that while OpenAI and Microsoft have reached licensing agreements with some large international publishers, many local and regional outlets have been left out of the negotiations. The plaintiffs argue that this creates an uneven playing field where only the largest media conglomerates receive compensation for their intellectual property, while local outlets are expected to provide their work for free under the guise of "fair use."

A Growing Legal Front Against Generative AI Models

This latest litigation adds to a mounting pile of legal challenges facing OpenAI and its primary financial backer, Microsoft. The most prominent of these is the ongoing lawsuit filed by The New York Times, which similarly alleges that the tech companies used millions of its articles to train their AI systems. The outcome of these cases will likely hinge on the interpretation of "fair use" in the age of machine learning.

Under U.S. copyright law, fair use allows for the limited use of copyrighted material without permission for purposes such as criticism, news reporting, teaching, and research. OpenAI has consistently maintained that its training processes fall under this category because the resulting models are "transformative." They argue that the AI does not simply store and regurgitate articles but learns the underlying patterns of language to create something entirely new.

However, the publishers in the current lawsuit argue that the use is not transformative but "extractive." They claim that because ChatGPT and Copilot can reproduce entire passages or closely paraphrase reporters’ work, they are functioning as a substitute for the original product rather than a new creative work. The Southern District of New York has become the primary battleground for these arguments, with multiple judges currently weighing the technical nuances of how LLMs process and store information.

Federal Intervention and the National Security Argument

In a surprising turn of events, the legal battle has drawn the attention of the federal government. The Department of Justice (DOJ), under the Trump administration, recently filed a Statement of Interest in the New York Times case, which has significant implications for the Seattle Times and Newsday lawsuit. The DOJ’s brief argued that a legal victory for the publishers could inadvertently harm local newsrooms and national security.

The administration’s position is that AI technology can be a powerful tool for small, independent publishers, allowing them to automate certain tasks and stay competitive with larger mainstream outlets. More controversially, the DOJ argued that stalling the advancement of American AI through restrictive copyright rulings could give foreign adversaries, such as China, a strategic edge in the global technology race. This "national security" defense suggests that the government views the rapid development of AI as a priority that may occasionally supersede traditional intellectual property protections.

This intervention has been met with skepticism from news industry advocates. Critics argue that the government’s stance prioritizes the growth of trillion-dollar tech companies over the survival of the independent press. They contend that national security is actually bolstered by a robust, well-funded local news ecosystem that can combat misinformation and hold local governments accountable—a task that AI chatbots are currently unable to perform reliably.

The Technological Defense: OpenAI and Microsoft Respond

OpenAI has publicly responded to the wave of litigation by asserting that its practices are both legal and beneficial to the broader information ecosystem. The company maintains that it only uses publicly available materials for training and that it provides "opt-out" mechanisms for publishers who do not want their sites crawled by its bots. OpenAI executives have often characterized their mission as democratizing information, arguing that AI tools help users synthesize and understand the vast amount of data available on the internet.

Microsoft, for its part, expressed surprise at the filing of the lawsuit by the Seattle Times and Newsday. In a statement, the company indicated a willingness to engage in dialogue with publishers to find mutually beneficial solutions. Microsoft has previously pointed to its "Bing for Publishers" program and other initiatives as evidence of its commitment to supporting the news industry, though these programs do not typically involve direct payments for training data.

The defense strategy for these tech giants often focuses on the "black box" nature of neural networks. They argue that it is technically impossible to "unlearn" specific data once a model has been trained. Therefore, the publishers’ demand that the companies destroy existing models could effectively reset the progress of generative AI by several years, a consequence the defendants argue would be disproportionate to the alleged harm.

Potential Consequences for the Artificial Intelligence Industry

If the courts side with the local newspapers, the repercussions for the AI industry could be profound. A requirement to license all training data would fundamentally change the economics of AI development. Currently, the industry relies on the ability to ingest massive datasets at little to no cost. If every piece of copyrighted text requires a licensing fee, the cost of developing a competitive LLM could skyrocket, potentially entrenching the dominance of the few companies wealthy enough to afford such fees.

Furthermore, a ruling that necessitates the destruction of models trained on infringing data—often called "algorithmic disgorgement"—would be a catastrophic setback for OpenAI and Microsoft. It would set a precedent that could lead to a cascade of similar demands from authors, artists, musicians, and other copyright holders whose work has been used to train AI without their explicit consent.

On the other hand, a victory for the tech companies could signal the end of the traditional subscription model for digital news. If AI systems can legally provide the "essence" of a paywalled article for free, the incentive for consumers to pay for news will continue to dwindle. This could lead to a future where local news is either heavily subsidized by the state or disappears entirely, leaving "news deserts" across large swaths of the country.

The Path Forward in the Southern District

As the lawsuit proceeds, the discovery phase will likely provide a rare glimpse into the internal workings of OpenAI and Microsoft. Lawyers for the newspapers will seek to uncover exactly how much paywalled content was used and whether the companies intentionally designed their systems to circumvent security measures.

The case of the local newspapers suing OpenAI and Microsoft over the use of paywalled articles is more than just a dispute over copyright; it is a fundamental test of how society values original content in an era of automated synthesis. The legal community, the tech industry, and journalists alike are watching closely, as the eventual ruling will define the boundaries of intellectual property for the next generation of technological innovation.

For now, the Seattle Times and Newsday join a growing coalition of plaintiffs—including the Authors Guild and various visual artists—who are demanding a seat at the table. Whether through a landmark court ruling or a series of massive licensing settlements, the relationship between those who create information and those who process it is being fundamentally rewritten. The outcome in the Southern District of New York will determine if local journalism remains a viable business or becomes an unpaid fuel source for the engines of artificial intelligence.

Tagged:

Leave a Reply

Your email address will not be published. Required fields are marked *