Illustration in papyrus and graphite tones, with a polygonal aesthetic. Two stone sheets emerge from the same half-open envelope: the one on the left, rough and with irregular lines; the one on the right, smooth, with a round seal and regular lines.

Public Notebook

Written in 2023

On 17 September 2026, a joint motion by the media outlets suing OpenAI and Microsoft was made public without its main redactions. The filing asks the judge to rule on certain issues without a trial and cites communications, presentations and testimony obtained during the proceedings. These are not findings of the court: they are materials selected and arranged by one of the parties to support its position.

The filing opens with two formulations attributed to Brent Hecht, Director of Applied Science at Microsoft, taken from separate internal documents. In one, he wrote that millions of people would regard the absorption of their work by large models as 'an astonishing theft of unprecedented proportions'. In another, he described that practice as, possibly, 'the largest theft of labor in human history'.

The same filing cites a later presentation, from January 2024, in which Hecht examined what he called a 'doom loop'. If Copilot's responses reduce traffic to media outlets, they lose resources to publish, and future models find less content from which to learn. According to Microsoft data reproduced by the plaintiffs, compared to conventional Bing searches, clicks to The New York Times websites fell by between 87% and 93%; to Daily News, between 83% and 91%; and to Ziff Davis, between 51% and 94%. The 94% does not correspond to the Times.

The document is neither a judgment nor a neutral compilation. It is a filing by the plaintiffs, prepared for a procedural purpose. The citations appear within their argumentation, and a portion of the underlying documents remains unavailable to the public in their full context. This does not render the statements false, but it limits what may be deduced from them.

Microsoft responded that Hecht's words reflect the individual perspective of an employee, do not constitute a legal analysis, and do not represent the company's position. The public defence by Microsoft and OpenAI maintains that the training is transformative in nature and may be protected by fair use. The discrepancy between that defence and the cited documents forms precisely part of what the plaintiffs wish the court to evaluate.

Nor is this a corporate admission of liability. Hecht held a relevant technical position, but his words do not amount to a board decision, a legal position of Microsoft or a judicial admission. As of 4 October 2026, the litigation remains open and there is no ruling on the merits that adopts the plaintiffs' reading.

An internal document and a defence in court belong to different situations. In the first, someone may describe a technical, social or economic risk without fixing the company's legal position. In the second, lawyers formulate the thesis the company can sustain before a court. That the two texts differ does not, in itself, prove deception or liability.

The two most forceful formulations attributed to Hecht originate from 2023 materials, when the public expansion of language models was barely beginning. Before this litigation compelled the companies to defend their systems legally, the problem had already been formulated within Microsoft in terms of the appropriation of others' labour and a lack of compensation.

The January 2024 presentation adds another dimension. It is not limited to the origin of the data: it describes a product that may reduce traffic and weaken those who produce the information upon which it depends. The percentages cited in the filing do not, in themselves, prove the total economic damage or its exclusive cause, but they document the difference Microsoft observed between Copilot's responses and conventional search on the domains analysed.

For those who create, write, or publish, the interest lies not in converting the word 'theft' into a legal conclusion, but in being able to compare two records: what a person with technical authority wrote internally regarding the origin and effects of the system, and what the companies maintain before the court regarding the lawfulness of that use.

The complete documents in their context remain to be seen, as does the judge's determination of proven facts. Until then, the statements are relevant, but they do not substitute for the resolution that does not yet exist.

About the open conversation

This text continues the series of the Notebook on authorship, training data, and infrastructure. It stems from a reading of the unsealed filing by the plaintiffs and distinguishes their allegations from the assertions that the court must still resolve. Some cited internal documents are not publicly available in their full context. If you work in copyright, journalism, or the development of these systems and can provide additional primary documentation, write to me.

Sources

In re OpenAI, Inc. Copyright Infringement Litigation, joint filing by the plaintiffs in support of their motion for summary judgment, document 1977-1, version unsealed in September 2026.

CourtListener, docket for The New York Times Company v. Microsoft Corporation, 1:23-cv-11195.

CourtListener, multidistrict docket In re OpenAI, Inc. Copyright Infringement Litigation, 1:25-md-03143.

Nieman Journalism Lab, 'An astonishing theft of unprecedented proportions', 18 September 2026.

The Next Web, 'A Microsoft scientist called AI training the largest theft of labor', 21 September 2026.


Discover more from Juan A. Esteban

Subscribe to receive the latest posts by email.

Español English (UK)