When the surplus lies not in the output, but in the substrate
In the preceding text, I maintained that the artist working with big data does not produce the content nor execute the form: they modulate. They decide the threshold, the friction, the taxonomy, the rule. The work was an operation upon an external flow that the artist had not generated—CERN data, personal logs, newspaper archives. The decision arrived after the data, upon it.
This text examines the case that reverses this scheme, and does so in a manner that compels a reformulation of the question. What occurs when the artist does not receive the data, but fabricates it? When the dataset is not found raw material, but a work in itself, produced by hand, before any algorithm intervenes? If in the first text the surplus appeared at the end of the chain—in the modulation of the output—here it appears at the beginning, sealed within the substrate. And that alters the status of everything that follows, including the algorithm.
1. Ten thousand tulips
The case is Myriad (Tulips) and Mosaic Virus, by Anna Ridler, from 2018. And it is worth recounting it for its procedure, because the procedure is the thesis.
Ridler wished to train a generative system using tulips. The standard path would have been to download thousands of images from the web, or to pay precarious workers to label them en masse. She did the opposite. She purchased, manipulated, and photographed ten thousand tulips herself, one by one, over the course of three months—the exact duration of a flowering season. And then she labelled the ten thousand images by hand, with a bespoke taxonomy, writing beneath each one its categories: the precise tone, the phase of decay, the degree of streaking on the petals caused by the mosaic virus infection.
That dataset—ten thousand hand-labelled photographs—is Myriad: it is exhibited as an installation, a physical grid of fifty square metres. Only afterwards does she introduce this material into a generative network to produce Mosaic Virus, a video of mutating tulips, whose blooms are linked to the real-time price of a cryptocurrency—a nod to the tulip mania of the 17th century, the first speculative bubble in history.
The decisive factor is the order of operations. Ridler does not modulate the output of a system trained on external data. She determines, from the root, what the system can possibly generate, because she has produced and classified every element from which the system will learn. The algorithm does not discover tulips on the internet: it interpolates between the tulips that Ridler photographed and named. The machine conceives nothing that escapes the physical and conceptual framework she sowed by hand.
2. The inversion of the chain
It is prudent to formulate the distinction with precision, for it is structural.
The standard chain of data art proceeds as follows: massive, external data, typically extracted by scraping the web, feeds an algorithm, and the artist intervenes by modulating the output—filtering, cropping, shaping the result. Human decision-making occurs at the end. This is the scheme of the four cases in the previous text, and it is the dominant one in theory: Manovich, Christiane Paul, and the majority of new media theorists assume that the raw material of data art is, by definition, a pre-existing, massive, automated, and incorporeal flow. The artist is a choreographer of flows they did not produce.
Ridler's chain runs in reverse: manual construction of the dataset, then algorithm, then work. Human decision-making is at the beginning, and it is of a different nature. It is not modulation of an output; it is fabrication of the substrate. And that has a consequence which degrades the status of the algorithm: if every input datum is a human decision—which flower, which light, which label, which category—then the generative network loses its aura as an autonomous creative agent. It is reduced to what it has technically always been: an interpolation engine that calculates averages and transitions between a set of prior human decisions. The algorithm does not add; it averages what was already sown.
Here the surplus—the difference that reorganises the system without allowing itself to be absorbed by its function—does not appear, as in the first text, by modulating the output. It appears encrypted in the input. It resides in the gesture of having fabricated by hand every drop of the channel, so that the machine cannot compute anything outside the physical and conceptual experience of the one who made it. The biological time of a flowering season, the muscular labour of photographing ten thousand flowers, the subjective criteria of each label: all of this remains sealed in the data, and the algorithm can do nothing but reproduce it, transformed.
3. Three variations on decision in the substrate
Ridler is the most radical case, but not the only one that shifts the decision towards the origin. Three other artists operate upon the substrate, each in their own way.
Memo Akten, in Learning to See, does not manufacture the data but designs the bias. He takes a live video feed of everyday objects—a sheet, some cables, a pair of hands—and has it interpreted by a network trained on disparate thematic datasets: waves, fire, clouds, flowers. The moving red sheet becomes the advance of a fire. Akten’s decision lies in the deliberately dissonant coupling between the physical stimulus and the model’s prejudice. He does not modulate the output: he configures in advance how the machine will misinterpret, and in doing so, he exposes that algorithmic vision is never neutral, that it is a strict reflection of the visual memory imposed upon it. The surplus lies in the design of the bias, not in the retouching of the result.
Mario Klingemann, in Memories of Passersby I—his own work, an autonomous system housed in a wooden console that generates faces that never repeat—intervenes in the training. He built a binary interface, Tinder-style, to manually review thousands of variants generated by the networks and select those that aligned with his criteria, strongly marked by the surrealism of Max Ernst and Breton’s “convulsive beauty”. He does not touch the pixel in real time: he acts as a breeder of the system’s mathematical genetics, sowing an aesthetic deviation towards deformation to prevent the network from converging on perfect, sterile photorealistic portraits. The surplus lies in the selection that guides the evolution of the system.
Trevor Paglen, in ImageNet Roulette—with Kate Crawford—makes the inverse movement to all the others: he does not sculpt the data to produce an image, he dismantles it to expose it. His installation processed the faces of the public with the ImageNet engine and stamped upon them the categories that the developers had invisibly assigned to those patterns—frequently racist, misogynistic, heirs to nineteenth-century phrenology. Paglen’s decision lies in the epistemic framing: visually declassifying the ideological substrate of the data, forcing one to look at the political biases fossilised in the raw material that the military-commercial infrastructure uses to categorise people. He does not produce aesthetic surplus: he produces critical transparency. It is the necessary reverse of this line—it reminds us that data is never neutral, and that working with it is working with its politics.
Manufacturing (Ridler), biasing (Akten), selecting (Klingemann), dismantling (Paglen). Four ways for human decision to reside in the substrate, not in the output. Four ways for the algorithm to cease being a creative agent and return to what it is: an engine that operates on what a human has already decided.
4. Corporeal data: a gap and a programme
What the Ridler case opens, and theory has yet to consider, is the status of data when it is manufactured by hand.
Data art theory assumes, almost without exception, that data is an incorporeal flow: something that is extracted, downloaded, processed, but which exists before and outside the artist. Under that assumption, data has no body, no author, no biological time. It is pure available information. And Ridler’s practice refutes that assumption in act: her dataset has a body—ten thousand physical photographs—it has time—a flowering season—it has muscular labour and it has an authorship as deliberate as that of any work. The data, here, is not computer abstraction: it is matter conditioned by the physical labour, time, and criteria of the person who produced it.
There does not exist, as far as I have been able to trace, an aesthetic theory of corporeal data or slow data: of the dataset produced like a sculpture or an engraving, slowly, by hand, as a work rather than as an input. New media theory theorises found data; not manufactured data. And that gap is not minor, because it touches the heart of the question of this line. If human surplus can reside in the substrate—if it is in how the data is manufactured and not only in how the output is modulated—then a good part of the debate regarding the creative autonomy of generative systems is poorly framed. It is not about how much freedom the algorithm has, but about how much of the work was already decided before the algorithm was switched on.
That is the axis that Cosmic Data LAB aims to develop. The first text situated the surplus in the modulation of the output. This one situates it in the manufacturing of the substrate. Between the two, a hypothesis that runs through the line: that in data art, as in any art, the decisive factor is not who or what executes the final form, but where the difference that the system cannot reabsorb is produced. Sometimes at the end of the chain. Sometimes, as in Ridler’s ten thousand tulips, long before the chain begins.
This notebook remains open, as always, to anyone who wishes to intervene—from data art, software studies, computational aesthetics, or practice.
Sources
Ridler, A. (2018). Myriad (Tulips) and Mosaic Virus. Project documentation and artist’s methodological statements. https://annaridler.com/works/myriad-tulips
Akten, M., Fiebrink, R., & Grierson, M. (2019). Learning to See: You Are What You See. arXiv:2003.00902. https://arxiv.org/abs/2003.00902
Klingemann, M. (2018). Memories of Passersby I. Sotheby’s, Contemporary Art Day Auction (Lot 109, 2019), catalogue note and technical documentation. Onkaos.
Crawford, K., & Paglen, T. (2019). Excavating AI: The Politics of Images in Machine Learning Training Sets. AI Now Institute, New York University. https://excavating.ai
Manovich, L. (2001). The Language of New Media. Cambridge, MA: MIT Press.
Esteban Ruiz, J. A. What the artist does when they no longer make the work. Public notebook at juanesteban.art, 2026.
Esteban Ruiz, J. A. Art as Structural Surplus: Toward a Relational Ontology Beyond Human Authorship (V2.3). PhilArchive and Zenodo, 2026.
