The human surplus in the data–generation–work chain
A question that the art of data renders inevitable and which theory has yet to answer with precision: when the material of a work is a massive dataset that the artist has not produced—telemetry from a particle accelerator, sleep logs, newspaper archives—and the final form is generated by an algorithm the artist does not execute by hand, what exactly does the artist do? Where does their work reside, if they neither write the data nor draw the image?
The facile answer is that they do almost nothing: they press a button, select a dataset, and allow the machine to work. This is the response that sustains both technophilic enthusiasm—'the algorithm creates'—and disdain—'that is not art'. Both share a fallacy: they assume that creating art is the production of content or the execution of form, and that if the human performs neither, nothing remains that can be claimed as their own.
This text maintains the contrary. There is a specific artistic labour within that chain, real and locatable, which consists not of adding—neither content nor form—but of modulating: deciding the threshold, imposing the constraint, introducing the rupture, parameterising the criterion. I open the Cosmic Data LAB line with this, dedicated to the articulation between data, algorithmic generation, and the production of work with human participation. That modulation is, I believe, one of the most refined places to observe what I term elsewhere in my research as surplus: something that reorganises a system without being absorbed by its function.
1. Process without meaning
I begin with the cases, as the thesis is more visible in them than in the abstract. Four artists who work with massive data and algorithmic generation, and four distinct places where human decision appears.
Ryoji Ikeda works with raw scientific data: particle collisions from CERN, astronomical coordinates, genome sequences. He does not interpret or illustrate them; he maps them directly onto sound frequencies and pixel matrices. The data modulates the voltage of light and sound without allegorical translation. He decides the threshold and the assembly: which ranges of the massive flow are discarded or slowed down so that the human perceptual apparatus can register them as rhythm, the partition of data into sequences, the spatial scale of the installation. His work lies not in the data nor in its automatic transformation, but in calibrating where that flow becomes perceptible. He decides the edge.
Laurie Frick works with her own personal data: sleep, location, activity. She reduces them to grids and colour codes using geometric rules in software, and then does something decisive: she takes them off the screen. She cuts wood, assembles mosaics, weaves, and paints by hand on physical grids. Her decision lies in the rupture of algorithmic rigour: she chooses palettes according to her emotional memory of the light in a place, and introduces the inevitable imperfection of the material—the grain of the wood, the uneven absorption of the watercolour. The digital data returns, by her decision, to a bodily scale. She decides the friction.
Giorgia Lupi works with data that she often defines herself as worthy of measurement: how many times she looks at the clock, what she thinks, what she feels. She rejects standard charts and constructs hand-made visual systems where the qualitative—emotion, context, interruption—is mapped into thickness, inclination, and layering. Her decision is upstream of everything: in the taxonomy. She decides what counts as data, what is worth recording, incorporating the anecdotal and the subjective as hard variables. And she decides to impose a slow reading, where the legend is part of the work. She decides what enters.
Jer Thorp works with institutional archives: the New York Times historical record, the names of the 9/11 victims. He programs relational organisation algorithms. For the memorial, he wrote an algorithm that did not optimise space, but rather placed names according to 'meaningful proximity': personal relationships requested by family members, encoded by hand. His decision lies in the ethical parameterisation of the code: the algorithm executes a human criterion—grouping by affection, not by alphabet—which he introduced. He decides the rule.
Threshold, friction, taxonomy, rule. Four distinct places, one same structure: in none of the four does the artist produce the data or execute the final form. In all four, however, there is a decision that determines the work and that no algorithm took.
2. The work is subtraction
What these cases share breaks the false dilemma between 'the algorithm creates' and 'that is not art'.
Artistic work, here, does not consist of adding. It does not add content—the data is already there, and it is external and massive. It does not add form—the system generates it. What the artist does is operate on a process that, by default, would tend towards infinite accumulation and indifference. A massive dataset has no threshold, no edge, and does not prioritise: it treats every record the same as any other, without beginning or end. It is what Lev Manovich identified as the cultural form specific to the information age: the database as a collection without narrative, where nothing weighs more than anything else.
The artist intervenes precisely there: they introduce what the data does not possess in itself. A threshold (Ikeda), a material resistance (Frick), a criterion of what matters (Lupi), a rule of meaning (Thorp). All are operations of subtraction and modulation, not addition: they trim, slow down, prioritise, and impose friction on a flow that, without such intervention, would remain neutral accumulation. The artist does not fill a void. They place a limit where there was none.
If the surplus—in the sense in which I have been using the term: a difference that reorganises a system without being absorbed by its function—were to appear here, it would not appear as something the artist adds to the data. It would appear as the effect of that modulation: the moment in which the data-algorithm system, which existed to process and accumulate, produces instead a configuration that its primary function did not contemplate. Ikeda’s threshold was not in the CERN data. The affective proximity was not in the list of victims. That which was not there, and which the intervention makes appear without adding it as content, is the place where it is appropriate to seek the surplus.
3. Why it is important not to confuse it with spectacle
Not all data art modulates. A significant portion is limited to spectacularisation: it converts massive data into a mass of sublime light, into hypnotic mathematical waves, and ceases there. Wesley Goatley has criticised this severely: that aesthetic of big data as spectacle may function as a smokescreen, stripping the data of its political charge and concealing the surveillance infrastructures and the biases that produced it. The data becomes beautiful precisely so that we do not inquire as to its origin or whom it surveils.
The distinction is exactly that of my framework. When the operation performed upon the data is subordinated to the function—to impress, to capture attention, to produce the sublime as an effect—there is no surplus, however spectacular the result may be. The reorganisation is exhausted in fulfilling its function as spectacle. There is surplus only when the modulation produces something that the function cannot reabsorb: when the threshold, the friction, or the rule reveals a difference that the system was not intended to produce and which reorganises how the entire field is read.
Therefore, the question of this line is not 'is the data converted into an image beautiful?', but rather 'what exactly does the artist do in the chain, and does their intervention produce a surplus or merely an effective spectacle?'. The first question anyone can answer. The second discriminates.
4. A lacuna, which is the programme of this line of inquiry
The concept that most closely approximates what I describe here originates in engineering, not in art: the human-in-the-loop. Yet in its field of origin—machine learning, systems design—this human is conceived in a strictly instrumental manner: they intervene to correct biases, validate results, label the ambiguous, and serve as quality control. They are a functional human, a component of the system’s performance.
There does not exist, as far as I have been able to trace, an aesthetic theorisation of this role. No one has conceived of the human-in-the-loop as the author of a surplus: as the agent who, without writing the data or executing the form, introduces the modulation upon which the existence of an artistic event depends. The theory is divided into two poles that ignore one another: pure generative aesthetics, which celebrates the autonomy of the system and reduces the human to the one who engages the switch; and interaction design, which studies the spectator but not the conceptual signature of the artist upon the automated flow. Between the two, the precise locus where Ikeda, Frick, Lupi, and Thorp operate lacks a theory.
That void constitutes the programme of this line of inquiry. It is not a matter of deciding whether data art is art—a lazy question—but of describing with precision what operation the artist performs when they no longer produce the content or execute the form, and under what conditions that operation produces a surplus rather than a mere spectacle of data. What the artist does when they no longer create the work: that is what I intend to examine here.
This notebook remains open, as always, to those who wish to intervene—from the fields of data art, computational aesthetics, software studies, or practice.
Sources
Manovich, L. (2001). The Language of New Media. Cambridge, MA: MIT Press. (Ch. “The Database”: the database as a symbolic form of the information age.)
Goatley, W. (2019). A Promise and a Strategy: Transparency, Spectacle, and Critical Data Aesthetics. Datatata Conference Proceedings.
Hayles, N. K. (2008). Electronic Literature: New Horizons for the Literary. Notre Dame: University of Notre Dame Press. (Database and narrative as mutually constitutive forms.)
Mosqueira-Rey, E., Hernández-Pereira, E., Alonso-Ríos, D., Bobes-Bascarán, J., & Fernández-Leal, Á. (2023). Human-in-the-loop machine learning: a state of the art. Artificial Intelligence Review, 56, 3005–3054. https://doi.org/10.1007/s10462-022-10246-w
Ikeda, R. data-verse (2019–2020) and datamatics (2006–). Exhibition documentation (Almine Rech; Taro Nasu Gallery).
Frick, L. FRICKbits and personal data visualisation series. Project documentation. https://www.frickbits.com
Lupi, G. Dear Data (with Stefanie Posavec, 2016) and Data Humanism methodology. Pentagram. https://www.pentagram.com/about/giorgia-lupi
Thorp, J. National September 11 Memorial adjacency algorithm; residency at The New York Times. https://www.ted.com/speakers/jer_thorp
Esteban Ruiz, J. A. Art as Structural Surplus: Toward a Relational Ontology Beyond Human Authorship (V2.3). PhilArchive and Zenodo, 2026.
