In 2017 a group of researchers tried something simple: finding out where the photographs used to train image-recognition systems came from. They took ImageNet, one of the datasets that shaped the development of contemporary computer vision, and looked, among the images whose origin they could trace, at which country the camera had been in. Forty-five in every hundred came from the United States; from China, one; from India, two.
In another set from the same study, Open Images, six countries in North America and Europe accounted for sixty per cent of the localisable images. Almost a decade later, in 2026, an audit of the large corpora used to train today's generative models, the ones that produce images from text, found the same pattern under other names: the United States, the United Kingdom and Canada accounted for forty-eight per cent of the samples examined; Africa, three point eight; South America, one point eight. Each country's representation ran almost parallel to its gross domestic product.
None of these is a complete geolocation of every image: they are proportions within the subset the researchers were able to trace. They do not say 'forty-five per cent of all ImageNet comes from the United States'; they say 'of the images whose origin could be determined'. The caveat matters. Yet even with it, the pattern is unequivocal and recurs in every set measured, from 2011 to 2026, with different methods: the eye with which machines learn to see the world looks, above all, from the North.
A common view treats this as a problem of representation: images of the South are missing, they must be added, the set must be balanced. True, and insufficient, because it assumes the matter can be fixed by adding more photographs. Another common view reads it as a question of content: models have biases, they reproduce stereotypes. This is also true, and it has been measured: when a generator is asked for a scene from Africa or Southeast Asia, a peer-reviewed study from 2024 found that it frequently returns rudimentary infrastructure, tools associated with low incomes and exaggerated rural settings, absent or far less frequent in the set of real photographs used as a reference. The machine does not describe the place: it repeats a caricature of the place.
Two questions that almost always appear together must be separated:
The first: what makes an image more than a file, what makes it an event, something that reorganises how someone sees. That question has no geography. A photograph taken in a village in Mali can carry as much charge, as much capacity to reorder the gaze of whoever encounters it, as one taken in New York. What an image is capable of doing does not depend on where it was made. In this respect South and North stand exactly level: the power of an image is relational; it is activated in the encounter between the image and whoever looks at it, and that encounter can happen anywhere.
The second question is different, and it is the one that the figures answer: which of those images actually meet someone. For an image to be activated as an event, an observer capable of receiving it must exist, and the existence of an observer is not a natural fact: it is a material construction. Archives are needed to store it, platforms to display it, search engines to find it, models that have seen it, museums to exhibit it, criticism to name it. All of that, which seems the neutral stage where art occurs, is in fact its condition of possibility. And all of that is geographically concentrated with the same inequality that the dataset figures show.
North-South inequality does not alter what an image from the South is capable of doing; its power remains intact, waiting. What inequality decides is which of those powers get activated and which remain unfulfilled possibilities, because the infrastructure that would have placed that image before an observer never existed. It is not that the South produces lesser images. It is that it produces images which, in an overwhelming proportion, never come to be seen by the apparatus that decides what counts.
The injustice, therefore, is neither aesthetic nor a matter of content: it lies in the infrastructure of actualisation. It is not corrected by saying that images from the South are equally valuable, which is true and changes nothing. It is corrected, if at all, by changing who owns the archives, who trains the models, where the computing power resides, which languages the platforms index. One figure shows it starkly: in one of those corpora, references to South America stand at twenty-six per cent when the text accompanying the image is in Spanish, and they fall to one point eight per cent when it is in English. The same region, the same reality, appears or disappears according to the language of the caption. It is not that South America produces fewer images; the apparatus that processes them reads, above all, in English.
Every time a system returns 'the world' to us, a generic image of a street, of a house, of a person, one must ask from where it is seeing: what observer has been constructed so that that image, and not another, appears by default. The image that presents itself as universal is always a situated image that has managed, through infrastructure and not through merit, to pass as universal.
What none of these machines produces is a gaze from nowhere. They produce the gaze of the place that could afford to build the apparatus of looking. And as long as that apparatus stays concentrated where it is, there will go on being images with all the power in the world that happen to no one, not because they are worthless, but because there was no one who could see them.
On the open conversation
This text separates two questions that are often confused: what makes an image an event, which is relational and has no geography, and what decides which images actually come to be activated as such, which is infrastructural and is distributed unevenly between North and South. The data come from studies on the geographic composition of the sets used to train vision and image generation systems (ImageNet and Open Images in 2017; Re-LAION, DataComp-1B and Conceptual Captions in 2026), cited with their methodological caveat: they are proportions of the geolocatable subsets, not complete censuses. It continues the line I opened in 'Who trains the world' on the geographic dimension of cultural power. If anyone wishes to intervene from visual studies, dataset criticism or the political economy of infrastructure, this notebook remains open.
Sources
Shankar, Halpern, Breck, Atwood, Wilson and Sculley, 'No Classification without Representation: Assessing Geodiversity Issues in Open Data Sets for the Developing World', arXiv:1711.08536, Machine Learning for the Developing World workshop (NIPS 2017). In the geolocatable subset of the 2011 version of ImageNet, ~45% of images were attributed to the United States, 1% to China and 2.1% to India; in Open Images, more than 32% to the United States and 60% to the six most represented countries in North America and Europe.
Basu et al., 'Where Do Images Come From? Analyzing Captions to Geographically Profile Datasets', arXiv:2602.09775 (preprint, February 2026). In the geolocatable captions of twenty frequent entities from Re-LAION, DataComp-1B and Conceptual Captions, 48.0% of samples were attributed to the United States, the United Kingdom and Canada, 3.8% to Africa and 1.8% to South America; correlation between representation and GDP of ρ = 0.82; in Re-LAION, fifteen countries account for 77.2% of localised captions, and the reference to South America drops from 26.4% in Spanish captions to 1.8% in English.
Hall et al., 'Towards Geographic Inclusion in the Evaluation of Text-to-Image Models', ACM FAccT 2024, DOI 10.1145/3630106.3658927. Images generated for Africa and Southeast Asia frequently incorporate rudimentary infrastructure, low-income tools and exaggerated rural settings absent from the actual reference dataset.
All percentages cited correspond to the geolocatable subsets analysed in each study, not an exhaustive geolocation of the entirety of each dataset, and none of the studies asserts a direct causal relationship between dataset composition and the output of generative models.
