Chapter 13 A genealogy of stereotypes: the representation of Indigenous Peoples with generative AI
Se, nas cidades, temos dificuldade em fazer as pessoas entenderem que existe uma grande diversidade indígena, imagine explicar isso para um banco de dados.
[If we have trouble getting people in the big cities to understand that there is not just one way of being Indigenous, imagine explaining that to a database.]
(Kuenan Mayu, comment made during one of the AIAI focus groups, May 2023)
From the moment we first started generating images with Midjourney in the context of the AIAI project, we noticed that the images evidenced a great many preconceived ideas about Indigenous Peoples, representing them in ways that are over-simplified and that do not draw on the cultural specificity of the communities in question. This bias in the representation of Indigenous Peoples has its genealogy in how they have been portrayed throughout the history of photography. To analyse these recurrent forms of stereotypical representation, this chapter looks back at a selection of photographs of Indigenous Peoples that were taken in South America at the end of the nineteenth century. In this period, aspects such as the mise-en-scène, the pose, the use of a background screen, as well as the alteration of cultural norms in terms of clothing and jewellery, began to impact the way we perceive Indigenous Peoples. An enlightening example is the case of photographic images of the Mapuche people which played a foundational role in the way outsiders have viewed that community. Today, we see these visual conventions reappearing in images generated by Artificial Intelligence tools like Midjourney, thus suggesting that historical bias in the representation of Indigenous Peoples is not just ongoing, but that it is being reproduced and amplified by these new generative technologies.
Mise-en-scène
Photographs of people taken in the late nineteenth century followed the same strict conventions as portrait painting in terms of mise-en-scène. According to Peter Burke (2001), having one’s photo taken was about capturing a scene in which the person, dressed in their Sunday best, was surrounded by objects that characterised them and that served a symbolic function, ‘hiding a religious or moral message presented through the “disguised symbolism” of everyday objects’ (21). This use of mise-en-scène was also applied to the first photographic images taken of Indigenous Peoples, such as the Mapuche. In the book Mapuche: fotografías Siglo XIX y XX: construcción y montaje de un imaginario [Mapuches: Photographs from the Nineteenth and Twentieth Centuries: Construction and Mise-en-Scène of a Social Imaginary], Margarita Alvarado (2001) analyses the work of photographers, such as Gustavo Milet Ramírez and Odber Heffer Bissett, who worked in Chile at the end of the nineteenth century. According to Alvarado, their photographs formed the basis for the way in which Mapuche culture – or Auraucanian culture as it was referred to at the time – was understood at that point in history.
Alvarado’s observations are fundamental to understanding how a stereotypical view of the Mapuche people was constructed through photographic images. When analysing a series of portraits of Mapuche women taken around 1900 in Gustavo Milet Ramírez’s studio (Fig. 13.1), Alvarado highlights a recurrent practice: ‘podemos percibir que todas estas diferentes mujeres aparecen exhibiendo exactamente las mismas joyas de plata. Milet interviene así directamente una normatividad social y estética instalando sobre estas mujeres joyas de plata que probablemente eran de su propiedad’ [We can see that all these different women appear to be displaying exactly the same pieces of silver jewellery. Milet directly intervened in social and aesthetic norms by requiring these women to display silver jewellery that probably belonged to the photographer himself] (23). Furthermore, according to Alvarado, objects of cultural significance such as traditional Mapuche jewellery – pieces known as trapelakucha, sikil or trarulongko – are not presented in relation to their Indigenous symbolism or their role in the construction of Indigenous identity, but are used to suit the aesthetic choices of the photographer: ‘No son más que una parafernalia utilizada para una escenificación, son solo un subterfugio para mostrarnos una realidad que, al fin y al cabo, es un delicado y maravilloso montaje’ [They are just props used to set the scene, a trick to show us a reality that, at the end of the day, turns out to be a delicate and marvellous invention] (23).
Fig. 13.1. Mapuche women wearing silver jewellery, Gustavo Milet Ramírez’s studio, Traiguen, Chile (ca. 1900). Nationaal Museum van Wereldculturen, Leiden, the Netherlands. Collection Wereldmuseum, nos. RV-A299-11 and RV-A299-19.
This stereotypical use of cultural symbols is reproduced in images generated with Artificial Intelligence and was observed in the very first tests carried out by the AIAI project team with Midjourney. If we used prompts such as ‘amazon_woman_child_craftwork’ and ‘andes_woman_child_weaving’, we saw that, just as with the images of Mapuche women photographed at the end of the nineteenth century, the AI images of women and children from the Amazon (Fig. 13.2, upper images) and the Andes (Fig. 13.2, lower images) do not faithfully reflect the clothing, textiles or adornments that correspond to those regions. For people who belong to these cultures, this can be quite shocking, because the iconography or colours of textiles have always corresponded to specific locations and identities and still do so today.
In the generated images, instead of showing specific iconographic elements that allow us to clearly identify where a person is from or which Indigenous People they belong to, we are presented with a confusing mixture of signs belonging to different ethnic groups. There are even some unidentifiable elements that have nothing to do with the cultural codes associated with Amazonian or Andean communities. This phenomenon is reminiscent of the appropriation and random exchange of jewellery in Milet’s studio, where cultural symbols were exchanged between models – who may well have not even been sitting for their photographs of their own free will – without considering their original usage or meaning.
Basing his observations on these and other images generated during the first phase of the project, the artist and researcher Tadeu Kaingang singled out one of the key elements of the representation of Indigenous Peoples in Brazil: the cocar or headdress. In these AI-generated images, he argued, the headdresses seemed to be more like those used in Maya culture and not at all like those used by Indigenous Peoples in Brazil. This observation highlights how AI, because it lacks a contextual understanding of cultural symbols, recombines them seemingly randomly, generating representations that are not just incorrect but which can contribute to the reproduction of stereotypical ways of seeing Indigenous Peoples.
The pose and the screen
Alvarado also highlights that, in these early images, the poses adopted by the subjects of Mapuche origin did not really reflect their normal behaviour, cultural forms of expression or surroundings, but rather those demanded by the photographer. The Mapuche subjects are participating in the photography studio’s process of mise-en-scène, with poses and expressions imposed on them in order to correspond to an aesthetic vision that was quite foreign to their way of being (2001, 23). For example, in the images reproduced in Fig. 13.3, the conventions of mise-en-scène are easily identified. The image of the Longko, taken in Fernando Valck’s studio in 1880, portrays the Mapuche chief posing with a solemn attitude, dressed in a traditional poncho combined with Western accessories, such as the boots with spurs and the hat we see on the floor, in this case against a completely neutral background. And in the image of the Mapuche women and children, taken in Gustavo Milet Ramírez’s studio (ca. 1900), the figures are positioned on either side of a gnarled tree trunk, some standing still and staring impassively straight at the camera so their faces can be clearly seen, while others are bent over, engaged in food preparation and thus referencing what purports to be a typical scene of daily life. Nonetheless, they are all positioned quite improbably in front of a screen depicting a classical plinth, climbing plants and trees fading into the mist that undoubtedly served for other more high-society-type portraits.
Fig. 13.2. Thea Pitman, images of Amazonian and Andean women and children generated with Midjourney v.5.1, May 2023, as tests before starting the AIAI project. Prompts: ‘amazon_woman_child_craftwork’ and ‘andes_woman_child_weaving’, respectively.
Fig. 13.3. Left: Mapuche Longko [chief], photograph by Fernando Maximiliano Valck Wiegand, Santiago de Chile, 1880. Museo Histórico Nacional, Santiago de Chile. Collection no. Fb-4363. Right: Mapuche women and child preparing food, posing against screen, Gustavo Milet Ramírez’s studio, Traiguen, Chile (ca. 1900). Nationaal Museum van Wereldculturen, Leiden, the Netherlands. Collection Wereldmuseum, no. RV-A299-14.
Fig. 13.4. Lucian de Silenttio, Man on Mountain Top, image generated with Midjourney v.5.1, May 2023, in the first phase of the AIAI project. Prompt: ‘Man on the highland mountain, dressed in a red and black striped poncho, next to him a golden dog reclining like a sphinx, in the background a city between mist and smoke coming out of buildings. Realistic, 4k photography’.
This form of photographic representation, far from disappearing over time, is now reproduced in images generated by Artificial Intelligence. An example can be seen in Fig. 13.4 generated by Lucian de Silenttio, a Bolivian interdisciplinary artist and educator of Aymara origin, during the early phase of the AIAI project. Lucian’s intention was to reproduce the view of the urban landscape of his everyday life in the city of El Alto, overlooking La Paz, Bolivia; an area where the presence of an urban form of Aymara culture predominates. This landscape is seen as a pictorial backdrop against which the figure of the character stands out.
Of the four variations generated by Midjourney, Lucian preferred the one at the bottom left because the person has long hair, although he commented that he would have liked them to exhibit more Indigenous features. However, without including any reference to Indigeneity or geographical location in the prompt, there is virtually no chance that such phenotypical features would appear in an image generated by Midjourney v.5.1, indicating the extremely sparse presence of Indigenous Peoples from the region in the database used and the lack of any shadow-prompting for specific forms of diversity included by the designers.1 Another of Lucian’s observations was about the poncho, a traditional Aymara woven garment. In the generated image, it looks like a raincoat or cape. In a personal communication, Lucian commented, ‘Es que [la IA] no conoce’ [The AI simply doesn’t know what it is]. He also noted that he found all of the images to be ‘muy convencionales’ [very conventional], especially the top-right image.
With respect to the question of the pose, in the old black-and-white photographs, the rigidity and immobility of the figures was due to technical requirements, as the chemical processes involved in photography at that time required the sitter to remain motionless for several seconds and thus hampered the representation of naturalness. We clearly see here that, although this technical limitation no longer exists, the rigidity of the pose is still reproduced in the generated images due to its strong presence in the images that make up the database.
The ethnic scene
Tadeu Kaingang also observed the construction of what we might call ‘ethnic scenes’ in the images that he generated as part of the first phase of the AIAI project. According to Alvarado, the traditional ‘escena étnica’ in photography was intended to convey ‘una atmósfera que tiene el fin de producir una ambientación elaboradamente dramática’ [an atmosphere that aims to produce an elaborately dramatic setting] (2001, 25) in which an attempt is made to create the illusion of a real landscape and of nature, as well as a seemingly truthful representation of an exotic and distinct cultural reality.
This type of mise-en-scène reappears in AI-generated images. An example is the prompt created by Tadeu Kaingang: ‘A family of five South American Kaingang indigenous people reforesting a patch of forest while under the watchful gaze of ancestral nature gods in supernatural form’ (see Fig. 10.2). The generated image omits any obvious representation of the ‘ancestral nature gods’ explicitly mentioned in the prompt, showing instead a forest pieced together from generic elements, and it is also notable that the clothing does not correspond to Kaingang culture, despite the reference to that culture being clearly stated in the prompt. The frontal gaze and rigid pose also reappear. As Tadeu commented during the focus groups, these are ‘estereótipos, onde não conseguimos enxergar a pessoa indígena contemporânea’ [stereotypes, where we can’t see contemporary Indigenous people]. These images of Kaingang people reminded Tadeu of illustrations made by Joaquim José de Miranda in the eighteenth century that were ‘desenhados no estilo do conquistador; eles eram povos indígenas completamente diferentes’ [drawn from the perspective of the conquistador and looked like totally different Indigenous Peoples], and they stand in clear contrast to documentary images of the Kaingang that can be found in the Instituto Socioambiental’s Povos Indígenas no Brasil [Indigenous Peoples in Brazil] encyclopedia (2018), which show Kaingang families in their real everyday settings, far from the artifices of a photographic studio.
During the focus group sessions in the first stage of the AIAI project, several participants expressed difficulty recognising themselves or their communities in the generated images. Mariela Tulián, Casqui Curaca [community leader] of the Tulián people, wanted to represent herself as a female jaguar surrounded by carob trees, as she did not imagine herself as a person so much as a big cat, an important figure for the people of the Gran Chaco. However, Midjourney generated an image that did not fulfil her expectations, making it difficult for her to identify with the jaguar or recognise the carob forest. This case brings us to the question of Indigenous self-representation, which is not always centred on the individual self but rather on their relationship with the environment. As Lucian de Silenttio pointed out in a personal communication: ‘Las montañas, los árboles y hasta los animales son sujetos’ [mountains, trees and even animals are subjects]. This vision challenges the anthropocentric logic that underlies both AI and traditional visual representations.
Under-representation
A further observation about the lack of representativity in the generated images was made by Mapuche weaver and researcher of natural dyes and medicinal plants, Loreto Millalén, who, in reference to her research practice, created this prompt in Spanish: ‘La gente se identifica cada una con una planta y trabajan en equipos para realizar actividades que son consideradas medicina’ [Each person identifies with a plant and they work in teams to perform activities that can be considered medicine]. The resultant image (Fig. 13.5) shows a white man dressed in a white coat with a stethoscope, thus evidencing the association of the word ‘medicine’ with the figure of a male Western doctor. In the various other images generated in relation with this prompt, and with no information included in the prompt itself about aesthetics, the images are reminiscent of cartoons or caricatures, completely removed from the symbolic, spiritual and communal complexity of traditional medicine that Loreto had in mind.
Fig. 13.5. Loreto Millalén, People Working with Medicinal Plants, image generated with Midjourney v.5.1, May 2023, in first phase of the AIAI project. Prompt: ‘La gente se identifica cada una con una planta y trabajan en equipos para realizar actividades que son consideradas medicina’ [Each person identifies with a plant and they work in teams to perform activities that can be considered medicine].
Lucian de Silenttio, who has gone on to generate a great many images since the time of these first focus groups, has stated forcefully:
No es posible que las Inteligencias [Artificales] nos permitan auto-representarnos a nosotros como indígenas … porque la Inteligencia Artificial es un programa que utiliza una base de datos determinada, y solo puede responderte recurriendo a esa base de datos en el que no estamos incluidos.
[It’s impossible for AI tools to allow us to self-represent as Indigenous people … because an AI tool draws on a specific database, and it can only respond to you by drawing on that database, in which we’re not included.]2
Experiences in which AI-generated images result in heteronormative and racially standardised representations have already been observed by social scientists and activists, highlighting the fact that models trained with data extracted from the web fail to adequately represent all of humanity.3 This is not accidental but structural. Massive databases perpetuate, reproduce and enhance stereotypes generated by a dominant culture, and along with them, the logics of exclusion and domination where the flourishing of other visions and narratives is suppressed.
All of this suggests that behind the user-friendly interface of AI tools, biased logics persist, inherited from the colonial history of representation. While it was possible to work harder to come up with prompts that produced results closer to the desired self-representation, this process involved a double-translation effort for non-English speakers: translating very different cultural imaginaries into conventional European languages, such as Spanish or Portuguese, and then again into English, the only language that the ‘machine’ could really work with at the time. This was a double translation that forced us to translate ourselves in order to exist in a system that did not represent us.
That said, there is always an exception to any rule. Aymar Ccopacatty, an Aymara artist from Peru who moved to the USA as a child, drafted a prompt directly in English that included specific details about ethnicity, geographical location and local flora, and the results actually exceeded his expectations. In the generated images (Fig 13.6, top) Aymar recognised a distinct resemblance to photographs he has of his great uncle Cicilio Ccopacati Quispe who is from the Qullana Socca community in Puno, Peru (Fig 13.6, bottom). Furthermore, perhaps because of the focus on representing time in his prompt, albeit slow, the generated images are much more dynamic than the static pose of the original photographs that more closely relate to the conventions of traditional portrait photography. We can, of course, also speculate on geographical and temporal distance making Aymar more predisposed to experiencing an affective response to the generated images in comparison with the other Indigenous artists involved in the project. Nonetheless, this one exception does suggest that, for those who are so inclined and who can bypass linguistic difficulties with text prompts, there are interesting possibilities to be explored in AI image generation in relation to Indigenous Peoples.
Fig. 13.6. Top: Aymar Ccopacatty, Aymara Grandfather Walking by Lake Titicaca, set of four images generated with Midjourney v.5.1, May 2023. Prompt: ‘an ancient Aymara grandfather, walking his farm lands on lake titicaca, high altitude, time is slow, slow slow down, the green shore and the blue depths of the lake shimmer. His crops are Potato, Avas, Quinua, Jirwa, Achachilas guard the snow capped glacier mountains far away, slow slow slow down, an Apacheta spring gurgles its sacred and mystical life giving water that flows downhill towards the lake’. Bottom: Cicilio Ccopacati Quispe, Puno, Peru (ca. 1980). Photos: Aymar Ccopacatty.
Notes
1 As noted in Chapter 5, Midjourney v.5.1 and subsequent versions used for this book have reverted to an algorithm that returns results with less ethnic diversity, after the very brief experiment of increasing ethnic diversity in v.5.0.
2 Quotations from the focus groups held virtually in May 2023 are all taken from aruma’s notes. While excerpts of the focus groups are available in Sebastián’s short videos (2023–24b, also available on Manifold – read.uolpress.co.uk/projects/artificial-intelligence-art-and-indigeneity), recordings of the whole focus groups are not.