Provenance is “the fact of coming from some particular source,” as the Oxford English Dictionary puts it (OED, “provenance” #2, n.d.). Others define provenance just as “origin” (Merriam-Webster, “provenance” #1, n.d.). It seems simple: provenance is where something comes from. Yet, it is a term that gets animated in particular and contested ways, for particular and contested ends, according to context. Provenance for art, appraisal and museology refers to a piece’s history of ownership, a means to verify its authenticity and often its value (OED, “provenance, #3, n.d.). In archaeology, it refers to the location of an object at the time of excavation, marked by spatial coordinates (Millar Reference Millar2002). It is at the core of archival theory (SAA 2026), or as Jarrett Drake has written (Reference Drake2016, para. 3), “perhaps the most sacred principle in the archival field,” as it is concerned with the history of records and structures, archival arrangement and description. More recently, provenance is a theme in the wealth of scholarship addressing machine learning (ML) and artificial intelligence (AI).Footnote 1 Here, it has largely referred to data provenance, or where the training data for these models originates.
Provenance’s activation in these domains does not mean that the concept has stayed put in each area. There is a history of cross-field exchange. Museums’ 21st-century re-examination of collecting practices has caused them to draw on concepts from ethics, law, anthropology and sociology to inform their approach to provenance (Feigenbaum and Reist Reference Feigenbaum and Reist2012). Long-standing notions of archival provenance have been altered through input from other fields, including disability studies, decolonial thought, museology and archaeology (Brilmyer Reference Brilmyer2022; Millar Reference Millar2002; NYPL 2024). Data provenance frameworks developed in the context of AI and ML have incorporated perspectives from libraries and archives, feminist ethics and even the food industry (Holland et al. Reference Holland, Hosny, Newman, Joseph and Chmielinski2020; Jo and Gebru Reference Jo and Gebru2020; Luccioni et al. Reference Luccioni, Corry, Sridharan, Ananny, Schultz and Crawford2022).
This contribution focuses on the exchange that exists, and that could be built, between archival provenance and data provenance from AI/ML. I ask how insights about provenance developed more recently in AI could also return to inform archival practice. To set the ground for this exchange, I trace frameworks related to provenance, first from archival studies and then from AI/ML. I ask how these frameworks overlap and diverge, and also how they have “flowed,” or been assimilated and adopted into the other field. I show how archival theories have been animated in documenting data in the context of AI and ML (Gebru et al. Reference Gebru, Morgenstern, Vecchione, Vaughan, Wallach, Daumé III and Crawford2021; Goebel et al. Reference Goebel, Chander, Holzinger, Lecue, Akata, Stumpf, Kieseberg, Holzinger, Holzinger, Peter Kieseberg and Weippl2018), and then reverse that flow of knowledge, arguing that a practice of provenance developed in critical research on AI, especially around publicly accessible data documentation, could inform provenance in archives, as well.
1. Origins
This contribution has its own origin story – its own provenance. Like most scholarship, one origin point lies in the literature of the field, emerging from archival studies and responsible AI. The other origin lies in two conversations: one with a computer scientist, who challenged qualitative scholars to explore epistemological movement between fields, and the other with a literary scholar and archive user, who wondered why the origin stories of archival collections were not more visible to the researchers using them.
The first conversation happened years ago, when I was interviewed for a role where I would be actively collaborating with computer science researchers building AI tools. The interviewer asked a question that I, a critical and qualitative social scientist, struggled with. Often, the computer scientist noted, qualitative scholars analyze and critique computer science methods. She was not opposed to this paradigm. However, she asked: in what ways could critical or qualitative insights be used to forward the technical goals of systems she and her team were designing? The question has stuck with me since, and not just because I did not get the job. It at once raised seemingly unresolvable tensions about ways of knowing and the ethics of quantification (D’Ignazio and Klein Reference D’Ignazio and Klein2020; Drucker Reference Drucker2021). Yet, it also observed the direction in which research insights flow when AI is in conversation with qualitative fields, surfaced questions about why insights flow that way and opened space for thinking about what it would look like if they flowed in alternative directions.
The other conversation took place a couple of years later, during a Q and A session. I had presented scholarship based on interviews with archivists from a variety of US institutions, and how they had conceptualized and built rapid-response archival collections. A literary scholar, who used archives for research, brought up something about the interviews that surprised him: how conscientious the archivists were in forming these collections, and how attuned they were to issues of power and representation in conceptualizing and processing these materials. He had never had the story of a collection told to him when he used archives for his scholarship: how the materials arrived there, how they got to be that way, and what processes they had been subject to in making their way to folders in a box on a shelf with a finding aid.Footnote 2 While theoretically well aware of the “constructedness” of the documentary record, he had not had the provenance of a collection, broadly understood, presented to him. More importantly, he had never had the provenance of the collection explained to him through the voices of the archivists who had worked on it.
These cross-field and cross-method conversations, one on epistemological movement and the other on the visibility of an archival collection’s origin story, are the starting points that led here. These conversations joined existing threads I was following from research and pedagogy on socio-technical systems and archives. From previous scholarship, I was familiar with the paradigms being developed for documenting training data provenance in ML and AI – paradigms that emphasized standardized documentation about how datasets and models are constructed, which are intended to communicate with groups using that dataset or model. At the same time, I was teaching and researching within archival studies, observing the ways that archivists grappled with multiple decisions, constraints and opportunities when developing a collection, and the relative invisibility of that work to the everyday user of archives. I wondered if knowledge being developed in studies of AI could move to address that invisibility. Could frameworks from AI ethics flow in such a way to inform the conceptualization, practice and visibility of archival provenance?
2. Provenance in archives
In archives, the principle of provenance developed in Europe in the 19th century so that archival repositories could arrange records in a way that reflected the organization of a collection as it was acquired, known as respect des fonds (Bailey Reference Bailey2013; Brilmyer Reference Brilmyer2022). Rather than arranging records by other organizational principles like category or community, the traditional guidance on provenance dictated that records made by separate creators should be kept separately, and that the original order of these records be maintained (Caswell Reference Caswell2016; NYPL 2024). Since then, provenance has remained one of the foremost principles of the archival field, with the Society of American Archivists characterizing it either as “the origin or source of something,” or “information regarding the origins, custody and ownership of an item or collection” (SAA 2026).
Provenance’s traditional instantiation has been challenged and rethought within archival studies, especially over the last thirty years, as postmodern archival theorists showed how archival work shaped the meaning of records (Caswell Reference Caswell2016; Douglas Reference Douglas2016), and frameworks from diverse social movements, including decolonization, anti-racism and feminism, highlighted the traditional archive’s reproduction of dominant power relations (Bastian et al. Reference Bastian, Griffin and Lowry2024; Brilmyer Reference Brilmyer2022; Cifor and Wood Reference Cifor and Wood2017; Ghaddar Reference Ghaddar2025; Lapp Reference Lapp2023; NYPL 2024; Rayan Reference Rayan2024; Wurl Reference Wurl2005). Provenance was also rethought because the principle of respect des fonds did not cleanly reflect what happened in archival practice (Bailey Reference Bailey2013; Brilmyer Reference Brilmyer2018; Reference Brilmyer2022; Fenyo Reference Fenyo1966; Millar Reference Millar2002; Trace Reference Trace2020). These contributions argue that changing traditional Western notions of provenance may change how power relations are expressed in an archive and alter the ways in which collections reflect a given subject. Moreover, because provenance functions as both a principle and a practice, an evolution in the conceptual scoping of provenance affects the hands-on work related to archival arrangement and description (Maemura et al. Reference Maemura, Worby, Milligan and Becker2018), or, as Wurl (Reference Wurl2005, 67) has put it, alters how an archivist “confronts a body of archival information on a processing table.”
One theme emerging in the recent rethinking of archival provenance is the use of the concept not as a straightforward path to arrangement and description, but as a means of recognizing and reflecting the contextual construction of an archival collection (Bettivia et al. Reference Bettivia, Cheng and Gryk2022; Brilmyer Reference Brilmyer2018; Douglas Reference Douglas2016; Duff and Harris Reference Duff and Harris2002; Millar Reference Millar2002). This recognition includes transforming provenance into a foundational theory that better reflects the power dynamics inherent in record keeping, including encompassing multiplicity and multi-temporality, oppression and imagination, in the maintenance of cultural memory (Rayan Reference Rayan2024). In doing so, archival theorists propose a vision of provenance that recalls its foundational definition, if not as “the fact of coming from some particular source” (OED, “provenance,” #2) then as the story of coming from some particular source (or sources). Laura Millar (Reference Millar2002), for instance, draws on museology and archaeology to highlight traditional archival provenance’s limited scope. Looking to paradigms from these fields expands provenance to encompass (ibid, 12):
… not just the creation of the records but also their history over time and [archivists] role in their management. The question we need to ask is not “how did the records come to be?” The question, rather, is “how did these records come to be here?”
Jennifer Douglas (Reference Douglas2016) similarly argues that current modes of description obscure the “constructedness of the fonds,” painting the archive as an “un-self-conscious and more or less spontaneous output of its creator.” These perspectives evolve provenance conceptually, which could theoretically lead to practice-based evolution. As Douglas says, the conceptual evolution of provenance attempts to show the gap between “what is done and what could be done,” if not outlining the steps to close that gap. That is, what is often left implicit in this conceptual evolution are the mechanics through which this story, this “constructedness,” might become visible to an archival user – to become visible to the literary scholar from the Q and A, for instance.
3. Provenance in AI and ML
Discussions on provenance from ML and AI overlap and diverge from the centuries-long discussions of provenance in archives. Like museology and archaeology (Anderson Reference Anderson2024; Millar Reference Millar2002), approaches to provenance from AI/ML domains can evolve archival approaches to provenance. Within these domains, there is a robust, ongoing conversation around the importance of documenting training data, the models they create and the outputs from these models (Gansky and McDonald Reference Gansky and McDonald2022). Documentation in all stages of this pipeline is meant to promote transparency, or the ability to look into, understand, and, therefore, govern these complex systems and their socio-technical entanglements (Ananny and Crawford Reference Ananny and Crawford2018).
In AI/ML conversations, provenance typically refers to “data provenance,” or documentation about the “creation and use of datasets” for training computational models (Gebru et al. Reference Gebru, Morgenstern, Vecchione, Vaughan, Wallach, Daumé III and Crawford2021, 86). Researchers in the AI bias field have urged practitioners to “develop standards to track the provenance, development, and use of training datasets throughout their lifecycle,” such that computer and social scientists could better monitor issues in model bias and representation (quote from Campolo et al. Reference Campolo, Sanfilippo, Whittaker and Crawford2017; Gebru et al. Reference Gebru, Morgenstern, Vecchione, Vaughan, Wallach, Daumé III and Crawford2021; World Economic Forum 2018). In a literal sense, data provenance helps show the constructedness of datasets, revealing some of the organizational, cultural and technical mechanisms through which data are “cooked” (Gitelman Reference Gitelman2013). Documenting where data come from can also serve technical goals, for instance, so that researchers are able to monitor drift and contamination in datasets (Luccioni et al. Reference Luccioni, Corry, Sridharan, Ananny, Schultz and Crawford2022).
Work related to provenance in these domains has generated numerous publications, groups and frameworks (Gebru et al. Reference Gebru, Morgenstern, Vecchione, Vaughan, Wallach, Daumé III and Crawford2021; Holland et al. Reference Holland, Hosny, Newman, Joseph and Chmielinski2020; Hutchinson et al. Reference Hutchinson, Smart, Hanna, Denton, Greer, Kjartansson, Barnes and Mitchell2021; Pushkarna et al. Reference Pushkarna, Zaldivar and Kjartansson2022). Among the most well-known frameworks for documenting data provenance is “Datasheets for Datasets,” created by an interdisciplinary team of computer and social scientists (Gebru et al. Reference Gebru, Morgenstern, Vecchione, Vaughan, Wallach, Daumé III and Crawford2021). Datasheets, as I abbreviate this framework henceforth, represent a workflow for dataset creators to reflect on the creation, distribution and implications of a dataset. But its purpose goes beyond internal documentation: rather, it is a workflow that creators document so that the information can be communicated to dataset users, such that they “have the information they need to make informed decisions about using a dataset” (ibid, 87). As the authors note, “transparency on the part of dataset creators is necessary for dataset consumers to be sufficiently well informed” (ibid, 87). Datasheets consist of up to 57 questions, including sections on the motivation behind building the dataset, the composition of the dataset, how it relates to potentially sensitive human data, the collection process, data processing, how the dataset should be used, how it will be distributed and how it will be maintained. Related frameworks, like FactSheets (Arnold et al. Reference Arnold, Bellamy, Hind, Houde, Mojsilović, Mojsilović, Nair, Ramamurthy, Olteanu, Piorkowski, Reimer, Richards, Tsay and Varshney2019), Data Cards (Pushkarna et al. Reference Pushkarna, Zaldivar and Kjartansson2022) and Dataset Nutrition Labels (Holland et al. Reference Holland, Hosny, Newman, Joseph and Chmielinski2020), have encouraged provenance documentation similar to Datasheets while emphasizing legibility, conciseness and “user friendliness” to the dataset user. That is, these frameworks have further stressed the need for documentation of data provenance to serve as a communicative bridge to the dataset user. This appears to be a difference between provenance from AI and provenance from archives, and a difference that speaks to the concerns of the literary scholar who felt he had never heard the story of the collections he consulted. Through the work of AI bias scholars who have drawn on frameworks from archives to develop these concepts (e.g., Jo and Gebru Reference Jo and Gebru2020), provenance in AI has emerged as a reflective and communicative practice, if not necessarily a foundational one for the broader AI field. For archives, provenance is a foundational and long-standing practice, but traditionally a structural one informing arrangement that may only implicitly present itself to the archival user.
At the same time, the reconceptualization of archival provenance as a collection’s history and management over time addresses similar concerns to those of computer scientists when they first started conceptualizing data provenance (Millar, Reference Millar2002). Data provenance was first discussed by computer scientists working on database systems at the beginning of the 21st-century. In 2000, computer scientists Bunneman, Khanna and Tan outlined foundational issues related to scientific datasets that were becoming widely available on the web. They asked (87): “When you find some data on the Web, do you have any information about how it got there?” They offered data provenance as a term that could address scholars’ need to know about this data, conceptualizing it as “a description of the origins of a piece of data and the process by which it arrived in a database” (ibid, 88). While speaking to different audiences, their descriptions of provenance – and the necessity of documenting it – closely resemble how archival scholars like Millar advocate for a conceptual evolution of archival provenance. Millar (Reference Millar2002) argues for a provenance that asks, “How did these records come to be here?” Bunneman and his co-authors argue for data provenance because, when you encounter a piece of data, “you would like to know how it got there” (Reference Buneman, Khanna and Tan2000, 87). Strategies for addressing Millar’s question may draw on discussions of provenance that first emerged in these discussions of database systems and have since been taken up in conversations about AI.
4. Exchange: toward more transparent provenance for archives
This implicit overlap between data provenance and archival provenance is not the only time these domains have spoken to each other. In my suggestions for archival practice to look to AI/ML provenance frameworks, I build on this history.Footnote 3 When they do speak to each other, it is more common for insights about provenance to flow from archival studies to AI/ML. Scholars have noted how archival studies, and library and information science (LIS) generally, can inform the responsible construction and especially documentation of AI and ML systems. For instance, in their well-cited “Lessons from Archives” (2020), Eun Seo Jo and Timnit Gebru look to paradigms and practices from archival studies – including mission statements, appraisal processes, consortia and investing in community archives – to inform more ethical data collection, annotation and use in AI and ML. Drawing attention to parallels between collecting for archives and data collection for AI, they argue that the ML community should “take lessons from other disciplines that have longer histories of addressing similar concerns,” like archival studies (Jo and Gebru, Reference Jo and Gebru2020, 307). In my own collaborative work on documenting ML dataset deprecation, we have referenced archival studies work to advocate for a more serious consideration of the afterlives of ML and AI training data (Luccioni et al. Reference Luccioni, Corry, Sridharan, Ananny, Schultz and Crawford2022). Safiya Umoja Noble’s noted research on search algorithms, Algorithms of Oppression (Reference Noble2018), has influenced thinking on automated decision-making systems in part by drawing on her background in LIS, arguing that critical perspectives from LIS can inform more ethical information engagement, within and beyond AI.
To reprise the interview question that the computer scientist asked me, what happens if that epistemological flow is reversed – if knowledge that tends to move in a particular direction moves another way? What if insights about provenance developed in AI and ML contexts are brought back to the practice of provenance in archives? AI and ML provenance takes the form of a documentary practice that prompts reflection from dataset creators and is intended for communication to dataset users. The kind of provenance developed in AI and ML could be a model for a more public and perhaps transparent form of provenance in archives. It could be a model that actualizes the provenance described by archival scholars like Douglas (Reference Douglas2016) and Millar (Reference Millar2002), one that embraces a collection’s history and story, and emphasizes that story as one that is also told to an archival user.
Evidence is emerging of what this could look like. Emily Maemura and Helena Byrne (Reference Maemura and Byrne2024) have brought the Datasheets framework to web archiving, running workshops with web archivists and developing a toolkit based on the original set of 57 questions. The adapted Datasheet focuses on asking web archive creators to address questions about how data were collected and processed. This kind of epistemological experimentation could be expanded toward archives that are less traditionally seen as datasets (like physical collections), and especially to push in the direction of communicating provenance to the archival user.
Looking to provenance practice from AI/ML in archives could take the form of archivists adapting frameworks like Datasheets as they accession, describe and arrange a collection, as it does in Maemura and Byrne’s work (Reference Maemura and Byrne2024). While that instantiation emphasizes processing, there may be value in archival users hearing from archivists about the general “motivation” behind a collection, to borrow the first section from Datasheets, or why, how and for whom a collection was built. Documentary frameworks like this could also be adapted to incorporate ideas surfaced in the archival literature, including descriptions of the literal site where the documents were originally used (Lehane Reference Lehane2012). These documents could be publicly available and circulated to researchers who plan to engage with a collection, much like finding aids are. Perhaps provenance could look like an infographic, similar to Dataset Nutrition Labels: an at-a-glance, user-centered description of a collection and how it came to be. Perhaps this epistemological exchange could come in even less structured forms, for instance, inspiring organizations to tell users about the process – contested, messy and complicated as it might be – of building their collections. What this epistemological movement from data documentation frameworks to archival provenance might risk, of course, is flattening the robust, critical and imaginative ways that archival provenance has been rethought in recent years. At its most basic, though, having AI and ML provenance inform archival provenance is about moving provenance to incorporate communication and a considered visibility of the collection’s story for an archival user.
Beyond telling the story of a collection’s accessioning and processing, I wonder about the ways that this type of documentation could foster greater connection between archivists and the archival user. Michelle Caswell (Reference Caswell2016) has considered the humanities’ theoretical engagement with “the archive” and discussed how archival theory is often missing from these discussions but could enrich them. One way to meet this opportunity is to make the constructedness of the archive visible to scholars who rely on it, using provenance to do this work. This is to say that frameworks from AI and ML might push archival provenance closer to this vision and, in the process, closer to its expansive original definition. After all, the word provenance comes from the French word provenir, meaning “to come forth, arise” (OED, “provenance”). Perhaps frameworks from AI and ML might do that for archival provenance, acting as a means of making the archivist’s vital role in shaping collections more visible.
Acknowledgements
I am grateful to the interlocutors from the University of Pennsylvania’s Center on Digital Culture and Society Colloquium series, who helped inspire this article, and to the three anonymous peer reviewers for their thoughtful and generative feedback.
Funding statement
The author received no financial support for this article.
Competing interests
The author declares none.
Frances Corry is an Assistant Professor in the Department of Information Culture & Data Stewardship at the University of Pittsburgh. Her work employs critical-historical approaches to information, examining the prehistories and afterlives of data-intensive systems – from social media platforms to AI tools.