1. Introduction
Numerous scholars have argued for studying the development of artificial intelligence (AI) systems across disciplines and in accordance with ethical guidelines and principles (e.g., Hine and Barnaghi Reference Hine and Barnaghi2024; Mittelstadt Reference Mittelstadt2019). To achieve this, (Critical) AI and ethics research has increasingly called for interdisciplinary collaboration between Computer Science and the Social Sciences and Humanities (SSH) (e.g., Raji et al. Reference Raji, Scheuerman and Amironesei2021; Sloane and Moss Reference Sloane and Moss2019), and there is a strong need for interdisciplinary AI expertise in domains such as medicine (e.g., Rostamzadeh et al. Reference Rostamzadeh, Mincu and Roy2022; Thomas Reference Thomas2021). Bringing these disciplines together aims to “grapple with the wide[-]reaching impacts of AI on people” (Rakova et al. Reference Rakova, Yang, Cramer and Chowdhury2021, 6), including bias and the technological amplification of social inequalities (e.g., Obermeyer et al. Reference Obermeyer, Powers, Vogeli and Mullainathan2019; Thomas Reference Thomas2021). Interdisciplinary research thus approaches AI as a “complex, multifaceted” problem where monodisciplinary lenses are too narrow (Hendrickx and Smuha Reference Hendrickx and Smuha2023) while interdisciplinary approaches may lead to “innovative solutions” (ICAI 2024).
Approaching AI as a problem where “no single academic discipline can provide an adequate definition of such problems much less a clear and feasible resolution” (Wyatt Reference Wyatt, Werthner, Prem, Lee and Ghezzi2022, 329), I (the author) have been participating in a four-year interdisciplinary research project on Responsible AI for health decision-making at a Dutch University and its University Medical Center (hereafter, “the project”). Running between 2022 and 2026, this project brought together researchers in Computer Science (Medical AI Research), Humanities (Media Studies), Medicine (Cardiology and Intensive Care Medicine [ICU]) and Law (Health Law) and was specifically developed and funded through a university-internal funding structure to facilitate intensive interdisciplinary collaboration between these groups.Footnote 1 By bringing together these various disciplinary perspectives, the central aim was to foster more comprehensive forms of Responsible AI development and application in medical domains such as Radiology, Cardiology and Intensive Care Medicine. This included interdisciplinary emphasis on ethical, political as well as legal inquiry into AI development and application processes (see the “Project and analytical approach” section for a full project overview). This article mobilizes this project as a case to critically reflect on a series of day-to-day experiences throughout four years of interdisciplinary collaboration and focuses on the generative potential of the frictions, discomforts and disagreements that emerged within the project.
Despite considerable consensus and commitment of the team and across the university to steer interdisciplinary research for advancing “Responsible AI,” the project represents a case where the interdisciplinary collaboration was nevertheless still marked by disagreements and frictions as the different disciplines sought to collaborate. These frictions and disagreements occurred through the diverging expectations around the so-called “mode of interdisciplinarity” (Barry and Born Reference Barry and Born2013), relating to different types of research relationships and hierarchies across disciplines. As a junior scholar working at the intersection of Critical Data and AI studies and Science and Technology Studies (STS), I had, for example, first-hand experience with a so-called “service-subordination mode” (Barry and Born Reference Barry and Born2013, 11) of interdisciplinary collaboration in which one discipline would “make up for […] an absence or lack in the other, (master) discipline(s)” (Barry and Born Reference Barry and Born2013, 11). At the start of the project, my work in the humanities was expected to contribute to AI research and medicine as the two leading disciplines of the project in this manner.Footnote 2 Some of the AI research and medical colleagues assumed that my role in the project was to develop the ethical codes that would “locate the permissions and prohibitions of [AI’s] use” (Amoore Reference Amoore2020, 7) in health care. Even though my experience shifted over time, this also caused discomfort on my side, as it did not align with my initial expectations around the collaboration.
More specifically, frictions arose from fundamental differences in disciplinary research cultures, terminology and practices of knowledge production (see also Knorr-Cetina Reference Knorr-Cetina1999; Wyatt Reference Wyatt, Werthner, Prem, Lee and Ghezzi2022). As Moats and Seaver (Reference Moats and Seaver2019) argue, sets of “common-sense distinctions” (3) between, for example, AI and humanities research approaches, methods and conceptual understandings were taken for granted. This complicated efforts to move collaborative work beyond a division where AI researchers did “model development” and humanities scholars did “ethics,” and to instead generate unexpected research outcomes.
Over time, however, I also came to notice the productive potential of friction and disagreement in our collaborative work. Rather than treating moments of discomfort as obstacles, several team members engaged in disciplinary reflection and questioned the roles each discipline played. I came to recognize that such frictions could be generative. This insight connects to Moats and Seaver’s (Reference Moats and Seaver2019, 3) argument not to simply ignore or overcome issues emerging from differences in approaches, methods and concepts but to understand the relationships between fields more carefully. Such work helps surface distinctive conceptual articulations and epistemic commitments toward AI-related problems and the frictions these generate. Reflecting on four years of collective research on Responsible Health AI, this article contributes to critical scholarship on interdisciplinarity by presenting a novel conceptual approach. Using the project as an exemplary case, I argue that attending to the specific frictions and tensions that emerged around key AI concepts provide ways to generate new possibilities for (inter)disciplinary thinking in the field.
By examining interdisciplinary work in the project, this article focuses on specific moments of “disconcertment” (Verran Reference Verran2001) caused by interdisciplinary discussions on three central but interrelated concepts and their associated ethics debates: (1) bias, (2) ground truth and (3) error. According to postcolonial-STS scholar Helen Verran (Reference Verran2001), moments of disconcertment are common but often overlooked “fleeting experience[s]” (5) that grow from “seeing certainty disrupted” (Verran Reference Verran1999, 141) and may open new ways to think and act differently. I have chosen these three concepts for several reasons. First, they are central to Critical Data, AI and STS scholarship on the power and politics of AI development and its societal implications (e.g., Ananny Reference Ananny2022 on error; Henriksen and Bechmann Reference Henriksen and Bechmann2020 on ground truth; Jaton Reference Jaton2021 on bias and ground truth). In this literature, such terms are not treated primarily as computational measures, but as socially and materially constructed categories that shape AI knowledge production and legitimization. Second, throughout the project, I repeatedly encountered how these seemingly mundane terms carried divergent disciplinary meanings, requiring ongoing translational work between AI and medical researchers, legal scholars and our humanities practice. These differences generated friction. For example, “bias” is often invoked to address (potential) “injustice and harm produced by [AI] systems” (Miceli et al. Reference Miceli, Posada and Yang2022, 2). However, it is frequently framed as a system-based technical problem that can be mitigated without engaging underlying social dynamics (Galanty et al. Reference Galanty, Luitse, Noteboom, Croon, Vlaar, Poell, Sanchez, Blanke and Išgum2024).
Following Verran’s (Reference Verran2001) definition of disconcertment, I critically reflect on disciplinary discussions, misunderstandings and disagreements around bias, ground truth and error within the project. I focus on epistemic tensions and moments of disconcertment that arose from these disputes. I argue that such moments, emerging through discipline-specific approaches to bias, ground truth and error, can provide productive ways for participants to take action and create space to deepen interdisciplinary work on Health AI. This space enables reflection on epistemological differences and helps interrogate taken-for-granted research relationships between disciplines (e.g., medicine as the AI application domain). Building on this insight, I propose that moments of disconcertment can be mobilized as what Aradau and Blanke (Reference Aradau and Blanke2022, 155) call “little tools of friction” in interdisciplinary AI projects. Such tools offer new ways of practicing “doing difference together” (Verran Reference Verran2011, 422) and facilitate the creation of new connections (Smolka et al. Reference Smolka, Fisher and Hausstein2021). These little tools of friction may thus generate new possibilities for (inter)disciplinary thinking on AI, fostering new directions for interdisciplinary AI-based medicine and beyond.
Combining a Critical AI and STS research approach, the article is structured in three steps. First, I situate my contribution within critical literature on interdisciplinarity and examine how researchers have worked toward new research directions for challenging topics such as AI. This section also introduces “disconcertment” (Verran Reference Verran2001) and “little tools of friction” (Aradau and Blanke Reference Aradau and Blanke2022) as analytical lenses to attend to disciplinary tensions. Second, I discuss the project’s practical organization and explain the analytical approach taken toward moments of disconcertment around bias, ground truth and error as they appeared across three subprojects. The remainder of the article analyzes these moments and considers how disciplinary differences in the meaning of terms in Health AI research can be mobilized as frictional tools to open new avenues for interdisciplinary engagement.
2. Disconcertment and little tools of friction in AI and interdisciplinarity
Critical scholarship on interdisciplinarity has identified various ways in which researchers address specific disciplinary rules, methodological boundaries and knowledge subjectivities to explore new research directions for complex topics such as AI. Integration is of key interest in this literature, as it is considered both to underpin an interactive interdisciplinary research process and to constitute one of its major challenges, as there is never a one-size-fits-all approach to a project (Pohl et al. Reference Pohl, Klein, Hoffmann, Mitchell and Fam2021). To achieve sustainable, interdisciplinary knowledge integration, Bowker (Reference Bowker, Anand, Gupta and Appel2018) has argued that interdisciplinary projects demand three things. First, the objects of study as well as their underlying theoretical frameworks require “the triangulation and synthesis of multiple methodologies, both qualitative and quantitative” (207). Second, such projects call upon researchers’ abilities to consider multiple epistemic viewpoints. This means that scholars should be open to other knowledge practices and integrate them into their interdisciplinary projects. Third, interdisciplinary research requires scholars to interact in “trading zones” (Galison Reference Galison1997) – i.e., the spaces that build and sustain “a culture of [scientific] collaboration, facilitated by [partial] exchange of ideas, theories, beliefs, values, and data” (Pohl et al. Reference Pohl, Klein, Hoffmann, Mitchell and Fam2021, 23). These highlight that total agreement or shared foundational principles are not always necessary “for cooperation nor for the successful conduct of work” (Star and Griesemer Reference Star and Griesemer1989, 388). Others have pointed out that interdisciplinarity remains an interactive process in which participants conduct “boundary work” (Gieryn Reference Gieryn1983) and use “boundary objects” (Star and Griesemer Reference Star and Griesemer1989) to negotiate and collaborate in different forms, creating new languages and objects to interact around specific research problems and manage differences.
Contributing to this foundational work, this article reflects on the disciplinary tensions that emerge because the shared language for interaction that is necessary for operating trading zones and boundary work is missing and leads to moments of epistemic “disconcertment” (Verran Reference Verran2001). Disconcertment refers to “the sense of being put out in some way,” and when it is accompanied by the term epistemic, disconcertment “implies that our taken-for-granted account of what knowledge is has somehow been upset or impinged upon so that we begin to doubt” (Verran Reference Verran and Green2013, 144). Focusing on disconcertment by encountering differences in Nigerian versus English practices of quantification within postcolonial Nigerian (Yoruba) math education, Verran (Reference Verran2001) showed that disruptions of certainty occurred when these different epistemologies met in practice. This happened through the encounter of distinct knowledge traditions (Western math and Yoruba math) because they did not fit the specific lines of reasoning connected to Verran’s individual forms of knowledge production, shaped by postcolonial politics. Following Verran (Reference Verran2001), a specific kind of unsettling – disconcerting – experience can thus arise when someone encounters a practice of knowing that is coherent and effective on its own terms, but fundamentally challenges the categories, logic and assumptions of one’s own (politically) established ways of knowing and producing knowledge to understand the world. Building on this argument, I align with Hillersdal et al. (Reference Hillersdal, Jespersen, Oxlund and Bruun2020) and argue that it can also emerge by being confronted with different traditions of AI knowledge production in AI Research, Medicine, Humanities and Law as we encountered in our interdisciplinary project.
To pay attention to moments of disconcertment means to signal “different ways of knowing, and different ways of being in the world” (Smolka Reference Smolka2020, 13). This allows for recognizing the disciplinary differences that have so far been taken for granted. In this article, the recognition of difference means to be attentive to the tensions and disconcertments that occur as we participate in discussions, misunderstandings and disagreements around specific key concepts in interdisciplinary work on Health AI. At the same time, based on her research in Nigeria, Verran (Reference Verran1999, Reference Verran2001) has argued that even though it can easily be explained away as failure or inadequacy through stories of institutional power relations, attending to disconcertment may open new avenues for understanding other ways of knowing, and “might generate new possibilities for answering moral questions of how to live” (Verran Reference Verran1999, 136). I follow this argument in understanding that attending to such moments of disconcertment may facilitate seeing and creating new relations among different groups that work together in an interdisciplinary project setting. This is due to their disruptive qualities. Disconcerting moments – such as those involving the interpretation of key terms – can allow participants to become more sensitive to differences between disciplinary approaches within the field and to mobilize these differences in new, and possibly productive ways. Here, disciplinary differences should not be approached as problems to be solved but as conditions to be worked with. This process is what Verran has termed “doing difference together” (Verran Reference Verran2011), and what previous STS research has adapted to “facilitate collaboration between scholars and scientists from different disciplines who study the same research object” (Smolka et al. Reference Smolka, Fisher and Hausstein2021, 1081; Hillersdal et al. Reference Hillersdal, Jespersen, Oxlund and Bruun2020). From this perspective, I analyze the moments of disconcertment I identify through friction and disagreement around bias, ground truth and error within the project and show how they can be “a sure guide … in generating possibilities for new futures” (Smolka Reference Smolka2020, 4) in this rapidly developing research field.
Taking seriously that disciplinary difference should be a condition to be worked with in interdisciplinary collaboration, this article considers how the moments of disconcertment caused by difference in conceptual meaning and understanding of terms in Health AI can productively be mobilized as tools to create productive frictions, which possibly open new avenues for interdisciplinary collaboration. Here, I do not refer to tools in the form of material objects to help perform tasks such as gardening or scraping, but moments as tools in the sense of instrumental interruptions to normalizing disciplinary assumptions, make differences visible and open up spaces for reflection and potential reorientation. Originally developed through an analysis of practices of friction against military AI development and data extraction by big tech companies, such “little tools of friction,” Aradau and Blanke (Reference Aradau and Blanke2022) explain, can be understood as mundane but political “actions that slow down” (152) […] and “open up a democratic scene of dissensus” (153) around the production and implementation of AI technologies. Taking Google employee’s public petition against project Maven in 2018 as well as the organization of hackathons as examples, the authors show that little tools should not be seen as forms of action that completely stop the production of AI, but as small interventions that “slow down, hinder, or redirect the movements of technology” (Aradau and Blanke Reference Aradau and Blanke2022, 155). As such, they allow for the possibility of starting resistance through small interventions. Following this conceptual approach, I argue that the moments of disconcertment that I analyze in this article can be used as small, relatively mundane tools that can insert friction into interdisciplinary projects as they condition groups to actively pay attention to difference between disciplinary ways of understanding and doing research. They can be deployed for the possible disruption and resistance of existing practices, assumptions and taken-for-granted ideas of interdisciplinary collaboration and show how it can be done differently.
3. Project and analytical approach
This article draws and reflects on the experience of working in an interdisciplinary research project (2022–2026) on Responsible AI for health decision-making. To initiate collaborative scholarship on the topic, the core research team consisted of five principal investigators (PIs), five PhD-level researchers and two postdoctoral researchers. The PIs were based in Medical AI Research (2); Media Studies (1); ICU (1) and Health Law (1). The PhD-level researchers were based in Medical AI Research (2), ICU (1), Cardiology (1) and Media Studies (1), and the postdoctoral researchers were based in Medical AI Research (1) and Health Law (1). The team members worked on various subprojects, including the research and development of two medical AI applications. Application development focused on the research and design of two specific systems: (1) an AI-driven predictive application to guide anesthesiologists and intensivists in the choice and amount of coagulation agent (a substance to promote of blood clotting) in the care of patients that have a high-risk of bleeding after major surgery; (2) an application for AI-driven detection and prediction of atrial fibrillation (irregular heart rate) with ICU patients. Further individual and collaborative work addressed the ethics and politics of dataset and AI-system development in medicine more broadly, as well as a comparative case study on the implications of AI use in ICUs across Europe for decision-making and professional autonomy.
Alongside these projects and initiatives, the team organized biweekly AI-lab meetings consisting of presentations to foster knowledge exchange, domain perspectives and project ideas. As these meetings also allowed for broader interdisciplinary discussion on Medical AI system and application development and implementation, they functioned as the primary moments of interdisciplinary exchange with the entire project team. As such, they also emerged as key events where moments of disciplinary tension and epistemic disconcertment surfaced.
As a junior researcher from Critical Data, AI and STS, I initially felt somewhat of an outsider as I had no medical training and invested a lot of time learning about the different types of medical data and the algorithmic techniques used to develop AI applications in the field. Working my way into the project, I first noticed that tensions could arise through meeting discussions in which different research approaches to AI for specific domain applications were shared, and domain expertise rubbed against one another. For instance, a conversation on the quantifiability of “the most important” sources of error in training data would trigger tension and disconcertment. AI-research colleagues considered such errors to be indicators that AI systems are not yet “good enough,” but can be optimized using specific metrics and algorithmic techniques. For law colleagues as well as some medical researchers, however, AI errors were impossible to quantify in the first place, as they associated them with normative issues of individual patient safety and health equity. I approached AI errors as “sociotechnical constructs – as relationships between people and machines that have somehow failed, broken down, behaved in unexpected ways” (Ananny Reference Ananny2022, 3–4) whose impact should be critically scrutinized. Taken together, this lack of shared language and the unanticipated differences in how the central interrelated concepts of error, bias and ground truth were understood led these notions to become key sites of disconcertment that I discuss in this article.
I analyze moments of disconcertment, in relation to bias, ground truth and error through the lens of three “interdisciplinary stories” (Ananny and Hudson Reference Ananny and Hudson2025). Each moment relates to one exemplary subproject that has been part of the larger research priority area described earlier. Here, I generally refer to medical AI-research colleagues as AI researchers; colleagues in Cardiology and Intensive Care Medicine will be mainly referred to as medical researchers and team members in Health Law will hereafter be called law colleagues. As the stories are written from my position and perspective as a junior Critical Data, AI and STS researcher and collaborator in this project, the analysis is based on my (the author’s) personal experiences, individual reflective meeting notes and archived meeting materials such as power point presentations and minutes. It is, therefore, important to acknowledge that the reflective analysis provides a specific personal account of the interdisciplinary collaborative research. My analytic interest and orientation as a humanities researcher was shaped by particular sensitivities to epistemic difference, friction and disconcertment, which enabled particular insights that might not always be representative of the experiences of the overall group. Taking this personal analytical perspective, however, has also allowed me to provide particular in-depth insight into disconcertment and its potential as emerging from lived interdisciplinary practice within the project.
In analytically narrating the subprojects’ stories as the project is approaching its end, I follow Ananny and Hudson (Reference Ananny and Hudson2025) in their aim to not argue that these projects “all mean any one thing, that they offer a coherent story” (1773) or that they represent the full space of disciplinary difference in AI research in health care. The analytical stories presented in this article reflect on partial accounts of interdisciplinary collaboration, disciplinary difference and moments of disconcertment to challenge different disciplinary practices of knowing, and show how conceptual differences can be mobilized as generative “little tools of friction” for new pathways of interdisciplinary work in (Health) AI. However, these four years of collaboration did not go without failure and disappointment (see also Fitzgerald and Callard Reference Fitzgerald, Callard, Whitehead, Woods, Atkinson, Macnaughton and Richards2016; Mol and Hardon Reference Mol and Hardon2020). For example, a follow-up conversation on the issue of AI error discussed above led to a dead end on the potential for a project collaboration between AI researchers and lawyers on AI evaluation strategies, as their interpretation of error focused on different parts of the development and implementation process. For the AI researchers, this concerned the technical evaluation of AI systems according to certain metrics in the research and development phase, while for law, it concerned the clinical validation of AI systems that focuses on patient safety and tests an application for juridical implementation in the hospital. While such moments of failure in interdisciplinary work should certainly not be ignored in further research as many issues may stay unresolved (e.g., Callard and Fitzgerald Reference Callard and Fitzgerald2015), they fall outside of the analytical scope of this article. Instead, my approach in attending the selected situated narratives of disconcertment within the project builds on Law’s (Reference Law2004, 85) argument that “[o]ther possibilities can be imagined … if we attend to non-coherence.” It seeks to encourage researchers to act and reflect within and across moments of disconcertment – which may lead to dead ends – without fixing conceptual definitions, foreclosing disciplinary meanings or ignoring epistemic difference, to provoke new ways of thinking about AI-focused interdisciplinary research.
4. Doing bias differently
Like in many Responsible AI projects and debates, the notion of bias – defined in the field as “any systematic and/or unfair difference in how predictions are generated for different patient populations that could lead to disparate care delivery” (Hasanzadeh et al. Reference Hasanzadeh, Josephson, Waters, Adedinsewo, Azizi and White2025, 2) – played a key role within our project discussions. Particularly, the biases stemming from datasets developed in medical domains, such as radiology (Oakden-Rayner Reference Oakden-Rayner2020), ophthalmology (Khan et al. Reference Khan, Liu, Nath, Korot, Faes, Wagner, Keane, Sebire, Burton and Denniston2021) and cardiology (Noseworthy et al. Reference Noseworthy, Attia, Brewer, Hayes, Yao, Kapa, Friedman and Lopez-Jimenez2020), caught our interest early in the project, as data are the key infrastructure for AI research (Thylstrup Reference Thylstrup2022). The construction and maintenance of datasets introduce model biases, which may lead to discriminatory outcomes for certain patient groups (Arora et al. Reference Arora, Alderman, Palmer, Ganapathi, Laws, McCradden, Oakden-Rayner, Pfohl, Ghassemi, McKay, Treanor, Rostamzadeh, Mateen, Gath, Adebajo, Kuku, Matin, Heller and Sapey2023; Oakden-Rayner Reference Oakden-Rayner2020). For this reason, the project team focused on this specific subject area for one of its subprojects. Drawing on scholarly work that has highlighted the importance of dataset documentation to enhance data quality (e.g., Gebru et al. Reference Gebru, Morgenstern, Vecchione, Wortman Vaughan, Wallach, Daumé III and Crawford2021; Maier-Hein et al. Reference Maier-Hein, Reinke, Kozubek, Martel, Arbel, Eisenmann, Hanbury, Jannin, Müller, Onogur, Saez-Rodriguez, van Ginneken, Kopp-Schneider and Landman2020; Rostamzadeh et al. Reference Rostamzadeh, Mincu and Roy2022), we examined how publicly available medical datasets are documented and how this documentation – or lack thereof – impacts the detection and mitigation of AI biases for various medical tasks and domains. To do so, we developed a tool for comprehensive dataset documentation evaluation with the aim to support the responsible creation of new dataset documentation and the evaluation of existing dataset documentation. This tool included key information on data acquisition, curation, annotation practices as well as intended use and data limitations (Galanty et al. Reference Galanty, Luitse, Noteboom, Croon, Vlaar, Poell, Sanchez, Blanke and Išgum2024, Reference Galanty, Luitse, Vlaar, Sánchez, Blanke and Isgum2025).
While the subproject eventually resulted in a successful collective research article (Galanty et al. Reference Galanty, Luitse, Noteboom, Croon, Vlaar, Poell, Sanchez, Blanke and Išgum2024, Reference Galanty, Luitse, Vlaar, Sánchez, Blanke and Isgum2025), it was not developed without disciplinary friction and disconcertment around how to approach and possibly mitigate the problem of dataset bias, which meant different things to different team members. Drawing on their disciplinary assumption that technically “bias means a deviation from the standard” (Ferrer Reference Ferrer2021, np), a first key point of friction emerged as our dedicated AI-research colleagues discussed pursuing a quantitative interrogation into public data similar to Oakden-Rayner (Reference Oakden-Rayner2020). From this AI perspective, such a project would enable a critical focus on the statistical properties of the training datasets, measure their level of bias and – depending on the research outcome – call for a series of fairness measures and interventions to, for example, improve patient representation in data. Even though they could see the value of this type of quantitative approach to data scrutiny, the proposal triggered disconcertment and friction with medical colleagues. To them, the epistemic authority of AI researchers tended to align with only technical priorities and solutions, creating knowledge hierarchies in which the quantitative approach was framed as the most legitimate path forward. However, the medical scholars strongly disagreed with the narrow focus on “filling data gaps” or introducing fairness measures to attenuate the potentially negative consequences of biased data. Instead, their understanding of bias was more closely tied to the important legal and medical value of health equity (e.g., WHO, n.d.) rather than technical data representation, which is why they argued that health equity should be more prominently considered. From this standpoint, the medical researchers argued that it is important to recognize that, at times, deliberately constructing data biases may be necessary to benefit patients. Oversampling groups where diseases follow gendered or racialized patterns, for instance, can actively promote more equitable outcomes (see also Grote Reference Grote2025; Pot et al. Reference Pot, Kieusseyan and Prainsack2021).
In a different vein, the AI research proposal led to disconcertment among the humanities team members and their Critical Data and AI research practice. In this field, an increasing amount of scholarship has highlighted the limits of dominant technical bias- and fairness-oriented approaches put forward by AI scholars and called for more comprehensive analyses of social practices and power relations in AI development, including data production (e.g., Miceli et al. Reference Miceli, Posada and Yang2022). From this disciplinary perspective, a quantitative study of bias in the public datasets assumed the problem is to be located and mitigated through technical artifacts and processes. This would ignore and obscure the – to humanities research – important, but understudied, questions of how bias could have been inscribed in the data in the first place. In addition, it would overlook what authoritative disciplinary values, interests and power relations from AI research and medicine informed: “what counts as bias and what does not, what problems [this type of] debiasing initiatives address, and what goals they aim to achieve” (Miceli et al. 2022, 2). Attempting to challenge these dominant epistemic viewpoints led us in the humanities team to propose an alternative qualitative study. The suggested approach would focus on scrutinizing underlying practices and decisions of dataset creators – i.e., the choices that have been made in the construction of a dataset (Jaton Reference Jaton2017). Critical Data and STS has emphasized that data are inescapably biased (D’Ignazio and Klein Reference D’Ignazio and Klein2020.; Jaton Reference Jaton2021) and local (Timmermans and Berg Reference Timmermans and Berg1997), as it is shaped by factors such as the politics and power relations in institutional organization, computational infrastructures and means of individual collection and generation.
The team achieved consensus that individual practice and decision-making play a role in shaping (public) datasets. However, some medical colleagues quickly shared their doubts and disagreement about the value of a qualitative study on how data labels are produced. According to them, such a study would struggle to demonstrate scientific evidence and would offer few quantifiable results to strengthen its generalizability. These concerns reflected (implicit) disciplinary hierarchies within the subproject, in which quantitative and metric-driven forms of knowledge were assumed to carry greater epistemic authority. In their reaction, these medical colleagues positioned qualitative approaches as less rigorous and less actionable, generating friction that made it more difficult to include such perspectives. To these colleagues, it would therefore be effective to scrutinize a representative number of open datasets and evaluate the quality of annotations by using specific metrics such as Fleisch–Kappa to measure the variability between labels within a given dataset. This would be a good starting point for generalizable statistical intervention and bias mitigation. The discussion, in turn, prompted a moment of disconcertment, as we were confronted with the assumptions that variability in annotations and the bias that may emerge from it was something that can be located and statistically corrected from within the data itself by selecting the right metrics. Our medical colleagues deeply valued the medical domain’s “universality that is grounded in statistical inference […] made through an exact, quantitative weighing of the evidence” (Berg and Timmermans Reference Berg and Timmermans2000, 42). According to their standpoint, this would allow for the study’s results to be more generally applicable for different tasks and medical subfields. Even though they were not equally valued by everyone, either one of the proposals had its research potential, yet the difference in disciplinary understandings of data bias and mitigation as well as appropriate methods created difficulty to come together and agree on the project’s direction.
The discomforting conversations about our disciplinary hesitations and doubt regarding the project proposals, however, opened a generative discussion that focused on a slightly different and more practical question. Instead of seeking consensus on how to understand bias or which approach to adopt, my AI-research colleague and I, as project leads, decided to stay with these differences and frictions and to position ourselves as data users in practice. We did so after a conversation with one of the AI research PIs who was supervising the project. We located the research impasse and tasked the entire PhD-team (additionally consisting of two medical PhDs) to challenge ourselves and describe what we considered as most important to know about a public dataset intended for model development, while also accounting for the various biases that might arise from the data when training a model. This PI’s seemingly mundane and practical question functioned as a tool of friction and led to an intervention that forced us, the PhD team involved in the subproject, to “do difference” while responding to the task from our own disciplinary point of view.
The AI research PI’s intervention and my AI-research colleague’s and my response subsequently triggered the entire PhD-project group to act and explore another avenue for the subproject: data documentation. Following the development of existing guidelines for documentation over time led us to draw from these documents and develop a tool that would allow us to qualitatively investigate how existing documentation guidelines were being followed for existing public data, as limited research existed on that topic. Even though some of the team’s medical and AI research scientists were slightly uncomfortable, noting that their primary focus was on using data rather than documenting it for future users, they nonetheless acknowledged documentation as a useful tool to promote information sharing between data creators and users. Yet, others had to be convinced that documentation could enable critical reflection on practices of data collection, (pre)processing, distribution and maintenance, helping to detect potential bias risks and harms. This triggered friction, as the humanities scholars had assumed the implementation of reflexive practices in dataset construction to address bias had already been normalized. However, because much of this work has been developed in AI ethics, within fields like computer vision and natural language processing, it did not translate seamlessly into the team’s Medical AI Research practices. Being confronted by this mismatch became a moment of disconcertment, as it revealed how sensitivities toward reflexive forms of inquiry that feel established in some scientific communities may remain less considered in others through the uneven distribution of epistemic authority within interdisciplinary AI research projects.
Looking back, the disconcertments, the AI research PI’s intervention and the follow-up humanities–AI initiative to stay with, rather than resolve, disciplinary differences about bias and mitigation strategies enabled us to take up the task and “do difference” in a way that ultimately proved productive for the team. Throughout the design of the study on documentation of public medical imaging and signal data, the PhD team created space for collective discussion where we agreed our disciplinary “partial knowledges” (Suchman Reference Suchman2002, 94) could meet in parallel. My AI-research colleague and I, for example, negotiated including specific components in the data documentation tool that would pose sets of questions about the specific situated and dataset-specific practice of data annotation. These components covered how the data were annotated and by whom, using what type of annotation instructions, or how disagreement between annotations was accounted for. Yet, together with the entire research collective, we also formulated questions that would meet the medical and AI research interests, such as the information on the hardware for data acquisition or the quantification of sources of error. As such, the work that led to the development of the tool and the study we conducted to demonstrate its application became a particularly valuable starting point to further develop the subproject. This resulted in the development of a workshop (MIDL 2025) and various presentations (e.g., Luitse Reference Luitse2025) to facilitate implementation with and by wider (interdisciplinary) Medical AI Research communities.
5. Disconcerting ground truth(s)
Moments of disconcertment also emerged through friction and discomfort about the notion of ground truth. This term refers to the referential datasets that are considered to contain “the ‘true’ values of the phenomena that are computationally modelled” (Högberg Reference Högberg2025, 85). Friction around ground truth became particularly visible during conversations about developing an AI-driven application for predicting coagulation strategies in the hospital’s ICU department. This decision-support application is concerned with the medically complex task of guiding anesthesiologists and intensivists in the choice and amount of medicine that promotes blood clotting (coagulation) for patients who have a high-risk of bleeding during or after major surgery. For the development of this application, colleagues from intensive care medicine collected patient data from tests to assess the coagulation using whole-blood samples. These so-called Rotational Thermoelectrometries (ROTEM) represent a key site of patient datafication in critical care, transforming the dynamic process of blood coagulation into a set of data outputs (ROTEM data) that can be interpreted by anesthesiologists and intensivists. The goal of this AI system, however, is to automate the analysis of the data outputs using machine-learning techniques and to simultaneously make a prediction about a patient’s risk of bleeding and provide suggestions for the choice and amount of blood clotting medicine to prevent bleeding from happening.
To create a consistent ground truth ROTEM dataset, my medical colleagues explained that all data samples were manually annotated by a panel of expert cardiovascular anesthesiologists and intensivists. In this case, this group of selected experts was given clinical ROTEM blood clotting test result data along with important details from the patients’ surgery. For each data sample, they were asked to judge whether the result was normal or not. If that was not the case, participants were asked to identify the most likely reason from a pre-defined list, such as low platelet count, a shortage of clotting proteins or the effects of specific medication. Subsequently, participants were asked to indicate whether they would recommend a treatment and, if so, to select one or more treatment options such as clotting protein concentrates, plasma or platelet transfusions. (Noteboom et al. Reference Noteboom, Kho, Veelo, van der Ster, van Haeren, Viersen, Müller, Hermanns, Vlaar and Schenk2025).
The presentation of the annotation procedure led to disconcertment among AI researchers, as our medical colleagues referred to the development of their ground-truth dataset as the process of achieving a new “gold standard.” In today’s medical practice, driven by scientific evidence, the gold standard is a dominant medical term used “to describe definitive and decisive standards” (Timmermans and Berg Reference Timmermans and Berg2003, 27–28) that represent measures that define the “truth” and therefore constitute “the rock bottom to which new candidates for standards are compared.” For these medical researchers, the annotated data would thus serve as the new baseline standard to which AI systems for the prediction of patient bleeding and coagulation strategies could be optimized and compared. Even though they emphasized that the system required validations through clinical trials, this group argued that well-performing predictive systems ultimately promise to enhance clinical decision-making and improve outcomes for critical care patients. This framing carried strong epistemic authority within the medical domain, positioning clinical knowledge practices as the primary mediators of truth and legitimacy within this subproject. As a result, alternative understandings of ground truth were initially sidelined. As the clinical framing dominated the terms of discussion, this caused a moment of friction with my AI-research colleagues, as they would not have referred to this process in terms of aiming to set a new gold standard. To them, the ground-truth dataset would be termed a benchmark, representing a pragmatic “workable truth basis” (Högberg Reference Högberg2025, 96) to which various systems can be tested and compared to make them doable (Fujimura Reference Fujimura1987) for a specific medical task. As a technical-statistical measure used for training and evaluation, it could not be used to make claims about the potential of improving actual clinical outcomes or specific care standards, even though these may be the goal of the project at hand.
The discomfort that emerged through the confrontation with the different understandings of ground truth between the medical and AI-research colleagues revealed differing epistemic commitments and knowledge hierarchies toward AI development between these disciplines. On the one hand, guided by a dominant culture of evidence-based medicine (Timmermans and Mauck Reference Timmermans and Mauck2005), the medical researchers referred to ground-truth data as a “golden referent” that may improve clinical standards, and ultimately optimize patient care by increasing the effectiveness and efficiency of treatment. On the other hand, from the AI-research perspective, the creation of ground-truth data is part of a technical experimental workflow (Mackenzie Reference Mackenzie2015), in which the performance of predictive systems can iteratively be evaluated in a technical way and without immediate clinical commitments. During one of the biweekly project meetings, the humanities PI took the initiative and actively posed some questions to the team, which steered everyone to stay with these conceptual differences and engage with the frictions, providing ground for the disciplinary reflection on everyone’s project priorities. During this mediated conversation, it became prevalent that the implied relation of ground truth to the medical gold standard showed that my fellow medical researchers epistemically prioritized the direct impact of their system for clinical settings and outcomes. AI research scholars, instead, valued technical performance and optimization of a system itself first in advance of further developing a clinically relevant application. These collaborative reflections, in turn, made everyone in the project more attentive to the different value systems at play. This sensitivity subsequently showed and forced us to be more precise in articulating the kinds of technologies we aimed to develop, and in clarifying our immediate and long-term research goals (e.g., technological or clinical).
Disconcertment additionally emerged through another presentation on ROTEM data creation and annotation. Here, the medical team emphasized that establishing an accurate ground truth depends above all on the consistency of expert annotators’ work. This was, however, difficult to establish as different participants vary in medical interpretation and judgment of data samples, which may lead to unwanted label variations. In response, a group of project members from both AI-research as well as the humanities were interested in how our medical colleagues were dealing with issues of the so-called interobserver variability. This means that the medical assessment of data samples (e.g., high risk of bleeding or not) can vary significantly between different experts, as these medical professionals each have their individual educational background, clinical practice and experience and style of annotation.
The discussion stirred visible discomfort and friction with some of our medical teammates as well as a few other AI-research colleagues, asking us why we were interested in that question if the model showed good performance. Their reaction reflected implicit epistemic hierarchies in which technical performance metrics carried greater authority than reflexive (qualitative) concerns about the practices behind the production of the annotations. Questions about interpretive variability, subsequently, caused friction as this group of colleagues seemed to position them as secondary to evaluative standards of a “good” system – i.e., model performance. Within Critical Data, AI and STS, however, studies of ground-truth dataset construction (e.g., Henriksen and Bechmann Reference Henriksen and Bechmann2020; Jaton Reference Jaton2017; Kang Reference Kang2023) are seen as epistemically valuable because they interrogate the material politics of data production in terms of “what is [being] materialised (e.g., values, categorisations, etc.) in the production process but also what becomes invisible, such as the doubts and moral dilemmas” (Schjøtt, Reference Schjøtt2026). Even though such inquiries seem to occupy a less authoritative position to our colleagues, their reaction also triggered counter-disconcertment with us in the humanities. From our position, such work should be considered equally valid to measures of evaluation as it provides useful insights into dataset construction processes which shape model behavior and outcomes but remain difficult to trace once through standard metrics.
Taking a different and technical perspective on the issue, the group of AI-researchers that posed the same question on interobserver variability explained that they were interested in the issue because they were developing quantitative techniques to account for labeling ambiguities caused by label variability — or the idea that the same data labels can have multiple medical meanings. Even though this group acknowledged the value of examining how these variabilities could have emerged in practice, they aimed to substantiate this work by pre-empting potential data biases and mitigating potential model flaws when models would be applied in diverse clinical settings. By doing so, this group of AI-research engineers highlighted that, within their disciplinary practice, the strengthening of data quality and model robustness was given precedence in preparing AI systems for future clinical applications.
The disconcerting moment that arose from the question, and the ensuing disciplinary debate around interobserver variability in ground-truth dataset construction, was, however, also operationalized as a frictional tool opening a new research pathway. After a medical colleague and I took the initiative to motivate the group to attend to and reflect collectively on these seemingly ordinary conversations, we stayed longer with disciplinary differences. This collective discussion created a productive ground for the medical team to integrate everyone’s disciplinary concerns and generate a new research project that specifically focused on interobserver variability in medical practice. Recognizing the value of better understanding label differences, our colleague organized focus groups in which expert anesthesiologists and intensivists reflected on their annotation work, particularly on cases of disagreement. This allowed the team to better understand and account for individual judgment in annotation and provided a new avenue for producing consensus. Rather than only measuring interobserver variability technically using specific metrics, the focus groups enabled annotators to agree on a shared label alongside individual ones. As such, this new research direction showed that the initiative taken by my medical colleague and me to attend to friction and disconcertment allowed the other medical colleagues to consider the diverse concerns and develop new strategies for knowledge production on data labeling without trying to achieve a shared language for the research.
6. Encountering error with friction
In parallel to bias and ground truth, friction and unease around the questions of AI error gave rise to moments of disconcertment in the project as well, as it carries different disciplinary meaning and understanding (see also Ananny Reference Ananny2022). These differences became specifically prevalent during a presentation and follow-up discussion on a project that concerned the early detection and prediction of abnormal heart rhythm (atrial fibrillation) with ICU patients, conducted by Medical AI-Research colleagues and one intensivist. Atrial fibrillation is a common and serious arrhythmia that increases morbidity and mortality, and this subproject aimed to integrate real-time physiological monitoring data (e.g., heart rate, blood pressure, blood saturation) with electronic health records. This involves detecting predictive patterns and biomarkers associated with the initial onset of irregular heart rhythm. Through the early detection and prediction of atrial fibrillation, this AI application is set to provide decision-support to intensivists, thereby allowing prompt intervention, prevent further patient complications and guide the allocation of selected medical resources for treatment.
The presentation in question focused on a specific case study into sex bias in AI-driven electrocardiogram (ECG) classification for atrial fibrillation, as well as sinus rhythm (healthy heart rhythm) and myocardial infarction (a heart attack). To assess the impact of sex imbalance in the training data, a team of two AI-research colleagues explained that they had evaluated the performance of three different model types trained on the same dataset comprising ECG data containing examples of atrial fibrillation, sinus rhythm and myocardial infarction. Subsequently, as “designing a model evaluation often involves choosing one or more evaluation metrics combined with a choice of test data” (Hutchinson et al. Reference Hutchinson, Rostamzadeh, Greer, Heller and Prabhakaran2022, 1860), the models’ performance was measured and compared between test sets consisting of ECG data of male and female populations using two types of standardized evaluation metrics. However, instead of using the metrics to summarize the performance of an ECG classification system into one single number or score (e.g., the system presenting a 97 percent accuracy score), our colleagues shared visual bar plots containing error bars to report their results. Error bars are graphical representations of confidence intervals that provide insight into the range of performance scores observed across multiple evaluation cycles, alongside a confidence level range that the true average performance falls within (e.g., 95 percent). As such, they are predominantly being used to indicate the statistical error or uncertainty around the performance estimate of the specific system at hand, as in this case, the AI-driven ECG classifier.
This specific AI-research approach to the concept of error as a quantitative measure to provide insight into the varying range of the models’ performance scores prompted disconcertment among medical as well as legal team members. When they began to pose questions about what the reported error bars actually meant for the models’ outcomes in practice, it became clear that these colleagues were working with a very different notion of error. For them, error refers to the “myriad glitches, breakdowns, and failures” (Lin and Jackson Reference Lin and Jackson2023, 4) that risk leading to potential patient harms, especially when a system is being deployed in real-world settings (e.g., the hospital’s ICU). Driven by the medical discipline’s inherent commitments to ensuring patient safety (Jerak-Zuiderent Reference Jerak-Zuiderent2012), such errors (as failures) should therefore be reduced and preferably eliminated to secure good hospital care. In addition, from a medical ethical-legal standpoint, this notion of error as failure that is to be eliminated is inherently tied to the attribution of values such as responsibility and accountability if something goes wrong (Grote Reference Grote2025). If an error cannot be fully eliminated, who or what is to be held responsible and accountable if errors with AI systems occur and negatively affect quality care for patients? Being able to be held accountable in this way has become an increasingly important and “good thing” in current health-care settings (Jerak-Zuiderent Reference Jerak-Zuiderent2015). The disconcertment with our medical and legal colleagues was thus triggered by an implicit hierarchical way of understanding error by the AI researchers, where statistical measures of error were accorded greater legitimacy than experiential, clinical and legal perspectives on patient risk and responsibility.
For our AI-research scholars, however, the emerging friction caused another moment of disconcertment as the error rates they referred to in their presentation had neither to do with the potential for the ECG classification system to fail at all, nor were they in this particular case concerned with the question of patient harm. In contrast, this group tried to push back on the epistemic focus on error as failure and safety outcomes by the medical and legal colleagues. From the AI research point of view, the error rates they were discussing provided evidence that the systems were behaving as intended and properly enabling a particular methodology to explore sex bias in ECG classification. Here, errors were presented as “integral to the [AI system’s] form of being and intrinsic to its experimental and generative capacities” (Amoore Reference Amoore2020, 23) within research and development settings, not the failures that someone or something should be held accountable for in actual clinical environments in which AI is being applied.
The disconcerting moment that surfaced as we were faced with these varying disciplinary understandings of error, in turn, shed light on some of the differing epistemic focus points toward different stages of AI evaluation within AI Research, Medicine and Law. This time, I took the lead and stayed with the friction for a moment and asked the group to reflect on the fact that AI researchers discussed the issue of error in the context of a “learner-centric” evaluation (i.e., model-centric evaluation) from which one can draw conclusions “about the quality of the learner or its environment based on the evaluation of the learned model” (Hutchinson et al. Reference Hutchinson, Rostamzadeh, Greer, Heller and Prabhakaran2022, 1860). This type of evaluation includes testing the performance of specific models, yet also sheds light on the quality of the training data and its impact on model performance. In the case of my colleagues’ project, the impact of sex imbalance in ECG training data was assessed through evaluations of the three different model types for ECG classification. Within this process, error is “part of the way a machine learns by itself” (Aradau and Blanke Reference Aradau and Blanke2021, 7). It is something to be tamed, as my fellow team members try to keep the size of the error bars to a minimum, but, as Aradau and Blanke (Reference Aradau and Blanke2021) have argued, this “taming is at the same time an optimi[s]ation” (7) to improve model performance in an AI-lab setting. In contrast, by associating the error with system failure that has potential harmful consequences, our medical and legal colleagues approached the topic of error in the context of an “application-centric evaluation,” where they “are concerned with how the model will operate within an ecosystem consisting of both human agents and technical components” (Hutchinson et al. Reference Hutchinson, Rostamzadeh, Greer, Heller and Prabhakaran2022, 1860) such as the hospital’s ICU. In validating whether a specific AI application can be used responsibly in clinical settings, any error must be controlled. In addition, it needs to adhere to patient safety requirements and medical principles like nonmaleficence, as well as broader values of AI accountability and transparency if something goes wrong in clinical practice.
My initiative to collectively attend to this moment of disconcertment between the team members from the medical, legal and AI research disciplines, not just allowed the entire team to become more attentive to the different ways of “seeing like an algorithmic error” (Ananny Reference Ananny2022), depending on the development stage or site of application. Similar to what had occurred after frictional discussions around bias and ground truth, attending to disconcertment through the collective discussion on the issue of error prompted me to mobilize the moment as a little tool of friction that could be acted upon. This action concerned taking a new research approach, changing direction and doing difference. The mundane interdisciplinary discussion and the disconcertment it caused prompted my colleagues in the humanities and me to initiate and develop a new subproject on the broader politics of machine-learning evaluation, a key but underexplored area within Critical Data, AI and STS research. Based on the discussions we encountered, we took the initiative to approach evaluation – and error analysis and reporting as part of this process – as a “system of arbitrary structure” (Hutchinson et al. Reference Hutchinson, Rostamzadeh, Greer, Heller and Prabhakaran2022, 1860) in which situated evaluation practices and micro-political decisions influence how a system’s performance is understood measured and validated. Following this approach, we ultimately brought together a group of international interdisciplinary scholars developing new conceptual frameworks and empirical case studies into machine-learning evaluation beyond AI in health care and across domains. In a themed journal issue on the topic (Luitse et al. Reference Luitse, Schjøtt and Blanke2024), the contributions highlight emerging approaches and understandings of the politics and social implications of these evaluative processes (e.g., Campolo Reference Campolo2025; Ravn et al. Reference Ravn, Galanos, Archer and Shanley2025).
Together, these analytical approaches illustrate the need for more historical awareness, attention toward wider attempts of standardized evaluation infrastructures, and situated accounts of machine-learning evaluation in practice as highly disciplinary, domain and context-dependent. My intervention to stay with the differences and disconcertments within the project team and ask everyone for reflection had thus allowed us in the humanities to see the different understandings of error within the broader context of evaluation practices in AI system and application development. In addition, as I pushed pause when friction and disagreement on the notion of error occurred, I took action and forced myself to come together with other colleagues within Critical Data and AI to reflect and develop new articulations of the research problems present in the underdeveloped area of scholarship on AI evaluation (see also Moats et al. Reference Moats, Holtrop, Van Eck, Varga, Dechesne and Waltman2025). As such, this example highlights how the moments of disconcertment we encountered not just had a generative role within the project itself. Instead, the interdisciplinary discussion and friction provided fruitful ground to do and think differently beyond our research group, creating ideas and discussion to develop new focus areas for knowledge production in (Health) AI.
7. Concluding discussion
This article has considered the generative potential of disconcertment in interdisciplinary collaboration through the case of a four-year research project on Responsible AI for health decision-making. Critical research on interdisciplinarity has largely addressed the complexities of such work by emphasizing knowledge integration. Concepts such as “trading zones” (Galison Reference Galison1997) and “boundary work” (Gieryn Reference Gieryn1983) have been central in showing how new collaborative spaces, languages and objects can be created to manage disciplinary differences (e.g., Hine and Barnaghi Reference Hine and Barnaghi2024). I advance this literature by focusing instead on what happens when such shared language is missing. In the project under study, the absence of common terms and frameworks across the Humanities, AI Research, Medicine and Law repeatedly prompted moments of “disconcertment” (Verran Reference Verran2001) between participants.
Following Verran’s (Reference Verran2001) foundational observations in Nigerian maths education, disconcertment is a specific kind of unsettling experience that arises when one encounters a mode of knowing that “works” on its own terms yet disrupts the categories, logics and assumptions of dominant knowledge practices. Mobilizing disconcertment in the context of interdisciplinary collaboration on Responsible AI in health, my analysis of three interdisciplinary stories has shown how such unsettling encounters – and related experiences of friction and discomfort – are tied to epistemic hierarchies, where some ways of knowing are treated as more authoritative than others. In the project, the different disciplines of AI Research, Medicine, Humanities and Law lacked a shared understanding of seemingly common AI concepts – bias, ground truth and error – which were anchored in divergent methods, epistemological commitments and research objectives. These differences produced moments of contention that threatened to derail collaboration or lead to dead ends.
However, Verran (Reference Verran2001) also showed that moments of disconcertment can be used to open new (productive) avenues for understanding other ways of knowing in the world. Drawing on this approach and on moments of disconcertment around the concepts of bias, ground truth, and error in the project I have been part of since 2022, I advance the argument that small, mundane but active efforts to engage with frictions, tensions and discomfort can be made productive when different disciplinary ways of knowing collide. Such efforts can prompt reflection and open new directions for interdisciplinary research collaboration. They can be mobilized as “little tools of friction” (Aradau and Blanke Reference Aradau and Blanke2022) in interdisciplinary projects, which encourage groups to actively pay attention to different disciplinary ways of knowing and doing (interdisciplinary) research and show how it can be done differently. Yet this requires two interrelated steps of follow-on action by multiple participating researchers, which can be observed across all three analytical reflections.
First, the three interdisciplinary stories I have put forward present different ways in which active interventions by project members to attend to moments of disconcertment provided ways for the interdisciplinary team to discuss and reflect on their own practices of knowing and acknowledging novel ways of approaching various research problems. Yet these analytical narrations do so in a highly fragmented way, which reflects the fragmented, uneven and situated nature of the interdisciplinary project itself. For example, even though epistemic hierarchies between AI Research, Medicine and Law in the story regarding AI error did not suddenly disappear, my initiative to attend to disconcertment steered our diverse group of interdisciplinary scholars to remain with difference between disciplinary encounters and made it productive instead of ignoring or smoothing it. However, such intervention required (hard) work throughout the project’s diverse trajectories. This article has shown that this work occurred in different forms and was initiated by diverse project members from all involved disciplines.
This becomes particularly visible in the subproject on dataset bias. The AI research PI’s intervention, followed by my AI-research colleague’s and my own decision (as project co-leads) to stay with our divergent understandings of bias and approaches to bias mitigation in public data, allowed AI, medical and humanities scholars to collectively reposition themselves. Rather than seeking quick consensus, we asked what kinds of information about public data are most important for AI model development. Similarly, the humanities PI’s questions and suggestions to remain with the different disciplinary understandings of ground truth – and to engage the frictions they generated – created space for team members to more articulate their priorities and values, distinguishing research goals without ranking them. In the case of error, staying with disconcertment made epistemological distinctions between practices of AI evaluation more visible. It clarified the different roles error was expected to play in these processes. This, in turn, required AI and medical team members to specify their analytical goals, the development stage of the AI system and its intended site of application.
Second, the analysis shows that remaining with disciplinary differences in approaches to bias, ground truth and error can help interdisciplinary teams to make difference productive. In the data bias subproject, the AI PI’s intervention and the follow-up conversations we initiated shifted the focus. Rather than clinging to discipline-specific approaches, the PhD team turned toward data documentation as a shared, productive focal point. For our medical colleagues working on inter-rater variability in data annotation, the questions raised by humanities and AI-research colleagues, and the disconcertment that followed an intervention by one medical researcher and myself, prompted them to explore something new. To better understand label differences, they convened focus groups with expert annotators to collectively reflect on annotation work, especially on cases of disagreement. These discussions helped them foreground individual judgment in annotation and offered an alternative way of building consensus in data. Lastly, my intervention to pause for a moment and steer collective reflection on disciplinary difference in discussions of AI and error between AI, medical and law researchers, created a productive opportunity for the humanities team to think and act differently beyond the project boundaries. Here, difference and disconcertment were used as a tool of friction – a productive force to start a new research strand on the practices and politics of AI evaluation more broadly, across domains and societal sectors.
Reflecting on these project experiences, I finally stress the need to actively formalize moments of disconcertment as part of practicing “doing difference together” (Verran Reference Verran2011) in interdisciplinary research. Importantly, disconcertment does not become productive on its own. As the analysis demonstrates, it requires actively attending, facilitating and mediating work that was largely, though not exclusively, carried out by humanities researchers in the project, including myself. In practice, this involved naming discomfort, convening collective reflection sessions and prompting collaborators to articulate and trace the roots of disciplinary disagreement. Disconcertment functioned as a little tool of friction only when someone effectively “pushed the pause button” and created a space where differences could be examined rather than smoothed over.
I therefore propose that interdisciplinary (Health) AI projects integrate attending to disconcertment as a deliberate method. It can be a little tool that humanities researchers in particular can mobilize to place themselves “in the loop” (Goodlad Reference Goodlad2023) and support critical, reflexive, discipline-aware collaboration that sustains epistemological difference in AI instead of avoiding it. This does not require institutional overhaul, but it does ask interdisciplinary teams mediating across different disciplines to take disconcertment seriously and act upon it. It requires scholars to assume explicit project roles (Balmer et al. Reference Balmer, Calvert, Marris, Molyneux-Hodgson, Frow, Kearnes, Bulpin, Schyfter, MacKenzie and Martin2015) to intervene and slow down when friction, disagreement and discomfort arise, and to facilitate conversations that surface disciplinary power relations, make epistemic authority visible, and enable collaborators to acknowledge each other’s ways of knowing. In this way, doing difference together becomes not only a theoretical aspiration but a practical intervention in interdisciplinary AI scholarship, particularly in contexts where some forms of knowledge risk becoming subordinate to others. Recognizing and working with difference, then, is important to rendering epistemic power relations visible and fostering more symmetrical forms of collaboration.
Acknowledgements
First, thank you to my colleagues in the Research Priority Area AI for Health Decision-making at the University of Amsterdam, without whom this interdisciplinary research project would not have been possible. In particular, to Maria Galanty, Clarisa Sanchez, Ivana Isgum and Alexander Vlaar for the inspiring research collaborations throughout the project. Second, I would like to thank the two anonymous reviewers for their thorough engagement and feedback, which helped greatly to revise and improve the text. Third, I thank the journal editors, Georgina Born and Tobias Blanke, for their guidance and encouragement to submit to this themed issue. Lastly, thank you to Thomas Poell for his highly valuable comments on earlier drafts of this article.
Funding statement
This work was supported by the University of Amsterdam Research Priority Area Artificial Intelligence for Health Decision-Making.
Competing interests
The authors declare none.
Dieuwertje Luitse is a PhD Candidate at the Amsterdam School for Cultural Analysis and part of the Department of Media Studies at the University of Amsterdam (UvA). Her research explores the ethics, politics and power of data, artificial intelligence (AI) systems and application development in health care. This project is part of the UvA’s interdisciplinary research priority area on AI for Health decision-making that brings together researchers from Computer Science, Medicine, Law and the Humanities. In addition, her research interest focuses on the study of computational infrastructures and the (historical) development of AI systems in relation to their socio-economic and political implications. Furthermore, she co-organizes the online Critical AI Studies Seminar and is currently co-editing the topical collection on the Politics of Machine Learning Evaluation in Digital Society.