Hostname: page-component-76d6cb85b7-6jg5l Total loading time: 0 Render date: 2026-07-21T08:04:10.847Z Has data issue: false hasContentIssue false

Scalable, context-sensitive psychiatric assessment with large language models and brief diaries

Published online by Cambridge University Press:  21 July 2026

Whitney R. Ringwald*
Affiliation:
University of Minnesota , USA
Aman Taxali
Affiliation:
University of Michigan, USA
Mike Angstadt
Affiliation:
University of Michigan, USA
Colin E. Vize
Affiliation:
University of Pittsburgh , USA
Chandra Sripada
Affiliation:
University of Michigan, USA
Aidan G. C. Wright
Affiliation:
University of Michigan, USA
*
Corresponding author: Whitney R. Ringwald; Email: wringwal@umn.edu
Rights & Permissions [Opens in a new window]

Abstract

Background

Accurate psychiatric assessment requires understanding a person’s unique experience within their psychosocial context. Clinical interviews have been the gold standard for assessment as the only methods capable of this complex task, but they are time and resource-intensive. Consequently, psychiatric assessment typically relies on patient report surveys that are decontextualized and narrow in scope. This comprehensiveness-scalability tradeoff is a major bottleneck in studying and treating psychopathology. We propose using large language models (LLMs) to score psychopathology from brief personal narratives as a low-burden, context-sensitive solution.

Methods

Participants (N = 108) completed brief (~1 minute), freeform audio diaries daily for 2 weeks. We used six LLMs to score wide-ranging psychopathology (Internalizing, Detachment, Disinhibition, Antagonism, Anankastia) from the diary transcripts. Leveraging an array of self-report and clinical interview measures, we tested the convergent, discriminant, concurrent, and clinical validity of LLM ratings for between-person differences and within-person fluctuations in psychopathology.

Results

Supporting convergent and discriminant validity, LLM ratings correlated most strongly with corresponding self-report domains at the between (average convergent r = .42) and within-person (r = .28) levels. LLM and self-report ratings had similar patterns of associations with external variables, except for Anankastia and Antagonism. Further, every LLM-rated domain related to psychopathology ascertained by clinical interview.

Conclusions

Across multiple forms of validity, we showed that LLMs can assess most major forms of psychopathology from mere minutes of audio. These results support scoring open-ended narratives with LLMs as a scalable, portable method to translate idiographic diagnostic data into standardized psychiatric assessments.

Information

Type
Original Article
Creative Commons
Creative Common License - CCCreative Common License - BY
This is an Open Access article, distributed under the terms of the Creative Commons Attribution licence (http://creativecommons.org/licenses/by/4.0), which permits unrestricted re-use, distribution and reproduction, provided the original article is properly cited.
Copyright
© The Author(s), 2026. Published by Cambridge University Press
Figure 0

Figure 1. Within-person correlations between LLM ratings and self-reports of psychopathology. Note: LLM and self-report psychopathology reflect daily fluctuations from a person’s average levels. Credibility intervals are in parentheses.Figure 1. long description.

Figure 1

Figure 2. Between-person correlations between LLM ratings and self-reports of psychopathology. Note: LLM and self-report psychopathology reflect averages estimated from daily ratings. Credibility intervals are in parentheses.Figure 2. long description.

Figure 2

Table 1. Nomological net similarity between LLM-rated psychopathology and self-report psychopathologyTable 1. long description.

Figure 3

Figure 3. Correlations between LLM-rated and self-reported psychopathology and interview-rated psychopathology. Note: SIDP-IV, Structured Interview for DSM-IV Personality; SCID-5, Structured Clinical Interview for DSM-5 Axis I Disorders. LLM and self-report psychopathology reflect averages estimated from daily ratings. * = credibility interval does not contain zero.Figure 3. long description.

Figure 4

Table 2. Rationales and transcript quotes selected by ChatGPT-5 to support its ratingsTable 2. long description.

Supplementary material: File

Ringwald et al. supplementary material

Ringwald et al. supplementary material
Download Ringwald et al. supplementary material(File)
File 1.3 MB