Introduction
Attention deficit hyperactivity disorder (ADHD) is a common heterogeneous neurodevelopmental condition with an estimated prevalence of 5–7.2% in children and adolescents, and 2.5% in adults [Reference Dobrosavljevic, Larsson and Cortese1]. Although symptoms are heterogeneous, the common core symptoms are varying degrees of inattention and hyperactivity/impulsivity symptoms leading to significant impairments in important life activities. Previous studies show that successful ADHD management and pharmacologic treatment can prevent undesired outcomes, such as academic impairments, injuries, criminality, suicide, and improve patients’ quality of life and long-term outcomes [Reference Harpin2–Reference Boland, DiSalvo, Fried, Woodworth, Wilens and Faraone4].
Adult ADHD is increasing and demanding more health care resources [Reference Adler, Faraone, Spencer, Berglund, Alperin and Kessler5]. The considerable diversity of ADHD presents diagnostic difficulties for healthcare professionals, potentially resulting in either overlooked cases or incorrect positive diagnoses [Reference Johnson, Morris and George6]. A delayed diagnosis and treatment of ADHD symptoms can ultimately result in escalated healthcare utilization. This increase in the use of health care resources may be fueled by the impact of co-occurring psychiatric conditions [Reference Du Rietz, Jangmo, Kuja-Halkola, Chang, D’Onofrio and Ahnemark7, Reference Garcia-Argibay, Pandya, Ahnemark, Werner-Kiechle, Andersson and Larsson8]. Undiagnosed adults with potential ADHD symptoms experience significantly greater challenges, highlighting the importance of accurate and early diagnosis to support at-risk individuals or those with overlapping symptoms [Reference Naya, Tsuji, Nishigaki, Sakai, Chen and Jung9].
Machine learning (ML) approaches have shown growing promise in healthcare by identifying data-driven patterns in routinely collected electronic health records (EHRs) that may not be readily apparent during standard clinical review. When applied at scale, such models can function as clinical decision-support tools, flagging individuals whose recorded symptom trajectories or service use patterns suggest elevated risk and who may benefit from earlier or more comprehensive clinical assessment [Reference Swinckels, Bennis, Ziesemer, Scheerman, Bijwaard and de10]. This has the potential to improve service efficiency by supporting targeted case-finding and triage, rather than relying solely on referral-based or clinician-initiated identification [Reference Stephenson, Eadie, Holmes, Asadpour, Gutierrez and Kumar11].
In the context of neurodevelopmental and psychiatric disorders such as attention-deficit/hyperactivity disorder (ADHD), such approaches may help address the persistent gap between symptom onset and diagnosis. Importantly, ML-based tools can be used to augment rather than replace existing clinical decision-making, supporting early risk stratification and prioritization for assessment [Reference Dwyer, Falkai and Koutsouleris12]. This aligns with emerging models of precision psychiatry, where data-driven tools complement clinical expertise to improve individualized care.
The potential of ML in automating the identification of ADHD has garnered significant research interest, with numerous studies demonstrating promising results through various computational approaches. However, the majority of these studies have primarily focused on utilizing structured behavioral assessments [Reference Tachmazidis, Chen, Adamou and Antoniou13–Reference Öztekin, Finlayson, Graziano and Dick16], neuroimaging data [Reference Peng, Debnath and Biswas17–Reference Khan, Waheeb, Riaz and Shang20], or physiological signals [Reference Tosun21–Reference Catherine Joy, Thomas George, Albert Rajan and Subathra23]. These types of data, while valuable, are often not readily available or require specialized, often expensive, procedures. In contrast, EHR are already available to healthcare providers for all their patients, making them a more convenient, cost-effective, and practical option for ML-based ADHD identification. Despite the growing accessibility and richness of EHRs, which provide comprehensive patient histories, few studies have explored the application of ML techniques to EHR data for ADHD identification. This gap represents a critical opportunity for advancing ADHD diagnosis, as EHR data can provide real-world, longitudinal insights that are often absent in other types of data, potentially enhancing the accuracy and scalability of ML-driven diagnostic tools.
Method
Data source
This study is based on retrospective data from the Regional Health Care Information Platform in Region Halland (RHIP), an integrated data platform located in southwestern Sweden [Reference Ashfaq, Lönn, Nilsson, Eriksson, Kwatra and Yasin24]. The platform incorporates information from both primary and secondary healthcare levels, including prescribed medications, clinical findings such as laboratory tests, and medical imaging used in patient care. Although clinical history-taking, self-reports, structured interviews, and neuropsychological testing are commonly used in the diagnostic assessment of ADHD, such detailed information is not codified in the available data. Instead, diagnoses, procedures, and medications are recorded using standardized coding systems, namely the International Classification of Diseases, 10th Revision, Swedish Edition (ICD-10-SE), and the Anatomical Therapeutic Chemical (ATC) classification system.
Study design
This study retrospectively analyzed data from RHIP for adults (18 years or older) identified as ADHD between 2015 and 2022. ADHD diagnosis was determined based on the ICD-10-SE code “F90” or at least one pharmacy dispensation of ADHD medications listed in Table 1. The control group consisted of adults referred to the outpatient clinics in psychiatry between 2015 and 2022, who had no ADHD-related diagnosis and had not been prescribed ADHD-related medications.
Considered ATC codes to identify ADHD patients

When including the entire control group, the ML model struggled to effectively distinguish the ADHD group. To address this, we explored excluding certain subgroups from the control group. The model only achieved strong performance when patients with depression or anxiety diagnoses were removed. A total of 62.8% of the initial control cohort met criteria for depression or anxiety and were therefore excluded. This left 5,126 control participants, which, although reduced, still represents a cohort size larger than the ADHD group (n = 3,570). Therefore, this exclusion did not adversely affect model performance or the stability of the classification task.
This exclusion is clinically justified, given the significant overlap in cognitive and behavioral symptoms – such as difficulties with concentration and memory, as well as variations in activity levels – between ADHD, depression, and anxiety disorders. Some studies suggest that depression and anxiety disorders are among the most frequently co-occurring psychopathologies with ADHD [Reference Quenneville, Kalogeropoulou, Nicastro, Weibel, Chanut and Perroud25, Reference Katzman, Bilkey, Chokka, Fallu and Klassen26], further supporting the rationale for their exclusion.
For each patient, we considered only the visits, or other contact with a health care provider that occurred between 2011 and 2022; patients who did not have any type of contact in this time frame were excluded.
To facilitate early diagnosis of adult ADHD, in this study, we explored the feasibility of predicting ADHD diagnoses at three different intervals: 6, 12, and 18 months before the index date. The index date was defined as the date of the first recorded ADHD diagnosis for individuals in the ADHD cohort and the date of the last recorded medical contact for individuals in the control group. To ensure that predictions were based solely on historical data, all patient records from the index date onward, as well as those within the designated prediction period, were excluded from the analysis, as illustrated in Figure 1. This approach aimed to simulate real-world diagnostic scenarios by focusing on information available before the onset of ADHD diagnosis in the affected group or the last medical contact in the control group. We collected variables from the Regional Health Care Information Platform, encompassing patient demographics, diagnoses, procedures, and prescribed medications.
Using patient historical health records for early diagnosis of adult ADHD.

Modeling strategies
To capture the temporal dynamics of medical records and manage the challenges posed by very long sequences and the significant variability in the number of contacts, we aggregated all variables over a fixed time window. In our experiments, we used an aggregation period of 2 months.
The modeling approach evaluated in this study utilized transformers, a model architecture introduced by Vaswani et al. [Reference Vaswani, Shazeer, Parmar, Uszkoreit, Jones, Gomez, Kaiser and Polosukhin27] that enhances natural language processing through self-attention mechanisms, allowing for efficient capture of long-range dependencies. Transformers process sequences in parallel and overcome the challenge of lacking inherent token order recognition through positional encoding, which adds positional information to token embeddings, enabling the model to understand sequence order. This architecture has shown great promise in modeling the sequential nature of EHRs [Reference Moore, Orset, Yassaee, Irving and Morelli28, Reference Li, Rao, Solares, Hassaine, Ramakrishnan and Canoy29]. Compared to other approaches that use traditional models on EHR data, such as logistic regression or random forests applied to manually engineered tabular features, our transformer-based approach can directly model the sequential structure of clinical events. This enables the model to capture temporal patterns and long-range dependencies that conventional models cannot represent without extensive feature engineering.
To apply the transformer architecture, a text-based representation of patient medical records was created, as shown in Figure 2. In this representation, clinical codes (such as diagnosis ICD codes and medication ATC codes) were used as text tokens. Additionally, two tokens representing patient age and gender were included. The age token consisted of the word “Age” followed by a digit representing the patient’s age group; for example, “Age_3” indicates a patient in their 30s. The gender token was represented by a single character (“F” or “M”). The time period when each clinical code was recorded was used for positional encoding to capture the temporal aspects of the EHR data.
Generating a text representation of the patient’s healthcare records.

Model implementation details
We implemented a customized BERT-based architecture tailored for the classification of structured event-sequence data. The model consists of a single encoder layer with four self-attention heads, using a hidden size of 512 and an embedding size of 256. Input representations are derived from a joint embedding module that combines token embeddings and event-index (time period) embeddings, followed by layer normalization. The encoder block includes multihead self-attention, a position-wise feed-forward network that expands to 2,048 units before projecting back to 512 units, Gaussian Error Linear Unit (GELU) activation, and layer normalization, with dropout set to 0.1. On top of the encoder, a multilayer classifier produces the final prediction through a sequence of linear layers (20, 10, 2), each followed by Rectified Linear Unit (ReLU) activation and dropout.
For optimization, the model was trained using the Adam optimizer with a learning rate of 5 × 10−7 and weight decay of 0.015. To address class imbalance, we employed a weighted binary cross-entropy loss (BCEWithLogitsLoss) with the positive class weight set according to the empirical distribution of the training data. Training was conducted for up to 5,000 epochs with a batch size of 50 and an early-stopping patience of 250 epochs. Dropout in the BERT module and classifier was set to 0.1 and 0.5, respectively.
Evaluation metrics and interpretability
The predictive performance of the model was assessed using the following metrics: area under the curve of the receiver operating characteristic for discrimination between patients, f1-score, sensitivity, specificity, and number needed to screen (NNS). The NNS is the inverse of the positive predictive value and represents the number of patients the model needs to flag in order to correctly identify one true positive case [Reference Kipnis, Turk, Wulf, LaGuardia, Liu and Churpek30, Reference Liu, Bates, Wiens and Shah31].
The Shapley Additive Explanations (SHAP) technique was used to offer insights into model decisions by revealing the most important features for early adult ADHD prediction. SHAP is a model-agnostic explanation method that provides both local and global explanations [Reference Lundberg and Lee32]. In our formulation, the global explanations show which key factors (demographics or clinical codes) increase or reduce the risk of adult ADHD. To explore model interpretability, we applied the approach using a 6-month prediction period on a randomly selected subset of 500 patients.
Results
A comprehensive statistical analysis was performed to compare demographic and clinical characteristics between ADHD and non-ADHD groups, as illustrated in Table 2. The cohort consisted of 8,696 participants, with 3,570 in the ADHD group and 5,126 in the non-ADHD group. Significant differences were observed in several variables. The ADHD group had a lower proportion of females (48.43%) compared to the non-ADHD group (53.06%), with a statistically significant difference (P = 0.003). The mean age of ADHD patients was notably younger (31.64
$ \pm 11.67 $
years) compared to the non-ADHD group (52.83
$ \pm $
22.77 years), yielding a highly significant P-value (P < 0.001). These demographic patterns are consistent with the findings of other studies [Reference Chung, Jiang, Paksarian, Nikolaidis, Castellanos and Merikangas33–Reference Hutt Vater, DiSalvo, Ehrlich, Parker, O’Connor and Faraone35].
Demographics and 2-month resource utilization variables (data after applying inclusion criteria)

Regarding healthcare utilization, ADHD patients had more frequent primary care visits (mean = 1.62, standard deviation = 1.66) compared to the non-ADHD group (mean = 1.25
$ \pm $
1.39, P < 0.001). Conversely, non-ADHD individuals exhibited a higher mean number of specialist visits (mean = 0.76
$ \pm $
1.33) compared to ADHD patients (mean = 0.67
$ \pm $
1.10, P = 0.001). Psychiatrist visits were significantly more common among ADHD patients (mean = 1.11
$ \pm $
3.00) than non-ADHD individuals (mean = 0.52
$ \pm $
1.30, P < 0.001).
Additionally, hospital admissions were lower in the ADHD group (mean = 0.10
$ \pm $
0.21) compared to the non-ADHD group (mean = 0.13
$ \pm $
0.21, P < 0.001). The average hospital length of stay was also shorter for ADHD patients (mean = 0.45
$ \pm $
4.03 days) than for non-ADHD patients (mean = 0.80
$ \pm $
2.24 days, P < 0.001). No significant difference was observed in the number of emergency visits between the two groups (P = 0.30). These findings highlight notable differences in healthcare utilization and demographic characteristics, which can inform predictive modeling for early ADHD diagnosis.
Table 3 presents the model’s performance across three prediction horizons: 6, 12, and 18 months before the index date. Overall, the results are fairly consistent across these timeframes, with only modest variations in performance metrics. The 6-month prediction window yielded the highest scores, with an area under the receiver operating characteristic curve (AUC) of 0.79 (95% confidence interval [CI]: 0.76–0.81), F1-score of 0.79 (95% CI: 0.76–0.81), sensitivity of 0.80 (95% CI: 0.77–0.84), and specificity of 0.77 (95% CI: 0.73–0.81). At 12 months, performance remained comparable, with an AUC of 0.76 (95% CI: 0.73–0.79), F1-score of 0.76 (95% CI: 0.73–0.79), slightly higher sensitivity at 0.74 (95% CI: 0.70–0.78), and a somewhat lower specificity of 0.78 (95% CI: 0.73–0.82). The 18-month results showed a further modest decline, with an AUC of 0.75 (95% CI: 0.71–0.78), F1-score of 0.75 (95% CI: 0.71–0.78), sensitivity of 0.74 (95% CI: 0.70–0.78), and specificity of 0.76 (95% CI: 0.71–0.80). While the 6-month horizon offered slightly stronger performance, the differences across all timeframes were relatively small, indicating stable predictive capability over different lead times.
Model performance results on the holdout dataset

The SHAP analysis of the 6-month prediction period model identified several clinical codes as influential in predicting the probability of an ADHD diagnosis, as shown in Figure 3. Among these, ‘F158: Other specified mental and behavioral disorders due to use of other stimulants, including caffeine’ (SHAP value: 0.26) highlights the well-documented link between ADHD and substance use disorders, with research suggesting that over half of adults with ADHD may meet criteria for such disorders during their lifetime [Reference Dunne, Hearn, Rose and Latimer36]. Similarly, ‘Y903: Blood alcohol level 0.60–0.79 mg per 100 ml’ (SHAP value: 0.24) underscores the association between ADHD and alcohol use disorders, as studies estimate that 15–25% of adults with alcohol use disorders have ADHD [Reference Wilens37]. Additionally, the codes “O648: Obstruction of labor caused by other specified abnormal fetal position and other specified abnormal fetal presentation” and “G02AB01: Metylergometrin” point to an increased risk of childbirth complications among mothers with potential ADHD. While these codes are not specific to ADHD diagnosis, they likely capture broader clinical and behavioral contexts that tend to co-occur in the healthcare records of individuals subsequently identified with ADHD.
SHAP values results on the holdout dataset.

The fairness analysis across genders for 6-month prediction period, as depicted in the Figure 4, reveals a notable difference in the true positive rate (TPR), with the model achieving 75.2% TPR for females compared to 66.7% for males. This indicates that the model identifies ADHD cases slightly more effectively in females than in males. The false positive rate and true negative rate are almost consistent across genders, suggesting balanced performance in non-ADHD predictions. However, the disparity in TPR highlights the need for further evaluation to ensure equitable diagnostic performance across genders.
Model fairness across genders.

Discussion
This predictive model using ML, where we extracted data only based on existing patient records, is a cost-effective method that could possibly be used to supplement existing procedures to screen for adult ADHD, thus improving accuracy. Our results indicate that demographic information, as well as healthcare-related data, may predict later ADHD diagnoses. Individuals who were later diagnosed with ADHD, compared to those who were not diagnosed, tended to be male, younger, with more primary care and psychiatrist visits, and with slightly fewer visits to non-psychiatry specialists at other hospitals, as well as fewer hospital admissions and shorter length of stay. Clinical codes indicate a higher risk of substance use among adult ADHD individuals, as well as a higher risk for mothers with ADHD experiencing birth-related complications. Below, we will discuss our results more in-depth, related to other findings, and the implications for real-world use. When disseminating these results, some are expected from clinical lore as well as from other studies, while others were more unexpected and warrant closer consideration.
Gender
The gender skewness toward more males diagnosed with ADHD is reflected in earlier studies where a disproportionate number of boys receive the ADHD diagnosis [Reference Dobrosavljevic, Larsson and Cortese1, Reference Harpin2]. However, this proportion has been found to be less pronounced with later age of being diagnosed and with studies of individuals receiving their diagnosis as adults, as being close to the same proportions [Reference Faheem, Akram, Akram, Khan, Siddiqui and Majeed38–Reference Cortese, Faraone, Bernardi, Wang and Blanco40]. Although our results on gender differences carry high statistical significance, the difference is not at all as pronounced as in the pediatric samples, but is found around 50%, corresponding to other studies and clinical impressions.
Age and health care utilization
Since general clinical utilization increases with age, except within specific groups with somatic or psychiatric issues, the large age discrepancy makes sense. Also, although more older individuals are now diagnosed with ADHD, prevalence decreases with age, while other psychiatric problems do not show the same trends, and in some cases, such as depression, increases in older age groups. We also found that individuals with subsequent ADHD diagnoses had more primary care and psychiatrist visits, although this group is younger, thus defying the general trend of higher health care utilization as a result of older age.
Adult individuals, without prior history in psychiatry and having exited the school system, with issues relating to suspected ADHD, are encouraged to first visit primary care units to exclude somatic issues or milder, non-ADHD psychiatric conditions. If screened and with symptoms related to ADHD without other somatic and psychiatric explanations, these individuals are referred to a psychiatric unit for further evaluation. Relating to the increased healthcare utilization with age described above, it follows that despite higher utilization during a specific time period in order to explore medical explanations for the functional impairments that are consequential to ADHD, healthcare utilization for nonpsychiatric issues is lower in this group, thus resulting in fewer specialist visits and hospital stays.
Substance use
Our findings on stimulant use were less expected and warrant closer examination. Elevated rates of substance use, including alcohol, among individuals with ADHD are well documented; however, such patterns are also observed in other psychiatric populations. Notably, the diagnostic code F15.8 (“Other stimulant-related disorders, including caffeine”) raises interpretive challenges. It remains unclear whether these cases reflect self-medication with stimulants before formal ADHD diagnosis, or whether they primarily capture excessive caffeine consumption, such as high coffee intake, rather than more severe stimulant use disorders. Clinical observations have long noted unusually high caffeine or stimulant use among individuals with ADHD, suggesting that such behaviors may represent compensatory strategies or informal indicators of the disorder. Further research is needed to clarify whether caffeine-related coding reflects benign consumption patterns, attempts at self-medication, or undetected substance use disorders.
Pregnancy and birth complications
Research indicates that mothers with ADHD are more likely to experience preterm births, cesarean deliveries, and neonatal complications. These associations may reflect the broader clinical and psychosocial challenges often observed in individuals with ADHD, including higher rates of depression, social isolation, and unplanned pregnancies [Reference Hesselman, Wikman, Skoglund, Kopp Kallner, Skalkidou and Sundström-Poromaa41, Reference Walsh, Rosenberg and Hale42]. These findings collectively highlight the multifaceted impact of ADHD on both behavioral and physiological outcomes, emphasizing the importance of early diagnosis and management.
The appearance of such codes in SHAP analysis suggests that data-driven EHR models may capture indirect but informative clinical contexts preceding ADHD recognition, rather than relying solely on explicit diagnostic indicators.
Comorbidity
The results that did not appear in this study may be just as interesting as the ones that did appear. First, we could not elicit useful models when depression and anxiety disorders were included in the analysis. As we pointed out above, many of the symptoms of depression and anxiety overlap with core ADHD symptoms, such as concentration difficulties, motor activity, and motivational issues [Reference Forbes, Neo, Nezami, Fried, Faure and Michelsen43]. However, that is also true with other psychiatric conditions, such as personality disorders, dissociative disorders, psychotic disorders, neurocognitive disorders, substance-use disorders, and so forth. It may be that those diagnostic codes are not as ubiquitous or predictive in this sample as anxiety and depression, and thus not appear in the material.
Teasing out profiles early that predict ADHD, rather than diagnoses that are more related to depression and anxiety disorders, can be very helpful in a clinical context where these differential diagnostic dilemmas are common in the later assessment process. Since the psychiatric diagnostic process is not based on finding clearly demarked underlying disease processes or any distinguishable specific causal mechanisms, we need to find ways to triangulate several different pieces of information. Predictive ML models may be a useful addition to improving accuracy in the diagnostic process and clarifying when ADHD may be a relevant explanation to the individual’s difficulties.
Importantly, the proposed model is not intended to replace standard diagnostic protocols or detailed clinical assessments for ADHD. Instead, it should be viewed as an early-stage risk identification tool that leverages existing EHR data to flag individuals who may benefit from further evaluation. Such predictive models can be integrated into clinical workflows as decision-support aids, operating in the background of existing care pathways, guiding clinicians toward patients whose health trajectories or comorbidities suggest higher ADHD likelihood.
By flagging individuals with elevated ADHD likelihood, the model may help clinicians prioritize patients for structured diagnostic assessment. In this way, the model’s primary clinical value lies in its potential to improve service efficiency and consistency in case identification, support more equitable access to diagnostic evaluation, and contribute to more proactive management of ADHD in adults.
Conclusion
Although this and similar studies are not intended to substitute clinical reasoning intrinsic to psychiatric diagnostic evaluations, including ADHD, they represent a step toward augmenting such processes through data-driven, objective insights. By integrating routinely collected EHR information, these models illustrate how computational tools can complement traditional assessment methods and contribute to a more evidence-based, equitable diagnostic workflow.
Future research should focus on evaluating real-world performance, replicability across healthcare systems, and integration with dimensional frameworks such as transdiagnostic or neurodevelopmental constructs. Together, these efforts could advance the development of clinically informed, generalizable models that enhance equitable access to diagnostic evaluation and support more proactive management of ADHD in adults.
Finally, aligning the model with emerging dimensional frameworks in psychopathology, such as transdiagnostic or neurodevelopmental constructs, could extend its relevance beyond categorical diagnoses and provide insights more consistent with contemporary theoretical models of mental health [Reference Kotov, Krueger, Watson, Achenbach, Althoff and Bagby44].
Data availability statement
The real-world data used in the analysis are protected to comply with Swedish integrity-protecting regulations, and thus, data cannot be accessible to the public due to the requirements of the Swedish Health and Medical Services Act.
Acknowledgments
Portions of this manuscript’s text were refined using ChatGPT (OpenAI, GPT-4, accessed between April 2025 and August 2025 to improve clarity and grammar. The authors reviewed and edited all AI-assisted text, and the final content is entirely the authors’ own work.
Financial support
This work was financially supported by Takeda Pharma AB. FE and OH were supported by the CAISR Health research profile funded by the Knowledge Foundation (grant 20200208 01H).
Competing interests
The authors declare none.
Ethical considerations
The study was approved by the Swedish Ethical Review Board (registration number 2022-07287-02). Informed consent was waived because this was a retrospective study that the Swedish Ethical Review Board approved. All the methods in this study were performed in accordance with relevant guidelines and regulations.







Comments
No Comments have been published for this article.