I. Introduction
Adoption of large language models in education and knowledge work has reached a scale that, three years ago, no one predicted. By the end of 2025, generative AI assistants had become routine fixtures in university classrooms, in software-engineering organisations, in legal offices and in public administration across the European Union, used daily by populations counted in the hundreds of millions. The headline trajectory is one of accelerating uptake, productivity claims, and self-reported gains in capability. The trajectory the present paper concerns is a quieter one, visible underneath the headline: a growing body of evidence that the same technologies producing those productivity gains are also producing, at the level of the individual cognitive trajectory and at the level of organisational practice, a systemic form of deficit that accumulates silently and reveals itself only when the substrate it has eroded is later required.
The phenomenon has acquired a name in the recent literature: cognitive debt. The term, proposed by Kosmyna et al. on the basis of electroencephalographic measurements during essay-writing tasks, captures a pattern in which sustained reliance on LLM mediation reduces the engagement of the neural circuits associated with memory, attention, and idea generation, and leaves the user unable to integrate into their own cognitive substrate the work that the tool has helped them produce.Footnote 1 Independent peer-reviewed studies have converged on compatible findings using different methods and different populations: cognitive ease at a cost, depth compromised, critical engagement reduced, the asymmetry between confidence in the tool and confidence in oneself pulling thinking in opposite directions.Footnote 2 The convergent picture is not yet exhaustive, and the empirical work has limitations we examine in detail below, but it is sufficient to make the question of cognitive debt relevant from a regulatory perspective rather than merely speculative.
The European Union is among the first jurisdictions to enact a comprehensive horizontal legal instrument addressing high-risk AI: Regulation (EU) 2024/1689, the Artificial Intelligence Act.Footnote 3 The regulation classifies AI systems used to evaluate learning outcomes, including when those outcomes steer the learning process, as high-risk under Annex III, point 3(b), triggering an extensive set of obligations on providers and deployers. Article 14, the keystone of the high-risk regime, requires that such systems be designed to enable meaningful human oversight, with explicit recognition of the cognitive bias of automatic over-reliance on AI outputs. The regulatory architecture is, by international standards, ambitious and substantively serious. It is not, we argue in this paper, well-designed for the cognitive-debt mechanism that the recent neurocognitive evidence has begun to document.
This paper makes a tri-disciplinary argument. We synthesise the convergent neurocognitive evidence on cognitive debt and identify the four mechanisms (cognitive offloading, atrophy through disuse, transfer-appropriate processing failure and engagement asymmetry) through which the debt accumulates. We supplement that evidence with longitudinal practitioner observations gathered by the first author over twelve years of management and architectural roles in software-engineering organisations spanning the pre- and post-LLM transition, suggesting that the patterns identified in controlled experimental settings reproduce at the scale of professional practice. We then analyse the regulatory architecture of the AI Act as it applies to LLM-mediated learning environments and identify what we term the cognitive blind spot of Article 14: the assumption, structurally embedded in the provision, that the natural person assigned to oversight retains the cognitive substrate that the supervised activity, performed sustainedly under the regulatory regime, progressively erodes. The blind spot is not a drafting error. It is a consequence of designing oversight as a synchronic competence, operative at the moment of decision, rather than as a diachronic capacity built and maintained over the lifecycle of the supervisor’s professional engagement with the system.
Our contribution sits at the intersection of three literatures that, to our knowledge, have not previously been brought into direct dialogue. The neurocognitive literature on LLM-mediated cognition has expanded rapidly since 2023 but rarely engages with regulatory analysis. The legal literature on the AI Act, including the closely related work of Laux and Ruschemeier on automation bias published in this journal,Footnote 4 has begun to address the cognitive limitations of human oversight but has not yet incorporated the cognitive-debt evidence specifically. The practitioner literature on LLM-assisted professional work has documented productivity effects with increasing rigour but has not connected those findings to the regulatory framework now in force. Each of these literatures, taken alone, supports only a partial argument. The neurocognitive literature establishes that cognitive debt is real but says nothing about whether any legal regime should respond to it. The legal literature establishes that human oversight under the AI Act has cognitive limits but has not connected those limits to the specific, cumulative erosion that the neurocognitive evidence documents. The practitioner literature establishes that LLM mediation reshapes professional work but has not asked what the regulatory framework now in force does or does not capture. Their integration, we suggest, supports a substantive one: that the regulatory regime governing high-risk educational AI in the European Union, as it stands at the moment the high-risk obligations are due to enter into force, addresses the wrong temporality of risk for the developmental harm the recent evidence makes visible.
The paper proceeds as follows. Section II establishes the neurocognitive case for cognitive debt by articulating the four mechanisms and reviewing the convergent experimental evidence supporting each. Section III reports practitioner observations from software-engineering organisations across the LLM-adoption transition, identifying two patterns (the asymmetry between expressed confidence and verifiable understanding, and the convergent local solutions accompanied by senior resistance) that mirror, at the scale of professional practice, the experimental findings of Section II. Section IV analyses the AI Act’s high-risk regime for educational AI under Annex III, point 3, identifies a coverage problem in the regulatory perimeter, and articulates the cognitive blind spot of Article 14. Section V addresses the limitations and counterarguments to which our argument is open, derives operational implications of the developmental and substitutive distinction for institutions, providers and regulators, and identifies the empirical work required to settle the questions we leave open. Section VI concludes.
II. The neurocognitive case for cognitive debt
The proposition that sustained reliance on large language models produces measurable cognitive costs has, in the past two years, moved from speculation to a position supported by converging lines of empirical evidence. We organise that evidence around four identifiable mechanisms, each of which is documented in at least one peer-reviewed study and corroborated across methodologies. The mechanisms are not independent (they reinforce one another over time), but separating them analytically clarifies what is at stake when we speak of cognitive debt, the term we adopt from Kosmyna et al. and extend, in later sections, beyond the individual scale at which it was originally proposed.Footnote 5
A note on the evidence base is in order before we proceed. The most prominent recent contribution to this literature is the preprint of Kosmyna et al., which combines electroencephalographic measurement, natural language processing, and human and AI essay scoring across four experimental sessions over four months.Footnote 6 The study has not, at the time of writing, completed peer review. We treat its findings as preliminary in that strict sense, and we do not rest the argument of this section on its results alone. The mechanisms we describe below are each independently supported by peer-reviewed work, and we use Kosmyna et al. as the most recent and methodologically integrative example of a convergence already visible in the literature.
1. Cognitive offloading as the baseline mechanism
The first mechanism is the oldest and best-documented. Cognitive offloading is the use of physical action or external tools to alter the information-processing requirements of a task in ways that reduce the cognitive demand on the agent.Footnote 7 The phenomenon is not new and is not specific to artificial intelligence: writing is offloading, calculators are offloading, navigation apps are offloading. What is new with LLMs is the breadth of cognitive activity that can be offloaded in a single tool, including activities (argument construction, synthesis, hypothesis formation) that previously had no equivalent external substitute.
Stadler, Bannert and Sailer, in a randomised study of ninety-one university students published in Computers in Human Behavior, provide one of the cleanest demonstrations of the trade-off involved. Participants assigned to gather information about a socio-scientific issue using ChatGPT 3.5 reported genuinely lower cognitive load than those using Google. The reduction was real, not perceived: the LLM condition required less mental effort to reach a usable answer. But the same participants produced lower-quality reasoning and weaker justifications in their final recommendations. The conclusion the authors draw is that LLMs reduce germane load, the load associated with the constructive cognitive work that builds durable understanding, alongside the extraneous load they were intended to reduce. Cognitive ease, in their formulation, comes at a cost.Footnote 8
Lee et al., in a survey of three hundred and nineteen knowledge workers presented at CHI 2025, extend this finding into professional practice. Their participants reported decreased cognitive effort across the full spectrum of activities associated with critical thinking (knowledge retrieval, comprehension, application, analysis, synthesis, and evaluation) when using generative AI compared to working without it. The reductions were not marginal: examples reported as requiring “much less effort” or “less effort” reached 72% for analysis tasks and 76% for synthesis tasks. The data does not, on its own, distinguish between offloading that supports the same depth of thinking with less friction and offloading that substitutes for thinking that is no longer occurring. That distinction becomes the central question of the remaining mechanisms.Footnote 9
2. Atrophy through disuse
The second mechanism is the gradual weakening of capacities that are no longer regularly exercised. Bainbridge’s classic formulation of the ironies of automation anticipated the problem decades before its present manifestation: when a routine task is handed over to a machine, the human loses the regular practice that hones the skill, and is then less equipped to intervene when the machine fails or needs supervision. The irony is that automation increases the importance of human capacity at precisely the moment it reduces opportunities to develop that capacity. In other words, the supervisor is expected to be most competent exactly where the system has given them the least practice: the conditions under which the human is called upon to intervene are, by construction, the conditions the human has had the fewest occasions to rehearse.Footnote 10
In the context of LLM use, atrophy through disuse becomes particularly consequential because the activities being offloaded are themselves constitutive of intellectual development. A student who uses an LLM to draft essays does not merely save time on essay writing; they bypass the recursive process by which essay writing produces the cognitive structures that make later, more complex thinking possible. The offloading and the development are the same activity. Removing one removes the other.
The neurophysiological evidence in Kosmyna et al. is consistent with this account. Their EEG measurements showed that brain-only participants exhibited the strongest and most distributed neural connectivity during the writing task, search-engine users showed moderate engagement, and LLM users showed the weakest connectivity. The gradient is what matters: the more the cognitive work was externalised, the less the brain engaged. Repeated over four sessions, this pattern translated into a behavioural measure that is striking in its directness: participants in the LLM group could not, immediately after submission, accurately quote from their own essays. The work had been produced; it had not been integrated.Footnote 11
3. Transfer-appropriate processing failure
The third mechanism is the most subtle and, for educational and regulatory purposes, the most consequential. The principle of transfer-appropriate processing, established in cognitive psychology, holds that knowledge is most effectively retrieved and applied in contexts that resemble the conditions under which it was originally encoded. If the learning context differs substantially from the application context, transfer fails, not because the knowledge was never acquired, but because it was acquired in a form that does not generalise.
Applied to LLM-mediated learning, this principle predicts an asymmetry that the experimental data confirms. Participants who first learned without AI assistance and only later encountered LLMs (the Brain-to-LLM condition in Kosmyna et al.) showed enhanced neural activation when subsequently using the tool, consistent with effective augmentation of pre-existing capacities. The reverse condition, in which participants relied on LLMs from the start and were later required to work without them (LLM-to-Brain), showed reduced alpha and beta connectivity, consistent with under-engagement and ineffective independent performance. The same individuals, the same tasks, the same tool: the only variable was the order of acquisition.Footnote 12
This finding has implications that extend well beyond the experimental setting. It suggests that the cognitive value of LLM mediation depends on the existence of a substrate that LLM mediation does not, by itself, build. Where the substrate exists, the tool augments. Where it does not, the tool substitutes, and the substitution prevents the substrate from forming. The temporal ordering matters because the substrate is built only through the activities that the tool is being used to avoid.
4. Engagement asymmetry and mechanised convergence
The fourth mechanism concerns the relationship between user confidence and the quality of cognitive engagement. Lee et al. found, in their survey of knowledge workers, that higher confidence in the AI was associated with less critical thinking, while higher self-confidence was associated with more critical thinking. The two confidences point in opposite directions. The user who trusts the tool more thinks less; the user who trusts themselves more thinks more, even when using the same tool for the same task. This asymmetry is significant because it locates the mechanism not in the technology but in the user’s posture toward it.Footnote 13
A related phenomenon, identified by Sarkar and corroborated across multiple studies of AI-assisted knowledge work, is what he terms mechanised convergence.Footnote 14 When AI tools are used widely, the work they produce becomes more homogeneous and less diverse. The effect is observable at the linguistic level: within the LLM condition of Kosmyna et al., named-entity recognition, n-gram patterns, and topic ontology all showed marked within-group homogeneity, far higher than in the search-engine and brain-only groups.Footnote 15 We will argue, in the practitioner observations of Section III, that the same phenomenon is observable at the level of code: convergent micro-decisions across engineers who do not share style guides or codebases, producing local optimisations that fail to compose with the surrounding system. Mechanised convergence is the sociotechnical signature of an environment in which the tool, rather than the user, is doing the meaningful selection.
5. The four mechanisms in interaction
These four mechanisms are not parallel hazards from which a designer might pick one to mitigate. They form a sequence. Cognitive offloading, when habitual, produces atrophy through disuse. Atrophy through disuse weakens the substrate on which transfer-appropriate processing depends, so that capacity acquired with the tool fails to generalise to contexts without it. The eroded substrate, in turn, leaves the user more dependent on the tool’s confidence than on their own, completing the loop and making the next round of offloading more likely. The accumulation of these costs over time is what Kosmyna et al. name cognitive debt.Footnote 16
The metaphor of debt is useful, and we adopt it for the remainder of the paper, but it should be read with the precision its authors intended. The debt accumulates silently. Each individual offloading event is small and locally rational: the task is completed faster, the immediate cost is lower, the immediate output is acceptable. The debt is the difference between the cognitive substrate that would have been built had the activity been performed unaided and the substrate that is in fact built. That difference does not appear on any individual occasion. It appears only when the substrate is later required and is found to be missing. By that point, the deficit cannot be repaid retrospectively, because the conditions under which it would have been built have passed. The debt, unlike a financial one, admits of no later settlement: the developmental window in which the substrate forms is specific to the moment the activity is first performed, and once that activity has been delegated rather than performed, the window does not reopen. What is lost is not the output, which the tool supplied, but the formation the output would have produced in the person who made it.
In Section III, we present practitioner observations that suggest this dynamic is now visible at the scale of organisational practice. In Section IV, we argue that the European Union’s principal regulatory instrument for high-risk AI in education (the AI Act, and specifically the obligations triggered by Annex III, point 3(b)) was drafted on a model of risk that the cognitive-debt mechanism does not fit, and that this misfit constitutes a regulatory blind spot of substantive consequence (Figure 1).
The four mechanisms of cognitive debt and their temporal interaction. Cognitive debt accumulates through a self-reinforcing cycle of four mechanisms: (1) cognitive offloading, where external tools replace internal cognitive effort; (2) atrophy through disuse, where unexercised capacities weaken; (3) transfer-appropriate processing failure, where capacities acquired with the tool fail to generalise to contexts without it, evidenced by the Brain-to-LLM versus LLM-to-Brain asymmetry observed by Kosmyna et al. (n 3); and (4) engagement asymmetry, where trust shifts from self to tool, producing the linguistic and cognitive homogenisation Sarkar (n 16) terms mechanised convergence. The cycle is unidirectional and self-reinforcing. Cognitive debt is the cumulative deficit relative to the substrate that unaided practice would have built; it appears not in any individual cycle but only when the substrate is later required and found absent.

III. Cognitive debt in professional practice: practitioner observations
Note on voice and authorship of this section. The observations reported in §3.2 and §3.3 were gathered by the first author (Ferreira) in the course of management and architectural roles in software-engineering organisations between 2012 and 2026. To keep voice clear, the first person singular (“I”) is used in this section when reporting those observations directly. The first person plural (“we”) is reserved for the interpretive frame and the integration of the observations with the literature, both of which are joint work.
1. Method and observational frame
The empirical material in this section consists of observations I gathered over a continuous twelve-year period, spanning eight software-engineering organisations of varying mandate, size and geography. The observation is participant rather than detached: in each role I held positions with managerial or architectural responsibility over the teams under observation. Engineering management constitutes a privileged observational vantage in this respect, since code reviews, architecture discussions, post-mortems, hiring loops and one-on-one conversations expose patterns of technical reasoning that are not visible from outside the team.
Software engineering is the focus of this section for three reasons, and it is worth stating them before the observations themselves, because the choice shapes how far the observations transfer. The first reason is evidential: software engineering is the professional domain in which the transition from pre-LLM to post-LLM practice has been fastest, most complete, and most nearly universal. Code-generating assistants moved from novelty to default infrastructure across the industry within roughly two years, which compresses into an observable window a shift that, in most other professions, is still partial and gradual. The second reason is methodological: software engineering externalises cognitive work to an unusual degree. A teacher’s reasoning about a student is largely private, but an engineer’s reasoning is inscribed in code, in review comments, in commit histories, and in architecture documents, where it is open to inspection by anyone with managerial access. The cognitive substrate, and its erosion, leave a documentary trace. The third reason is simply access: the first author held managerial and architectural positions in this domain across the relevant period, and the participant vantage described above was available here and not elsewhere. We are explicit about this because it bounds the claim. Software engineering is offered as a domain in which the cognitive-debt mechanism is unusually visible, not as a domain in which it is unusually severe. The teaching professions, which the AI Act’s Annex III(3) provisions most directly govern, are precisely the setting to which these observations would need to be carried by future, domain-specific study; Section V returns to this limitation. The claim here is one of structural analogy, not of established generality.
Several of the organisations involved sit before the public release of ChatGPT in November 2022 and form the pre-LLM baseline. The remaining four, namely a global education platform, a cross-border product organisation operating across Brazil, the United States, and Europe, an industrial digital-twin firm in the oil and gas sector, and a European energy utility undergoing platform modernisation under formal governance constraints, straddle or follow that release. The transition was therefore not observed from a single vantage but reconstructed across overlapping contexts. The diversity is methodologically useful, since patterns that recur across all four warrant more confidence than those visible in only one.
I should also state my own position. I am not a non-user of these tools. I have used LLMs in my own technical work, both before this paper began and during its preparation. The question I have asked myself daily, and have not always answered well, is when to use them and when to refrain. The discipline of asking the question, more than abstinence from the tool, is what I take to be the relevant practice. The observations that follow should be read as those of a participant who has himself navigated, imperfectly, the same incentives that produced the patterns described.
We claim no formal ethnographic protocol. Field notes were not systematically maintained, observations were not coded against an a priori scheme, and no inter-rater reliability was sought. The analytical move performed in this section is therefore the retrospective synthesis of recalled observation rather than prospective measurement, and we use the phrase descriptively rather than as the name of an established method. This carries known limitations: confirmation bias in selective recall is plausible, the samples within each organisation are small, and the observer’s progressive engagement with this question may have shaped subsequent attention. We have nevertheless chosen to include this material because the observational window of twelve years across eight organisations, four of them spanning the post-LLM period, is unusual, because the patterns recur with sufficient consistency across contexts, and because the observed patterns align closely enough with formal experimental findings reported elsewhere in the literature to warrant articulation as falsifiable propositions for future formal study. We present two such patterns below.
2. The asymmetry between expressed confidence and verifiable understanding
The most pervasive shift across the post-2022 organisations was not a decline in apparent technical competence but a change in the relationship between expressed competence and underlying understanding. Engineers and technical managers, including peers, not only juniors, increasingly entered discussions with positions stated fluently and with high confidence. The positions were often correct, sometimes nuanced, and frequently delivered in the cadence of someone who had thought deeply about the matter. What had changed was what happened when the position was probed for its foundations.
A characteristic exchange came to recur. A position was advanced: that a specific caching strategy was preferable for a given workload, that a given authentication pattern was more secure than the alternative, that one observability approach scaled better than another. The position was correct in the sense of matching current best practice. When asked why, under what conditions the recommendation held, what trade-offs it implied, what it shared with or differed from related options, the discussion did not deepen. It deflated. The respondent could often restate the position in different words but could not trace it backward to the considerations that produced it. The confidence was real; the understanding behind it was fragile.
This pattern closely mirrors the most striking finding in the Kosmyna et al. experimental data. Participants in the LLM group could produce essays that human teachers and AI judges scored highly, but could not, minutes after submission, quote from those same essays. They had produced the artefact without integrating it into memory. The interpersonal version of this pattern, namely being able to state and defend a technical position one cannot trace to its foundations, is the workplace projection of the same phenomenon. It is not deception. The speakers are not pretending. The fluency is genuine. What is absent is the substrate that, in the pre-LLM baseline, accompanied the same fluency as a matter of course.Footnote 17
The consequence in technical discussion is a degradation of the discussion’s epistemic value. Disagreement, when it occurred, became harder to resolve productively, because the disagreement could no longer be traced to the divergent reasoning chains that produced each position. Two engineers asserting incompatible positions, neither of whom can articulate the foundations of their own assertion, cannot move toward synthesis through argument. The discussion either collapses to authority, with whoever is more senior or more insistent prevailing, or to majority, where whichever LLM-mediated formulation happens to be more widely repeated wins by frequency rather than by force of reason. My repeated experience of this dynamic, recurring across teams that did not share members or organisational culture, was the original motivation for the present inquiry.
3. Convergent local solutions and the senior asymmetry
A second pattern, distinct from but related to the first, emerged consistently across the post-2022 contexts. It manifests differently along the engineering experience curve, and the asymmetry between its junior and senior expressions is itself the most analytically interesting feature.
Among engineers in the early-to-mid stages of their careers, code began to acquire what can only be described as identical vices across people who had not previously produced similar work. The pattern is not plagiarism in any direct sense; the code is novel, written by the engineer in their own files. The pattern is one of convergence in micro-decisions. Similar variable-naming conventions appear across teams that did not share style guides. Similar refactoring strategies are deployed for problems that admit a wider range of approaches. More substantively, a recurring failure mode emerged in which a specific change, typically a targeted modification suggested by an LLM, solved the immediate problem at hand while creating new problems at the global level of the system. The local optimisation was correct; the global model that would have flagged its incompatibility with adjacent components was absent.
This pattern echoes, at the level of code, the homogenisation that Kosmyna et al. observed at the linguistic level in LLM-assisted essays: consistent homogeneity across named-entity recognition, n-gram distribution, and topic ontology within the LLM group.Footnote 18 The mechanism is plausibly the same. When the suggestion-generating tool is shared, suggestions converge. When the user lacks the global model that would let them recognise that a locally-correct suggestion is globally damaging, the suggestion is implemented. The consequence is technical debt of a particular shape: not the familiar pattern of cumulative shortcuts under time pressure, but the less familiar pattern of correct-looking changes that fail to compose with the rest of the system.
The mirror pattern emerged among more senior engineers. Their reaction to the tools was, predominantly, resistance, not on ideological grounds but on operational ones. The repeated complaint, articulated in slightly different forms across organisations, was that arriving at a usable result from the LLM required more work than simply doing the task. The reason can be stated precisely. Senior technical work draws heavily on tacit knowledge: the substrate of accumulated context, prior decisions, and integrative judgment that cannot be fully verbalised in the form of a prompt. Translating that substrate into the input format the tool requires is itself a substantial cognitive operation, and one that tends to lose precisely the elements of the underlying knowledge that gave the original judgment its value. For senior engineers facing problems whose solution depends on this substrate, the prompt-construction overhead consistently exceeded the marginal utility of the suggestion that resulted.
The asymmetry produces a counter-intuitive prediction that aligns with the Brain-to-LLM result of Kosmyna et al. the tool is least useful to those whose cognitive substrate is most developed, and most attractive to those whose cognitive substrate is least developed. This is the opposite distribution from the one we would want for cognitive scaffolding. In the experimental data, participants who built independent cognitive substrate before encountering the LLM showed enhanced neural connectivity when later using it, while those who relied on the LLM from the start showed degraded connectivity when later required to work without it. The workplace projection is that the LLM functions as effective augmentation in the hands of the senior engineer who already knows, but the senior engineer rarely chooses to use it, because the activation cost exceeds the gain. Meanwhile, it functions as substitute scaffolding in the hands of the junior engineer who does not yet know, and is used precisely there: at the point in the experience curve where substitution most damages the development of independent capacity. The phenomenon is not a misuse of the technology by individuals; it is a structural property of where the tool is most economically attractive, and that location is the worst possible location from a cognitive-development standpoint (Figure 2).Footnote 19
The experience-curve asymmetry of LLM utility. Two qualitative curves are plotted against a cognitive-substrate axis ranging from novice to expert. Adoption/use intensity decreases monotonically with experience, as activation costs and self-confidence rise. Augmentation value increases monotonically with experience, as the cognitive substrate that lets the tool function as augmentation rather than as substitute scaffolding accumulates. The two curves cross at mid-career. To the left of the crossover, the tool is most adopted where its augmentation value is lowest: the substitute-scaffolding region in which cognitive debt accumulates with use. To the right of the crossover, the tool’s augmentation value is highest but adoption is lowest, because the activation cost of compressing tacit knowledge into prompts often exceeds the marginal gain. The asymmetry mirrors, at the scale of professional practice, the Brain-to-LLM versus LLM-to-Brain finding of Kosmyna et al. (n 3), and yields a prediction of regulatory significance: LLM mediation is most economically attractive at exactly the position on the experience curve where its developmental harm is greatest.

IV. The regulatory gap: AI act Annex III(3) and the cognitive dimension
Regulation (EU) 2024/1689, the Artificial Intelligence Act, is the first comprehensive horizontal legal instrument on artificial intelligence in any major jurisdiction. Its risk-based architecture (prohibited practices, high-risk systems with extensive obligations, limited-risk systems with transparency duties, and minimal-risk systems left to the market) has become a reference point for legislators internationally and a binding compliance regime for any provider or deployer placing AI systems on the European single market or whose output is used in the Union. This section concerns one specific application of that architecture: the classification, in Annex III, point 3, of AI systems used in education and vocational training as high-risk, and the obligations triggered by that classification under Articles 9 to 27.
We argue that the regulatory model the AI Act constructs for educational AI, while substantively serious and considerably more demanding than any pre-existing instrument, was drafted on a conception of risk that the cognitive-debt mechanism does not fit. The misfit is not a drafting error and is not unique to the educational context (comparable misfits have been identified for automation bias more generally by Laux and RuschemeierFootnote 20 ), but its consequences in education are particularly significant, because the affected interest is not a discrete decision but the developmental trajectory of the learner, and because the regulatory beneficiary and the regulatory subject are, in many cases, the same person.
1. Scope and obligations: high-risk AI in education under Annex III, point 3
Annex III, point 3 of the AI Act lists four categories of educational and vocational-training applications as high-risk: (a) systems used to determine access or admission, (b) systems used to evaluate learning outcomes including when those outcomes steer the learning process, (c) systems used to assess the appropriate level of education a person will receive or access, and (d) systems used to monitor and detect prohibited behaviour during tests. Of these, point (b) is the most consequential for our argument and the focus of what follows. Its scope is, on its face, broad. It appears to capture not only formal summative assessment, such as the algorithmic grading of an exam, but any AI system whose evaluation of learning outcomes is used to direct what the learner will encounter next. On this reading, the formulation would extend to a wide range of LLM-based educational tools that grade student work and adjust subsequent instruction on that basis. The breadth of that reading is, however, contested, and the boundary between summative assessment and the formative, adaptive personalisation of learning has become the central interpretive question for the scope of point (b). We turn to that question, and to the Commission’s draft guidance on it, in the next subsection.
The obligations triggered by Annex III classification are substantial. Under Article 9, providers must establish, document, and maintain a risk management system across the entire lifecycle of the system, identifying foreseeable risks and adopting risk-management measures proportionate to those risks. Article 10 imposes data governance requirements on training, validation and testing datasets, including representativeness and the management of biases. Article 13 requires that high-risk AI systems be designed to be sufficiently transparent to enable deployers to interpret outputs and use them appropriately. Article 14, to which we return in detail in §4.3, requires that systems be designed to allow effective human oversight, with explicit attention to automation bias. Article 26 places parallel duties on deployers, including the duty to assign oversight personnel with the necessary competence, training, authority and support. Article 27 mandates a fundamental rights impact assessment by certain categories of deployer, including public bodies and entities providing public services, before putting a high-risk system into use. Above these system-specific obligations sits Article 4, in force since 2 February 2025, requiring providers and deployers to ensure a sufficient level of AI literacy among their staff and other persons dealing with the operation and use of AI on their behalf.
The temporal architecture of the regime is staggered, and was revised significantly during the drafting of this paper. Provisions on prohibited practices and AI literacy entered into force on 2 February 2025, and obligations on general-purpose AI models on 2 August 2025. The obligations applicable to high-risk systems under Annex III were originally due to apply from 2 August 2026, but the Digital Omnibus on AI, the package of amendments advanced by the European Commission in late 2025, postponed that date: under the revised timetable, the Annex III high-risk obligations under Article 6(2) are due to apply from 2 December 2027, and those for high-risk systems embedded in regulated products under Article 6(1) from 2 August 2028. The Commission published its draft guidelines on high-risk classification under Article 6(5) on 19 May 2026, opening a stakeholder consultation that closed in June 2026; the final guidelines, and the harmonised standards under the joint mandate to CEN and CENELEC, remain in development. Compliance is therefore being built against a regulatory text that is substantially fixed but whose operational meaning, and even its date of application, has continued to move.
2. Application to LLM-mediated learning environments
The AI Act’s Annex III(3)(b) classification captures, on its face, the standard cases for which it was drafted: an automated essay scorer that grades a written submission and adjusts the difficulty of subsequent assignments; an adaptive testing system that estimates competence and selects the next question; a personalised tutor that tracks student responses and chooses the next learning activity. The compliance question for these cases is largely settled. The provider conducts the conformity assessment under Annex VI, the system is registered in the EU database, the deployer institution conducts its fundamental rights impact assessment under Article 27 where applicable, and human oversight is implemented as required by Article 14. These cases are difficult but tractable.
The boundary of that classification is, however, less settled than the standard cases suggest, and the Commission’s draft guidelines on high-risk classification, published under Article 6(5) on 19 May 2026, address it directly. The guidelines do not alter the classification rules; they set out the Commission’s interpretation of Article 6 and offer non-exhaustive examples. Two features are relevant here. The first is the filter mechanism of Article 6(3), which the guidelines confirm exempts from the high-risk regime systems that fall within an Annex III use case but do not materially influence the outcome of decision-making, including systems that perform only a narrow procedural task or that merely improve the result of a previous human activity. The second is the contested status of adaptive, personalised learning. Education-technology bodies have argued in consultation that personalised learning tools, which adjust pace and content to the individual student on an interim and ongoing basis, perform a preparatory task to assessment rather than an evaluation of learning outcomes, and should therefore fall outside points 3(b) and 3(c); on this reading, only summative assessment, the evaluation of learning at the end of a course, would attract the high-risk classification. Whether the final guidelines adopt that narrower reading is, at the time of writing, unresolved. But the direction of interpretive travel matters for our argument, and not in the direction that might be expected.
A narrower reading of point (b), confined to summative assessment, would not weaken the argument of this paper; it would sharpen it. If formative, adaptive personalisation is read out of the high-risk perimeter, then the very category of system in which sustained, iterative, learner-facing LLM mediation occurs, and therefore the category in which cognitive debt has the greatest opportunity to accumulate, sits outside the high-risk regime altogether. The narrower the perimeter, the larger the share of cognitively consequential LLM use that the high-risk obligations never reach. The coverage problem we develop in the remainder of this subsection is, in other words, a structural feature of the regime that the draft guidelines, if anything, widen rather than close.
The cases that emerged in the course of 2024 and 2025, and that will dominate practice once the high-risk obligations apply, are different in structure. They involve recursive loops between LLM-based systems and human users in which the boundaries between learner work, evaluation and steering become difficult to identify in operational terms. A representative configuration is the following. A student uses an LLM to draft a written submission. The submission is uploaded to a learning management system whose grading layer is itself LLM-based and produces a numerical mark together with structured feedback. The same platform recommends next learning activities on the basis of the mark, the feedback, and a model of the student’s evolving performance. The instructor reviews aggregate dashboards rather than individual submissions. At which point in this loop does Annex III(3)(b) attach?
The answer, on the face of the regulation, is that it attaches at the grading and recommendation layer: that layer is an AI system used to evaluate learning outcomes, and the outcomes steer the learning process. The provider of that layer therefore bears the Article 9 to 15 obligations, and the educational institution deploying it bears the Article 26 to 27 obligations. The student’s use of the LLM at the drafting stage is, formally, outside the regulated perimeter, since the LLM in that role is not being used to evaluate learning outcomes, but to produce them. It is the upstream tool, regulated, where applicable, under the general-purpose AI provisions of Articles 51 onwards rather than the high-risk regime. The asymmetry is deliberate by design: the AI Act regulates the system that makes the consequential decision, not every tool that touches the inputs to that decision.
In practice, the asymmetry produces a coverage problem. The cognitive-debt mechanism described in Section II operates predominantly at the upstream stage, at the point where the student offloads the constructive work to a tool that is not, in its own right, classified as high-risk under Annex III(3)(b). The downstream evaluation system is regulated; the upstream production system, where the developmental harm accumulates, is not regulated under the same regime. A deployer that achieves full compliance with Articles 9 to 27 on its grading and recommendation layer can still be operating an environment in which cognitive debt accumulates at scale, because the regulatory instrument is correctly aimed at a different risk altogether: the risk that the evaluative decision is wrong, biased, or inadequately overseen.
The empirical study by Manganello et al. of early-adopting Italian universities under Annex III(3)(b) is illuminating in this regard.Footnote 21 The authors find what they describe as a systematic misalignment between regulatory expectations and institutional capacity: the AI Act’s classification of educational assessment AI as high-risk implies technical safeguards comparable to those required for medical devices or critical infrastructure, but the universities studied lack the organisational structures for algorithmic accountability at that level. The misalignment they identify is one of capacity. The misalignment we identify here is conceptually prior to capacity: it concerns what the regulatory regime is designed to detect at all. An institution can build the auditing capacity Manganello and colleagues call for, achieve compliance with every applicable provision, and still remain blind to the dynamic that the cognitive-debt evidence reviewed in Section II makes visible.
3. The cognitive blind spot of Article 14
Article 14 of the AI Act is the operational keystone of the high-risk regime. Where the substantive obligations of Articles 9 to 13 set out what providers must do in the design and documentation of the system, Article 14 sets out what the system must enable the human to do during operation. The provision is structured around a normative ideal, namely meaningful human oversight, and a recognition, unusual in EU regulatory text, of a specific cognitive limitation that threatens that ideal. Article 14(4)(b) requires that the system be provided to deployers in such a way that the natural persons assigned to oversight are enabled to remain aware of the possible tendency of automatically relying or over-relying on the output produced by a high-risk AI system (automation bias).Footnote 22
The reference to automation bias is, as Laux and Ruschemeier observe in their analysis published in this journal, the first explicit recognition of a cognitive bias in EU law. They argue persuasively that the legal text gives a scientifically simplified account of the phenomenon, that the asymmetric division of responsibility between providers (who must enable awareness) and deployers (who must exercise it) does not adequately address the design and contextual factors that produce the bias, and that mandating awareness is a weaker regulatory tool than directly regulating the conditions under which the bias arises. Their critique is sound, and its implications extend, in our view, beyond automation bias as the paper defines it.Footnote 23
Article 14 assumes a particular cognitive subject. The natural person assigned to oversight is presumed to possess, at the moment of oversight, the cognitive capacities required to perform that oversight meaningfully: domain understanding sufficient to evaluate the system’s output, the analytical capacity to identify when that output is wrong or unreliable, and the independent judgement required to decide, in the language of Article 14(4)(d), not to use the high-risk AI system or to otherwise disregard, override or reverse the output.Footnote 24 The provision specifies the supervisor’s duties without specifying the cognitive substrate that makes those duties performable. That substrate is treated as an exogenous given.
The cognitive-debt evidence reviewed in Section II destabilises the assumption. The substrate the supervisor requires, namely domain understanding, analytical capacity and independent judgment, is the same substrate that sustained interaction with the supervised system progressively erodes. The teacher who oversees an LLM-based grading system over five years, while themselves using LLMs to draft lesson plans, generate feedback and prepare assessments, is not the same cognitive subject at year five as at year one. The supervisor’s capacities are not preserved by the existence of the oversight obligation; they are shaped, like any cognitive capacity, by the activities the supervisor regularly performs and refrains from performing. Where the activities being offloaded are constitutive of the very judgment the supervisor is later required to exercise, the oversight obligation becomes circular: it requires capacities that the supervised activity, performed under the regime, undermines.
This is what we mean by the cognitive blind spot of Article 14. The provision recognises automation bias as a synchronic phenomenon: a tendency operative in the moment of decision, addressable through awareness training and design features that prompt the supervisor to engage. It does not recognise cognitive debt as a diachronic phenomenon, namely an accumulation of capacity loss across the lifecycle of the supervisor’s professional engagement with the system. The two phenomena are related but distinct. Automation bias asks: at this moment, is the supervisor over-relying on the output? Cognitive debt asks: across the years of this supervisor’s practice with this category of system, has the substrate that would let them detect over-reliance been built or eroded? An oversight regime that addresses only the first question is not therefore failing on its own terms; it is not addressing the second question at all.
An objection must be met at this point. It might be said that the AI Act already addresses what we have called the diachronic dimension, through the AI literacy obligation of Article 4, in force since February 2025, which requires providers and deployers to ensure a sufficient level of AI literacy among the persons who operate and use AI systems on their behalf. If staff must be kept AI-literate on an ongoing basis, the objection runs, then the regime does reach beyond the moment of decision and does attend to the capacities the supervisor brings to oversight over time. The objection is serious and, in our view, only partly answerable, which is itself an instructive result.
Article 4 is the provision of the AI Act that comes closest to the concern of this paper, and we do not dispute that it has a temporal dimension the moment-of-decision provisions lack. But AI literacy, as Article 4 frames it and as the emerging interpretive practice understands it, is a competence directed at the system: knowing what an AI system does, how to operate it, where its outputs are unreliable, and what its limits are. Cognitive debt is not a deficit in competence directed at the system; it is an erosion of the domain substrate directed at the task that the system is mediating. These are different objects. A teacher can be highly AI-literate, fully aware of how an LLM-based grader works and where it fails, and nonetheless have lost, through years of delegating the underlying disciplinary work, the independent command of the subject matter that would let them recognise a plausible-looking but wrong evaluation. Literacy in the tool does not preserve the substrate the tool is displacing. Indeed, the relationship may run the other way: greater fluency with the system can make delegation easier and more habitual, and so accelerate the very erosion that meaningful oversight requires the supervisor to resist. Article 4 therefore mitigates one part of the problem, the part concerning the supervisor’s understanding of the system, while leaving untouched, and possibly worsening, the part concerning the supervisor’s understanding of the domain. The scope and content of the literacy duty have, moreover, become harder to state with confidence following the Digital Omnibus revisions, which makes it an unstable foundation on which to rest a claim that the diachronic risk is already covered. The existence of Article 4 narrows the blind spot we identify; it does not close it.
The point is not that the AI Act should have addressed cognitive debt and failed to do so. The empirical evidence on which a regulatory response could be drafted has emerged largely after the political agreement on the regulation was reached, and the specific neurophysiological findings of Kosmyna et al. post-date the regulation by a full year.Footnote 25 Our point is rather that the regulatory architecture, as it now stands, contains a structural feature, namely its conception of human oversight as a synchronic competence rather than a diachronic capacity, that the present and foreseeable evidence base shows to be empirically inadequate for the educational context. The questions this raises are several: whether the harmonised standards now under development by CEN and CENELEC can incorporate diachronic considerations; whether the Commission’s Article 6(5) implementation guidelines can clarify oversight obligations to require monitoring of supervisor capacity over time; whether the Article 4 AI literacy duty, which applies horizontally and to deployers as well as providers, can be interpreted to cover cognitive-debt awareness in addition to operational competence; and, more fundamentally, whether direct regulation of the cognitive-debt risk, along the lines Laux and Ruschemeier propose for automation bias,Footnote 26 is both legally feasible and politically achievable. We do not attempt to resolve these questions here. We argue that they are now on the regulatory agenda, whether or not they have yet been recognised as such, and that the evidence base is sufficient to make their consideration overdue.
The mechanisms identified in Section II and the practitioner observations of Section III do, however, support a tractable interim distinction that institutions and providers can begin to apply within the existing regulatory architecture without waiting for the harmonised standards or implementation guidelines to clarify the position. The distinction separates developmental uses of LLM mediation from substitutive uses by reference to a single criterion: whether the activity being mediated is one whose performance, by the human, is constitutive of the capacity the educational programme is designed to build. Where the activity is constitutive, for instance drafting an argument in a writing course, deriving a proof in a mathematics course, or formulating a clinical hypothesis in a medical course, substitutive use erodes the substrate the programme exists to develop, and should be designed against. Where the activity is ancillary, such as formatting references, translating between human languages of equivalent fluency, or generating routine boilerplate, substitutive use is harmless and sometimes helpful. The criterion is verifiable in practice because it is internal to the institution’s own learning objectives: an institution that cannot say which activities are constitutive of its programme cannot articulate the programme at all. We return to the operational consequences of this distinction in Section V (Figure 3).
What the AI Act regulates and what cognitive-debt evidence suggests it should regulate. Three pairs of risks are aligned across a structural axis of opposition. Output errors, addressed by Articles 9, 10, and 13, concern incorrect or biased decisions at the moment of inference; substrate erosion, identified by Kosmyna et al. (n 3) and Stadler, Bannert, and Sailer (n 4), concerns the cumulative weakening of the user’s cognitive substrate across interactions with the system. Automation bias, addressed by Article 14(4)(b) and analysed in this journal by Laux and Ruschemeier (n 6), concerns over-reliance at the moment of decision; cognitive debt, identified by Kosmyna et al. (n 3) and corroborated by Lee et al. (n 4), concerns the silent accumulation of capacity deficit that reveals itself only when the supervisor is called upon. Oversight as competence, the design assumption embedded in Article 14(1) and (4), conceives of human supervision as a capacity operative at the moment of supervision; oversight as capacity (the dimension of supervision the AI Act does not currently address) concerns what is built and maintained across the supervisor’s professional engagement with the system. The opposition is not exhaustive but is structural: the regime addresses what the system does at the moment of decision, while cognitive-debt evidence concerns what the supervisor becomes across the lifecycle of practice.

V. Discussion: limits, counterarguments and implications
The argument advanced in this paper rests on three legs (neurocognitive evidence in Section II, practitioner observation in Section III, and legal analysis in Section IV), each of which is open to objection. We address the most consequential objections in turn, draw out the practical consequences of the developmental and substitutive distinction introduced at the close of Section IV, and identify the empirical work that would be required to settle the questions we have left open.
1. The fragility of the empirical base
The strongest objection to the cognitive-debt thesis concerns the quality and quantity of the evidence supporting it. The most prominent recent contribution, Kosmyna et al., is a preprint with fifty-four university-aged participants in essay-writing tasks across four sessions over four months.Footnote 27 The sample is not representative of the populations to which the regulatory implications of our argument would attach, the task is one of many cognitive activities to which LLM mediation could be applied, and the experimental window is short relative to the multi-year educational and professional trajectories that the cognitive-debt mechanism, as we describe it, requires to manifest fully. Marrone and Kovanovic, in their cautionary commentary on the public reception of the study, are correct to warn that the popular reading of “ChatGPT rots your brain” overstates what the data shows.Footnote 28
The convergence with Stadler, Bannert and Sailer and Lee et al. addresses some of these limitations.Footnote 29 Stadler and colleagues studied a different population (German university students) using a different experimental paradigm and a different LLM, and reached compatible conclusions about the trade-off between cognitive ease and depth of inquiry. Lee and colleagues sampled a much larger population (three hundred and nineteen knowledge workers across nine hundred and thirty-six real-world AI-use instances) and found compatible patterns of reduced cognitive effort and the confidence-direction asymmetry. The replication is not exact and the methods are different; what is converging is not a single experimental result but a family of findings whose joint pattern is harder to dismiss than any of the individual studies. The argument we have built rests on this convergence rather than on Kosmyna et al. alone, and we believe that distinction is sufficient to carry the regulatory analysis. It is not sufficient, and we do not claim it is, to support specific quantitative predictions about how much cognitive substrate is lost over what time horizon under what conditions of use. Those questions remain open, and represent the most urgent direction for further empirical work.
2. The augmentation counterargument
A more substantive objection holds that we have selected the evidence that supports the cognitive-debt thesis and underweighted the evidence for cognitive augmentation. The literature contains genuine examples of the latter. Suriano et al. report that students who actively engage with ChatGPT, by asking follow-up questions, prompting for justifications, and critiquing AI-provided arguments, show improvements in critical, creative and reflective thinking compared to peers who do not use the tool.Footnote 30 Cui et al., in field experiments at Microsoft and Accenture, found productivity gains from Copilot adoption that were larger for less experienced developers, suggesting that under some conditions LLM mediation accelerates capacity acquisition rather than retarding it.Footnote 31 The Brain-to-LLM finding within Kosmyna et al. itself shows enhanced neural connectivity when LLM mediation is added to an established cognitive substrate.Footnote 32
The augmentation evidence is real, and our argument does not contradict it. The distinction we drew at the close of Section IV, namely developmental versus substitutive use, is the distinction these findings occupy. Augmentation effects appear where the user retains primary cognitive responsibility for the activity and uses the tool to extend, refine or accelerate that activity. Substitution effects appear where the user offloads the cognitive activity itself and consumes the output. The same tool, the same task, the same user can produce either pattern depending on which posture is adopted. The claim of our paper is not that LLM mediation is always damaging. It is that the educational and professional environments now being constructed tend, as a default, toward substitution rather than augmentation, that this default produces cognitive debt at scale, and that the regulatory architecture under which these environments operate has not been designed to detect or address this dynamic. Whether a given individual user, in a given moment, achieves augmentation rather than substitution is a question for that user. Whether populations of users, under the incentive structures that the technology produces, achieve augmentation in aggregate is the question for the regulator. The evidence we have reviewed suggests the answer to the first question is sometimes and the answer to the second is not by default.
3. The single-observer limitation in Section III
The practitioner observations in Section III carry a methodological limitation distinct from those of the experimental literature. They were gathered by a single observer, without a pre-registered coding protocol, across organisations whose selection was determined by employment trajectory rather than by sampling design. We acknowledged these limitations in §3.1 and did not present the observations as empirical evidence in the formal sense. We presented them as falsifiable propositions for future structured study, anchored in the convergent experimental literature reviewed in Section II. The work that would test those propositions (multi-site studies of technical discussion patterns in software-engineering teams across the LLM-adoption transition, ideally with archival material such as code review records and meeting transcripts) has not been done at the time of writing. That gap is itself a finding: the rapid normalisation of LLM-assisted technical work has outpaced the empirical apparatus needed to assess its developmental consequences. The practitioner observations we have offered are best read as a hypothesis that the patterns identified in controlled experimental settings reproduce at the scale of professional practice, and as a call for the structured observational work that would test that hypothesis.
4. Operational implications of the developmental–substitutive distinction
The distinction introduced at the close of Section IV has consequences that institutions, providers, and regulators can act on without waiting for further evidence or further legislation. We identify three.
For educational institutions deploying LLM-based assessment systems within the scope of Annex III(3)(b), the distinction reframes the Article 26 oversight obligation. The deployer is required to assign oversight personnel with the necessary competence; the question we raise is what the necessary competence must include. We argue it must include the capacity to identify, in the institution’s own programmes, which learning activities are constitutive of the capacities the programme is designed to build, and to design the use of LLM mediation accordingly. This is a curriculum-level competence rather than an operational one, and is properly the responsibility of academic governance rather than of compliance staff. The Article 4 AI literacy duty, read in conjunction with Article 26, supports this reading: the staff dealing with the operation of the system on the institution’s behalf must be in a position to make these judgements, and this requires literacy in the cognitive consequences of substitutive use, not only in the operational features of the tool.
For providers of LLM-based educational systems, the distinction has design consequences. A system that supports developmental use is designed to scaffold the learner’s own cognitive activity, for instance by providing hints rather than answers, requiring justification rather than accepting unargued conclusions, and surfacing the reasoning behind a recommendation rather than only the recommendation itself. A system that supports substitutive use is designed to minimise friction between request and output. The two design philosophies are not compatible at the level of default behaviour, and the compliance documentation required by Articles 11 to 13 should distinguish between them. A provider whose risk management system under Article 9 has not addressed which mode of use the system defaults to has not, in our view, complied with the spirit of the obligation, even if the documentation in formal terms is complete.
For regulators, the distinction suggests an interpretive direction for the harmonised standards now under development and for the Article 6(5) classification guidelines, published in draft in May 2026 and not yet finalised. Neither instrument is constrained to address only synchronic risks. The CEN-CENELEC mandate as amended in 2025 explicitly contemplates state-of-the-art research on human-AI interaction, and the Commission’s guidelines have latitude to clarify the interpretation of human oversight under Article 14 in ways that incorporate diachronic considerations. The cognitive-debt evidence, while still developing, is now sufficient to warrant explicit treatment in both instruments. Whether such treatment will materialise is a political question we do not pretend to answer. That it should is the regulatory implication of the argument we have made.
5. What remains open
We close this section with an explicit register of what the present paper does not establish. It does not establish the magnitude of cognitive debt under any specific pattern of LLM use; that requires longitudinal studies that have not yet been conducted. It does not establish that the regulatory measures we suggest would, if adopted, produce the outcomes we describe; that requires regulatory experiments and impact evaluation that the AI Act’s implementation phase will, in time, generate. It does not establish that the cognitive-debt mechanism is the dominant risk in educational AI; competing risks (algorithmic bias in assessment, privacy violation, surveillance overreach) may, on closer analysis, prove more urgent in particular contexts. What it does establish, we believe, is that cognitive debt is a phenomenon of genuine regulatory relevance, that the AI Act in its present form does not address it, and that the convergent evidence base is now sufficient to make its consideration overdue. The work that follows from this position, across empirical, doctrinal and political dimensions, is, in our view, the more interesting work.
VI. Conclusion
The empirical and regulatory landscape on which we have written this paper is moving rapidly. The neurocognitive evidence on LLM-mediated cognition is at an early stage and will, in the year following submission, likely be substantially extended by studies that we have not been able to incorporate. The harmonised standards under development by CEN and CENELEC, the European Commission’s Article 6(5) implementation guidelines, and the secondary instruments that will shape the AI Act’s operational meaning in education are, at the moment of writing, taking the form they will hold for the coming years. The question of how the regulation governs the technology is being decided now, in implementation rather than in legislation, and the conditions under which that decision is made will determine for some time whether the cognitive consequences of educational AI fall within the regulatory frame at all.
Our argument has been that they should. The neurocognitive evidence on cognitive debt, while provisional in many of its details, has converged across methods and populations sufficiently to make the phenomenon one of clear regulatory relevance. The AI Act, as it stands, addresses the synchronic risks of the high-risk educational systems within its perimeter (output errors, biased training data, automation bias at the moment of decision), and does so seriously. It does not address the diachronic risk that sustained engagement with such systems progressively erodes, in the persons assigned to oversee them, the cognitive substrate that meaningful oversight requires. The blind spot is built into the design of the regime, and the evidence base for closing it is now available.
What we have not attempted in this paper is to specify the policy instrument by which the closure should occur. We have indicated three candidates (the harmonised standards, the implementation guidelines, and the Article 4 AI literacy duty interpreted with reference to cognitive consequences rather than only operational competence), and we believe each is feasible, but we have not argued that any is sufficient. The harder questions are political rather than analytical, and concern whether the European institutions and the educational sector are prepared to recognise that the developmental consequences of AI mediation are part of what high-risk regulation in this domain must address, or whether the regime will continue to address only the discrete decision and leave the trajectory unattended.
The choice between those positions is, in our view, the more interesting work. We hope this paper contributes to the conditions under which it is taken seriously.
Acknowledgments
The authors thank the anonymous reviewer of the European Journal of Risk Regulation for an exceptionally careful, generous, and constructive reading of the manuscript, and for comments that materially strengthened the final version.
Financial support
This research received no specific grant from any funding agency in the public, commercial or not-for-profit sectors.
Competing interests
The authors declare no competing financial or non-financial interests relevant to the subject matter of this manuscript.
Use of generative AI in manuscript preparation
In accordance with current editorial practice on the transparent acknowledgement of generative AI, the authors disclose that large language model assistance was used during the drafting and editing of this manuscript, primarily for prose refinement, structural clarification and bibliographic formatting. All substantive arguments, factual claims, citations and conclusions are the intellectual product of the authors, who take full responsibility for the content of the manuscript.
Bibliography
Journal articles and conference papers
Bainbridge L, “Ironies of Automation” (1983) 19(6) Automatica 775
Lee H-P and others, “The Impact of Generative AI on Critical Thinking: Self-Reported Reductions in Cognitive Effort and Confidence Effects from a Survey of Knowledge Workers” in Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems (ACM 2025)
Manganello F and others, “Testing the Applicability of a Governance Checklist for High-Risk AI-Based Learning Outcome Assessment in Italian Universities under the EU AI Act Annex III” (2025) 8 Frontiers in Artificial Intelligence
Risko EF and Gilbert SJ, “Cognitive Offloading” (2016) 20(9) Trends in Cognitive Sciences 676
Stadler M, Bannert M and Sailer M, “Cognitive Ease at a Cost: LLMs Reduce Mental Effort but Compromise Depth in Student Scientific Inquiry” (2024) 160 Computers in Human Behavior 108386
Suriano R and others, “Student Interaction with ChatGPT Can Promote Complex Critical Thinking Skills” (2024) Learning and Instruction (forthcoming)
Laux J and Ruschemeier H, “Automation Bias in the AI Act: On the Legal Implications of Attempting to De-Bias Human Oversight of AI” (2025) European Journal of Risk Regulation (advance access)
Working papers and online sources
Cui Z and others, “The Effects of Generative AI on High-Skilled Work: Evidence from Three Field Experiments with Software Developers” (SSRN paper, 2024)
Kosmyna N and others, “Your Brain on ChatGPT: Accumulation of Cognitive Debt when Using an AI Assistant for Essay Writing Task” (arXiv preprint, June 2025) <https://arxiv.org/abs/2506.08872>
Marrone R and Kovanovic V, “MIT Researchers Say Using ChatGPT Can Rot Your Brain. The Truth Is a Little More Complicated” (The Conversation, 23 June 2025)
Sarkar A, “When Copilot Becomes Autopilot: Generative AI’s Critical Risk to Knowledge Work and a Critical Solution” (arXiv preprint, December 2024) <https://arxiv.org/abs/2412.15030>
Legislation
Regulation (EU) 2024/1689 of the European Parliament and of the Council of 13 June 2024 laying down harmonised rules on artificial intelligence (Artificial Intelligence Act) [2024] OJ L 2024/1689


