1. Introduction and problem
Before the bell, in a secondary history class, the screen assembles three objects. A chat answers about an episode from the past in the tone of a lesson. In the margin, a model confidence score. Below, a plan already sequenced for the next hour. The contract calls the bundle an autonomous agent and leaves the teacher an accept button. The scene is not field data for this article: it is the display ceiling the market already treats as support. Nothing is written, before use, about what counts as an acceptable answer. There is no trace with which to read the frame. There is no guarantee of stopping the explanation if the attribution of responsibility is false. There is no name for who answers, and no threshold at which the decision returns to a person.
The thesis does not fit in the short title. Those four objects do not constitute professional support if the supervision protocol is missing. The craft is the judgment that delegates, corrects, or stops. Fisher (2024) argues, as a philosophical thesis and not as a classroom trial, that a model can issue propositional content in response to a factual question without having assessed its truth: fluency is not a norm of veracity. Papaioannou (2026) shows, in a comparative qualitative study of 1922 in Asia Minor, that the language of the prompt is associated with terminology, with the attribution of responsibility, and with historiographical framing. Status: argument in one case, empirical finding of framing in the other. Neither proves that a chat “already teaches.” They ask something else: whether a teaching act was delegated, or whether a surface that speaks like a teacher was displayed.
The text rests on two distinctions already made in this series and does not remake them. On 16 September 2026 the harness was understood as the layer that turns generative capacity into work: bounded tools, state, verifiers, a system stop. On 18 September the evaluation of agents was understood as the layer that measures whether that work is reliable against a criterion external to the model’s prose. Neither is the judgment of someone who has a group, a curriculum, and a responsibility. Human-in-the-loop supervision (HITL) is the third layer: governance. A system can produce work and still not deserve delegation. It can have been measured on another distribution and still fail in this explanation, with this group. Design inference, not a finding of our own: without the third layer there is a product and there is a metric; there is no professional support.
Six illegitimate substitutions organize the problem. Calling autonomy what still requires a responsible person: Nemoto (2026) proposes agents whose ontological status is unresolved, yet which enter a sustained relation with the person who uses them. Treating the copilot as a substitute: Kim et al. (2026) show that human–AI collaboration, as a condition, does not homogenize how machine teachers are valued. Confusing literacy with clicks. Reading an agent evaluation as situated judgment. Believing that the harness already governs the educational act. Promoting the score to a criterion: Zhai et al. (2024) define over-reliance as accepting the artificial dialogue without questioning it when reliability cannot be assessed. Zhang and Tur (2024) register plans as a cited affordance and, in the same movement, register concerns. Affordance is not the craft. This article separates artifacts and judgment, does not pretend to protocols the corpus did not measure, and proposes four tests that are not a standard. No sample sizes, d, r, AUC, percentages, or DOIs are invented.
2. State of the art: situated practice vs artifact
The discourse of the “AI teacher” crushes four strata that should not be mixed: human-centered design (Alfredo et al., 2024; Brusilovsky, 2024; Dakshit et al., 2026; Kotsis and Stylos, 2026); the craft of those who form teachers (MacPhail et al., 2026; Moorhouse and Kohnke, 2024; Wan and Gu, 2026); critical literacy, not interface skill (Nabhan and Habók, 2026; Baran et al., 2026; Zhang and Samsudin, 2026; Anders and Dux Speltz, 2025; Brünner et al., 2025; Traga Philippakos and Rocconi, 2025; Haroud and Saqri, 2025); and the limit of the model when it is made to speak as a teacher (Papaioannou, 2026; Nemoto, 2026; Kudina and de Boer, 2025; Fisher, 2024; Zhai et al., 2024; Yan et al., 2024; Kim et al., 2026; Zhang and Tur, 2024). The artifact lives in the advertisement. Situated practice lives in identity, in professional learning, and in the protocol.
Alfredo et al. (2024) review human-centered learning analytics and educational AI: excluding teachers and students from design feeds mistrust, and the balance between human control and automation remains open. Brusilovsky (2024) recalls that such control was pioneered in education and later fell behind analogous work on recommender systems. Dakshit et al. (2026) propose PEARL —personalization, explainability, attribution, representation, and situated agency— illustrated with a simulated tutor, not with a classroom veto. Kotsis and Stylos (2026) read the European AI Act alongside STEM: transparency, accountability, and human oversight condition pedagogy and agency, and they treat AI as an epistemic actor. The statute is not yet a school protocol; it does make it illegitimate to call autonomous what the norm wants under the watch of persons.
The craft does not begin at the button. MacPhail et al. (2026) describe, through interviews in five jurisdictions, school-based teacher educators: role responsibility, heterogeneity, a disposition to work with those in formation, and development needs the post does not resolve. It is not an AI study. It is the subject a copilot does not download. Moorhouse and Kohnke (2024) listen to English teacher educators in initial language teacher education in Hong Kong: they anticipate effects on curriculum, teaching, and assessment, and most report a lack of confidence and competence. Wan and Gu (2026) review human–machine dialogic learning and tie it to pedagogy, environment, and platform. Dialogue is not, by itself, professional capacity. There is situated practice when the educator remains the subject of the curriculum. There is an artifact when the generated plan takes that place.
Literacy in this corpus is already judgment, not a click. Nabhan and Habók (2026) validate the ED-AI framework for language teachers: knowledge, evaluation, collaboration, contextualization, teacher autonomy —not the model’s— and ethics. Coefficients are not imported. Baran et al. (2026) design critical professional learning without an effect this text cites. Zhang and Samsudin (2026) compare an AI-integrated genre pedagogy with a generic program, from familiarity to criticality, without imported effect sizes. Anders and Dux Speltz (2025) join functional, critical, and creative literacy to planning, iterating, and evaluating, on the student side. Brünner et al. (2025) make probability, context, and manipulation visible. Traga Philippakos and Rocconi (2025) separate the tool, reported confidence, and the need for professional development. Haroud and Saqri (2025) separate support and replacement. Yan et al. (2024), Kudina and de Boer (2025), Lemasters and Hurshman (2025), and Balducci (2024) set the floor —ethical caution, functionalized language, orality, tasks with process— without themselves installing a veto protocol.
3. Review method
A critical narrative review was conducted, not a meta-analysis and not a rate of “copilot success.” The argument is one of category: what counts as support when a teacher supervises a system, and what remains an artifact. The window is 2021–2026. Where print year and online year diverge, the print year is cited: Wan and Gu (2026); Lemasters and Hurshman (2025); Kudina and de Boer (2025); Yan et al. (2024); Zhang and Tur (2024). Inclusion: peer-reviewed work with a Crossref DOI verified on 22 September 2026, slot 09:02 America/Mexico_City, on human supervision, teacher AI literacy, human-centered educational AI, teacher agency, or the limits of the model in a teaching role. Early-childhood axes already published in this series, catalogs without a paper, and any figure absent from the sources were excluded. The twenty-four entries of fuentes.md were used. No authors, journals, or DOIs were added.
Each source is marked with one of three statuses. Empirical finding: what was observed. Framework: an argument or review that does not trial the four tests. Design inference: what this article concludes and cannot attribute as a result of the source. Without a study that equates them to support, chat, score, plan, label, or dashboard are a category ceiling. There is no PRISMA of our own. The tests in section 7 are hypotheses, not a standard. Producing work is not governing the teaching act.
4. Axis 1. The artifacts do not constitute support
The support this axis denies is not that of an adult in an early-childhood room. It is the support a vendor declares by saying the teacher “already has a copilot.” The chat that “teaches” is the first artifact. Fisher (2024) sets the limit: faced with a question of fact, the model can release content without having assessed its truth. If the speaker is not under that norm, there is theater of explanation. Papaioannou (2026) adds a finding of framing: the language of the prompt is associated with how the episode is named, with who bears responsibility, and with the historiography. The chat inherits a frame; it does not “cover the topic.” Zhang and Tur (2024) map plans and materials as cited uses, not as a verdict that the plan is the act of teaching. Support begins where someone with the craft accepts or rejects that frame.
The confidence score is not a criterion written before use. It is a signal from the system about its own output, or a platform metric: this corpus does not audit a product and does not invent a threshold. Zhai et al. (2024) show the shortcut. Over-reliance is accepting the dialogue’s recommendation without a question, especially when assessing reliability is hard. A high score reduces the question precisely when the question is the craft. Yan et al. (2024) leave automated feedback and grading among the practical and ethical challenges: that a system scores does not mean the teacher has fixed which error is not delegated. Model confidence and teacher criterion are not convertible. One speaks of the system. The other speaks of this situation, this source, and this consequence if the frame is false.
The generated plan does not inherit the craft either. Wan and Gu (2026) tie the affordances of human–machine dialogue to pedagogy, environment, and platform. A file downloaded outside that design does not inherit the review. Moorhouse and Kohnke (2024) hear educators who see curriculum, teaching, and assessment already affected and who, for the most part, do not feel competent: the plan would move the gap to the accept button. Haroud and Saqri (2025) keep support and replacement apart. A plan that occupies the criterion has crossed sides even if the brochure says support. Balducci (2024) recalls that the security of a task depends on authenticity, collaboration, and process, not on fluency. The plan can be an input. Exhibited as the class, it is an artifact.
The “autonomous” agent, and the catalog AI teacher or machine teacher, close agency falsely. Nemoto (2026) proposes indeterminacy, not resolved intentionality. Kim et al. (2026), in two experiments with undergraduate students in the United States, report that valuation differs by gender and that the difference is sharper under the label of human–AI collaboration. Status: student perception, not a protocol audit. Kudina and de Boer (2025) add that functionalizing language thins judgment. The dashboard without a protocol and the button that does not stop the act complete the series. Alfredo et al. (2024) ask for real control, not a graphic that hides the trace. Dakshit et al. (2026) ask for traceable decisions. Fluency for veracity, score for criterion, plan for craft, label for agency, dashboard for supervision, and button for veto: none of these substitutions is in the corpus.
5. Axis 2. Relational craft in the garden
The heading belongs to the series template. The referent does not. There is no early-childhood classroom here. The ground is the relation between a responsible person and a system that produces text, plans, and scores. There is craft when that person fixes the criterion before delegating, reads the trace, can veto, names who is responsible, and knows the threshold at which the decision ceases to be the model’s. MacPhail et al. (2026) describe, without speaking of AI, the subject: role responsibility, heterogeneity, a disposition to work with those in formation, and professional development the post does not resolve. A copilot does not recruit that identity. HITL protects it; it does not turn it into interface operation.
The first gesture is the prior criterion, not the impression left by the demo. Anders and Dux Speltz (2025) show, on the student side, that planning includes domain knowledge and criteria before iterating. That taxonomy of human-in-the-loop practice is not a teacher protocol: the learner holds the cycle. Nabhan and Habók (2026) include evaluation and ethics alongside knowledge, collaboration, contextualization, and professional autonomy: to be literate, in that framework, is to judge, not to press. Baran et al. (2026) place criticality in the professional learning of teacher educators. Traga Philippakos and Rocconi (2025) separate knowing the tools and reporting confidence from the need for development. The criterion is written as an intolerable error —a false datum, an unjust attribution, a task the students were supposed to do, an assessment that is not delegated— not as a score.
The second gesture is the trace. Papaioannou (2026) shows that frame and attribution sit in the output and change with the language of the question: whoever does not read accepts the mirror. Dakshit et al. (2026) ask for decisions that are traceable for students, teachers, and administration. PEARL is demonstrated by simulation; the principle is not thereby validated as school practice. Alfredo et al. (2024) tie trust to participation in design and deployment. The minimum trace is what was said, under which prompt, and what was left out, not a participation summary. Yan et al. (2024) recall that questioning, feeding back, or grading with a model does not become harmless by adoption. Without a trace there is faith in fluency, not review.
The third gesture is the veto in the moment, not an administrator setting. Brusilovsky (2024) recovers learner control as a research debt: it neighbors the veto and is not the veto. A student may want to direct a trajectory, and the teacher must still be able to cut a false explanation. Kotsis and Stylos (2026) place human oversight under the AI Act in European STEM: the norm asks for the watch of persons and does not design a button. Zhai et al. (2024) describe the cost of being unable to refuse the recommendation. Kim et al. (2026) show that the collaboration label does not equalize perceptions. The veto is not an average of satisfaction.
Attribution and escalation travel together. Nemoto (2026) blocks loading professional blame onto an indeterminate agency. Fisher (2024) blocks treating the model as a speaker that has already assessed the truth. Kudina and de Boer (2025) describe the thinning of the person who was supposed to judge. The name of whoever adopts the output, and of whoever reviews whether it enters assessment or sensitive content, is written outside the model. Balducci (2024) and Lemasters and Hurshman (2025) do not describe a copilot threshold: they push tasks with a human process and, in philosophy, an orality the model does not sustain in silence. Grading, certifying, handling harm, leaving the curriculum, or contradicting a binding source escalates to a named person, with the cut decided beforehand and not when the score looks high.
The sixth gesture is critical literacy. Zhang and Samsudin (2026) set discipline and textual genre against a generic program. Brünner et al. (2025) make probability, context, and manipulation visible. Haroud and Saqri (2025) keep replacement distinct from support. Moorhouse and Kohnke (2024) locate the gap in teacher educators’ confidence and competence: the response is professional learning, not abdication of the curriculum. Knowing what not to delegate, how to read a trace, when to stop, and to whom responsibility returns is the literacy that counts. A prompt course can raise reported confidence and leave the protocol absent. The teacher uses, corrects, or switches off. The teacher is not a customer of a function.
6. Contrast. Category boundaries
Human supervision is not a timid autonomy. Nemoto (2026) leaves agency indeterminate. Kotsis and Stylos (2026) require oversight as a condition of teacher agency under the Act, not as the small print of an agent that “no longer needs” the teacher. Kim et al. (2026) study machines that teach as a matter of student perception: that role is not a teaching post. Calling the copilot autonomous and adding “with a human in the loop” in a footnote erases the boundary instead of crossing it. HITL denies that the system is the subject of the act.
Copilot and substitute are separated by who can say no. Haroud and Saqri (2025) treat support and replacement as categories that teachers and students do not live in the same way. Moorhouse and Kohnke (2024) leave curriculum, teaching, and assessment with educators who are affected and often underprepared: the subject remains the educator. Zhang and Tur (2024) may register help with planning without turning the model into the responsible planner. The copilot produces inputs under someone else’s criterion. The substitute occupies the criterion.
Literacy and click training do not coincide either. Nabhan and Habók (2026) disaggregate knowledge, evaluation, collaboration, contextualization, teacher autonomy, and ethics: “I know how to use the tool” does not exhaust the construct. Baran et al. (2026), Zhang and Samsudin (2026), and Brünner et al. (2025) seek situated criticality, probability, context, and manipulation, not generic familiarity. Anders and Dux Speltz (2025) require criteria set by the person before iterating. Ending a workshop at “I can already generate the plan” can be functional literacy and still be governance illiteracy.
Agent evaluation is not professional judgment, and the harness is not governance. In the sense fixed on 18 September, evaluation measures reliability on a distribution of tasks, not in this hour. Yan et al. (2024) accumulate challenges that a row of results does not close: being well measured at generating items does not authorize grading where the threshold has not been written. In the sense of 16 September, the harness produces work with tools, state, verification, and a technical stop. That stop obeys a budget or a verifier. It does not know a binding source, nor a relation in which it is better not to delegate even if the verifier is green. One can have a harness and an evaluation and still not have support.
The score is not the acceptance criterion. Zhai et al. (2024) tie over-reliance to the difficulty of assessing reliability. Dakshit et al. (2026) ask for explainability and justifiable decisions, not a scalar that replaces judgment. A score may be a clue; it is not the criterion that should have been written first. Other internal boundaries are crushed with the same haste. The human in the loop in Anders and Dux Speltz (2025) is the self-regulating student. Brusilovsky’s (2024) learner control is not the teacher’s veto. PEARL’s agency is not institutional attribution of responsibility. The dialogue in Wan and Gu (2026) may form teachers; it does not govern. The orality of Lemasters and Hurshman (2025) and the task principles of Balducci (2024) are human cuts of another grain: they do not, alone, install trace review. The category is conjunctive. One piece missing, and the system falls back into artifact, even if the slide says copilot.
7. Four tests of support (not an artifact)
What follows is a design inference of this article, anchored in the corpus and delivered by no source as a standard. It is not a norm, not a validated rubric, and it does not score products. If a chat, a score, a plan, an “autonomous” agent, a dashboard, or a button fails the four tests, it cannot be declared professional support.
7.1. Test of an acceptance criterion prior to use, not of the score or the demo. Anders and Dux Speltz (2025) place criteria before iteration. Nabhan and Habók (2026) treat evaluation as a dimension of literacy, not as an accessory. Baran et al. (2026) inscribe criticality in the learning of teacher educators. Traga Philippakos and Rocconi (2025) separate knowing the tool from being prepared to judge it. The test is met when, before delegation, it is written what output is acceptable, what error is not delegated, and what uses are out of bounds, for example certifying or closing a dispute of fact. If the only criterion is a high score or an impressive demo, the test fails. The criterion speaks of the curriculum and of possible harm, not of the model’s subjective assurance.
7.2. Test of traces the teacher can review, not of the engagement dashboard. Dakshit et al. (2026) require traceable decisions. Papaioannou (2026) shows that the frame travels in the output. Alfredo et al. (2024) tie control to real participation, not to an aggregate graphic. Yan et al. (2024) leave open the challenges of questioning, feedback, and grading: inspection is part of the ethical response. The test is met when the teacher can read what was asked, what the system answered, which tools intervened if any, and which version reached the students. If the teacher sees only participation, screen time, or “mean confidence,” there is display surveillance, not review.
7.3. Test of the right to veto and to stop, not of the collaboration label. Brusilovsky (2024) recalls that control is not given by the word collaborative, and that learner control does not exhaust the teacher’s veto. Kotsis and Stylos (2026) make human oversight a condition of agency, not an optional mode. Zhai et al. (2024) describe the cost of being unable to refuse the recommendation. Kim et al. (2026) show that labeled collaboration does not erase differences in valuation: the taste of part of the group is not a license to continue. The test is met when the teacher can prevent an output from arriving, or withdraw it, without waiting for another purchasing cycle and without the system translating the cut into a style preference. If stopping requires a privilege the classroom teacher does not have, or if the button only confirms what was already delivered, there is choreography.
7.4. Test of explicit attribution of responsibility and of an escalation threshold. Nemoto (2026) offers no subject on whom to load professional blame. Fisher (2024) offers no speaker who has assessed the truth. Kudina and de Boer (2025) describe the thinning of the person who was supposed to judge. MacPhail et al. (2026) recall that role responsibility is an identity, not a signature at the foot of a plan. Moorhouse and Kohnke (2024) leave teacher educators as subjects of the curriculum: the gap calls for learning, not for surrendering the name. Balducci (2024) and Lemasters and Hurshman (2025) keep human cuts in the task and in orality. The test is met if it is written who answers and when the decision escalates to a named person —grading, harm, leaving the curriculum, a binding source, a disagreement the trace does not resolve—. “The AI is responsible,” without a name or a threshold, is a disclaimer.
The four are read together. A criterion without a trace cannot be applied. A trace without a veto is an archive. A veto without attribution does not say who will stand by the cut tomorrow. A threshold without a prior criterion is invented under pressure, when the score most invites one not to think. A plan, a formative dialogue (Wan and Gu, 2026), a student cycle (Anders and Dux Speltz, 2025), explainability in the manner of PEARL (Dakshit et al., 2026), or a reliability measurement made in another layer may sit inside the craft as subordinate pieces. None, alone, is support. Catalog autonomy does not turn them into a detail for later.
8. Discussion
Three tensions organize the discussion. The first opposes displaying artifacts and exercising the craft. Fisher (2024), Papaioannou (2026), Nemoto (2026), and Kim et al. (2026) do not say the same thing: absence of truth-assessment, frames that move with language, unresolved agency, collaboration that does not equalize perceptions. The market crushes them into “the AI teacher has arrived.” The piece may be true in its domain. “Piece = professional support” is not. Zhang and Tur (2024) may list the plan among affordances without authorizing that equation. Wan and Gu (2026) may show dialogue useful for teacher learning without turning the chat into the educator.
The second opposes the series’ previous layers to governance. The harness produces checkable work. Evaluation measures reliability on a distribution. Both can be well made and this hour’s act can still be non-delegable: because the frame misattributes a historical responsibility (Papaioannou, 2026), because the person using the system does not detect the failure (Zhai et al., 2024), because assessment has not been redesigned (Balducci, 2024; Lemasters and Hurshman, 2025), or because the Act asks for oversight and teacher agency that an autonomous mode contradicts (Kotsis and Stylos, 2026). Buying a harness and buying an evaluation does not buy HITL. Omitting the third layer —“the model is already reliable, therefore it already supports”— is a display bias, not a conclusion of the reviews.
The third opposes critical literacy and click confidence. Moorhouse and Kohnke (2024) find educators who anticipate the impact and do not feel competent: that is a datum, not a moral failure. The response coherent with Baran et al. (2026), Zhang and Samsudin (2026), Brünner et al. (2025), and Nabhan and Habók (2026) is learning that includes evaluation, ethics, context, and situated criticality. The incoherent response is a prompt workshop that raises reported confidence (Traga Philippakos and Rocconi, 2025) and calls that literacy. Haroud and Saqri (2025) show that openness to adoption is not distributed equally between students and teachers, and that support and replacement are not synonyms. Confusing enthusiasm for use with a protocol leaves the school governed by whoever most wants to delegate.
The tests read these tensions as hypotheses, not as a classroom finding. PEARL is a simulation (Dakshit et al., 2026). MacPhail et al. (2026) did not study copilots. The human in the loop in Anders and Dux Speltz (2025) is a student, and Brusilovsky’s (2024) control is a research agenda, not a veto button. Criterion, trace, veto, and a named threshold are a candidacy for support. Chat, score, plan, label, dashboard, or button, alone, are consumption.
9. Limits
The review is narrative. It does not estimate pooled effects and does not apply a PRISMA of its own. Sample sizes, d, r, AUC, and percentages are not invented, even where some sources report them. Nabhan and Habók (2026) validate a construct; they do not rank teachers. Zhang and Samsudin (2026) are read for the move from familiarity to criticality, not for an imported effect. Kim et al. (2026) speak of undergraduates, not of the teaching profession. Moorhouse and Kohnke (2024) gather perceptions in Hong Kong. Papaioannou (2026) is a historiographical case. Brünner et al. (2025) and Baran et al. (2026) design professional learning; they do not audit products.
Other sources are floor or neighbor, not direct evidence of the protocol. MacPhail et al. (2026) study the identity of school-based teacher educators. Fisher (2024), Nemoto (2026), and Kudina and de Boer (2025) argue. Balducci (2024) offers an opinion on tasks. Lemasters and Hurshman (2025) propose a turn to orality. Dakshit et al. (2026) simulate PEARL. Anders and Dux Speltz (2025) document student practices. Brusilovsky (2024) reviews learner control. Alfredo et al. (2024), Yan et al. (2024), Zhang and Tur (2024), Zhai et al. (2024), and Wan and Gu (2026) map fields; they do not certify a copilot. Kotsis and Stylos (2026) analyze statute and agency in European STEM, not a school protocol. Haroud and Saqri (2025) and Traga Philippakos and Rocconi (2025) describe perceptions and needs, not compliance with the four tests.
In this kit there is no trial that equates chat, score, plan, or an autonomy label with support that has a criterion, a trace, a veto, and attribution. That absence is a category ceiling, not a magnitude of harm. Contracts are not audited and brands are not condemned. The four tests are not a validated instrument. The word “garden” in heading 5 does not transfer early-childhood findings: this article does not study that level. The corpus is in English and this synthesis is an editorial translation. The stack harness, evaluation, and HITL is a distinction of the series, not an empirical result of the twenty-four sources.
10. Conclusions
A chat that “teaches,” a score, a plan, or an “autonomous” agent does not constitute professional support if teacher supervision is missing. Neither does a dashboard or an accept button. The craft is the judgment that delegates, corrects, or stops. Fisher (2024), Papaioannou (2026), Nemoto (2026), Zhai et al. (2024), Kim et al. (2026), and Zhang and Tur (2024) fix the ceiling: fluency without truth-assessment, framing in the output, indeterminate agency, over-reliance, collaboration that does not equalize perceptions, and plans that are an affordance in the literature, not the craft.
Where there is craft there is a subject and a protocol. MacPhail et al. (2026) and Moorhouse and Kohnke (2024) leave responsibility with those who form teachers, even when they feel unsure: the response is learning, not surrender. Nabhan and Habók (2026), Baran et al. (2026), Zhang and Samsudin (2026), Brünner et al. (2025), Anders and Dux Speltz (2025), Traga Philippakos and Rocconi (2025), and Haroud and Saqri (2025) move literacy from the click to evaluation, ethics, criticality, and the distinction between support and replacement. Alfredo et al. (2024), Brusilovsky (2024), Dakshit et al. (2026), and Kotsis and Stylos (2026) ask for control, oversight, and traceable decisions. Yan et al. (2024), Kudina and de Boer (2025), Lemasters and Hurshman (2025), Balducci (2024), and Wan and Gu (2026) prevent closing the problem with a catalog role.
The stack is not substitutable from the inside. The harness produces work. Evaluation measures reliability. Human supervision governs delegation in situation, with four conjunctive tests: a prior criterion, traces the teacher can review, a right of veto, and explicit attribution with an escalation threshold. Where the sources did not measure that protocol, this article does not treat it as measured. Where they measured frames, perceptions, constructs, or arguments, these are not translated into “the agent already supports.” Turning a generative system into support is exercising judgment about when it serves, when it is corrected, and when it is stopped. The rest is chat, score, plan, or marketing autonomy. It is not the craft, and it should not be presented as what it is not.
Laboratorio Editorial de NEXTECH.IA / Ingeniero Mitre.
References
- Alfredo, R., Echeverria, V., Jin, Y., Yan, L., Swiecki, Z., Gašević, D., y Martinez-Maldonado, R. (2024). Human-centred learning analytics and AI in education: A systematic literature review. Computers and Education: Artificial Intelligence, 6, Article 100215. https://doi.org/10.1016/j.caeai.2024.100215
- Anders, A. D., y Dux Speltz, E. (2025). Developing generative AI literacies through self-regulated learning: A human-centered approach. Computers and Education: Artificial Intelligence, 9, Article 100482. https://doi.org/10.1016/j.caeai.2025.100482
- Balducci, B. (2024). AI and student assessment in human-centered education. Frontiers in Education, 9, Article 1383148. https://doi.org/10.3389/feduc.2024.1383148
- Baran, E., Dilek, M., Ziba, M., y Xiao, X. (2026). Human-centered AI for teacher educators: Designing professional learning for critical AI literacy. Computers and Education Open, 11, Article 100399. https://doi.org/10.1016/j.caeo.2026.100399
- Brünner, B., Schön, S., y Ebner, M. (2025). From Gretel to Strudelcity: Empowering teachers regarding generative AI for enhanced AI literacy with CollectiveGPT. Education Sciences, 15(2), Article 206. https://doi.org/10.3390/educsci15020206
- Brusilovsky, P. (2024). AI in education, learner control, and human-AI collaboration. International Journal of Artificial Intelligence in Education, 34(1), 122–135. https://doi.org/10.1007/s40593-023-00356-z
- Dakshit, S., Mokhtari, K., y Khalid, A. (2026). Designing understandable and fair AI for learning: The PEARL framework for human-centered educational AI. Education Sciences, 16(2), Article 198. https://doi.org/10.3390/educsci16020198
- Fisher, S. A. (2024). Large language models and their big bullshit potential. Ethics and Information Technology, 26(4), Article 67. https://doi.org/10.1007/s10676-024-09802-5
- Haroud, S., y Saqri, N. (2025). Generative AI in higher education: Teachers’ and students’ perspectives on support, replacement, and digital literacy. Education Sciences, 15(4), Article 396. https://doi.org/10.3390/educsci15040396
- Kim, J., Spence, P. R., y Kelly, S. (2026). Machine teachers and human-AI collaboration: Student differences in perceptions of AI-based education. Interactive Learning Environments. Advance online publication. https://doi.org/10.1080/10494820.2026.2667455
- Kotsis, K. T., y Stylos, G. (2026). The AI Act and the future of STEM education in Europe: Rethinking pedagogy, assessment, and teacher agency. Frontiers in Education, 11, Article 1845045. https://doi.org/10.3389/feduc.2026.1845045
- Kudina, O., y de Boer, B. (2025). Large language models, politics, and the functionalization of language. AI and Ethics, 5(3), 2367–2379. https://doi.org/10.1007/s43681-024-00564-w
- Lemasters, R., y Hurshman, C. (2025). A shift towards oration: Teaching philosophy in the age of large language models. AI and Ethics, 5(2), 1203–1215. https://doi.org/10.1007/s43681-024-00455-0
- MacPhail, A., Vanassche, E., Guberman, A., Czerniawski, G., y Boks-Vlemmix, J. (2026). School-based teacher educators’ professional identity and professional needs. Teaching and Teacher Education, 176, Article 105523. https://doi.org/10.1016/j.tate.2026.105523
- Moorhouse, B. L., y Kohnke, L. (2024). The effects of generative AI on initial language teacher education: The perceptions of teacher educators. System, 122, Article 103290. https://doi.org/10.1016/j.system.2024.103290
- Nabhan, S., y Habók, A. (2026). Language teachers’ AI literacy: A psychometric study based on the ED-AI framework. Computers and Education: Artificial Intelligence, 10, Article 100583. https://doi.org/10.1016/j.caeai.2026.100583
- Nemoto, R. (2026). Continuous intentionality and indeterminate agency in large language models. AI and Ethics, 6(3), Article 322. https://doi.org/10.1007/s43681-026-01181-5
- Papaioannou, C. (2026). Language and the framing of historical narrative in large language models: The case of Asia Minor (1922). AI and Ethics, 6(3), Article 262. https://doi.org/10.1007/s43681-026-01138-8
- Traga Philippakos, Z. A., y Rocconi, L. (2025). AI literacy: Elementary and secondary teachers’ use of AI-tools, reported confidence, and professional development needs. Education Sciences, 15(9), Article 1186. https://doi.org/10.3390/educsci15091186
- Wan, P., y Gu, X. (2026). Developing teachers’ professional abilities: A systematic review of human-machine dialogic learning for teacher education. Interactive Learning Environments, 34(2), 615–639. https://doi.org/10.1080/10494820.2025.2507280
- Yan, L., Sha, L., Zhao, L., Li, Y., Martinez-Maldonado, R., Chen, G., Li, X., Jin, Y., y Gašević, D. (2024). Practical and ethical challenges of large language models in education: A systematic scoping review. British Journal of Educational Technology, 55(1), 90–112. https://doi.org/10.1111/bjet.13370
- Zhang, P., y Tur, G. (2024). A systematic review of ChatGPT use in K-12 education. European Journal of Education, 59(2), Article e12599. https://doi.org/10.1111/ejed.12599
- Zhang, Y., y Samsudin, M. A. (2026). From familiarity to criticality: Cultivating EFL teachers’ AI literacy through an AI-integrated genre-based pedagogy. Education Sciences, 16(1), Article 150. https://doi.org/10.3390/educsci16010150
- Zhai, C., Wibowo, S., y Li, L. D. (2024). The effects of over-reliance on AI dialogue systems on students’ cognitive abilities: A systematic review. Smart Learning Environments, 11(1), Article 28. https://doi.org/10.1186/s40561-024-00316-7