1. Introduction and problem
In basic-education schools an alarm circulates that is symmetrical to the promise of personalization: generative artificial intelligence “cheats,” “writes the essay,” and “forces” the school to hunt the student with a detector. That alarm usually silences its condition of possibility. To hunt, a text is needed. To grade, an artifact. To accuse, a statistical trace—perplexity, burstiness, an “AI” percentage—presented as proof of authorship. The artifact is not abstract. It is a ten-year-old’s composition, an eight-year-old girl’s science paragraph, a dictated story in the final stretch of early childhood education. Those who sign it are learning to write, not defending a doctoral thesis. If the school evaluates the product as if it were transparent with respect to the subject, and a model can produce a plausible product, integrity ceases to be a honor code and becomes a problem of evidence: what counts as proof that that child learned?
The problem is not hostility to the tool nor indulgence toward fraud. It is confusion among three objects. The first is an empirical finding: what is observed when detectors are tested against human and generated texts; what primary pupils report about homework and copying; how teachers and students rank “learning” and “cheating” with ChatGPT. The second is a normative framework: what UNESCO, the European integrity network, the International Baccalaureate or a school district prescribe when they speak of age, attribution, prohibition or the presumption of innocence. The third is a pedagogical inference: what should not be sold as a solution—the infallible detector, the ban that restores authorship, the idea that the model “makes the child cheat”—if the source does not measure it. Mixing them produces a pedagogy of the hunt: the school “uses AI” or “bans it” without being able to say whether the text it grades has a subject, whether the child thought, or whether only a linguistic surface was scored.
The thesis of this article is restrictive. In basic education, the problem is not hunting cheats with a detector, but what counts as evidence that a child learned when a model can produce the artifact. Where the source does not measure that detectors work in primary classrooms, that is not asserted. Where it does not measure that AI “makes the child cheat,” that is not promised. Where it does not measure that banning ChatGPT restores authorship, that is not invented. The gap is the text without a subject. This article does not treat adaptive tutors (16 August), teacher well-being (21 August, 9:02), curricular literacy about AI (20 August evening), UNESCO teacher competence (19 August), privacy (21 August, 1:02), inclusion, parental mediation, play/PopBots, multilingualism, or gaps in “learning AI.” The object is integrity: authorship, evidence, detection and false positives.
2. State of the art: integrity, authorship and the assessable artifact
It is useful to separate three strata. The first is conceptual: what academic integrity means when the text is no longer a reliable index of a subject. The second is normative: which frameworks fix attribution, age, prohibition or redesign of assessment. The third is empirical: what has been documented on detectors, on primary or secondary pupils using generative models, and on school policies, not only university ones.
In the conceptual stratum, Tauginienė et al. (2018), in the glossary of the European Network for Academic Integrity (ENAI), define academic integrity as compliance with ethical and professional principles, standards and practices by individuals or institutions in education, research and scholarship. Foltýnek, Bjelobaba, Glendinning, Khan, Santos, Pavletić, and Kravjar (2023) introduce “unauthorized content generation”: academic work, in whole or in part, for credit, progression or award, with unapproved or undeclared human or technological assistance. That definition does not presuppose that using a model is, in itself, misconduct. Eaton (2023) names a horizon of “postplagiarism”: it is conceptual, not a classroom trial. Cotton, Cotton, and Shipway (2023) debate honesty risks in higher education; it is not a primary-school RCT. Lodge, Howard, Bearman, Dawson, and Associates (2023)—an expert document published by Australia’s Tertiary Education Quality and Standards Agency, not a regulation of that agency—propose multiple, contextualized judgments and emphasis on process, not only on the product. Dawson and Bearman (2020) called “future-authentic assessment” the representation of a discipline’s likely realities; that framework is university-based and is retained as grammar, not as a finding about basic education.
In the normative stratum, Miao and Holmes (2023) orient a human-centred use of generative AI and propose, among other measures, an age limit for independent conversations with genAI platforms (UNESCO’s institutional communication places it at thirteen) and ethical and pedagogical validation of tools by institutions. That text is cited here as a framework of integrity and agency, not as an axis of teacher competence. UNESCO’s (2021) Recommendation on the ethics of AI elevates human oversight and proportionality. The International Baccalaureate (2023) states that it will not ban AI software because banning is an ineffective way to deal with innovation, and that it does not regard any work produced—even only in part—by such tools as the student’s own: it must be credited in the body of the text and in the bibliography. Eaton (2024) insists on the presumption of innocence and that not all AI use constitutes misconduct. Those prescriptions do not measure children’s learning; their authority is normative.
In the empirical stratum, detector evidence is abundant and, overwhelmingly, university-based or drawn from adult and adolescent texts (Liang et al., 2023; Elkhatat et al., 2023; Weber-Wulff et al., 2023; Perkins et al., 2024). Evidence of primary pupils using models for homework is scarce; Oreški et al. (2025) is one of the few named studies of that band. Mah et al. (2024) document tensions between “learning” and “cheating” with writing teachers who include Pre-K and with linguistically diverse secondary students. District and school-program policies (Elsen-Rooney, 2023; Banks, 2023; International Baccalaureate, 2023) document bans, reversals and attribution rules, not effects on children’s authorship. The gap is not the absence of alarm: it is the scarcity of evidence that the written product, in basic education, remains proof of a subject.
3. Review method
A critical narrative review was conducted, not a meta-analysis. The purpose was not to estimate a homogeneous effect size—nonexistent across a fourteen-detector test, a questionnaire of 301 pupils and a district communiqué—but to articulate evidence of learning with verified sources. Inclusion: (a) integrity, authorship, detection, school use of models or school policies; (b) preferably 2021–2026, with still-current frameworks (Tauginienė et al., 2018; Dawson and Bearman, 2020; UNESCO, 2021); (c) journal or proceedings with DOI, organizational report, school-system statement or the classifier vendor’s technical announcement; (d) verifiable DOI or official URL. Adaptive tutors, teacher well-being, curricular AI literacy, UNESCO teacher competence, privacy, inclusion, parental mediation, play/PopBots, multilingualism and “learning AI” gaps were excluded as axes.
The search was run on 21 August 2026 (around 13:06, America/Mexico_City) on DOI pages, Springer, Cell/Patterns, MDPI, Taylor & Francis, UNESDOC, IB, Chalkbeat, TEQSA, arXiv, OpenAI and ERIC. Each source was verified against at least one of those pages. Nuclear cases: (4) detection; (5) primary school and the cheat/learn tension; (6) school policies. Sadasivan et al. (2023) saturate theoretical unreliability; Cotton et al. (2023) and Lodge et al. (2023) saturate the university frame and are marked as such.
The analysis distinguished three enunciative statuses. Empirical finding: what was observed in the sample or in the tool-testing record. Normative framework: what an organization or institution prescribes. Pedagogical inference: the translation into the basic-education classroom, marked as such. The limits are those of any narrative review (section 9).
4. Case 1. Detectors: false positives, evasion and the risk of accusing those who write as children or as non-natives
Liang, Yuksekgonul, Mao, Wu, and Zou (2023) published in Patterns an evaluation of seven GPT detectors on 91 TOEFL essays from a Chinese forum (human writing by non-native English speakers) and 88 U.S. eighth-grade essays from the Hewlett Foundation’s ASAP set. Empirical finding: the detectors accurately classified the U.S. eighth-grade essays but incorrectly labeled more than half of the TOEFL essays as “AI-generated” (mean false-positive rate: 61.3%). All unanimously identified 19.8% of the human TOEFL essays as AI-authored, and at least one flagged 97.8%. Those unanimously misclassified had lower text perplexity: that metric penalizes restricted linguistic ranges. Simple prompts reduced the bias and, at the same time, bypassed detectors. The authors caution against use in evaluative settings, particularly with non-native speakers. It is not a finding about early childhood or the first cycle: children’s text was not measured. Pedagogical inference, marked as such: if predictable L2 writing is confused with a machine, the writing of a child acquiring vocabulary is a plausible candidate for the same error. That inference is not measured; the paper’s only child-adolescent empirical anchor is that native eighth-grade essays were not massively misclassified.
Elkhatat, Elsaid, and Almeer (2023) tested, in the International Journal for Educational Integrity, classifiers from OpenAI, Writer, Copyleaks, GPTZero and CrossPlag on fifteen GPT-3.5 paragraphs, fifteen GPT-4 paragraphs and five human controls, all on an engineering topic (cooling towers). Empirical finding: the tools were more accurate with GPT-3.5 than with GPT-4; on human controls they showed inconsistencies, false positives and uncertain classifications. The object is university-technical. It does not measure children. It does not authorize saying that GPTZero or Copyleaks “detect cheating” in a primary notebook.
Weber-Wulff, Anohina-Naumeca, Bjelobaba, Foltýnek, Guerrero-Dib, Popoola, Šigut, and Waddington (2023) tested fourteen tools, including Turnitin and GPTZero, on 54 documents of known ground truth (human; human machine-translated; ChatGPT-generated; generated and hand-edited; generated and paraphrased with Quillbot): 756 tests. Empirical finding: they are neither accurate nor reliable (all below 80% accuracy; only five exceeded 70%). There is a bias toward classifying as human. Accuracy on original human text was 96% and dropped about twenty points with machine translation into English. Manual editing of generated text left accuracy around 42%; automatic paraphrasing, 26%. About 20% of unobfuscated AI texts would be attributed to humans; with obfuscation, about 50% or more. Six of fourteen tools produced false positives; with GPTZero, half of the positive classifications would be false accusations in the authors’ frame. Turnitin obtained the best accuracy in that battery and zero false positives on the sample’s humans, but failed against editing and paraphrasing. The authors conclude that those systems should not be used as the sole basis of an accusation, because they do not provide verifiable evidence. Perkins, Roe, Postma, McGaughran, and Hickerson (2024) inserted 22 GPT-4 submissions (with prompts to reduce detectability) among real university assessments in Vietnam, marked by fifteen faculty members with Turnitin. Empirical finding: the tool identified 91% as containing some AI content, but only 54.8% of the content; faculty referred 54.5% to misconduct; mean grades were 52.3 (AI) and 54.4 (students). It is not basic education. “Flagging as AI” is not equivalent to the real proportion nor to unanimous judgment.
OpenAI (2023) announced a classifier of AI-written text and, on 20 July 2023, withdrew it “due to its low rate of accuracy.” In its own evaluation on an English challenge set, it correctly identified 26% of AI-written text (true positives) and incorrectly labeled 9% of human text as AI (false positives). The company warned that it had not thoroughly assessed effectiveness on content written in collaboration with human authors. Sadasivan, Kumar, Balasubramanian, Wang, and Feizi (2023), in an arXiv preprint, argue empirically and theoretically that reliable detection is fragile: paraphrasing attacks degrade detectors (including watermarking schemes and neural classifiers) and, as models better emulate human text, even the best possible detector tends toward performance barely above a random classifier. Status: preprint and theoretical impossibility result, not a school trial.
Accumulated status. Empirical finding: 2023–2024 detectors, including products used in schools and universities, produce false negatives that are easy to obtain with editing or paraphrasing, and false positives that disproportionately hit non-native and machine-translated writing. The vendor that most visibly offered its own classifier withdrew it for inaccuracy. It is not a finding that Turnitin, GPTZero or another detector “works” as proof of a child’s authorship. Pedagogical inference: deploying an “AI” percentage on a basic-education child’s composition is to risk accusing someone who writes as they are learning to write—with low perplexity, with formulas, with translation—and, at the same time, letting the obfuscated artifact through. That is not integrity: it is a statistical lottery over a subject who cannot defend themselves with a detection methodology.
5. Case 2. Primary homework and the tension between scaffolding and shortcut: a subject who does use the model
Oreški, Oreški, and Ružić (2025) published in Informatics in Education a self-administered questionnaire to 301 students in primary grades 5–8 (ages 11 to 14) from two schools—one urban and one rural—in northwestern Croatia (10 February to 14 March 2025), in informatics classes and with informed consent. Convenience sample: 154 boys and 147 girls; 49 in grade 5, 115 in 6, 101 in 7 and 36 in 8. Empirical finding: 57.8% (174) report using AI tools for homework. There is no sex difference (χ² = 0.66, p = .417); there is a grade difference (χ² = 9.53, p = .023): from 40.8% in grade 5 to 66.7% in grade 8. Spearman correlation between daily social-media time and AI use: r = 0.34, p < .001. Mode of use: 36.2% read, think and use the answer as a guide; 25.5% adapt it; 15.0% copy parts; 12.3% copy the answer. Only 36.5% consider copying without modification unacceptable; 38.2% say it depends on the assignment; 16.3% that it is acceptable if the teacher does not notice. 13.0% call any homework use cheating; 31.2%, only if the whole answer is copied; 43.5% that it is not if it serves as help or ideas. History, Croatian language, Geography and Mathematics lead use by subject. The authors call for ethics content in informatics: that is a recommendation, not a measured curricular effect.
Status. Empirical finding from upper primary (not early childhood or the first cycle): homework use is majority, rises with grade, and the norm on copying is lax. It is not a finding that AI “makes the child cheat,” nor of learning gained or lost: there is no pretest. Pedagogical inference, marked as such: if at age eleven more than half use the model for the task that is graded, the product may have no subject. Distinguishing that is not a detector problem; it is evidence design (process, orality, visible revision). Transferring the 57.8% to the first cycle would be illegitimate: Oreški et al. (2025) did not measure that band, and Miao and Holmes (2023) mark an age threshold that advises against that extrapolation.
Mah, Walker, Phalen, Levine, Beck, and Pittman (2024), in Education Sciences, asked sixteen Pre-K through postsecondary writing teachers (Bay Area Writing Project) and twelve secondary students from a linguistically diverse charter to rank vignettes of ChatGPT use by how much was learned and how much was cheated. Empirical finding: teachers and students used similar criteria; they reached similar conclusions about learning and divergent ones about cheating. Disagreements clustered in four tensions: shortcut versus scaffold; generating ideas versus generating language; ChatGPT support versus analogous supports (tutor, family, internet); learning with the model versus learning without it. The study does not measure detection or achievement. It includes Pre-K teachers, but participating students are secondary, not early childhood. Pedagogical inference: even when there is agreement on whether learning occurred, there is none on whether cheating occurred. In basic education, where the child is learning to be an author, that dissociation is the core: one can grade a “clean” text and not know whether there was a subject, or accuse a “suspicious” text and not know whether there was learning. Mah et al. (2024) do not resolve that dissociation with a numerical threshold; they document that the threshold is normative and disputed.
6. Case 3. School policies: banning, attributing, not restoring authorship
On 3 January 2023, the New York City Department of Education blocked ChatGPT on school devices and networks. Elsen-Rooney (2023) reported the official justification: negative impacts on student learning and concerns about safety and accuracy; the block did not cover devices outside the department. In May, Chancellor David C. Banks (2023) called that response a “knee-jerk” fear and shifted the accent toward exploration with teacher support. Documentary-empirical finding: the largest U.S. school system banned institutional access and, months later, reversed the panic frame. It is not a finding of reduced cheating or restored authorship: those variables are not measured. Pedagogical inference: a network ban does not return the subject to the text written at home or on a phone.
The International Baccalaureate (2023), on 1 March 2023, states that it will not ban AI software because banning is ineffective in the face of innovation. It does not regard any work produced, even only in part, by such tools as the student’s own: it must be credited in the body and in the bibliography. Status: a framework of a program that includes the Primary Years Programme and the Middle Years Programme, not a detector trial or a measure of learning. Foltýnek et al. (2023) ask for policies, training, transparent acknowledgment and human responsibility; AI use is not automatically misconduct and a system should not be listed as co-author. Eaton (2023, 2024) adds postplagiarism and the presumption of innocence. Miao and Holmes (2023) ask for institutional validation and an age limit for independent conversations. Lodge et al. (2023)—experts, not a TEQSA mandate—propose process and multiple judgments.
Accumulated status. Policy finding: there are schools and programs that banned, that reversed, and that, instead of banning, require attribution. Normative framework: ENAI, UNESCO/Miao and Holmes, the IB and Eaton coincide, with nuances, that integrity does not reduce to hunting software and that the human remains responsible. It is not a finding that those policies restored children’s authorship. Pedagogical inference: in basic education, copying the university policy of “declare the prompt” may be unintelligible to a seven-year-old and, at the same time, copying the “ChatGPT forbidden” policy does not prevent the home artifact. What can be designed, as a hypothesis of evidence and not as a measured result, is an assessment in which the subject is visible: process, orality, production in presence, revision. That is consistent with Lodge et al. (2023) and Weber-Wulff et al. (2023) when they ask to prioritize prevention and process over detection; it is not a demonstration that such a design already works in early childhood or the first cycle.
7. Inferential framework: evidence of learning, not hunting artifacts
The framework that follows is this article’s pedagogical inference, anchored in the cases and in the verified instruments. It is not a new international standard. It distinguishes five tests. If a proposal of “integrity with AI” in basic education does not pass them, it is not deployed as a solution.
7.1. Test of the visible subject. Weber-Wulff et al. (2023) conclude that detectors do not provide verifiable evidence and that written assessment should focus on the process of skill development, not the final product. Lodge et al. (2023) propose emphasis on process. Inference: in basic education, grading only the file or the clean notebook is to accept a text that may have no subject. Evidence of learning is the traceability of the child’s thinking—drafts, dictation, conversation, correction in view—not the polish of the paragraph. This test does not claim that process is already measured in early childhood; it claims that, without it, the product is ambiguous.
7.2. Test of the false positive as asymmetric harm. Liang et al. (2023) document bias against non-native writing; Weber-Wulff et al. (2023) document increased false positives with machine translation and, for GPTZero, a high risk of false accusation in their battery; OpenAI (2023) withdrew its classifier for inaccuracy. Eaton (2024) requires a presumption of innocence. Inference: in a basic-education classroom, a false positive is not a minor technical error. It is accusing a child who is learning to write, or a child who writes in a second language, or a child whose text passed through a translator. The power asymmetry—the teacher holds the percentage; the child has no expert method—replicates the imbalance that integrity norms say they avoid. A detector is not used as sole proof. It is not promised that “with another threshold” the harm disappears: no nuclear source demonstrates that for primary school.
7.3. Test that the model is not the cheater. Oreški et al. (2025) show heterogeneous uses: guide, adaptation, partial copy, full copy, and divided beliefs about whether that is cheating. Mah et al. (2024) show that “learning” and “cheating” do not coincide. Foltýnek et al. (2023) hold that AI use is not automatically misconduct. Inference: saying that “AI makes the child cheat” attributes moral agency to a token statistic and withdraws it from school design (what is asked, under what conditions, with what evidence). The child may copy; the model does not “cheat.” Where the source does not measure intent, this article does not invent it.
7.4. Test that banning does not restore authorship. Elsen-Rooney (2023) and Banks (2023) document a ban and its reversal without evidence of restored authorship. The International Baccalaureate (2023) declares prohibition ineffective as a way of dealing with innovation. Inference: blocking a URL on the classroom network does not reconstruct the subject in the text submitted from home. It may have other aims (safety, age, data); those aims do not become a finding of integrity. Miao and Holmes (2023) mark an age limit for independent conversation: that is a protection framework, not proof that the ban makes six-year-olds into authors.
7.5. Test of age-proportionate evidence. Miao and Holmes (2023) and the thirteen-year threshold for independent conversation place early childhood education and the first cycle below autonomous use of generative chat. Oreški et al. (2025) measure ages 11–14, not 5–8. Inference: the younger the age, the less sense it makes to hunt an essay and the more sense it makes not to ask for an artifact a model can produce in silence. In early childhood, language evidence is oral, gestural, graphic and situated. Transplanting university-essay panic to the writing corner is an error of object. In the second primary cycle, where Oreški et al. (2025) do observe use, evidence has to become process, not a detector.
Operationally, the framework admits: (a) in-presence tasks and oral production; (b) dated drafts and visible revision; (c) simple, taught attribution when use is authorized, in the line of the IB and ENAI; (d) dialogue with the child about how the work was done, in the line of Weber-Wulff et al. (2023). It rejects: (e) the detector as proof of misconduct in basic education; (f) accusation by percentage; (g) the thesis that the model “cheats”; (h) the thesis that banning ChatGPT restores authorship; (i) evaluating a polished text as if it were, in itself, learning.
8. Discussion
Three tensions organize the discussion. The first is between product and evidence. Detectors try to save the faith that the text “shows” the pupil. Liang et al. (2023), Elkhatat et al. (2023), Weber-Wulff et al. (2023) and Perkins et al. (2024) show, in corpora that are not early childhood, that that statistic over-accuses those who write with a narrow range and lets through those who obfuscate. OpenAI (2023) withdrew its classifier. Sadasivan et al. (2023) argue a theoretical ceiling. Inference: clinging to the product as proof is, in 2026, a pedagogical decision. Lodge et al. (2023) and Weber-Wulff et al. (2023) shift the weight to process from higher education; carrying that to basic education is a hypothesis, not a writing-corner trial.
The second tension is between cheating and learning. Mah et al. (2024) document that one can agree on the latter and diverge on the former. Oreški et al. (2025) document that most of their sample do not consider unmodified copying unacceptable, and that a plurality do not call help-use cheating. That does not prove that “nothing is happening”; it proves that the norm is in dispute in the band where the child is becoming an author. Cotton et al. (2023) warn, in university, of honesty risks and the difficulty of detection. Transferring that alert to basic education without passing through process evidence turns the child into a suspect by default. Eaton (2024) recalls the presumption of innocence precisely because institutional practice tends to invert it: the search is how to prove misconduct, not how to sustain a judgment of learning.
The third tension is between prohibition and attribution. New York illustrates panic and its reversal (Elsen-Rooney, 2023; Banks, 2023). The IB illustrates attribution without a ban (International Baccalaureate, 2023). Neither measures whether basic-education children became more authors. Miao and Holmes (2023) add age and validation: read in early childhood, that pushes against seating a five-year-old to converse alone with a model, and against pretending that a network filter resolves evidence. Foltýnek et al. (2023) ask for policies and training, not an oracle.
If the object is basic education, the teacher does not need a better detector. They need the authority to say that an impeccable paragraph is not, without more, a child who learned, and that a “suspicious” paragraph is not, without more, a child who cheated. Replacing that judgment with a percentage is to abdicate assessment. Replacing it with a ban is to abdicate design. The subject is not restored on the classroom network; it is made visible in the evidence the school accepts.
9. Limits
This review is narrative. It does not apply a full PRISMA protocol nor estimate combined effects on “cheating” or “learning.” The bulk of detector evidence is university-based or from adult and adolescent texts (Weber-Wulff et al., 2023; Elkhatat et al., 2023; Perkins et al., 2024; OpenAI’s classifier, 2023). Liang et al. (2023) include native eighth-grade essays—not early childhood or the first cycle—and TOEFL essays by non-natives. Sadasivan et al. (2023) is a preprint. Oreški et al. (2025) is a convenience sample of two schools, ages 11–14, self-report, with no measure of learning or detection. Mah et al. (2024) is qualitative, with Pre-K-to-postsecondary teachers and secondary students, not with early-childhood children using ChatGPT. Cotton et al. (2023) and Lodge et al. (2023) are, respectively, a higher-education debate and an expert document for the Australian tertiary sector: they must not be cited as a finding about basic education nor as a TEQSA regulation. The IB, New York, ENAI, Eaton and Miao and Holmes statements are prescriptive or institutional. Geography is biased toward the United States, Europe, Vietnam (Perkins) and Croatia; no Latin American detector trial or primary-homework genAI study meeting the criteria was located with the same degree of verification. That absence is a gap, not proof of nonexistence. Section 7 inferences are ethical-pedagogical threshold hypotheses for basic education, not evidence of national implementation. This article does not claim that all AI in the task is misconduct, nor that detectors “already work,” nor that banning ChatGPT restores authorship, nor that the model makes the child cheat: no nuclear source measures those promises in early childhood education or in the first cycle.
10. Conclusions
In basic education, academic integrity in the face of generative AI is not resolved by hunting cheats with a detector. It is resolved by asking what counts as evidence that a child learned when a model can produce the artifact. Three families of verified evidence support the argument. 2023–2024 detectors produce false positives—acutely against non-native and translated writing—and are evaded by editing or paraphrasing; OpenAI’s own classifier was withdrawn for inaccuracy (Liang et al., 2023; Elkhatat et al., 2023; Weber-Wulff et al., 2023; Perkins et al., 2024; OpenAI, 2023). In Croatian upper primary, homework use is already majority and the copying norm is lax; teachers and students do not fully agree on what cheating is, though they agree more on what learning is (Oreški et al., 2025; Mah et al., 2024). School policies oscillate between the ban that does not measure restored authorship and the attribution that does not measure children’s learning (Elsen-Rooney, 2023; Banks, 2023; International Baccalaureate, 2023; Foltýnek et al., 2023; Miao and Holmes, 2023).
Law and policy are not mute. Miao and Holmes (2023) mark human agency, institutional validation and an age limit for independent conversation. ENAI (Foltýnek et al., 2023) distinguishes authorized use from unauthorized generation and requires human responsibility. Eaton (2023, 2024) requires not treating all use as misconduct and not inverting the presumption of innocence. The IB (2023) denies that machine output is, without more, the student’s work. Lodge et al. (2023) ask for multiple judgments and process. Where the sources do not measure that detectors work in basic education, that AI makes the child cheat, or that banning ChatGPT restores authorship, this article does not invent it. The specific risk of basic education is evaluating a text without a subject. Integrity, then, is not software. It is the design of evidence in which the child—not the model, not the percentage—remains the author of the learning.
Ingeniero Mitre / Laboratorio Editorial de NEXTECH.IA
References
- Banks, D. C. (2023, May 18). ChatGPT caught NYC schools off guard. Now, we’re determined to embrace its potential. Chalkbeat New York. https://www.chalkbeat.org/newyork/2023/5/18/23727942/chatgpt-nyc-schools-david-banks/
- Cotton, D. R. E., Cotton, P. A., & Shipway, J. R. (2023). Chatting and cheating: Ensuring academic integrity in the era of ChatGPT. Innovations in Education and Teaching International, 61(2), 228–239. https://doi.org/10.1080/14703297.2023.2190148
- Dawson, P., & Bearman, M. (2020). Concluding comments: Reimagining university assessment in a digital world. In M. Bearman, P. Dawson, R. Ajjawi, J. Tai, & D. Boud (Eds.), Re-imagining university assessment in a digital world (pp. 291–296). Springer. https://doi.org/10.1007/978-3-030-41956-1_20
- Eaton, S. E. (2023). Postplagiarism: Transdisciplinary ethics and integrity in the age of artificial intelligence and neurotechnology. International Journal for Educational Integrity, 19, 23. https://doi.org/10.1007/s40979-023-00144-1
- Eaton, S. E. (2024). Future-proofing integrity in the age of artificial intelligence and neurotechnology: Prioritizing human rights, dignity, and equity. International Journal for Educational Integrity, 20, 8. https://doi.org/10.1007/s40979-024-00175-2
- Elkhatat, A. M., Elsaid, K., & Almeer, S. (2023). Evaluating the efficacy of AI content detection tools in differentiating between human and AI-generated text. International Journal for Educational Integrity, 19, 17. https://doi.org/10.1007/s40979-023-00140-5
- Elsen-Rooney, M. (2023, January 3). NYC education department blocks ChatGPT on school devices, networks. Chalkbeat New York. https://www.chalkbeat.org/newyork/2023/1/3/23537987/nyc-schools-ban-chatgpt-writing-artificial-intelligence/
- Foltýnek, T., Bjelobaba, S., Glendinning, I., Khan, Z. R., Santos, R., Pavletić, P., & Kravjar, J. (2023). ENAI recommendations on the ethical use of artificial intelligence in education. International Journal for Educational Integrity, 19, 12. https://doi.org/10.1007/s40979-023-00133-4
- International Baccalaureate. (2023, March 1). Statement from the IB about ChatGPT and artificial intelligence in assessment and education. https://www.ibo.org/news/news-about-the-ib/statement-from-the-ib-about-chatgpt-and-artificial-intelligence-in-assessment-and-education/
- Liang, W., Yuksekgonul, M., Mao, Y., Wu, E., & Zou, J. (2023). GPT detectors are biased against non-native English writers. Patterns, 4(7), 100779. https://doi.org/10.1016/j.patter.2023.100779
- Lodge, J. M., Howard, S., Bearman, M., Dawson, P., & Associates. (2023). Assessment reform for the age of artificial intelligence. Tertiary Education Quality and Standards Agency. https://www.teqsa.gov.au/guides-resources/resources/corporate-publications/assessment-reform-age-artificial-intelligence
- Mah, C., Walker, H., Phalen, L., Levine, S., Beck, S. W., & Pittman, J. (2024). Beyond CheatBots: Examining tensions in teachers’ and students’ perceptions of cheating and learning with ChatGPT. Education Sciences, 14(5), 500. https://doi.org/10.3390/educsci14050500
- Miao, F., & Holmes, W. (2023). Guidance for generative AI in education and research. UNESCO. https://doi.org/10.54675/EWZM9535
- OpenAI. (2023). New AI classifier for indicating AI-written text (update of 20 July 2023: the classifier is no longer available due to low accuracy). https://openai.com/index/new-ai-classifier-for-indicating-ai-written-text/
- Oreški, P., Oreški, T., & Ružić, I. (2025). Primary school students’ awareness of the ethical aspects of using artificial intelligence tools. Informatics in Education, 24(3), 559–585. https://doi.org/10.15388/infedu.2025.18
- Perkins, M., Roe, J., Postma, D., McGaughran, J., & Hickerson, D. (2024). Detection of GPT-4 generated text in higher education: Combining academic judgement and software to identify generative AI tool misuse. Journal of Academic Ethics, 22, 89–113. https://doi.org/10.1007/s10805-023-09492-6
- Sadasivan, V. S., Kumar, A., Balasubramanian, S., Wang, W., & Feizi, S. (2023). Can AI-generated text be reliably detected? (arXiv:2303.11156). arXiv. https://doi.org/10.48550/arXiv.2303.11156
- Tauginienė, L., Gaižauskaitė, I., Glendinning, I., Kravjar, J., Ojstršek, M., Ribeiro, L., Odineca, T., Marino, F., Cosentino, M., & Sivasubramaniam, S. (2018). Glossary for academic integrity. European Network for Academic Integrity. https://www.academicintegrity.eu/wp/wp-content/uploads/2018/02/GLOSSARY_final.pdf
- UNESCO. (2021). Recommendation on the ethics of artificial intelligence. https://unesdoc.unesco.org/ark:/48223/pf0000381137
- Weber-Wulff, D., Anohina-Naumeca, A., Bjelobaba, S., Foltýnek, T., Guerrero-Dib, J., Popoola, O., Šigut, P., & Waddington, L. (2023). Testing of detection tools for AI-generated text. International Journal for Educational Integrity, 19, 26. https://doi.org/10.1007/s40979-023-00146-z