1. Introduction: from suspicion to case file

A percentage arrives before the case file. The person who grades sees the proportion classified as generated, without a draft, without the writing conversation, and without what the student can explain. The number may remain in the classroom or enter a procedure, with effects on the grade, the record, and trust. The false positive falls on a person; the false negative lets through a text that perhaps cannot be sustained. They are not the same harm.

The problem does not begin with that interface. Cotton et al. (2024), as a conceptual framework, see in conversational interfaces and in GPT-3 a double face: greater participation, collaboration, and accessibility, and also dishonesty that is difficult to detect. There is no trial with known ground truth in their text; there is a proposal of policies, training, and more than one method of prevention. Birks and Clare (2023) add that misconduct is not new and that prevention frameworks exist, rooted in crime prevention, which remain useful when a language model facilitates the misconduct. Their close is a research agenda, not a model sanction.

The review question takes PCC form. The population is the academic texts of higher-education students and, as transfer, other academic writing already tested with detectors. The concept is the automatic classification of human authorship as against generated authorship. The context is assessment and integrity procedures. “Do detectors work?” mixes products, units, and definitions of error that the studies do not share. The question that organizes the article is what a score can prove when a case has to be decided.

The thesis is restrictive. Under control, accuracy improved and, in intact texts, several false positives became rare; inconsistency persists, as do the drop under paraphrase, the difficulty with hybrids, and a possible bias toward non-native writers. The score opens a conversation and does not close a case file. What follows marks each claim as an empirical finding, a conceptual framework, or a design inference.

2. Background: three waves of research on detectors

There is no straight line, but three waves: that of 2023 asks whether human text can be distinguished from generated text; the second attends to bias and to evasion; that of 2025–2026 returns to commercial products and to text that is hybrid, humanized, or edited. Two planes must not be confused. Benchmark accuracy is a frequency whose provenance is known. Probative value is what a score authorizes one to say of a person, without that truth and with a consequence. An excellent test set may not be proof of a violation.

Weber-Wulff et al. (2023) test 12 public tools plus Turnitin and PlagiarismCheck, and ask whether they separate human text from ChatGPT text and whether machine translation and obfuscation alter detection. The finding: they are neither accurate nor reliable, they are biased toward saying human, and they worsen with obfuscation. Elkhatat et al. (2023), with fifteen paragraphs from model 3.5 and fifteen from model 4 on cooling towers, five controls, and the tools OpenAI, Writer, Copyleaks, GPTZero, and CrossPlag, are more often correct on 3.5 than on 4 and produce false positives and uncertainty on the human material. Walters (2023), with 16 detectors and three groups of 42 essays, finds high accuracy for Copyleaks, Turnitin, and Originality.ai, and failure of most of the other 13 against model 4; payment improves little. Dalalah and Dalalah (2023) enter by the title alone, with neither an abstract nor a result.

The second wave changes the question. Liang et al. (2023) report, as an empirical finding, that detectors frequently classify non-native English writing as generated, with a risk of marginalizing it in assessment. The abstract does not bring a repeatable rate. Perkins et al. (2024b) set six detectors against content modified to evade detection (n = 805) and find a drop in accuracy of 17.4% under simple techniques. They conclude that the detectors cannot be recommended for determining violations, because of accuracy and because of the risk of false accusation, although they may support learning if the use is not punitive.

Van Vlasselaer et al. (2026), Malik and Amjad (2025), and Hyatt et al. (2025) measure accuracy again on text that is no longer pure; the figures come later, not as a slogan. Barrot and Aranda (2025) enter by the title alone, without an abstract: distinguishing what was produced by the student, what was edited, and what was generated. As a reading inference, not as a meta-analysis, no design observes provenance, permission, understanding, and consequence at once.

3. Method: critical PRISMA-lite review with an evidential reading

This is a narrative synthesis, not a meta-analysis: there is no common unit of text, of model, or of threshold. On 9 October 2026 the Crossref REST API was queried, without Scopus, Web of Science, ERIC, or PsycInfo: 16 queries of 60 rows, journal articles from 2021 onward. Of 960 records, 892 remained after deduplication. Eligibility required authors, a year, a generative-AI term in the title and another term of detection, integrity, or writing, and it set aside image, deepfake, code, audio or video, and disinformation. The exclusions — 45, 292, 101, and 17 — and the 437 records eligible for screening are in Table 1.

StageRecordsCriterion applied
Identification in the Crossref REST API96016 queries, 60 rows; journal article; publication from 2021
After deduplication892Duplicate records across queries removed
Automatic exclusion: no authors45No registered authorship
Automatic exclusion: no generative-AI term in the title292The title does not name generative AI, an LLM, ChatGPT, or generated text
Automatic exclusion: no term of integrity, detection, or writing101Neither the title nor the journal places detection, integrity, plagiarism, assessment, writing, or students
Automatic exclusion: other modality17Image, code, audio or video, disinformation
Eligible for screening437Result of the automatic filter
Excluded in manual screening414Title and abstract, under the criteria stated in the text
Included from the seed23Screening of the eligible set
Included by targeted search1Perkins et al. (2024b), verified in Crossref on 9 October 2026
Corpus included2420 with an abstract via Crossref, OpenAlex, or Semantic Scholar; 4 with the title only
Table 1. PRISMA-lite flow of identification and selection (real record of 9 October 2026).

Manual screening, a single pass and without double screening, excluded 414 of 437: corrections and errata, duplicate preprints, conference proceedings, journals without verifiable peer review or of doubtful indexing, adoption without integrity or detection, classifiers without student writing, and opinion without an argument. 23 remained from the seed, plus Perkins et al. (2024b) by targeted search and verification in Crossref. They are 24: 20 with an abstract and 4 with the title only — Jiang et al. (2024), Dalalah and Dalalah (2023), Barrot and Aranda (2025), and Zhang and Lv (2026). There was no formal risk-of-bias tool.

The evidential reading asks three things: what text entered — human, wholly generated, edited, hybrid, or humanized — or whether the work is a framework, voices, or policy; what error the abstract declares, without inventing a percentage; and whether that reaches an individual case. It does not reach when the ground truth is fabricated, and still less when the real stock of work does not have that truth. No study observes permission, understanding, and consequence at once. Behavioural health and admissions enter as transfer, not as a course task.

4. Accuracy under control: what changed between 2023 and 2026

Table 2 gathers the studies that report performance in the abstract that was recovered. It is not a league table: the rows do not share length, model, year, or a definition of a correct classification. They are read as designs, and the interpretation separates that accuracy from probative value.

StudyToolsTexts evaluatedWhat it reports (according to its abstract)
Weber-Wulff et al. (2023)12 public tools, plus Turnitin and PlagiarismCheckHuman text and ChatGPT text; also machine translation and obfuscationNeither accurate nor reliable. A bias toward classifying text as human. Obfuscation worsens performance in a significant way
Elkhatat et al. (2023)OpenAI, Writer, Copyleaks, GPTZero, and CrossPlag15 paragraphs from model 3.5, 15 from model 4, and 5 human controlsGreater accuracy with model 3.5 than with model 4. False positives and uncertain classifications on the human controls
Walters (2023)16 detectors; Copyleaks, Turnitin, and Originality.ai stand out42 essays from ChatGPT-3.5, 42 from ChatGPT-4, and 42 human essays without AIThose three show high accuracy on all three sets. Most of the other 13 fail against model 4. Payment improves only slightly
Gosling et al. (2024)The AI detector associated with Turnitin160 student responses and 160 from ChatGPT, with 16 promptsScores for the generated text significantly higher. The human scores, all at zero
Malik and Amjad (2025)Turnitin, ZeroGPT, GPTZero, and Writer AIChatGPT, Perplexity, and Gemini, in the original and after Grammarly, Quillbot, and human editing of 10% to 20%Turnitin, the most accurate and consistent, with an AI score of 100% even under adversarial techniques. Quillbot affects above all ZeroGPT, GPTZero, and Writer AI. The same model and the same tool yield different scores depending on the file. Perplexity is detected better; Gemini scores lower
Van Vlasselaer et al. (2026)GPTZero, Pangram, Copyleaks, and Turnitin160 synthetic documents with known ground truth (human, AI, hybrid, and humanized) and 1163 master’s theses from 2024–2025 without known ground truthPangram is superior on pure, hybrid, and humanized AI. The others underestimate, especially with the more advanced model. All correctly classify human text. False positives are rare. On the theses, Pangram flags 45.5%, generally at low or moderate levels
Popkov and Barrett (2025)One free detector and Originality.AI100 human articles from 2016–2018 and 200 chatbot texts (100 from ChatGPT and 100 from Claude)Problematic false positives and false negatives, whether paid or free. A median of 27.2% of the human text flagged as AI by the free detector. Originality.AI performs better and remains limited against Claude
Zhao et al. (2024)Domain-specific models, not a universal detector3755 letters of recommendation and 1973 graduate statements of purpose from Fordham, with GPT-3.5 Turbo counterpartsA general and universal detector appears beyond current reach. Specialized models can achieve high accuracy on that admissions material
Hyatt et al. (2025)Online detectors, alone and by consensusSTEM student writing, human and generatedEach detector varies. Together they inform when human intuition fails. Consensus reduces false positives almost to zero
Table 2. Accuracy studies in the corpus (the authors’ synthesis from the abstracts; not a standardized comparison across tools).

The improvement, as an empirical finding, is real and is not linear. Weber-Wulff et al. (2023) do not describe the zero of Gosling et al. (2024), the 100% for Turnitin in Malik and Amjad (2025), the rare false positives of Van Vlasselaer et al. (2026), or the consensus of Hyatt et al. (2025). Nor does a singular detector follow: Walters (2023) set three capable tools against a majority of the other thirteen, ineffective against model 4. The false positive is not closed. Malik and Amjad (2025) see different scores in files from the same model. Popkov and Barrett (2025), in behavioural-health journals, find a median of 27.2% of human text flagged by a free detector and a better Originality.AI, limited against Claude. Zhao et al. (2024) achieve high accuracy with domain models — 3755 letters and 1973 statements, counterparts from GPT-3.5 Turbo — and place a universal detector out of reach. It remains a benchmark.

The limit is one of design. Except for the 1163 theses, ground truth is known, but neither the assignment brief nor understanding is observed. Pangram flags 45.5% of those theses, at levels that are generally low or moderate, without known ground truth: they are flags, not prevalence. The design inference is that classifying a pure text better does not decide a case. The score measures a resemblance; the violation joins a person, an assignment brief, and understanding.

5. False positives, non-native writers, and base rates

Equity is not an appendix to accuracy. Liang et al. (2023) document, as an empirical finding, that detectors frequently classify non-native English writing as generated, with a risk of marginalizing those writers. The abstract does not give a repeatable figure. Even so, it forbids reading “rare false positive” as if the error were evenly distributed: what is rare in the aggregate can be frequent for someone who writes in a language that is not the first.

Jiang et al. (2024) formulate the question where it would matter most to close it: whether, when ChatGPT essays are detected in a large-scale writing assessment, there is bias against non-native speakers of English. There was no abstract. To attribute to them a yes or a no would be to invent a finding. The title keeps the debate open. The silence neither refutes Liang et al. (2023) nor confirms it.

There is counter-evidence, and it does not dissolve the problem. Gosling et al. (2024) leave all 160 human texts at zero. Van Vlasselaer et al. (2026) see correct classification of wholly human text and rare false positives, with an improvement relative to previous studies. That prevents the claim that every detector over-flags. It does not annul Liang et al. (2023): those abstracts do not separate native prose from non-native prose. Popkov and Barrett (2025), with a median of 27.2% on articles from 2016–2018, show that “rare” depends on the tool, the genre, and the year. The evidence is in tension.

The example shows the base rate, not a university. With a low false-positive rate, the non-generated majority produces innocent people whose score cannot be distinguished from a correct detection, and no one indicates which flag belongs among the 9.5. Sanctioning by the number reaches them. The 45.5% in Van Vlasselaer et al. (2026) is not the 5% of the assumption, and the near zero in Hyatt et al. (2025) does not put a name to the case either. The design inference: the number opens the review and does not close it, all the more so if the error falls on non-native writers (Liang et al., 2023).

6. Evasion, hybrid texts, and the shifting boundary of authorship

A detector can be right against text that resembles its training and fail against text that is modified in minutes. Perkins et al. (2024b) measure that fragility: six detectors, n = 805, and a drop of 17.4% under simple techniques, with the refusal, already cited, to use them to determine violations and with the door left open to a non-punitive use. Weber-Wulff et al. (2023) had already located the same fragility when they asked about machine translation and obfuscation, and when they concluded that obfuscation worsens performance inside a verdict of lack of accuracy and of reliability. Evasion is not a late finding.

Malik and Amjad (2025) leave Turnitin at 100% after Grammarly, Quillbot, and editing of 10% to 20%; ZeroGPT, GPTZero, and Writer AI fall above all with Quillbot, and Gemini scores lower than Perplexity. The same model and the same tool do not give the same number on different files. In 160 documents, Van Vlasselaer et al. (2026) include the hybrid — generated passages inside human text — and the humanized — rewritten with a prompt so as to look like student writing. Pangram sustains its accuracy and the other three underestimate, especially with the more advanced model. That false negative is not erased because one tool, on that test set, does see the mixture.

Barrot and Aranda (2025) contribute no result; their object names three classes that a percentage flattens: produced by the student, edited with AI, and generated. Wright (2026), in an analysis of policy and not of accuracy, holds that since 2023 many rules do not separate the generation of assessed content from non-generative uses — optical recognition, speech-to-text, and handwriting recognition. As reported, 94% of the 50 leading universities in the United States had guides for faculty. Having a guide is not the same as distinguishing transcription from generation.

The design inference is that authorship is no longer binary. The hybrid, the three classes in the object of Barrot and Aranda (2025), and the distinction drawn by Wright (2026) between generation and mere conversion of format show more than two components. To read the percentage as one of two boxes is to impose on the case file an ontology that this literature is abandoning. The score, when it says something, says a resemblance: it does not say who decided, with what permission, and with what understanding.

7. Teacher judgement versus the score

If the software does not close the case, what remains is the eye of the person who grades. Perkins et al. (2024a) test that eye with 22 ChatGPT submissions designed to make detection difficult and with 15 teachers, alongside genuine submissions. The empirical finding cuts through the optimism: Turnitin flags 91% of the submissions, but only 54.8% of the content; staff report 54.5% to the procedure; the means are 52.3 and 54.4. The flag on the document, the coverage of the text, and the grade are not the same fact, and almost half of the experimental material never reaches the case file.

Neither the eye nor the software is a gold standard. Hyatt et al. (2025) show that, in the aggregate, the detectors give warning when intuition fails and that consensus reduces false positives almost to zero. Guan and Han (2025), with n = 156 STEM students in two groups — with a language model, or independent writing — evaluate a detector and a survey of ethical awareness. The abstract gives no figures for accuracy and none for the survey. The finding is the limitation of these tools in the writing course, where style is learned and a classifier may confuse learning with substitution.

The design inference is that neither the eye nor the program, alone, sustains a determination. The discordance among the flag, the coverage of the text, and the teacher’s report is a reason to ask a question, not a tie to be broken in favour of the sanction. Consensus helps where intuition fails, and it does not know the assignment brief (Hyatt et al., 2025; Guan and Han, 2025).

8. From detecting to validating: epistemic authorship and oral defence

Ebrahimzadeh et al. (2026) change the unit of analysis. As a validity framework, they ask what happens if, without supervision, a generated text that is not understood is submitted. The integrity of co-authorship, which they propose as a source of validity evidence, is violated in that act. The development is a prototype, “AI Viva”: an oral defence together with questions of comprehension, with quantitative and dialogic feedback. The validation is preliminary and by experts, not a cohort trial. The turn — from hunting the file to verifying epistemic ownership — is taken here as a framework, not as demonstrated efficacy.

Hutson et al. (2026) measure that turn in a pilot. Integrevise pairs the written work with a brief viva: it neither detects nor grades. At a private liberal-arts university in the Midwest, the autumn of 2025 and the spring of 2026 are distinct cycles, not a single cohort, and the evidence is preliminary. There were 52 vivas. Of seven respondents, 14.3% agreed that the oral component helped them think more deeply and 57.1% disagreed. In 20 of 52, that is 38.5%, there were no comments. The qualitative material suggests gaps that the written work does not show. They ask that research continue, not that a policy be closed.

The alternative also has a cost, and its evidence is weak. Disagreement at 57.1% against agreement at 14.3%, on seven voices, warns that the oral defence may not be experienced as thinking more deeply and may become a second examination. Perkins et al. (2024a) recommend fewer tasks that a tool can imitate, or assessments that include AI, together with more training. That recommendation is coherent with the framework of Ebrahimzadeh et al. (2026) and is not a demonstrated efficacy. The design inference is modest: verifying understanding answers to what the score does not name, and today it is a hypothesis, not an automatic improvement.

9. Institutional policy: purpose, proportionality, and due process

Taylor and LaCroix (2026), as a conceptual framework, hold that whether the use of generative AI counts as misconduct depends on the purpose of the university. They read the apparent rise in cases as incoherence — technological enthusiasm, corporate influence, and a norm that push in different directions — in which the student is held to account for conduct that the institution itself shaped. The analysis of mission statements finds a gap between the ideal and the practice. Without an alignment of mission, pedagogy, and technology there is no credible demand for integrity, and a detector inherits that incoherence.

Ying (2026) contributes the finding of a scoping review of 38 empirical studies. There is an ethical grey zone: the student is neither a passive recipient nor an offender without restraint, but an agent who interprets what counts as lawful use. At the same time there is a gap between awareness and practice, and there are rationalizations of that gap. The voices depend on demographic, psychological, and contextual factors. As a reading of that finding, a single policy falls short: to sanction the grey zone without hearing it is to punish negotiations that students experience as ambiguous.

Zhang and Lv (2026) enter by the title alone: a comment, without an abstract, according to which detection can undermine integrity. There is no datum to attribute to them. The statement gives warning in the direction of the risk of false accusation (Perkins et al., 2024b) and of bias (Liang et al., 2023). Birks and Clare (2023) link misconduct facilitated by AI to prevention frameworks already used for classic misconduct: the response begins with the opportunity and the assignment brief, not with the punishment. Situational prevention is here a design inference, not a finding measured by them.

From that there emerges no validated norm, but a design inference. The score is not sole proof. The person who is flagged may explain versions, format, transcription, editing, help, and the assignment brief they believed to be in force. The record of versions, if the task is a process task, is provided for in the assessment and is not improvised after the suspicion arises. The conversation precedes the sanction, and the sanction, if it comes, answers to understanding, permission, process, and text. The package has not been tried in the 24 studies. To call it proven good practice would exceed the corpus.

The inference for Mexican and Latin American universities is one of transfer: there is no datum, and no regulation is invoked. No source evaluates detectors in Spanish, and Table 2 describes texts in English, whereas in the region one writes in Spanish and, in graduate study, often in English as a second language. To transfer the threshold is to import the bias of Liang et al. (2023) onto ground that has not been measured; Jiang et al. (2024) leave the question open even in English. It authorizes neither enthusiasm nor prohibition. The percentage does not travel alone as far as the sanction.

10. Discussion: an indication is not proof

A score depends on the year, on the product, on the model, and on whether the text is pure, edited, hybrid, or humanized. In the laboratory, with known ground truth, some products classify correctly and others flag human prose or underestimate what was generated. In the case file the ground truth is missing and often the process as well, and the consequence is there in excess. The indication is a reason not to file the doubt away; the proof would be a sufficient reason for a violation. The score does not get there. Figure 1, an inferred diagram and not an evaluated protocol, orders the step so that the decision does not jump from the percentage to the sanction. Each rung answers a question that the rung before it does not answer.

RungQuestion it answersEvidence from the corpusDefensible use (design hypothesis)
Detector scoreDoes the text resemble, according to the tool, a generated class?Improvement on pure text and rare false positives in recent designs; failures from 2023 and inconsistency across files (Gosling et al., 2024; Van Vlasselaer et al., 2026; Hyatt et al., 2025; Weber-Wulff et al., 2023; Malik and Amjad, 2025)Open a review. Do not determine the violation
Contrast with history and draftsDoes the text fit the writing trajectory and the versions of the task?The test sets judge the finished text, almost always with fabricated ground truth (Walters, 2023; Zhao et al., 2024). Wright (2026) separates format conversion from generationAsk for an explanation of process before reading the score as a substitution of authorship
Conversation with the studentCan the student account for the decisions in the text and for the permission they believed they had?A grey zone and student agency (Ying, 2026). Staff reported 54.5% of the experimental submissions (Perkins et al., 2024a)Hear the explanation before any sanction
Verification of understanding or oral defenceDoes the student understand, and can they sustain, what they put in their own name?Co-authorship integrity is violated when what is not understood is submitted; the viva has preliminary expert validation (Ebrahimzadeh et al., 2026). A mixed pilot: 14.3% agreement and 57.1% disagreement, with seven respondents (Hutson et al., 2026)Check epistemic authorship without taking the oral defence to be effective
Proportional decisionWhat response fits the formative purpose and the doubt that remains?Misconduct depends on purpose (Taylor and LaCroix, 2026). Prevention precedes the sanction (Birks and Clare, 2023). Detection can harm integrity (the object of Zhang and Lv, 2026)Sanction only if the whole sustains the violation. The score in isolation does not suffice
Figure 1. Conceptual diagram, not empirical data: decision ladder from indication to determination, inferred from the review.

If the task forbade generation, to ignore a high score would be as unserious as to obey it, and Gosling et al. (2024) warn that the technical route is not enough. The rungs that follow cover the history the text does not tell (Wright, 2026), the unstable report (Perkins et al., 2024a), the norm the student is already interpreting (Ying, 2026), verification that still has no demonstrated efficacy (Ebrahimzadeh et al., 2026; Hutson et al., 2026), and the purpose that defines the offence (Taylor and LaCroix, 2026). To sanction on the indication alone is to realize the risk that Zhang and Lv (2026) state.

An indication is not a proof, for four reasons. The benchmark has known ground truth and the case does not; not even in 2023 was accuracy uniform (Elkhatat et al., 2023; Walters, 2023). The base rate flags innocent people, and the score does not say which flag is the innocent one. Paraphrase, the hybrid, and transcription make the same percentage cover different facts. Certifying learning requires that the person sustain the work. Nothing obliges anyone to switch the detectors off; what is required is that the number not be the last word. Cotton et al. (2024), as a framework, called for more than one method, and the ladder does not pretend that each rung has been measured.

11. Limitations

The corpus is from Crossref alone: Scopus, Web of Science, ERIC, and PsycInfo did not enter, and the Anglophone predominance belongs to what was retrieved, not to the reading. Generalization to Mexico and Latin America is inferential, and no source evaluates detectors in Spanish. The tools move faster than the journal: to use Weber-Wulff et al. (2023) or Elkhatat et al. (2023), which evaluate generations such as GPT-3.5 and GPT-4, in order to describe a product of 2026 would be as mistaken as to use Van Vlasselaer et al. (2026) in order to rehabilitate the stock of tools from 2023. Table 2 is a series of windows, not a sheet that is still in force.

The synthetic sets permit a statement of accuracy; the 1163 theses, without known ground truth, permit flags to be described and do not permit them to be verified. Four sources enter with the title only — Jiang et al. (2024), Dalalah and Dalalah (2023), Barrot and Aranda (2025), and the comment by Zhang and Lv (2026) — and nothing that is said of them exceeds that object. There was neither double screening nor a risk-of-bias tool. Behavioural health and admissions are transfer, not the classroom of an undergraduate programme. The synthesis forgoes a single accuracy figure for the field: with these designs such a figure would be more precise in form than in meaning.

12. Conclusions

The question that remains is not whether, on pure texts, a detector separates what is human from what is generated: some can, and others cease to be able to do so if the model, the paraphrase, or the hybrid changes. The question is what a university does with the number when the number points to a person. The answer of the corpus is a narrow one. The number is an indication: it opens a review, it does not declare the violation, it does not point out the innocent person, and it does not certify that the person who submits understands what was submitted.

There is an improvement under control that it would be disloyal to deny, and there is at the same time instability across files, fragility before the hybrid and before paraphrase, flags on non-native prose and, outside the classroom, human text flagged by detectors that in other designs almost do not fail. To cite only the improvement is to sell a product. To cite only the failure is to deny the improvement. Neither of the two series turns the benchmark into proof of a case file.

The person who grades does not correct that field alone: on the evasive submissions, the flag, the coverage, and the report did not coincide, and the grades remained almost level with one another. To check that the person can sustain what they sign is coherent with the harm that matters, and it arrives with weak evidence — a prototype validated by experts, and a pilot of mixed perceptions. To recommend it is a design inference. To treat it as an evaluated solution would repeat the shortcut of the detector that “already works.”

Without a purpose, the detector watches conduct that the institution helps to produce (Ying, 2026). That the score not be sole proof, that the process be open to explanation, and that the conversation precede the sanction is a condition of design, not a law of the corpus. In Mexico and Latin America the point weighs more heavily for lack of data: a threshold built on English, applied to Spanish or to English as a second language, repeats a bias that has not been measured. If the indication does not survive the process and the explanation, there is no sanction attributed to the number.

NEXTECH.IA Editorial Lab / Ingeniero Mitre

References

  1. Barrot, J. S., & Aranda, M. R. R. (2025). Efficacy of AI-Text Detection Tools in Distinguishing Student-Produced, AI-Edited, and AI-Generated Essays. Technology, Knowledge and Learning. https://doi.org/10.1007/s10758-025-09884-0
  2. Birks, D., & Clare, J. (2023). Linking artificial intelligence facilitated academic misconduct to existing prevention frameworks. International Journal for Educational Integrity, 19(1), Article 20. https://doi.org/10.1007/s40979-023-00142-3
  3. Cotton, D. R. E., Cotton, P. A., & Shipway, J. R. (2024). Chatting and cheating: Ensuring academic integrity in the era of ChatGPT. Innovations in Education and Teaching International, 61(2), 228–239. https://doi.org/10.1080/14703297.2023.2190148
  4. Dalalah, D., & Dalalah, O. M. A. (2023). The false positives and false negatives of generative AI detection tools in education and academic research: The case of ChatGPT. The International Journal of Management Education, 21(2), 100822. https://doi.org/10.1016/j.ijme.2023.100822
  5. Ebrahimzadeh, M., Shibani, A., & Shum, S. B. (2026). Coauthorship integrity: Reconceptualising assessment validity for the age of generative artificial intelligence. Computers and Education: Artificial Intelligence, 10, 100609. https://doi.org/10.1016/j.caeai.2026.100609
  6. Elkhatat, A. M., Elsaid, K., & Almeer, S. (2023). Evaluating the efficacy of AI content detection tools in differentiating between human and AI-generated text. International Journal for Educational Integrity, 19(1), Article 17. https://doi.org/10.1007/s40979-023-00140-5
  7. Gosling, S. D., Ybarra, K., & Angulo, S. K. (2024). A widely used Generative-AI detector yields zero false positives. Aloma: Revista de Psicologia, Ciències de l'Educació i de l'Esport, 42(2), 31–43. https://doi.org/10.51698/aloma.2024.42.2.31-43
  8. Guan, Q., & Han, Y. (2025). From AI to authorship: Exploring the use of LLM detection tools for calling on “originality” of students in academic environments. Innovations in Education and Teaching International, 62(5), 1514–1528. https://doi.org/10.1080/14703297.2025.2511062
  9. Hutson, J., Poyer, K., Ogoe, E., & Atologun, K. A. (2026). Beyond AI Detection: A Pilot Study of IntegreviseTM and Viva-Based Verification of Student Understanding in AI-Mediated Assessment. Trends in Higher Education, 5(3), 59. https://doi.org/10.3390/higheredu5030059
  10. Hyatt, J.-P. K., Bienenstock, E. J., Firetto, C. M., Woods, E. R., & Comus, R. C. (2025). Using aggregated AI detector outcomes to eliminate false positives in STEM-student writing. Advances in Physiology Education, 49(2), 486–495. https://doi.org/10.1152/advan.00235.2024
  11. Jiang, Y., Hao, J., Fauss, M., & Li, C. (2024). Detecting ChatGPT-generated essays in a large-scale writing assessment: Is there a bias against non-native English speakers? Computers & Education, 217, 105070. https://doi.org/10.1016/j.compedu.2024.105070
  12. Liang, W., Yuksekgonul, M., Mao, Y., Wu, E., & Zou, J. (2023). GPT detectors are biased against non-native English writers. Patterns, 4(7), 100779. https://doi.org/10.1016/j.patter.2023.100779
  13. Malik, M. A., & Amjad, A. I. (2025). AI vs AI: How effective are Turnitin, ZeroGPT, GPTZero, and Writer AI in detecting text generated by ChatGPT, Perplexity, and Gemini? Journal of Applied Learning and Teaching, 8(1), 91–101. https://doi.org/10.37074/jalt.2025.8.1.9
  14. Perkins, M., Roe, J., Postma, D., McGaughran, J., & Hickerson, D. (2024a). Detection of GPT-4 Generated Text in Higher Education: Combining Academic Judgement and Software to Identify Generative AI Tool Misuse. Journal of Academic Ethics, 22(1), 89–113. https://doi.org/10.1007/s10805-023-09492-6
  15. Perkins, M., Roe, J., Vu, B. H., Postma, D., Hickerson, D., McGaughran, J., & Khuat, H. Q. (2024b). Simple techniques to bypass GenAI text detectors: implications for inclusive education. International Journal of Educational Technology in Higher Education, 21(1), Article 53. https://doi.org/10.1186/s41239-024-00487-w
  16. Popkov, A. A., & Barrett, T. S. (2025). AI vs academia: Experimental study on AI text detectors’ accuracy in behavioral health academic writing. Accountability in Research, 32(7), 1072–1088. https://doi.org/10.1080/08989621.2024.2331757
  17. Taylor, T. B., & LaCroix, T. (2026). Purpose before policy: academic integrity, generative AI, and rhetorical stance. Higher Education. https://doi.org/10.1007/s10734-026-01706-1
  18. Van Vlasselaer, M., Van Droogenbroeck, F., & Spruyt, B. (2026). Who wrote this? Evaluating the reliability of AI detection tools in higher education. International Journal for Educational Integrity, 22(1), Article 16. https://doi.org/10.1007/s40979-026-00226-w
  19. Walters, W. H. (2023). The Effectiveness of Software Designed to Detect AI-Generated Writing: A Comparison of 16 AI Text Detectors. Open Information Science, 7(1), Article 20220158. https://doi.org/10.1515/opis-2022-0158
  20. Weber-Wulff, D., Anohina-Naumeca, A., Bjelobaba, S., Foltýnek, T., Guerrero-Dib, J., Popoola, O., Šigut, P., & Waddington, L. (2023). Testing of detection tools for AI-generated text. International Journal for Educational Integrity, 19(1), Article 26. https://doi.org/10.1007/s40979-023-00146-z
  21. Wright, C. (2026). Transcription is not generation: distinguishing non-generative AI tool use from academic misconduct in higher education assessment. International Journal for Educational Integrity, 22(1), Article 24. https://doi.org/10.1007/s40979-026-00234-w
  22. Ying, J. (2026). Academic Integrity in the Age of Generative AI: a Scoping Review of Research on Higher Education Student Voices. Journal of Academic Ethics, 24(3), Article 78. https://doi.org/10.1007/s10805-026-09752-1
  23. Zhang, S., & Lv, X. (2026). AI detection risks undermining academic integrity. Nature Human Behaviour, 10(8), 1395–1396. https://doi.org/10.1038/s41562-026-02528-y
  24. Zhao, Y., Borelli, A., Martinez, F., Xue, H., & Weiss, G. M. (2024). Admissions in the age of AI: detecting AI-generated application materials in higher education. Scientific Reports, 14(1), Article 26411. https://doi.org/10.1038/s41598-024-77847-z