1. Introduction and problem

In basic and early childhood schools a dangerous equation circulates: becoming literate in artificial intelligence would mean sitting children in front of a chatbot. That equation confuses three distinct objects. The first is the use of a tool. The second is understanding a set of ideas—data, pattern, prediction, limit, human agency—that makes it possible to recognize a system, explain where an error comes from, and decide when not to delegate a judgment. The third is independent conversation with a generative model, which UNESCO’s Guidance on generative AI reserves to a minimum age of thirteen and subjects to age verification and supervision (Miao & Holmes, 2023). Mixing them produces a pedagogy of the interface: the student “uses AI” without being able to say what a training datum is, why a classifier errs, or who is accountable for a decision.

The problem is not small. Long and Magerko (2020) defined AI literacy as the set of competencies that enables people to critically evaluate AI technologies, communicate and collaborate with them, and use them as a tool. That definition does not require programming; nor does it reduce to pressing a prompt. Touretzky, Gardner-McCune, Martin, and Seehorn (2019) organized what every K–12 student should know into five big ideas—perception, representation and reasoning, learning, natural interaction, and societal impact—across bands K–2, 3–5, 6–8, and 9–12. UNESCO’s student framework—distinct from the teacher framework—proposes twelve blocks in four aspects and three levels: understand, apply, and create (Miao, Shiohira, & Lao, 2024). The 2026 AILit framework adds four domains—engage, create, manage, and shape—as a common reference, not as a kindergarten method (European Commission & OECD, 2026).

This article’s thesis is restrictive. AI literacy in basic and early childhood education is not “using ChatGPT in class.” It is a curriculum of ideas compatible with age thresholds. That curriculum can begin well before age thirteen—recognizing a sensor, labeling examples, seeing that a model errs when the data are unbalanced, affirming that a person decides—and, precisely for that reason, it does not need independent conversation with a generative system. Where a source does not measure child development, academic achievement, or “readiness for the future,” this article does not promise it.

The gap that organizes the work is triple. First, K–12 literacy reviews document abundant proposals and a scarcity of evaluations of whether students understood the concepts (Casal-Otero et al., 2023). Second, the early childhood literature often mixes tools that use AI to teach something else with teaching about AI (Su & Yang, 2022). Third, Miao and Holmes’s (2023) age threshold leaves a pedagogical interval that schools often fill with improvisation: what is taught about AI when one does not yet converse with it alone? This article answers that question for students, not for teachers (the 19 August morning axis), not for parental mediation (19 August extra), not for inclusion and disability (20 August, 5:00 a.m.), and not with PopBots or personalized tutors as the spine.

2. State of the art: competencies, big ideas, and thresholds

Three strata should be kept apart. The first is conceptual: how AI literacy is defined. The second is normative: what UNESCO, UNICEF, AI4K12, and AILit prescribe. The third is empirical: what interventions with young children and primary students measure when the object is understanding AI, not using it as a tutor.

In the conceptual stratum, Long and Magerko (2020) synthesized interdisciplinary literature into seventeen competencies and sixteen design considerations. Digital literacy is a prerequisite; computational literacy—knowing how to program—is not. The framework groups what AI is, what it can do, how it works, how it should be used, and how people perceive it. Ng, Leung, Chu, and Qiao (2021), from thirty articles, proposed four aspects: know and understand, use and apply, evaluate and create, and ethical issues. That quartet is useful as a map, not as classroom evidence. Casal-Otero et al. (2023) reviewed 179 Scopus documents and found two broad families—learning experiences and theoretical perspectives—and a repeated scarcity: “there were hardly any experiences that assessed whether students understood AI concepts after the learning experience.” That statement is a synthesis finding: it licenses skepticism toward catalogs of activities that do not measure understanding.

In the normative stratum of big ideas, Touretzky et al. (2019) and AI4K12.org—a joint AAAI–CSTA initiative—fix five ideas and four grade bands. In K–2, perception is introduced as the difference between a simple sensor and a machine that extracts meaning; learning, as pattern recognition; societal impact, as a conversation about fairness at a child’s scale. In grades 3–5, students collect and label data to train simple classifiers and discuss that data quality affects the result (AI4K12 Initiative, 2020; Touretzky & Gardner-McCune, 2021). Those progressions are curriculum-design frameworks, not trials. Their value is that they allow teaching about AI without generative conversation: a six-year-old can sort stones and see that the human “model” also errs; a chatbot is not required.

In the global normative stratum, Miao, Shiohira, and Lao (2024) anchor the student framework in human rights, inclusion, and equity. The Understand level is intended for all and includes unplugged options. The Create level, especially system design, is elective deepening. By 2022, only about fifteen countries had included AI objectives in the national curriculum. The framework warns against competencies dictated by private platforms. That warning is the reverse of this thesis: literacy is not brand training.

Miao and Holmes (2023) are not a student competency framework; they are cited for the threshold, not as a teacher axis. The Guidance proposes “defining and enforcing an age limit for the use of GenAI” and states that “the minimum threshold should be thirteen years of age” for independent conversations, in dialogue with COPPA (thirteen) and the GDPR (sixteen without parental permission). Untested tools are not appropriate for primary education. UNICEF (2021) formulates nine requirements; supporting well-being, protecting data, ensuring safety, providing transparency, and preparing children for a world with AI are nuclear here. “Prepare” is a policy mandate, not an effect measurement. The 2021 Recommendation requires human oversight and non-discrimination (UNESCO, 2021). AILit organizes nineteen competencies and a tool-agnostic foundation (European Commission & OECD, 2026). OECD (2025) asks what students should learn in a future of powerful AI; it does not prescribe a chatbot for the first cycle.

In the early-childhood empirical stratum, Su and Yang (2022) reviewed seventeen studies from 1995 to 2021 and concluded that most reported improvements in concepts of AI, machine learning, computing, and robotics, and also in creativity, emotion control, collaborative inquiry, literacy, and computational thinking. That claim is a synthesis of a small, heterogeneous corpus, not a meta-analysis of developmental effects. This article retains the finding that concept experiences exist; it does not adopt the inference that AI “develops” the child. Vartiainen, Tedre, and Valtonen (2020) and Sanusi et al. (2024)—detailed in section 6—show, with small samples in informal settings, that children aged three to thirteen can infer input–output relations when they “teach” a classifier. They do not measure school performance.

The gap, then, is not the absence of frameworks. It is the confusion among framework, finding, and marketing. Basic and early childhood schools already have, in the verified documents, a sufficient lexicon: big idea, grade band, Understand level, data, bias, agency. What is missing is the discipline of not replacing that lexicon with a chatbot license.

3. Review method

A critical narrative review was conducted, not a meta-analysis. The purpose was not to estimate a comparable effect size—none exists homogeneously across a thirty-hour workshop, a life-science curriculum, and an informal case study with six children—but to articulate a curricular argument with verified sources. Inclusion criteria were: (a) publication between 2021 and 2026, with a justified exception for still-current frameworks (Touretzky et al., 2019; Long & Magerko, 2020; Vartiainen et al., 2020; UNESCO, 2021; UNICEF, 2021); (b) focus on student competencies in early childhood or basic education (approximately ages 3 to 14), not teacher competence or parental mediation; (c) document type of peer-reviewed journal, proceedings with a DOI, or a report from UNESCO, UNICEF, OECD, the European Commission, AAAI, or ACM; (d) access to a DOI page, UNESDOC, an institutional repository, or a publisher site confirming authors, year, title, and findings. Sources whose DOI did not open to verifiable metadata were excluded, as were the UNESCO teacher framework of 2023–2024 as an axis, PopBots and play scaffolding with generative AI as a nuclear case, personalized tutors and adaptive assessment, and inclusion/disability as the central problem.

The search was executed on 20 August 2026 (around 20:00, America/Mexico_City) on DOI pages, the ACM Digital Library, the AAAI Open Journal, Springer, ScienceDirect, Taylor & Francis, UNESDOC, OECD iLibrary, UNICEF Innocenti, ai4k12.org, and repositories (MIT DSpace, UCL Discovery, UEF). Each cited source was verified against at least one of those pages. The corpus was organized into three nuclear cases of distinct types: an early-secondary curriculum with pre–post and, later, a teacher comparison (DAILy); an upper-primary curriculum in a life-science context, complemented by a conceptions study (PrimaryAI); and two Teachable Machine studies in early childhood and early primary, outside the PopBots axis. As saturation, Ng et al. (2021), Casal-Otero et al. (2023), Su and Yang (2022), Touretzky and Gardner-McCune (2021), and OECD (2025) were retained.

The analysis distinguished three enunciative statuses. Empirical finding: what was observed or measured in the sample. Normative framework: what an organization or expert framework prescribes. Pedagogical inference: the translation this article proposes for the early childhood and basic classroom, marked as such. The limits are those of any narrative review: there is no full PRISMA protocol, there is a language bias toward English, and Latin America is underrepresented (section 9).

4. Case 1. DAILy: a curriculum of technical, ethical, and career-future ideas

Lee, Ali, Zhang, DiPaola, and Breazeal (2021) described, in SIGCSE, a virtual summer 2020 workshop with the thirty-hour Developing AI Literacy (DAILy) curriculum. Thirty-one students aged 10.4 to 14.7 (mean 12.57) participated; 61% were female and 87% came from groups underrepresented in STEM and computing. The curriculum interweaves technical concepts (logic systems, supervised learning with Teachable Machine, neural networks via a participatory game, generative adversarial networks), algorithmic bias, and career exploration. The authors report conceptual learning and an increase in interest that did not reach statistical significance, with a possible ceiling effect (high prior interest: 3.59 out of 5; post 3.71). That experience report is not a controlled trial.

Zhang, Lee, Ali, DiPaola, Cheng, and Breazeal (2023) published in the International Journal of Artificial Intelligence in Education (vol. 33, pp. 290–324) the exploratory study of the same program with twenty-five students in an urban summer STEM program, mostly grades 7–9, 32% African American, 56% Hispanic/Latino, and 56% female. The concept inventory (AI-CI) moved from a mean of 23.69 (SD = 4.61) to 26.05 (SD = 4.34), t(24) = 3.37, p < .01, Cohen’s d = 0.53. There were gains in general concepts, logic systems, machine-learning concepts, and supervised learning. There were none in neural networks (pre 2.89; post 2.87). The prediction item—distinguishing whether a technology predicts—did not improve; about half persisted in confusions. Awareness of AI-related careers did increase significantly (3.15 to 3.48, p < .001). At exit, nearly half explained AI not only as a technical subject but as one with personal, career, and societal implications. In interviews with nineteen students, almost all articulated benefits and harms, and sixteen proposed diverse datasets as bias mitigation (Zhang et al., 2023).

Status of the evidence. Empirical finding: in that self-selected sample, mediated by those who designed the curriculum and in a virtual format, concept scores and career awareness improved; difficulties in prediction and neural networks persisted; interest did not change significantly. It is not a finding that DAILy improves performance in mathematics, language, or general “development”: the study does not measure that. It is not a finding of population causality: there is no randomization or control group in this paper. The authors declare recruitment limits and immediate effects.

Zhang, Lee, and Moore (2024) published in AAAI a comparison study with teacher implementation during school hours. The experimental group had 89 middle-school students (51 sixth graders, 38 eighth; 42% female; 84% racial or ethnic minorities); the comparison group, 69 (41 sixth, 20 seventh, 8 eighth; 39% female; 77% minorities). Both completed the same pre- and post-test. The authors report that the experimental group developed a deeper understanding of AI concepts and more positive attitudes toward AI and its impact on future careers than the comparison group. Status: empirical finding from a non-randomized comparison on concepts and attitudes, not on school grades. The added value is that the curriculum was not confined to a researcher camp.

Pedagogical inference, marked as such: DAILy illustrates a curriculum of ideas—algorithm, label, unbalanced set, bias, stakeholder, prediction—for students approaching or crossing the age-thirteen threshold. Teachable Machine appears as a means to see the datum, not as a chatbot. The finding that prediction remains a difficult concept is precisely an argument against prompt pedagogy: whoever cannot distinguish an automatic door from a system that infers is not literate, even if they have chatted. For the first cycle of basic education, DAILy is not copied: the principle of interweaving technique and ethics is extracted, and the abstraction is lowered (section 7).

5. Case 2. Upper primary: everyday conceptions and PrimaryAI

Ottenbreit-Leftwich, Glazewski, Jeon, Jantaraweragul, Hmelo-Silver, Scribner, Lee, Mott, and Lester (2023a) published in the same IJAIED issue a qualitative study of the everyday ideas of grade 4 and 5 students (ages 9 to 11) about AI, juxtaposed with their teachers’ reflections, to inform curriculum co-design. Themes of conceptions, examples, and ethics emerged. The empirical finding is about an entry point, not a gain: students already bring examples (voice assistants, recommendations, “smart robots”) and incipient ethical judgments; teachers note the need to anchor content in prior knowledge and to feel able to teach it. The article does not claim that an AI curriculum improves science learning or socioemotional development.

The same team implemented PrimaryAI, a grades 3–5 curriculum that integrates machine learning, computer vision, AI planning, and ethics in an authentic life-science context—the decline of New Zealand’s yellow-eyed penguin—with unplugged activities and an immersive environment. Ottenbreit-Leftwich et al. (2023b) reported, in a SIGCSE poster, the spring 2022 implementation by two teachers (one fourth and one fifth grade in different schools). Unit pre- and post-tests, in a one-group design, showed statistically significant knowledge gains in machine learning and computer vision, with no significant gender differences. Teachers indicated that the curriculum engaged students and offered sufficient scaffolding to teach with little prior AI knowledge; they asked for greater alignment between AI and life-science concepts, and more local connections (Ottenbreit-Leftwich et al., 2023b).

Status of the evidence. Empirical finding: in two upper-primary classrooms, with regular teachers and a problem-based curriculum, students improved concept scores in machine learning and vision. It is not a finding of systemic inclusion, nor of improvement in biology beyond what the units assessed, nor of transfer to other subjects. The poster format limits available statistical detail; it is retained as evidence of feasibility in an ordinary classroom, not as a confirmatory trial. The conceptions study (2023a) and the implementation (2023b) complement each other: the first says which ideas already circulate; the second, that a curriculum anchored in a living problem can move those ideas toward normative AI concepts.

Pedagogical inference: upper primary (ages 9–11) is a band in which the curriculum of ideas can become disciplinary without becoming generative conversation. Collecting images, labeling, training a species or trait classifier, and asking who is missing from the dataset is literacy. Asking a language model to “explain the penguin” is not, and collides with Miao and Holmes’s (2023) threshold. PrimaryAI also shows that the vehicle can be school science, not a new “AI” subject. That is a design inference, not a finding of superiority over other vehicles.

6. Case 3. Early childhood and early primary: Teachable Machine, not the chatbot

Vartiainen, Tedre, and Valtonen (2020) published in the International Journal of Child-Computer Interaction a sociocultural case study with six children aged 3 to 9 who, in non-school settings, “taught” and explored Google’s Teachable Machine. Fine-grained analysis of video and interviews illustrates an embodied, fast-paced process: children produced datasets and models, observed, explored, and explained their own interaction. The empirical finding is that that interaction supported reasoning about the relationship between their own bodily expressions and the tool’s output: an incipient understanding of input and output in machine learning. The authors discuss agency and participation. They do not report tests of academic performance or standardized socioemotional development. The sample is minimal and the context informal and Finnish.

Sanusi, Sunday, Oyelere, Suhonen, Vartiainen, and Tukiainen (2024) replicated the question in an informal African setting, with eighteen children aged 3 to 13 and the same type of tool. The empirical finding, according to the authors, is that Teachable Machine contributed to data literacy and conceptual understanding across K–12, irrespective of background, and that children could infer the relationship between their expressions and the system’s output. The study is presented as a baseline for the African context, not as a school-efficacy trial. Online publication dates from 2023; the Computer Science Education volume is 2024.

Status of the evidence. Empirical finding: in two small-sample qualitative studies, children from age three, with adult mediation and a visible classifier, can narrate that “if I show this, the machine says that” and, at times, that the error changes if the example changes. It is not a finding that Teachable Machine develops language, executive function, or achievement. It is not a finding that a kindergarten should install webcams. Su and Yang (2022) add that the ECE AI corpus is scarce and heterogeneous; this article does not convert that scarcity into a license for mass adoption.

Why this case and not PopBots. On 18 August the scaffolding of play with generative AI and conversational robots was examined. Here the object is different: a classifier the child trains, not an agent that speaks to the child. The pedagogical difference is the threshold. Teaching a model with gestures or fruit photos makes the datum visible. Conversing alone with a language model hides the datum and feigns an interlocutor. Miao and Holmes (2023) advise against the latter under age thirteen. Vartiainen et al. (2020) and Sanusi et al. (2024) show that the former is, at least, researchable.

Pedagogical inference: in early childhood education, the curriculum of ideas is bodily and material. Data is “the photos we put in.” Pattern is “all of these are apples.” Prediction is “now guess this one.” Limit is “it erred because we only showed red apples.” Human agency is “we decide which photos and whether we trust.” That pentagram requires no chatbot account and no personal data in a general-purpose model (UNICEF, 2021; Miao & Holmes, 2023). It requires an adult who does not celebrate the machine’s success as magic.

7. Inferential framework: a curriculum of ideas below the independent-conversation threshold

The framework that follows is this article’s pedagogical inference, anchored in the cases and the verified instruments. It is not a new international standard. It distinguishes five axis-ideas and three tests of admissibility. If a “literacy” proposal fails the tests, it is not deployed as a curriculum for early childhood or basic-education students.

7.1. Five axis-ideas, not five applications. Data: something that is collected, labeled, and can be biased; Long and Magerko (2020) and DAILy place it at the center of understanding supervised learning (Zhang et al., 2023). Pattern: what a system extracts from examples; AI4K12 introduces it from K–2 as recognition, not as algebra (Touretzky et al., 2019). Prediction: an inference from what was learned, distinct from a sensor that opens a door; DAILy shows that the concept is difficult even at age twelve (Zhang et al., 2023). Limit: the system errs, does not understand the world, and does not replace a human judgment; Miao and Holmes (2023) insist that generative models do not understand objects or social relations. Human agency: someone decides the goal, the datum, the use, and the responsibility; it is the first block of UNESCO’s Understand level for students (Miao, Shiohira, & Lao, 2024) and a domain of AILit (European Commission & OECD, 2026). Inference: those five ideas are taught with object classification, decision games, unbalanced sets, and debate about “who is missing,” not with a language-model account.

7.2. Age-threshold test. Miao and Holmes (2023) set thirteen as the minimum for independent conversations with generative platforms and declare untested tools inappropriate for primary education. That is a normative framework, not a developmental trial. Inference: below that threshold, the school may teach about AI; it must not offer the student a generative interlocutor alone. A teacher may, in professional learning, use a model with non-identifiable child data; that is another article. The kindergarten or first-cycle student does not “ask the AI.” The student asks the teacher, a peer, or the material. Teachable Machine, used to classify gestures or images without uploading identities to a general-purpose model, can be a means of the data–pattern–limit axis (Vartiainen et al., 2020); a homework chatbot cannot.

7.3. Understanding test, not use test. Casal-Otero et al. (2023) document the rarity of assessing whether understanding occurred. DAILy and PrimaryAI assess concepts; Vartiainen et al. (2020) assess explanations. Inference: an activity counts as literacy if the student can say, in their language, what datum was used, what pattern was sought, what was predicted, how the system erred, and who decides. If it only produces a flashy output, there is use, not understanding. UNESCO’s Understand level (Miao, Shiohira, & Lao, 2024) and AI4K12’s K–2 and 3–5 bands (Touretzky et al., 2019) suffice to design that assessment; the Create level is not needed.

7.4. No-development-promise test. None of the nuclear cases measures global cognitive development, literacy, or mathematics as an effect of AI literacy. Su and Yang (2022) report heterogeneous syntheses that this article does not reify. Inference: the argument for teaching these ideas is civic and epistemic—recognizing systems that already operate on childhood, protecting data, not attributing a mind to a machine (UNICEF, 2021; UNESCO, 2021)—not a performance shortcut. Whoever sells “AI to raise PISA” has, in this corpus, no backing.

7.5. Band correspondence, inferred, not prescribed as a standard. Early childhood (approximately 3–6): bodily data and pattern; limit as visible error; agency as “I choose the photos”; no account, no generative conversation, with an unplugged option (Miao, Shiohira, & Lao, 2024, Understand level). Lower primary (6–8): sensors versus perception; small labeled sets; “who is not in the photos”; still no independent chatbot. Upper primary (9–11): a classifier in a science or language problem; bias as missing data; “do no harm” ethics and proportionality in everyday cases (PrimaryAI; Miao, Shiohira, & Lao, 2024). Early secondary (12–14): DAILy or equivalent: prediction, bias, stakeholders, careers; generative conversation only if the threshold is met, permission exists, and mediation is present, and even then as an object of critique, not as the author of the assignment (Miao & Holmes, 2023; Zhang et al., 2023). UNESCO’s Create level remains elective.

Operationally, the framework admits: (a) unplugged classification, decision, and “AI or not AI” activities; (b) visible image, sound, or gesture classifiers with local data and no identities in general-purpose models; (c) bias discussion with everyday examples; (d) assessment of explanations, not of outputs. The framework rejects: (e) ChatGPT or homologues as an independent interlocutor under age thirteen; (f) literacy reduced to prompt engineering; (g) the claim that teaching AI develops the child or improves performance if it was not measured; (h) replacement of teacher or student judgment by a model’s output.

8. Discussion

Three tensions organize the discussion. The first is between use and understanding. It is a finding that students aged 10 to 14 can improve a concept inventory and articulate bias (Zhang et al., 2023; Zhang, Lee, & Moore, 2024), that students aged 9 to 11 can gain in machine-learning and vision concepts (Ottenbreit-Leftwich et al., 2023b), and that children aged 3 to 9 can verbalize a classifier’s input and output (Vartiainen et al., 2020). It is a framework that literacy is defined as critical evaluation, not as an interface skill (Long & Magerko, 2020; Ng et al., 2021; Miao, Shiohira, & Lao, 2024). It is not a finding that “using AI” produces that understanding. DAILy shows, in fact, that one can interact with tools and still confuse prediction with automatism (Zhang et al., 2023). Confusing use with understanding is the category error this article names in its title.

The second tension is between precocity and threshold. Qualitative early-childhood evidence does not authorize advancing generative conversation; it authorizes advancing the ideas of data and error. Miao and Holmes (2023) and UNICEF (2021) protect privacy, safety, and agency. A kindergarten that installs a chatbot “so they get familiar” violates the normative threshold and does not gain Vartiainen et al.’s (2020) finding, which was obtained with a classifier and with adults present. Legitimate precocity is conceptual and ethical, not conversational.

The third tension is between global framework and local classroom. UNESCO (2024), AI4K12, and AILit (2026) offer progressions. Casal-Otero et al. (2023) warn that a competency framework is needed to guide modular didactic proposals adjusted to schools. Inference: adopting UNESCO’s student framework is not adopting a product. It is deciding, with local bands, which Understand-level block is taught this year and with what evidence of understanding it is closed. AILit adds the PISA 2029 horizon; that horizon does not justify training seven-year-olds on a chatbot to “do well on the exam.” It justifies, if anything, teaching children to manage and shape AI as citizens, beginning by not attributing a mind to it.

If the object is understanding, the early childhood and basic-education teacher need not be an engineer. They need a shared lexicon and the authority to stop the activity when the system asks for data that UNICEF (2021) obliges them to protect. Replacing that judgment with a license is, again, use without understanding—this time, institutional.

9. Limits

This review is narrative. It does not apply a full PRISMA protocol or estimate pooled effects. The nuclear cases have small samples or non-randomized designs: n = 25 and n = 31 in exploratory DAILy (Zhang et al., 2023; Lee et al., 2021); n = 89 versus 69 without randomization (Zhang, Lee, & Moore, 2024); two classrooms in PrimaryAI (Ottenbreit-Leftwich et al., 2023b); n = 6 and n = 18 in Teachable Machine (Vartiainen et al., 2020; Sanusi et al., 2024). Geography is biased toward the United States, Finland, and an informal African setting; no Latin American AI-literacy trial in early childhood or primary that met this article’s criteria was located with the same degree of DOI openness. That absence is a gap, not proof of non-existence. The PrimaryAI poster limits statistical detail. Su and Yang (2022) add studies that mix tools-for-learning with teaching-about-AI; this article does not reanalyze those seventeen items one by one. The UNESCO, UNICEF, AI4K12, and AILit frameworks are prescriptive: their authority is normative, not empirical. Section 7 inferences are curriculum-design hypotheses, not evidence of national implementation. The age-thirteen threshold is a policy recommendation, not a developmental-psychology finding measured in this corpus.

10. Conclusions

Making basic and early childhood students literate in artificial intelligence is not sitting them down to converse with a generative model. It is teaching a curriculum of ideas—data, pattern, prediction, limit, human agency—that current frameworks already name and that three families of verified studies show, modestly, as learnable. In early secondary, DAILy associates a thirty-hour, mediated, ethically interwoven curriculum with gains in concepts and career awareness, not with general performance gains, and leaves prediction as a difficult concept (Zhang et al., 2023; Zhang, Lee, & Moore, 2024). In upper primary, PrimaryAI associates a life-science problem with knowledge gains in machine learning and vision, in a teacher-led classroom, with no reported gender differences (Ottenbreit-Leftwich et al., 2023b). In early childhood and early primary, Teachable Machine allows, in small informal samples, verbalizing input and output (Vartiainen et al., 2020; Sanusi et al., 2024). None of those findings authorizes a promise of child development or a better grade.

Law and policy are not ambiguous. UNESCO’s student framework prioritizes a human-centred mindset and reserves the Understand level for all, with unplugged options (Miao, Shiohira, & Lao, 2024). AI4K12 distributes the five ideas by band (Touretzky et al., 2019). UNICEF requires protecting data, safety, and civic preparation, not premature exposure (UNICEF, 2021). Miao and Holmes (2023) set thirteen years for independent conversation. The 2026 AILit speaks of engaging, creating, managing, and shaping, not of chatting ahead of time (European Commission & OECD, 2026). Where the sources do not measure development or achievement, this article does not affirm them. Use is not understanding. Understanding, at these ages, fits in a classroom without a chatbot.

Ingeniero Mitre / Laboratorio Editorial de NEXTECH.IA

References

  1. AI4K12 Initiative. (2020). Five Big Ideas in AI and grade-band progression charts. AAAI & CSTA. https://ai4k12.org/
  2. Casal-Otero, L., Catala, A., Fernández-Morante, C., Taboada, M., Cebreiro, B., & Barro, S. (2023). AI literacy in K-12: A systematic literature review. International Journal of STEM Education, 10, 29. https://doi.org/10.1186/s40594-023-00418-7
  3. European Commission & OECD. (2026). Empowering learners for the age of AI: An AI literacy framework for primary and secondary education. OECD Publishing. https://www.oecd.org/en/publications/empowering-learners-for-the-age-of-ai_65cd27d4-en.html
  4. Lee, I., Ali, S., Zhang, H., DiPaola, D., & Breazeal, C. (2021). Developing middle school students’ AI literacy. In Proceedings of the 52nd ACM Technical Symposium on Computer Science Education (pp. 191–197). ACM. https://doi.org/10.1145/3408877.3432513
  5. Long, D., & Magerko, B. (2020). What is AI literacy? Competencies and design considerations. In Proceedings of the 2020 CHI Conference on Human Factors in Computing Systems. ACM. https://doi.org/10.1145/3313831.3376727
  6. Miao, F., & Holmes, W. (2023). Guidance for generative AI in education and research. UNESCO. https://doi.org/10.54675/EWZM9535
  7. Miao, F., Shiohira, K., & Lao, N. (2024). AI competency framework for students. UNESCO. https://doi.org/10.54675/JKJB9835
  8. Ng, D. T. K., Leung, J. K. L., Chu, S. K. W., & Qiao, M. S. (2021). Conceptualizing AI literacy: An exploratory review. Computers and Education: Artificial Intelligence, 2, 100041. https://doi.org/10.1016/j.caeai.2021.100041
  9. OECD. (2025). What should teachers teach and students learn in a future of powerful AI? OECD Publishing. https://www.oecd.org/content/dam/oecd/en/publications/reports/2025/05/what-should-teachers-teach-and-students-learn-in-a-future-of-powerful-ai_4578ec74/ca56c7d6-en.pdf
  10. Ottenbreit-Leftwich, A., Glazewski, K., Jeon, M., Jantaraweragul, K., Hmelo-Silver, C. E., Scribner, A., Lee, S., Mott, B., & Lester, J. (2023a). Lessons learned for AI education with elementary students and teachers. International Journal of Artificial Intelligence in Education, 33, 267–289. https://doi.org/10.1007/s40593-022-00304-3
  11. Ottenbreit-Leftwich, A., Glazewski, K., Hmelo-Silver, C., Jantaraweragul, K., Chakraburty, S., Jeon, M., Scribner, J. A., Lee, S., Mott, B., & Lester, J. (2023b). Is elementary AI education possible? In Proceedings of the 54th ACM Technical Symposium on Computer Science Education (p. 1364). ACM. https://doi.org/10.1145/3545947.3576308
  12. Sanusi, I. T., Sunday, K., Oyelere, S. S., Suhonen, J., Vartiainen, H., & Tukiainen, M. (2024). Learning machine learning with young children: Exploring informal settings in an African context. Computer Science Education, 34(2), 161–192. https://doi.org/10.1080/08993408.2023.2175559
  13. Su, J., & Yang, W. (2022). Artificial intelligence in early childhood education: A scoping review. Computers and Education: Artificial Intelligence, 3, 100049. https://doi.org/10.1016/j.caeai.2022.100049
  14. Touretzky, D., Gardner-McCune, C., Martin, F., & Seehorn, D. (2019). Envisioning AI for K-12: What should every child know about AI? Proceedings of the AAAI Conference on Artificial Intelligence, 33(01), 9795–9799. https://doi.org/10.1609/aaai.v33i01.33019795
  15. Touretzky, D. S., & Gardner-McCune, C. (2021). Artificial intelligence thinking in K-12. AI4K12. https://ai4k12.org/wp-content/uploads/2021/08/Touretzky_Gardner-McCune_AI-Thinking_2021.pdf
  16. UNESCO. (2021). Recommendation on the ethics of artificial intelligence. https://unesdoc.unesco.org/ark:/48223/pf0000381137
  17. UNICEF. (2021). Policy guidance on AI for children 2.0. UNICEF Office of Global Insight and Policy. https://www.unicef.org/innocenti/reports/policy-guidance-ai-children
  18. Vartiainen, H., Tedre, M., & Valtonen, T. (2020). Learning machine learning with very young children: Who is teaching whom? International Journal of Child-Computer Interaction, 25, 100182. https://doi.org/10.1016/j.ijcci.2020.100182
  19. Zhang, H., Lee, I., Ali, S., DiPaola, D., Cheng, Y., & Breazeal, C. (2023). Integrating ethics and career futures with technical learning to promote AI literacy for middle school students: An exploratory study. International Journal of Artificial Intelligence in Education, 33, 290–324. https://doi.org/10.1007/s40593-022-00293-3
  20. Zhang, H., Lee, I., & Moore, K. (2024). An effectiveness study of teacher-led AI literacy curriculum in K-12 classrooms. Proceedings of the AAAI Conference on Artificial Intelligence, 38(21), 23318–23325. https://doi.org/10.1609/aaai.v38i21.30380