1. Introduction: the pedagogical problem of a mediation that speaks

Early childhood education has been built, in the contemporary pedagogical tradition, as a space in which play, oral language and the sensitive presence of the adult constitute the primary scaffolding of development. When a machine begins to speak, to reply, to recommend and to simulate understanding, that triad is challenged in a way that is not merely technical. Generative artificial intelligence —systems capable of producing text, image, audio or dialogue from models trained on large volumes of data— enters classrooms and homes at a speed that, as Miao and Holmes (2023) warn, outpaces regulation. The problem formulated here is not whether the technology “works”, but under what ethical and didactic conditions an algorithmic mediation can sustain, rather than supplant, the play and orality of early childhood. The question is deliberately restrictive: it is not a matter of transferring to early childhood the promises of productivity that circulate in higher education or professional training, but of interrogating an age range in which the attribution of intentionality, the confusion between companion and artefact, and the vulnerability of personal data acquire a specific weight.

That restriction is justified by the state of the evidence itself. Su and Yang (2022), in a scoping review of seventeen studies published between 1995 and 2021, show that research on artificial intelligence in early childhood education is still a narrow field, though no longer an empty one. The reviewed works indicate that artificial intelligence improved understanding of concepts of artificial intelligence, machine learning, computer science and robotics, and that it was also associated with the development of creativity, emotion control, collaborative inquiry, literacy and computational thinking. That list of effects does not authorise unrestricted enthusiasm: seventeen studies across more than a quarter of a century are, in themselves, a datum of caution. At the same time, they prevent the lazy scepticism that declares nonexistent everything that does not reach the scale of a massive randomised trial. The present article occupies that interval: it takes the available findings seriously, refuses to invent others, and draws from them a pedagogical thesis about scaffolding.

The thesis may be stated as follows. If artificial intelligence, in the forms already studied —curricula of construction and training of systems, robotic toys with an artificial-intelligence interface, domestic conversational agents— has shown a capacity to support concepts, emotions, inquiry and language, then generative artificial intelligence does not introduce an ontological rupture, but an intensification of verbal mediation. That intensification increases the potential for scaffolding play and oral language and, in the same measure, increases the risk of an undue surrender of adult presence. Yang (2022) formulates the conceptual core that children can explore: algorithms trained on large volumes of data identify patterns, predict and recommend, and they have limits. That idea, and not fascination with the fluency of a chatbot, should guide curriculum design. Miao and Holmes (2023) add the normative correlate: a human-centred approach, the current lack of protection of data privacy, and an age limit for independent conversations with generative artificial-intelligence platforms. Scaffolding, in this framework, is not a complacent metaphor: it is the condition of possibility of any defensible use.

2. Artificial intelligence in early childhood education: the scope of scarce and convergent evidence

The review by Su and Yang (2022) offers the most prudent map of what can be affirmed without overreach. Seventeen studies, a temporal arc from 1995 to 2021, and a set of results that do not disperse in contradictory directions, but converge on two families of effects. The first family is conceptual and disciplinary: the children who participated in the reviewed experiences improved their understanding of notions of artificial intelligence, machine learning, computer science and robotics. The second family is competency-based and socio-emotional: advances were observed in creativity, emotion control, collaborative inquiry, literacy and computational thinking. This double family is pedagogically eloquent. It is not only a matter of “teaching machines”, but of a displacement of the learning subject: artificial intelligence appears, in the available evidence, both as an object of knowledge and as a mediation that reorganises practices of creation, affective regulation, joint inquiry and language.

That eloquence, however, must not be confused with a robust causal generalisation. A scoping review is not a meta-analysis; seventeen studies do not constitute a saturated corpus; and the heterogeneity of devices —social robots, toys, voice agents, construction curricula— prevents treating “artificial intelligence” as a single variable. The value of Su and Yang (2022) lies precisely in that honesty of scale: the field exists, it is small, and what it shows is promising in directions that coincide with the classical aims of early childhood education. Anyone wishing to found a classroom policy or a public policy on this basis must know that they are arguing from emerging evidence, not from a consolidated doctrine. The present article accepts that condition and turns it into method: every empirical claim refers to a named study; every pedagogical inference is declared as such.

Across a broader age range, Su, Guo, Chen and Chu (2023) have mapped the teaching of artificial intelligence in K–12 classrooms through another scoping review. Its very existence is a contextual datum: what in early childhood is still a reduced set of experiences is inscribed, in later schooling, in a denser curricular debate. That unevenness does not authorise the unmediated import of secondary-school frameworks into the kindergarten. It does authorise recognition that the question of what is taught, how it is taught and with what ethics artificial intelligence is taught is not a whim of the generative juncture of 2023–2024, but a problem that precedes ChatGPT and that the irruption of large language models does not invent, but aggravates. Kasneci et al. (2023) have discussed the opportunities and challenges of those models for education in general; Yan, Greiff, Teuber and Gašević (2024) have examined the promises and challenges of generative artificial intelligence for human learning. None of those works replaces the specific evidence of early childhood; both prevent treating the phenomenon as a merely local or merely technical matter.

The articulation among these bibliographic strata allows a methodological claim. There is a small and valuable empirical core on artificial intelligence in early childhood education (Su and Yang, 2022; Williams et al., 2019; Kewalramani, Palaiologou, Dardanou, Allen and Phillipson, 2021; Kewalramani, Kidman and Palaiologou, 2021; Druga, Williams, Breazeal and Resnick, 2017). There is a curricular framework that translates that core into design principles (Yang, 2022). There is a broader debate on generative models and human learning (Kasneci et al., 2023; Yan et al., 2024). There is, finally, an international governance that attempts to check the asymmetry between innovation and regulation (Miao and Holmes, 2023; Miao and Cukurova, 2024; Miao, Shiohira and Lao, 2024). The article moves among those four strata without mixing their statuses: the empirical is cited as empirical; the curricular, as curricular; the normative, as normative; the argumentative, as argumentative.

3. Artificial intelligence literacy: what young children can explore and how the curriculum should be designed

Yang (2022) displaces the problem from technological fascination toward a classical question of curriculum theory: why, what and how to teach artificial intelligence to young children. The answer offered is not a catalogue of tools, but a redefinition of literacy. Artificial intelligence literacy is understood as a form of digital literacy: not training in a particular software, but the capacity to inhabit a world in which opaque systems identify patterns, predict behaviours and recommend actions. The conceptual core that children can explore is, in Yang’s (2022) formulation, notably sober and notably sufficient: algorithms are trained on large volumes of data; they identify patterns; they predict; they recommend; and they have limits. That fivefold structure —data, patterns, prediction, recommendation, limit— is pedagogically more fertile than any demonstration of conversational fluency, because it restores to children the right to interrogate the machine rather than to admire it.

Yang’s (2022) how is articulated in two principles that this article takes as the compass of scaffolding. The first is learning-by-making: artificial intelligence is understood when it is built, programmed, trained and put to the test, not when it is contemplated as an oracle. The second is pedagogy-as-relational: knowledge does not circulate in a closed circuit between child and system, but in a network of presences —peers, teachers, families, cultural objects— that give meaning to what the machine does and to what the machine cannot do. To those principles is added a demand for embodiment and cultural responsiveness: the curriculum is not a universal package downloaded onto any classroom, but a proposal that must be embodied in local bodies, languages, games and cosmologies. The “AI for Kids” curriculum appears, in that framework, as an illustration of design and not as an imperial template.

This conception has a direct consequence for the debate on generative artificial intelligence. A language model that produces fluent responses can seduce the adult with the illusion of an inexhaustible tutor and seduce the child with the illusion of a companion who “understands”. If Yang’s (2022) conceptual core is taken seriously, that double illusion is precisely what the curriculum must dismantle. Exploring the fact that a system identifies patterns in data, predicts and recommends, and that those operations have limits, is a form of literacy that protects play: the child may speak with the machine, but must not be left alone with the belief that the machine speaks as an adult speaks. Relational pedagogy is not a humanist ornament; it is the device that prevents algorithmic mediation from becoming a presence without responsibility.

In dialogue with Su and Yang (2022), Yang’s (2022) curricular proposal allows a rereading of the reported effects. If the scoping evidence shows improvements in concepts of artificial intelligence and machine learning, that is coherent with a curriculum that places at its centre the idea of algorithms, data and limits. If it shows advances in creativity, collaborative inquiry and literacy, that is coherent with learning-by-making and with relational pedagogy. If it shows emotion control and computational thinking, that suggests that the mediation does not operate only on the cognitive-declarative plane, but in practices of regulation and systematisation that early childhood education already recognises as its own. None of this proves that a generative chatbot produces, by itself, the same effects. It does prove that design matters more than the device, and that the defensible design is the one that keeps children in the position of exploring limits, not of consuming fluency.

4. Case study I. PopBots: building, programming, training and interacting as a path to literacy

Williams, Park, Oh and Breazeal (2019) designed PopBots as a toolkit and an artificial-intelligence curriculum for early childhood education. The study involved eighty Pre-K and Kindergarten children, aged four to seven, and situated a social robot as a learning companion. The central finding, in the terms this article is authorised to retain, is that the children learned artificial-intelligence concepts by building, programming, training and interacting. That fourfold sequence —build, program, train, interact— is not a minor methodological detail: it is the empirical translation of the learning-by-making that Yang (2022) formulates on the curricular plane. PopBots does not present artificial intelligence as a service to be consulted, but as an artefact that is fabricated and tested in the company of a robot that, precisely because it is a companion, does not cease to be an object of design.

The age of the participants —four to seven years— places the study at the very threshold of early childhood education and the beginning of schooling. Eighty children do not constitute a universal population, but they do constitute a scale that exceeds anecdote and that allows the claim to be taken seriously that, with an adequate curriculum and toolkit, early childhood can appropriate concepts of artificial intelligence without requiring a premature formalisation in the style of textual programming for adults. The social robot as learning companion introduces, at the same time, a tension that the rest of this article’s literature helps to name. A robotic companion can sustain motivation, conversation and play; it can also activate attributions of intelligence and identity that Druga et al. (2017) documented with other agents. The design of PopBots, by requiring children to build, program and train, offers a structural antidote to that uncritical attribution: whoever trains the system is in a better position to understand that the system does not “know” in a human way.

From the thesis of this article, PopBots illuminates the scaffolding of play in a specific way. Play does not appear as an extrinsic reward added at the end of a lesson, but as the medium in which construction and interaction become meaningful. A social robot that accompanies does not replace the teacher: it reorganises the space of joint activity. The adult remains the one who holds the question, who names the limit, who translates the experience into language. Algorithmic mediation, in this case, is material and programmable; it is not yet the opaque fluency of a large-scale generative model. Precisely for that reason PopBots serves as a counterweight: it shows that the most fertile path for early childhood is not unlimited conversation with an inscrutable system, but the situated fabrication of a system whose behaviour can be interrogated. If generative artificial intelligence is to enter the kindergarten, the precedent of Williams et al. (2019) suggests that it should enter as an object that is explored and limited, not as a voice that is consulted in solitude.

A reading in terms of literacy is also warranted. Su and Yang (2022) report improvements in literacy and computational thinking across the studies in their review; PopBots supplies a plausible mechanism for those improvements: by building and programming, children rehearse sequences, hypotheses and corrections; by training, they confront the idea that data leave a trace in the system’s behaviour; by interacting, they put oral language into play in a circuit that is not purely human. Scaffolding, here, is triple: material (the toolkit), curricular (the sequence of activities) and social (the robot as companion and the adult as guarantor). To remove any of those three pillars —especially the third— would turn the experience into something else: the consumption of novelty or the abandonment of the child before a speaking machine.

5. Case study II. Robotic toys, lockdown and emotional capital: five children in Australia, 2020

Kewalramani, Palaiologou, Dardanou, Allen and Phillipson (2021) investigated the use of artificial-intelligence robotic toys in the home during the COVID-19 lockdown in Australia in 2020. The study, oriented by design-based research, involved five children with diverse additional needs. The methodological device was not the measurement of academic performance, but the production of empathy-based dialogues and the analysis of what the authors conceptualise as emotional capital and as a form of imaginary togetherness —a shared companionship that is also sustained on the plane of the imaginary. In a context of confinement, interruption of in-person routines and forced reconfiguration of care, the robotic toy appears not as a didactic luxury, but as a mediator of emotional life and of joint presence.

The facts this article retains are deliberately austere and, for that reason, more demanding. Five children do not allow population inferences; the diversity of additional needs prevents treating the group as homogeneous; the 2020 lockdown is a historical context unrepeatable in its exact form and, at the same time, a revealer of what happens when the classroom dissolves into the home. What the study authorises one to affirm is that, under those conditions, AI robotic toys participated in practices of empathic dialogue, in the mobilisation of emotional capital and in a shared imaginary companionship. That does not prove that “robots improve inclusion”. It proves that, in a research design attentive to emotion and relationship, technical mediation can be inscribed in the work of sustaining the bond when the in-person bond of the educational centre has been interrupted.

Kewalramani, Kidman and Palaiologou (2021), in a convergent paper, make the case for AI-interfaced robotic toys in early childhood settings as a pathway for children’s inquiry literacy. The formulation is important: this is not digital literacy in the narrow sense of handling a device, but an inquiry literacy —asking, exploring, constructing meaning— mediated by an artefact that responds. Read together with Su and Yang (2022), this line reinforces the family of effects that includes collaborative inquiry and literacy. Read together with Yang (2022), it confirms that relational pedagogy is not an abstract postulate: it is the way in which a toy ceases to be a task-solver and becomes a situated interlocutor, provided that the adult sustains the dialogue.

Lockdown adds a lesson that generative artificial intelligence makes urgent. When the public space of early childhood education closes, technical mediation can colonise the home with an ease that the classroom, with its rituals and its multiple presences, usually moderates. A robotic toy, in the design of Kewalramani, Palaiologou, Dardanou, Allen and Phillipson (2021), is inscribed in empathic dialogues and in an imaginary togetherness; a domestic generative model, without that design and without that adult, can be inscribed in a conversational solitude. The difference does not lie in the “intelligence” of the system, but in the quality of the human scaffolding that surrounds it. Emotional capital is not produced by the algorithm: it is mobilised by relations. The article retains this distinction as a criterion for evaluating any proposal that seeks to bring the automatic generation of language into early childhood in care contexts.

6. “Hey Google, is it OK if I eat you?”: attribution of intelligence, identity and playfulness

Druga, Williams, Breazeal and Resnick (2017) conducted initial explorations of child-agent interaction with twenty-six children aged three to ten who used Alexa, Google Home, Cozmo and the Julie Chatbot. The very title of the paper —a child’s question to an assistant about whether it is permitted to “eat” it— condenses the hybrid status of these objects: they are spoken to as people are spoken to, they are granted an imaginary body, and they are submitted to the rules of play. The themes the study identifies —perceived intelligence, identity attribution, playfulness and understanding— constitute, for this article, the earliest and most honest map of what happens when an algorithmic voice enters the child’s world.

Perceived intelligence is not intelligence measured by a test, but that which the child grants the agent on the basis of its replies, its failures, its tone and its punctuality. Identity attribution names the next step: not only “this seems clever”, but “this is someone”. Playfulness recalls that early childhood does not relate to artefacts primarily as a user of services, but as a player who rehearses hypotheses, provocations, transgressions and affections. Understanding, finally, signals that there is cognitive work under way: children are not passive recipients of a magic, but interpreters who construct models of how that which speaks to them works. Druga et al. (2017) do not deliver, in the terms this article permits itself to use, a metric of curricular learning; they deliver something perhaps more decisive for the ethical problem: a phenomenology of the relation.

That phenomenology intersects with PopBots and with the robotic toys in a way that should be made explicit. In Williams et al. (2019), the robot is a learning companion, but the curriculum requires building, programming and training: identity attribution is tensed by fabrication. In Kewalramani, Palaiologou, Dardanou, Allen and Phillipson (2021), the toy participates in an imaginary togetherness and in empathic dialogues: attribution is not denied; it is worked on the emotional and relational plane. In Druga et al. (2017), interaction with voice assistants and a chatbot shows that attribution emerges even —and above all— when the system has not been built by the child, but presents itself as an already-made, domestic, always-available presence. Contemporary generative artificial intelligence resembles this third scenario more than the first two: it is an already trained, already fluent, already installed voice that does not need to be fabricated in order to be consulted.

From this follows a design criterion that this article regards as non-negotiable. If the attribution of intelligence and identity is a phenomenon documented from the age of three, and if playfulness is the primary mode of that attribution, then any introduction of generative systems into early childhood education must treat that attribution as pedagogical content and not as a side effect. One must be able to ask, in the language of Yang (2022), what patterns the system identifies, on what data it is fed, what it recommends and where its limits fail. One must be able to sustain, in the language of Miao and Holmes (2023), that that conversation cannot be independent: the adult is part of the circuit, not an optional supervisor. Play is not protected by forbidding all interlocution with machines; it is protected by preventing interlocution from becoming a duet without a human witness.

7. Generative artificial intelligence, oral language and the scaffolding of play

Generative artificial intelligence introduces, relative to the systems studied in early childhood between 2017 and 2021, a difference of degree that tends to become a difference of nature: fluency. A robot that is built and trained, a toy that replies with a limited repertoire, a voice assistant that misunderstands and provokes laughter, let their seams show. A large-scale language model conceals them. Kasneci et al. (2023) have examined the opportunities and challenges of those models for education; Yan et al. (2024) have analysed the promises and challenges of generative artificial intelligence for human learning. This article does not extract from those works findings that are not in the sources verified here; it extracts a pedagogical consequence for early childhood education: the more the machine resembles a competent interlocutor, the more it obscures the conceptual core that Yang (2022) wants children to explore —data, patterns, prediction, recommendation, limits— and the more it tensions relational pedagogy.

The oral language of early childhood is not a channel for the transmission of contents. It is the medium in which intersubjectivity is built, roles are rehearsed, symbolic play is negotiated and emotion is regulated. If Su and Yang (2022) report advances in literacy, collaborative inquiry and emotion control in artificial-intelligence experiences, that suggests that technical mediation can, under design, be inscribed in those practices. The generative temptation consists in inverting the vector: instead of inscribing the machine in play, inscribing play in the machine, as if an automatic dialogue could replace the idiosyncrasy of children’s speech, its silences, its repetitions, its productive errors. Scaffolding, in the tradition this article claims, is a transitory, contingent help, withdrawn progressively by an adult who knows the child. A generative system does not withdraw its help: it offers it in an unlimited, uniform manner and without pedagogical memory of the concrete subject, except that which its data logs —themselves problematic— may accumulate.

It is then possible to distinguish three possible uses, none of which is presented here as an empirical finding, but as a design hypothesis grounded in the sources. The first is the use of generation as an object of exploration: producing with the system, in the adult’s presence, a story, a play prompt or a variation of a ritual, and asking afterwards what data, what patterns and what limits are noticed. That use aligns with Yang (2022) and with the PopBots sequence (Williams et al., 2019). The second is the use of generation as support for inquiry and for the literacy of asking, in the line of Kewalramani, Kidman and Palaiologou (2021), provided that the question remains the child’s and not an adult prompt that imposes itself. The third, which this article discards as incompatible with Miao and Holmes (2023), is the use of generation as an independent interlocutor of childhood: conversation alone with the platform. That third use is not a rhetorical extreme; it is the commercial default of contemporary systems and, therefore, the structural risk that early childhood education must name.

Play, in this scheme, is not the place of the playful application of a technology, but the criterion of judgement. A use is defensible when it expands the repertoire of symbolic actions, speaking turns, shared hypotheses and emotional regulations that the group already practises. A use is not defensible when it reduces play to waiting for a brilliant reply, when it captures attention in a dyadic circuit with the screen or the loudspeaker, or when it displaces the adult from the position of scaffolding to the position of a technician who “turns the system on”. Kewalramani, Palaiologou, Dardanou, Allen and Phillipson (2021) show that, even in confinement, the value of the robotic toy is at stake in empathic dialogue and in emotional capital, not in the autonomy of the artefact. That lesson holds, a fortiori, for a model that speaks better than any toy.

8. Governance and competencies: UNESCO facing the asymmetry between innovation and regulation

Miao and Holmes (2023), in UNESCO’s guidance for generative artificial intelligence in education and research, establish three facts that this article treats as constraints, not as optional recommendations. The first: generative artificial intelligence advances faster than regulation. The second: data privacy is unprotected. The third: the approach must be human-centred, and there is an age limit for independent conversations with generative artificial-intelligence platforms. In early childhood education, those three constraints become almost tautological and, nevertheless, are transgressed daily. Systems reach the home before they reach educational policy; children’s voices are recorded in opaque infrastructures; independent conversation is, in practice, the simplest gesture —leaving the device within reach— and, at the same time, the most incompatible with scaffolding.

Miao and Cukurova (2024) propose an artificial-intelligence competency framework for teachers: fifteen competencies organised in five dimensions —human-centred mindset, ethics, foundations and applications, AI pedagogy, and AI for professional learning— and three progression levels: Acquire, Deepen and Create. The framework is not written as a kindergarten curriculum, but its architecture is revealing for whoever designs mediation in early childhood. The human-centred mindset and ethics precede technical foundations; AI pedagogy is a dimension in its own right, not an appendix to “tool use”; professional learning recognises that the teacher is also a subject who is formed, not an operator trained in an afternoon. The levels Acquire–Deepen–Create prevent both the immobility of whoever declares themselves forever incompetent and the productivism of whoever pretends to “create” without having acquired.

Miao, Shiohira and Lao (2024) offer the correlative framework for students: twelve competencies in four dimensions —human-centred mindset, ethics, techniques and applications, and system design— with levels of Understand, Apply and Create. The difference in architecture from the teacher framework is not trivial. Students are not formed, in this scheme, to “use the tool better”, but to inhabit a mindset, an ethics, a technicality and a capacity for design. In early childhood education, “system design” does not mean software engineering; it means, in the wake of PopBots (Williams et al., 2019) and of Yang (2022), that children can build, train and interrogate artefacts, and that they understand that those artefacts have limits. “Understand–Apply–Create” is a progression that, in this age range, must be read in the key of play and orality, not of early accreditation.

The articulation of the three UNESCO documents produces a minimum standard that this article adopts as the normative closure of the thesis. There is no defensible use of generative artificial intelligence in early childhood education without a teacher who, at least at the Acquire level, sustains a human-centred mindset and an explicit ethics (Miao and Cukurova, 2024). There is no defensible childhood literacy that omits ethics and reduces the experience to technical application (Miao, Shiohira and Lao, 2024). There is no independent conversation of young children with generative platforms that can be reconciled with the guidance of Miao and Holmes (2023). The scaffolding of play and oral language, far from being a romanticism, is the didactic translation of that standard: the adult remains, data are treated as a problem, the system is interrogated, the child is not left alone with a voice that simulates care.

9. Ethical limits of algorithmic mediation in early childhood

The limits this article defends are not derived from a pessimistic anthropology of technique, but from the conjunction of the facts already cited. If data privacy is unprotected (Miao and Holmes, 2023), then every vocal interaction of a child with a generative system is, potentially, an extraction. If children attribute intelligence and identity to agents and relate to them from playfulness (Druga et al., 2017), then that extraction is not an informed formality, but an asymmetrical relation in which consent cannot be presupposed. If the documented pedagogical value depends on building, programming, training and interacting (Williams et al., 2019), on empathic dialogues and emotional capital (Kewalramani, Palaiologou, Dardanou, Allen and Phillipson, 2021), and on a relational and culturally responsive pedagogy (Yang, 2022), then a system that operates as a domestic oracle without fabrication, without mediated dialogue and without cultural embodiment does not automatically inherit those values. The limit is not “do not use technology”; the limit is not to confuse fluency with justification.

A second limit concerns equity and the diversity of needs. The study by Kewalramani, Palaiologou, Dardanou, Allen and Phillipson (2021) involves five children with diverse additional needs in a lockdown home. That framing, far from being a marginal note, obliges distrust of universal solutions. One and the same conversational agent does not produce the same scaffolding in different bodies, languages, neurodivergences and care situations. The cultural responsiveness that Yang (2022) demands and the human-centred mindset of the UNESCO frameworks (Miao and Cukurova, 2024; Miao, Shiohira and Lao, 2024) coincide here: the design that does not start from difference produces exclusion with the additional eloquence of seeming innovative. The article does not have evidence to affirm which adaptations work in each case; it does have evidence to affirm that difference was already at the centre of one of the most sensitive investigations in the corpus, and that to ignore it would be a form of pedagogical bad faith.

A third limit is epistemic and is addressed to those who research and those who write —including this text. Su and Yang (2022) reviewed seventeen studies across twenty-six years. PopBots studied eighty children. The Australian 2020 study examined five. Druga et al. (2017) explored twenty-six. Those figures are not a defect to be hidden with ambitious prose; they are the honest contour of the available knowledge. Kasneci et al. (2023) and Yan et al. (2024) widen the debate toward generative models and human learning, but they do not fill the void of specific evidence on early childhood education and the automatic generation of language. Whoever promises, on the basis of this corpus, that generative artificial intelligence “develops oral language” or “improves symbolic play” in early childhood will be inventing a finding. Whoever, on the basis of the same corpus, argues that human scaffolding is the condition of any defensible mediation will be interpreting, not fabricating.

The fourth limit is institutional and concerns formation. The fifteen teacher competencies of Miao and Cukurova (2024) and the twelve student competencies of Miao, Shiohira and Lao (2024) are not improvised in a day of “digital updating”. Acquiring a human-centred mindset and an operative ethics is a work of professional culture. In early childhood education, that culture already possesses languages of its own —play, care, observation, documentation, family participation— that must not be replaced by the jargon of innovation. The competency framework should be translated into those languages, not the other way around. A kindergarten that “uses artificial intelligence” without being able to name what data are collected, what patterns are invoked, what is recommended to families and where the age limit for independent conversation lies is not in the vanguard: it is in ethical default.

10. Implications for the design of practices: scaffolding, not automation

The implications that follow are presented as design orientations, not as results of a new experiment. First: the object of the activity, when there is artificial intelligence in early childhood education, should be the system itself —its data, its patterns, its predictions, its recommendations and its limits—, in the line of Yang (2022), and not the obtaining of a brilliant verbal product. Second: the privileged sequence is that of PopBots —build, program, train, interact— (Williams et al., 2019), adapted to local materials and cultures, not the sequence of consult–copy–exhibit. Third: the evaluation of the experience cannot be reduced to the correctness of a reply; it must attend to inquiry literacy (Kewalramani, Kidman and Palaiologou, 2021), to emotional capital and imaginary togetherness (Kewalramani, Palaiologou, Dardanou, Allen and Phillipson, 2021), and to the ways in which children perceive intelligence, attribute identity, play and understand (Druga et al., 2017).

Fourth implication: the adult is not a generic facilitator, but the guarantor of the relational circuit. The UNESCO frameworks (Miao and Cukurova, 2024; Miao, Shiohira and Lao, 2024) require a human-centred mindset and an ethics that, in the early childhood classroom, translate into presence, interpretation and veto. The veto includes, explicitly, independent conversation with generative platforms (Miao and Holmes, 2023). Fifth: families are not an appendix of school innovation. The Australian case of 2020 shows that the home can be the principal scene of mediation, for good when there are designed empathic dialogues, for ill when the device remains as an algorithmic nanny. Sixth: the scale of the evidence obliges institutional modesty. Seventeen studies (Su and Yang, 2022) and an expanding K–12 debate (Su et al., 2023) do not suffice to adopt massive commercial platforms in early childhood; they do suffice to sustain projects of design, documentation and situated research.

Seventh implication, specifically linguistic. If oral language is the medium of scaffolding, then the language of the activity cannot be, by default, the dominant language of the models. Yang’s (2022) relational and culturally responsive pedagogy requires that play and conversation occur in the languages of the community, and that the fact become visible that many systems predict and recommend better in some languages than in others. That unevenness is a literacy content, not a technical detail. Eighth: creativity and emotion control, reported by Su and Yang (2022) as part of the effects of artificial intelligence in early childhood, are not “delivered” by the model; they are cultivated in joint activity. A generative system that effortlessly produces a poem or a prompt can, if the design is careless, impoverish precisely what the evidence associates with well-designed experiences.

Ninth implication, closing this section: generative artificial intelligence need not be absolutely prohibited in order to be rigorously limited. It needs to be redirected toward the forms of mediation that the corpus has already made intelligible —fabrication, relation, inquiry, emotion, limit— and away from the forms that the corpus and the norm have already made suspect —independent conversation, data opacity, unworked attribution, curricular universalism. Kasneci et al. (2023) and Yan et al. (2024) recall, on the plane of human learning in general, that promises and challenges travel together. In early childhood education, that joint travel is not a balanced metaphor: the challenge weighs more, because the subject who attributes identity to an agent (Druga et al., 2017) is not in symmetry with the system that collects their voice.

11. Conclusions

This article has sustained a restrictive and, for that reason, operative thesis. Generative artificial intelligence may participate in the scaffolding of play and oral language in early childhood education only if it is treated as a limited algorithmic mediation, visible in its operations of data, patterns, prediction and recommendation (Yang, 2022), inserted in a relational and culturally responsive pedagogy, and subordinated to an adult who does not delegate the conversation. The scoping evidence gathered by Su and Yang (2022) authorises the claim that, in the experiences studied between 1995 and 2021, artificial intelligence was associated with a better understanding of concepts of artificial intelligence, machine learning, computer science and robotics, and with advances in creativity, emotion control, collaborative inquiry, literacy and computational thinking. It does not authorise the claim that contemporary generative models automatically reproduce those effects.

The case studies specify the mechanism. PopBots shows that eighty children aged four to seven can learn artificial-intelligence concepts by building, programming, training and interacting with a social robot as a learning companion (Williams et al., 2019). The work of Kewalramani, Palaiologou, Dardanou, Allen and Phillipson (2021) shows that, in the Australian lockdown of 2020, five children with diverse additional needs mobilised emotional capital and an imaginary togetherness in empathy-based dialogues mediated by AI robotic toys, in a design-based research design. Kewalramani, Kidman and Palaiologou (2021) situate those toys in an argument about inquiry literacy. Druga et al. (2017) recall, with twenty-six children aged three to ten before Alexa, Google Home, Cozmo and a chatbot, that perceived intelligence, identity attribution, playfulness and understanding are the phenomenological ground of any mediation that speaks.

Upon that ground, UNESCO governance is not a diplomatic epilogue: it is the condition of legitimacy. Generative artificial intelligence is faster than regulation; data privacy is unprotected; the approach must be human-centred; and there is an age limit for independent conversations with the platforms (Miao and Holmes, 2023). Teachers need fifteen competencies in five dimensions, with progression from Acquire to Create (Miao and Cukurova, 2024). Students need twelve competencies in four dimensions, with progression from Understand to Create (Miao, Shiohira and Lao, 2024). In early childhood education, those frameworks are translated into a language that the profession already possesses: to care, to play, to converse, to observe, to document, to include. The algorithmic mediation that cannot be stated in that language does not deserve the classroom.

There remains, as an agenda and not as a finding, the research that this corpus still does not offer: situated studies, of honest scale, on what happens when a generative model enters —always with an adult, always with a limit, always with a question about the data— the symbolic play and the orality of early childhood. Until that evidence exists, the ethics of interpretation consists in not filling the void with promises. Scaffolding is a human practice. The machine can, in the best of designs, enlarge it. It cannot, without ceasing to be what it is, occupy it.

  1. Druga, S., Williams, R., Breazeal, C., & Resnick, M. (2017). “Hey Google is it OK if I eat you?”: Initial explorations in child-agent interaction. In Proceedings of the 2017 Conference on Interaction Design and Children (pp. 595–600). ACM. https://doi.org/10.1145/3078072.3084330
  2. Kasneci, E., Sessler, K., Küchemann, S., Bannert, M., Dementieva, D., Fischer, F., Gasser, U., Groh, G., Günnemann, S., Hüllermeier, E., Krusche, S., Kutyniok, G., Michaeli, T., Nerdel, C., Pfeffer, J., Poquet, O., Sailer, M., Schmidt, A., Seidel, T., Stadler, M., Weller, J., Kuhn, J., & Kasneci, G. (2023). ChatGPT for good? On opportunities and challenges of large language models for education. Learning and Individual Differences, 103, 102274. https://doi.org/10.1016/j.lindif.2023.102274
  3. Kewalramani, S., Kidman, G., & Palaiologou, I. (2021). Using Artificial Intelligence (AI)-interfaced robotic toys in early childhood settings: A case for children’s inquiry literacy. European Early Childhood Education Research Journal, 29(5), 652–668. https://doi.org/10.1080/1350293X.2021.1968458
  4. Kewalramani, S., Palaiologou, I., Dardanou, M., Allen, K.-A., & Phillipson, S. (2021). Using robotic toys in early childhood education to support children’s social and emotional competencies. Australasian Journal of Early Childhood, 46(4), 355–369. https://doi.org/10.1177/18369391211056668
  5. Miao, F., & Cukurova, M. (2024). AI competency framework for teachers. UNESCO. https://doi.org/10.54675/ZJTE2084
  6. Miao, F., & Holmes, W. (2023). Guidance for generative AI in education and research. UNESCO. https://doi.org/10.54675/EWZM9535
  7. Miao, F., Shiohira, K., & Lao, N. (2024). AI competency framework for students. UNESCO. https://doi.org/10.54675/JKJB9835
  8. Su, J., Guo, K., Chen, X., & Chu, S. K. W. (2023). Teaching artificial intelligence in K–12 classrooms: A scoping review. Interactive Learning Environments, 32(9), 5207–5226. https://doi.org/10.1080/10494820.2023.2212706
  9. Su, J., & Yang, W. (2022). Artificial intelligence in early childhood education: A scoping review. Computers and Education: Artificial Intelligence, 3, 100049. https://doi.org/10.1016/j.caeai.2022.100049
  10. Williams, R., Park, H. W., Oh, L., & Breazeal, C. (2019). PopBots: Designing an artificial intelligence curriculum for early childhood education. Proceedings of the AAAI Conference on Artificial Intelligence, 33(01), 9729–9736. https://doi.org/10.1609/aaai.v33i01.33019729
  11. Yan, L., Greiff, S., Teuber, Z., & Gašević, D. (2024). Promises and challenges of generative artificial intelligence for human learning. Nature Human Behaviour, 8(10), 1839–1850. https://doi.org/10.1038/s41562-024-02004-5
  12. Yang, W. (2022). Artificial intelligence education for young children: Why, what, and how in curriculum design and implementation. Computers and Education: Artificial Intelligence, 3, 100061. https://doi.org/10.1016/j.caeai.2022.100061