1. Introduction and problem

In the 3-to-6-year span —kindergarten, preschool, CENDI, infant school— a package of artifacts has been installed that claim to count as support for theory of mind (ToM) and perspective-taking: a detector/CNN/multimodal system that scores ToM, perspective-taking score, false-belief accuracy, mindreading index, mentalizing level or ToM developmental age; a GenAI of ToM worksheets, perspective scripts, false-belief lesson packs or empathy-ToM cards by prompt; a dashboard of ToM trials/hour or perspective compliance; a VR/app/robot as a closed treatment that «trains ToM»; and a ranking of ToM developmental level by gaze, proxemics, emotion or language without mediation. The artifacts make it possible to display that the center «already does ToM with AI». The leap —from score, script, count, VR-treatment or ranking to claiming support— is not authorized by the mappings of AI in ECE nor by pedagogy when there are real mismatches of seeing/knowing and of belief, scaffolding that names mental states, mentalistic language, role play that requires putting oneself in the other's place, and pedagogical interpretation, not a «ToM-score».

The thesis is restrictive: those artifacts do not constitute support for the development of theory of mind / perspective-taking in early childhood. In early childhood education that practice develops with real mismatches of perspective, adult co-presence that interprets the seeing/knowing and the belief of the other, mentalistic language and children's agency to attribute a mental state distinct from their own (OECD, 2021); not a mind detector, a GenAI of worksheets, a dashboard of trials, a closed VR-treatment or a ranking of ToM developmental level. AI may support the adult's preparation in a subordinate way; it must not substitute the relationship nor convert the child into a ToM-score vector. Wang, Wang, Qin, Wu y Wu (2026) contribute Social Story on cognitive and affective perspective-taking at 4–5 years: mediated practice, not a score nor a GenAI recipe. Liang, Sun, Xu, Lee y Siu (2026) situate visual perspective-taking (VPT): a seeing/knowing base, not a dashboard. Mulvihill, Armstrong, Casey, Redshaw, Scarinci y Slaughter (2023) examine educators' mentalistic language: the craft of MSL, not a mindreading index. Heise y Bowman (2025) propose the CAT for 3–8 years: a measure, not pedagogy. Süngü y Alıcı (2025), Witt, Seehagen y Zmyj (2022) and Tuglaci e Ilgaz (2025) show that false-belief performance depends on the object, on membership and on fantastical elements: a task finding, not an authorization of false-belief accuracy as assessment. Bhavna, Akhter, Banerjee y Roy (2024) —transfer: 3–12 years and adults; fMRI— decode states and predict false belief: a detector artifact, not a 3–6 classroom. That authorizes asking what was measured: ToM-score or developmental ranking —not the practice when a child attributes a belief that he or she does not share and an adult names the mismatch.

The problem is aggravated by seven category confusions. First: it is not prosocial helping/sharing/consoling as the axis —Shi, Zhang y Zhu (2024) link ToM and sharing; Rubio, Neira, Villacura-Herrera y Castillo (2022) predict helping and cooperation via ToM; the axis is seeing/knowing, desire/belief and simple false belief, not helping/sharing/consoling—. Second: it is not conflict negotiation / turn-taking. Third: it is not generic SEL —there may be emotion IN affective PT (Wang et al., 2026; Macheta, Gut y Pons, 2023); the axis is not a SEC auto-score—. Fourth: it is not participation/voice nor CLASS as the axis. Fifth: it is not joint attention. Sixth: it is not executive functions as the axis —Shahaeian, Haynes y Frick (2023) are read only as a peripheral language–ToM–FE contrast—. Seventh: it is not clinical emotion monitoring nor inclusion —Ari-Arat y Tutkun (2026) associate PT and intentions toward peers with NEE; this article is not rewritten as inclusion—. Mark-making, dialogic reading, literacy, orality, sand/water/playdough/loose parts/outdoor, well-being, formative assessment, documentation, STEM, motor skills, creativity, free play and symbolic play are not recycled as the axis: Wolf (2022) is cited for ToM (inconsistent representations and external support), not as a pedagogy of symbolic play. There is an economy of coordination: the artifacts fit on a slide; the episode of a child who attributes a simple false belief and an adult who names the mismatch does not. OECD (2021) requires meaningful interactions, with digitization subordinated. Contributions: separating craft and artifacts; examining marked cases; and offering four tests of support.

2. State of the art: from the situated practice of perspective-taking to the artifact that is displayed

It is useful to separate four strata that the market of «AI for ToM / perspective-taking in early childhood education» usually mixes. The first is the construct from 3 to 6 years as situated relational practice: attributing to another a seeing, a knowing, a desire or a belief that may not coincide with one's own; simple false belief; mentalistic language; role play that requires putting oneself in the other's place with mediation (OECD, 2021; Wang et al., 2026; Liang et al., 2026; Heise y Bowman, 2025). The second is the craft that cultivates it —an adult who scaffolds the mismatch, names mental states and sustains that the child attribute without reducing to «correct/incorrect»; co-presence that interprets seeing/knowing; narratives that make visible that the other did not see the same thing— (Mulvihill et al., 2023; Wang et al., 2026; Wolf, 2022). The third is the evidence of AI affordances in ECE, computational ToM and FB detectors —without equating them to situated pedagogy— (Chen, 2024; Ljungcrantz, 2026; Nikolopoulou, 2025; Matta, 2026; Bhavna et al., 2024; Tison y Zawidzki, 2025). The fourth is the framework of practice and of GenAI, which treats the young child as a subject who attributes perspectives with agency, not as a ToM-score vector nor as the exclusive addressee of empathy-ToM cards or developmental rankings (OECD, 2021; Miao y Holmes, 2023).

OECD (2021) anchors meaningful interactions: framework. The verified systems measure objects distinct from the craft —Social Story of PT (Wang et al., 2026), VPT (Liang et al., 2026), educators' MSL (Mulvihill et al., 2023), CAT as a measure (Heise y Bowman, 2025), FB sensitive to object, membership and fantasy (Süngü y Alıcı, 2025; Witt et al., 2022; Tuglaci e Ilgaz, 2025)—. Inference: a mindreading index does not name «he did not see it»; an adult who scaffolds the perspective mismatch does. Chen (2024) and Ljungcrantz (2026) map AI in ECE without equating it to situated ToM. Nikolopoulou (2025) and Miao y Holmes (2023) set GenAI caution. Matta (2026) reviews AI and ToM as an architecture of functional attribution: artificial mindreading, not kindergarten 3–6. Bhavna et al. (2024) saturate the ceiling of the detector. Tison y Zawidzki (2025) mark the normative difference between human and artificial social cognition.

3. Review method

A critical narrative review was conducted, not a primary meta-analysis. The purpose was not to estimate a homogeneous effect size of the artifacts, but to articulate an argument of pedagogical category with verified sources. Inclusion criteria: (a) 2021–2026, with marked transfer when the sample does not equal 3–6 in a classroom of perspective-taking, includes 7–8 years or 3–12 and adults, is fMRI/deep learning of brain states, is computational ToM or HRI/LLM of adult/general social cognition, or the object is a psychometric measure without equating to pedagogy; (b) ToM, perspective-taking, false belief, VPT, educators' mentalistic language, Social Story of PT, pretend play in its link with ToM, cooperation predicted by ToM, emotion–PT, or AI in ECE / computational ToM / FB detector with artifact/craft relevance; (c) kindergarten, preschool, CENDI or 3–6, or explicit transfer; (d) peer-reviewed journal, DOI or OECD/UNESCO GenAI report; (e) verifiable DOI or publisher page. Excluded as central object were axes already used in this series —prosocial helping/sharing/consoling, conflict negotiation/turn-taking, SEL, participation/voice, CLASS, joint attention, inclusion, well-being, formative assessment, documentation, mark-making, dialogic reading, literacy, orality, sand/water/playdough/loose parts/outdoor, free play, symbolic play, STEM, motor skills, creativity and inquiry, and executive functions as axis—. Shahaeian et al. (2023) only as a peripheral contrast. Shi et al. (2024) and Rubio et al. (2022) distinguish ToM from helping/sharing; they do not recycle the axis of 8 September. Wolf (2022) is read for ToM, not as symbolic play. Heise y Bowman (2025) = measure ≠ pedagogy; 3–8, partial transfer. Bhavna et al. (2024) = 3–12 and adults, fMRI: transfer. Matta (2026) and Tison y Zawidzki (2025) = artificial ToM, adult/general transfer.

The search was executed on 9 September 2026 (slot 09:02 America/Mexico_City, weekday routine) on DOI pages, Crossref, Springer, Elsevier, Wiley, Taylor & Francis, MDPI, Frontiers, JAIR, OECD iLibrary, UNESDOC and publisher sites. Each source was verified against fuentes.md and crossref_check.json. Empirical finding, framework and pedagogical inference were distinguished. No N, d, r, AUC or DOI was invented: when an artifact lacks a verified trial as a display package in CENDI 3–6, it is discussed as a category ceiling (Chen, 2024; Ljungcrantz, 2026; Nikolopoulou, 2025; Miao y Holmes, 2023; Matta, 2026; Bhavna et al., 2024; Tison y Zawidzki, 2025). Twenty-one sources from the verified corpus were used.

4. Case 1. A CV/multimodal detector of ToM / perspective-taking score / false-belief accuracy / mindreading index / mentalizing level / ToM developmental age, a GenAI of ToM worksheets/perspective scripts/false-belief lesson packs/empathy-ToM cards, a dashboard of ToM trials/hour, a VR/app/robot as a closed treatment that «trains ToM», or gaze/proxemics/emotion/language→ToM developmental level ranking do not constitute support for theory of mind / perspective-taking

Chen (2024) and Ljungcrantz (2026) saturate the ceiling of artifacts when AI in ECE is presented as if it were support for ToM / perspective-taking. Chen (2024) maps affordances —tutoring, analytics, content generation, agents—, not that a perspective-taking score cultivates the situated attribution of seeing/knowing. Ljungcrantz (2026) reviews AI–ECE interaction 2020–2024: state of the art, not situated perspective-taking. Inference: «high false-belief accuracy = there was support» is a ceiling of analytics. A mindreading index can coexist with an absence of scaffolding. The grammar of display inverts the sequence: first the detector is activated; then it is declared that «there is already ToM with AI».

Nikolopoulou (2025) and Miao y Holmes (2023) name the risk of a GenAI of ToM worksheets / perspective scripts / false-belief lesson packs / empathy-ToM cards by prompt that closes the episode: a framework of caution, not a trial versus an adult who names «he did not see it; where will he look?» (OECD, 2021; Wang et al., 2026). Inference: producing a worksheet by prompt may be subordinate preparation; support begins with a real mismatch of perspective and an adult who names without reducing to correctness. When the script speaks for the adult and the score speaks for the child, the threshold of human mediation is not met even if the dashboard looks «green». OECD (2021) requires meaningful interactions: a trial/hour is not an interaction.

Bhavna et al. (2024) saturate the anchor of the detector and must be read with marked transfer. They propose a framework of explainable deep learning to decode brain states (ToM and pain) and predict false-belief performance; a dataset on the order of 155 participants (122 children aged 3–12 years and 33 adults). Status: neuroinformatics finding (pass/fail/inconsistent in FB). Transfer: 3–12 plus adults; fMRI, not a CENDI 3–6 classroom; task prediction ≠ pedagogical assessment. Inference: a system that «detects ToM» or predicts false-belief accuracy does not equal scaffolding of seeing/knowing nor authorize CNN/vision/multimodal of gaze, emotion or language as ToM developmental age. Matta (2026) —scoping of 45 sources; neuro-symbolic architectures of functional attribution— is a framework of artificial mindreading, not 3–6 pedagogy. Tison y Zawidzki (2025) argue that human social cognition involves normative attitudes that LLMs do not implement: a normative difference, not a spec sheet of a robot that «trains mindreading». Restrictive inference: detector ≠ adult who names the mismatch; GenAI of empathy-ToM cards ≠ open episode; dashboard of trials/hour ≠ observation; VR/app/robot as a closed treatment ≠ co-presence; ranking by gaze/proxemics/emotion/language ≠ mediation. The artifacts lack, in the verified corpus, pedagogical-equivalence trials with the craft in CENDI 3–6; no d, r or AUC of «support» is invented. They share a grammar of substitution: the score speaks for the attribution; the GenAI closes; the dashboard substitutes observation; the VR/robot is sold as treatment; the ranking converts process into a level without mediation.

5. Case 2. What the kindergarten does do when there is support for perspective-taking: situated relational craft and what Wang Social Story, Liang VPT, Mulvihill MSL, Heise CAT, Süngü/Witt/Tuglaci false belief, Wolf pretend play, Rubio, Ari-Arat and Macheta measure —with Matta, Bhavna and Tison in transfer

The floor of the craft is not a score: it is real mismatches of seeing/knowing and of belief; adult scaffolding that names mental states (think, know, see, want, believe) and sustains that the child attribute a perspective distinct from their own; co-presence that interprets the episode —not only «got the trial right»—; and, when there is role play, the requirement of putting oneself in the other's place with mediation, not a correctness script (OECD, 2021; Wang et al., 2026; Mulvihill et al., 2023; Wolf, 2022). Inference: there is support when that practice is protected; not when a perspective-taking score, a GenAI of worksheets or a ranking of ToM developmental level is displayed. The craft is recognized in the episode: a peer did not see the object's displacement; the adult names («she was not there; where will she look?») and sustains that the child attribute; the robot or the VR, if they appear, are subordinate means, not substitutes.

Wang et al. (2026) saturate the anchor of mediated practice: a 12-week quasi-experiment with 60 Chinese children aged 4–5 years; Social Story versus business-as-usual; CPT and APT; SPT at pretest, posttest and two-month follow-up. Empirical finding: the training group achieved larger gains in total SPT, CPT and APT, maintained at follow-up. Marked inference: Social Story with an adult ≠ GenAI of perspective scripts as a closed recipe; gains in SPT ≠ pedagogical ToM-score nor VR/app as a treatment that substitutes mediation. Liang et al. (2026) —254 preschoolers aged 3–5 years (130 girls); photographer task— saturate the VPT anchor: development associated with stimulus features, age and SES. Inference: VPT is a seeing/knowing base («from there it cannot be seen»); not a VPT-accuracy dashboard.

Mulvihill et al. (2023) —13 educators; 77 children aged 3–5 years; MSL in a wordless picture book and in a construction task— find no significant relation between group MSL and children's ToM. Status: empirical finding, including the null. Restrictive inference: mentalistic language is material of the craft —naming seeing, wanting, knowing, believing in the episode—, not a count that «produces ToM» nor a surveillance mentalizing level. The craft requires co-presence that interprets the concrete mismatch, not only group instruction. Heise y Bowman (2025) —n = 206; 3–8 years; partial transfer because it includes 7–8— present the CAT (diverse desires, diverse beliefs, knowledge access, expertise, false belief, VPT and false sign; prediction, explanation and comprehension). Finding: a robust measure that replicates and nuances the Wellman y Liu scaling. Contrast inference: CAT is a research instrument, not a dashboard of ToM trials/hour nor an assessment that substitutes the adult.

Süngü y Alıcı (2025) —150 children aged 3–6 years; standard FB and three alternatives that manipulate presence and location of the object— find a predominance of reality reasoning and more successes when the object is removed. Inference: failing an FB trial does not equal «absence of ToM»; a false-belief accuracy misreads task strategies. Witt et al. (2022) examine at age 4 whether membership (accent; gender) influences false-belief attribution: they find no consistent influence. Inference: attribution is not a stable mindreading index; a uniform ranking erases situatedness. Tuglaci e Ilgaz (2025) —fantastical versus realistic FB, controlling FE— find better performance on fantastical FB; 3-year-olds benefit when fantastical FB precedes realistic FB. Inference: context changes performance; a ToM-score of realistic trials does not capture situated perspective-taking; it is not converted into an FE axis.

Wolf (2022) argues that early pretend play does not by itself constitute evidence of ToM —it does not require distinguishing one's own perspective from the other's—, but it does constitute evidence of handling inconsistent representations with external support; that capacity underlies belief attribution. Marked distinction: it is cited for ToM, not as a symbolic-play axis. Inference: role play that requires putting oneself in the other's place needs mediation; the external support is the craft, not a treatment-robot. Rubio et al. (2022) —40 children aged 3–7 years (M = 5,075; partial transfer because it includes 7)— propose that first-order ToM predicts basic helping and second-order ToM predicts more complex cooperation. Distinction inference: ToM may predict cooperation; that does not recycle the prosocial axis of 8 September nor convert cooperation into perspective compliance. Ari-Arat y Tutkun (2026) —193 children aged 4–6 years; Bayburt, Türkiye— find small associations between perceived capacities and intentions toward peers with NEE; perspective-taking was not independently associated with the intentions. Inference: PT is not allowed to be reduced to an inclusion-score; this article is not converted into an inclusion axis. Macheta et al. (2023) —99 Polish children aged 3–6 years; TEC and three ToM tasks— find that only the opacity task predicted emotion comprehension. Inference: accessing an object under one description does not guarantee accessing it under all; emotion comprehension ≠ clinical emotion monitoring nor an empathy-ToM card.

Matta (2026), Bhavna et al. (2024) and Tison y Zawidzki (2025) are read in transfer: computational ToM, an FB detector and a normative difference do not equal 3–6 pedagogy. Taken together: there are systems that train PT with Social Story, describe VPT, examine MSL, validate measures, show the sensitivity of FB, discuss pretend play and ToM, predict cooperation, associate PT and intentions, link opacity and emotion, or decode states with DL. None —by itself— signs a real mismatch of seeing/knowing, scaffolded mentalistic language, co-presence that interprets the episode and agency to attribute a distinct belief in the classroom. The technical evidence is real in its domain; «score / GenAI / VR-treatment / detector = support for ToM» is an illegitimate category inference.

6. Case 3. ToM / perspective-taking ≠ prosocial helping/sharing/consoling, conflict negotiation/turn-taking, SEL as axis, participation-voice, CLASS, joint attention, FE as axis or clinical emotion monitoring

Seven borders protect the axis. First: prosocial helping/sharing/consoling —Shi et al. (2024) situate ToM and emotions in sharing; Rubio et al. (2022) predict helping and cooperation via ToM; there may be belief attribution IN a sharing episode; the axis is seeing/knowing, desire/belief and simple false belief, not helping/sharing/consoling, already published on 8 September—. Inference: «we work on sharing = there is ToM» does not sign. Second: conflict negotiation / turn-taking —there may be a dispute IN a perspective episode; the axis is not the turn—. Third: generic SEL —Wang et al. (2026) train APT; Macheta et al. (2023) link opacity and emotion; the axis is not a SEC auto-score nor an empathy-ToM card—. Fourth: child participation/voice —there is mentalistic speech; this article is not converted into voice already published—. Fifth: CLASS / interactions as axis —OECD (2021) anchors meaningful interactions; it does not authorize reducing ToM to a CLASS score—. Sixth: joint attention —naming «he did not see it» is not joint attention as axis—. Seventh: FE as axis and clinical emotion monitoring —Shahaeian et al. (2023) —N = 142; 3–6; three waves— find that early FE influenced ToM, that the effect became bidirectional and that language influenced both; Tuglaci e Ilgaz (2025) control FE; they are read only as a peripheral contrast. Ari-Arat y Tutkun (2026) prevent recycling inclusion. Miao y Holmes (2023) and Nikolopoulou (2025) require GenAI mediation. Inference: coherent AI at 3–6 remains on the adult's side, not as a detector, GenAI of lesson packs, dashboard, VR-treatment, mindreading robot or ranking without mediation.

These borders protect the axis. ToM / perspective-taking may touch, in a lateral way, sharing, cooperation, emotion, language, FE or intentions toward peers with NEE; it is not allowed to be reduced to those already published axes. A center that «works on SEL», «scores helping» or «trains FE» has not demonstrated by that label alone the craft of scaffolding seeing/knowing, desire/belief and simple false belief. Shi et al. (2024) and Rubio et al. (2022) are read for distinction, not as a prosocial axis; Shahaeian et al. (2023) are read as peripheral, not as an FE axis; Wolf (2022) is read for ToM, not as a symbolic-play axis.

7. Inferential framework: four tests for claiming that there is support for theory of mind / perspective-taking, not an artifact

The framework that follows is a pedagogical inference of this article, anchored in the cases and in the verified instruments. It is not a new international standard. It distinguishes four tests. If a kindergarten, preschool, CENDI or infant school does not pass them, it cannot declare that the artifacts constitute support for the development of theory of mind / perspective-taking in early childhood.

7.1. Test of situated practice (real mismatch of seeing/knowing and of belief; simple false belief; desire/belief; agency to attribute a distinct perspective), not of the CV/multimodal detector nor of the VR/app/robot of «training ToM». OECD (2021) and the practice and task cases (Wang et al., 2026; Liang et al., 2026; Süngü y Alıcı, 2025; Witt et al., 2022; Tuglaci e Ilgaz, 2025; Wolf, 2022) define the craft/artifact contrast. Chen (2024), Ljungcrantz (2026), Matta (2026), Bhavna et al. (2024, transfer) and Tison y Zawidzki (2025) map analytics, computational ToM and detector as affordance or architecture, not as situated perspective-taking. Inference: if the «evidence» is ToM-score, false-belief accuracy, mindreading index, mentalizing level or ToM developmental age, the center did analytics or decoding, not support.

7.2. Test of adult co-presence that names mental states and scaffolds the mismatch —not only «got the trial right»—, not of the GenAI of ToM worksheets / perspective scripts / false-belief lesson packs / empathy-ToM cards. OECD (2021), Wang et al. (2026), Mulvihill et al. (2023) and Wolf (2022, external support) situate practice and mediation. Miao y Holmes (2023) and Nikolopoulou (2025) require GenAI mediation. Inference: a script by prompt or a closed lesson pack does not sign the classroom episode in which the adult names «she did not see it» and the child attributes.

7.3. Test of human mediation and of the interpretation of ToM as relational practice, not of the dashboard of ToM trials/hour / perspective compliance nor of ranking by gaze/proxemics/facial emotion/language. OECD (2021) requires meaningful interactions. Heise y Bowman (2025) anchor measure ≠ pedagogy. Süngü y Alıcı (2025), Witt et al. (2022) and Tuglaci e Ilgaz (2025) anchor the sensitivity of FB to the task, not surveillance of counts. Inference: high trials/hour do not by themselves raise scaffolding. Subordinate preparation of the adult is legitimate; the scorer that converts the child into a ToM-score vector is not.

7.4. Test of category distinction and of professional judgment, not of the product catalogue. ToM / perspective-taking ≠ prosocial helping/sharing/consoling, conflict negotiation/turn-taking, SEL, participation-voice, CLASS, joint attention, FE as axis or clinical emotion monitoring. Shi et al. (2024), Rubio et al. (2022), Shahaeian et al. (2023, peripheral), Macheta et al. (2023) and Ari-Arat y Tutkun (2026) prevent those confusions. Miao y Holmes (2023), Nikolopoulou (2025) and OECD (2021) require human mediation. Inference: support is fulfilled with a real mismatch of perspective, co-presence that names/scaffolds and agency to attribute —not by scoring ToM nor by substituting mediation with a robot.

The framework admits digital subordinate to the adult's preparation (Wang et al., 2026, read as mediated practice, not as a GenAI recipe; Heise y Bowman, 2025, as a measure for the researcher, not as a dashboard). It rejects declaring support by the artifacts (Miao y Holmes, 2023; Nikolopoulou, 2025; Bhavna et al., 2024; Matta, 2026). The four tests are read together.

8. Discussion

Three tensions. First: displaying the artifacts versus exercising support. The cases of sections 5–6 —with marked transfers and distinctions— sustain the contrast; Chen (2024) and Ljungcrantz (2026) map AI without equivalence to situated ToM. Second: automated assessment or VR-treatment versus relational pedagogy —declaring support by false-belief accuracy, mindreading index or by a detector of brain states is inverted pedagogy; Heise y Bowman (2025) show that even a comprehensive measure is not the craft—. Third: GenAI of empathy-ToM cards / a robot that «trains mindreading» / ranking by gaze versus craft (Miao y Holmes, 2023; Nikolopoulou, 2025; OECD, 2021; Tison y Zawidzki, 2025): selling computational ToM or FB prediction as «ToM with AI» in the kindergarten confuses product with practice. Mulvihill et al. (2023) confirm that group MSL is not allowed to be translated into a ToM-score; Süngü y Alıcı (2025) confirm that an FB trial misreads strategies; Miao y Holmes (2023) subordinate generation to professional judgment.

The four tests of section 7 read these tensions. The empirical contrast defines the floor that the artifacts do not reach on their own. Inference: support is eroded when they are treated as if they were the practice. Subordinate digital is legitimate toward the adult (preparation of mentalistic language; CAT for the researcher, not for the dashboard; Wang et al., 2026, as mediated practice); inverting the sequence is not (OECD, 2021). If the center shows real mismatches of seeing/knowing and of belief, an adult who named mental states, and children who attributed a distinct perspective with agency, it may speak of support; if it only shows detectors, GenAI scripts, trials/hour, VR-treatments, mindreading robots or rankings without mediation, it speaks of artifacts. Chen (2024) and Ljungcrantz (2026): AI affordances in ECE ≠ situated ToM.

9. Limits

This review is narrative. It does not apply its own PRISMA nor estimate combined effects. No d, r or AUC is invented. Wang et al. (2026) are a quasi-experiment (60; 4–5): mediated practice, not a GenAI recipe. Liang et al. (2026) measure VPT (254; 3–5): a visual construct. Mulvihill et al. (2023) measure 13 educators and 77 children: a null with caution. Heise y Bowman (2025) measure n = 206 at 3–8: partial transfer; CAT is a measure, not a classroom. Süngü y Alıcı (2025) measure 150 children in laboratory FB. Witt et al. (2022) delimit 4 years and membership. Tuglaci e Ilgaz (2025) manipulate fantasy in FB: they are not converted into an FE axis. Wolf (2022) is a philosophical argument. Rubio et al. (2022) measure 40 children aged 3–7: partial transfer; distinction, not a prosocial axis. Ari-Arat y Tutkun (2026) use convenience (193; 4–6). Macheta et al. (2023) measure N = 99. Shahaeian et al. (2023) are longitudinal (N = 142): peripheral of FE. Shi et al. (2024) are used only to distinguish sharing. Bhavna et al. (2024) are fMRI 3–12 plus adults: transfer. Matta (2026) and Tison y Zawidzki (2025) are AI–ToM frameworks, not ECE 3–6. Chen (2024), Ljungcrantz (2026) and Nikolopoulou (2025) map AI in ECE, not the display package in a Latin American CENDI. No verified trials of the artifacts as a single package of «ToM with AI» in a 3–6 classroom were located; they are discussed as a category ceiling. OECD and Miao y Holmes are framework. The inferences of section 7 are category hypotheses, not implementation evidence.

10. Conclusions

The artifacts —a CV/multimodal detector or scores of ToM / perspective-taking / false-belief accuracy / mindreading index / mentalizing level / ToM developmental age; a GenAI of worksheets/scripts/lesson packs/empathy-ToM cards; a dashboard of trials/hour; a VR/app/robot as a closed treatment; a ranking without mediation— do not constitute support for the development of theory of mind / perspective-taking in early childhood education. Chen (2024) and Ljungcrantz (2026) confirm affordances without equivalence to situated ToM. Nikolopoulou (2025) and Miao y Holmes (2023) set GenAI limits. The empirical cases (sections 5–6) require reading Social Story, VPT, MSL, CAT, FB, pretend play, cooperation, emotion–PT, language–ToM–FE and sharing by what they measure and by what they do not equal as situated pedagogy. Matta (2026), Bhavna et al. (2024) and Tison y Zawidzki (2025) are read as an artifact ceiling, not as 3–6 craft. When there is support there are real mismatches of seeing/knowing and of belief, co-presence that names mental states and agency to attribute a distinct perspective (OECD, 2021). ToM / perspective-taking is distinguished from helping/sharing/consoling, conflict negotiation/turn-taking, SEL, participation-voice, CLASS, joint attention, FE as axis and clinical emotion monitoring.

Where the sources do not measure a kindergarten, this article does not claim it. Where they measure mappings of AI, Social Story of PT, VPT, educators' MSL, CAT, FB sensitive to object/membership/fantasy, pretend play, cooperation, emotion–PT, language–ToM–FE, sharing, computational ToM or FB prediction with transfer, it does not translate them into support via ToM-score, worksheet by prompt, closed VR-treatment or substitute-robot. Accompanying girls and boys from three to six years in theory of mind / perspective-taking is to exercise real mismatches of seeing/knowing and of belief, simple false belief, desire/belief, mentalistic language, role play that requires putting oneself in the other's place with mediation, scaffolding that names mental states, and co-presence that interprets ToM as situated relational practice, not as a «ToM-score». The rest is a detector, a GenAI of worksheets, a dashboard of trials, a closed VR-treatment, a mindreading robot and a ranking. It is not support for theory of mind / perspective-taking in early childhood education, and it must not be presented as what it is not.

Laboratorio Editorial de NEXTECH.IA / Ingeniero Mitre.

References

  1. Ari-Arat, C., y Tutkun, C. (2026). Associations between preschool children’s perspective-taking skills and their perceived abilities and behavioral intentions toward peers with special educational needs. Current Psychology, 45(17). https://doi.org/10.1007/s12144-026-09997-4
  2. Bhavna, K., Akhter, A., Banerjee, R., y Roy, D. (2024). Explainable deep-learning framework: decoding brain states and prediction of individual performance in false-belief task at early childhood stage. Frontiers in Neuroinformatics, 18. https://doi.org/10.3389/fninf.2024.1392661
  3. Chen, J. J. (2024). A scoping study on AI affordances in early childhood education: Mapping the global landscape, identifying research gaps, and charting future research directions. Journal of Artificial Intelligence Research, 81, 701–740. https://doi.org/10.1613/jair.1.16882
  4. Heise, M. J., y Bowman, L. C. (2025). The Comprehensive Assessment of Theory of Mind (CAT): A novel measure of 3- to 8-year-old children’s theory of mind and an evaluation of mental-state scaling. Child Development, 96(5), 1787–1806. https://doi.org/10.1111/cdev.14263
  5. Liang, J., Sun, J., Xu, X., Lee, K., y Siu, C. T. S. (2026). Early development of visual perspective-taking: Its associations with stimuli feature, child age, and family socio-economic status. Early Childhood Education Journal, 54(2), 757–772. https://doi.org/10.1007/s10643-025-01866-2
  6. Ljungcrantz, L. (2026). The interaction of AI and early childhood education. A state-of-the-art review 2020–2024. Early Childhood Education Journal, 54(5), 3565–3581. https://doi.org/10.1007/s10643-025-02079-3
  7. Macheta, K., Gut, A., y Pons, F. (2023). The link between emotion comprehension and cognitive perspective taking in theory of mind (ToM): a study of preschool children. Frontiers in Psychology, 14. https://doi.org/10.3389/fpsyg.2023.1150959
  8. Matta, D. (2026). Artificial intelligence and theory of mind. Journal of Psychology and AI, 2(1). https://doi.org/10.1080/29974100.2026.2628373
  9. Miao, F., y Holmes, W. (2023). Guidance for generative AI in education and research. UNESCO. https://doi.org/10.54675/ewzm9535
  10. Mulvihill, A., Armstrong, R., Casey, C., Redshaw, J., Scarinci, N., y Slaughter, V. (2023). Early childhood educators' mental state language and children's theory of mind in the preschool setting. British Journal of Developmental Psychology, 41(3), 227–245. https://doi.org/10.1111/bjdp.12449
  11. Nikolopoulou, K. (2025). Child-centered integration of generative AI in early learning: Balancing promises and challenges. AI, Brain and Child, 1(1). https://doi.org/10.1007/s44436-025-00023-1
  12. OECD. (2021). Starting Strong VI: Supporting meaningful interactions in early childhood education and care. OECD Publishing. https://doi.org/10.1787/f47a06ae-en
  13. Rubio, F., Neira, C., Villacura-Herrera, C., y Castillo, R. D. (2022). First and second-order theory of mind as predictors of cooperative behaviors in preschool and school children. Psykhe, 31(SI 1). https://doi.org/10.7764/psykhe.2021.36317
  14. Shahaeian, A., Haynes, M., y Frick, P. J. (2023). The role of language in the association between theory of mind and executive functioning in early childhood: New longitudinal evidence. Early Childhood Research Quarterly, 62, 251–258. https://doi.org/10.1016/j.ecresq.2022.09.003
  15. Shi, Y., Zhang, M., y Zhu, L. (2024). Sharing and allocation in preschool children: The roles of theory of mind, anticipated emotions, and consequential emotions. Behavioral Sciences, 14(10), Article 931. https://doi.org/10.3390/bs14100931
  16. Süngü, M., y Alıcı, T. (2025). Discrimination of false response from object reality in false belief test in preschool children. Journal of Intelligence, 13(10), Article 124. https://doi.org/10.3390/jintelligence13100124
  17. Tison, R., y Zawidzki, T. (2025). Human versus artificial social cognition and metacognition: the normative difference. AI and Ethics, 5(5), 5255–5271. https://doi.org/10.1007/s43681-025-00774-w
  18. Tuglaci, E., y Ilgaz, H. (2025). The effect of fantastical elements on preschoolers’ false belief task performance. Journal of Experimental Child Psychology, 260, Article 106321. https://doi.org/10.1016/j.jecp.2025.106321
  19. Wang, X., Wang, J., Qin, L., Wu, Y., y Wu, J. (2026). Promoting cognitive and affective perspective-taking in 4-5-year-olds through social story training: A quasi-experimental study. Early Childhood Education Journal, 54(5), 3627–3638. https://doi.org/10.1007/s10643-025-02093-5
  20. Witt, S., Seehagen, S., y Zmyj, N. (2022). The influence of group membership on false-belief attribution in preschool children. Journal of Experimental Child Psychology, 222, Article 105467. https://doi.org/10.1016/j.jecp.2022.105467
  21. Wolf, J. (2022). Implications of pretend play for Theory of Mind research. Synthese, 200(6), Article 523. https://doi.org/10.1007/s11229-022-03984-5