1. Introduction and problem

Between ages 0 and 6—kindergarten, daycare, CENDI, early childhood school—a package of three artefacts has been installed as if it counted as pedagogical joint attention. The first is an eye-tracking system that marks “correct/incorrect” if the child follows the gaze: a pipeline that turns gaze matching into a binary verdict and delivers the dashboard as if it were response to joint attention (RJA) or initiation of joint attention (IJA) in the classroom. The second is a social robot that runs RJA/IJA prompts: a physical or virtual agent that directs gaze, points or requests following and presents the trial as if it were triadic adult–child–referent practice. The third is a computer-vision model that scores “joint visual attention” from video: a classifier that labels shared-gaze episodes and declares the score as mediation of shared focus. All three are visible and cheap in coordination time. They allow centres to exhibit that they “already do joint attention with artificial intelligence.” The leap—from accumulating correct gaze matching, completing robotic trials or obtaining a joint-visual-attention score to claiming that pedagogical joint attention exists—is authorised neither by evidence on social robots or eye-tracking correlates of joint attention nor by what early childhood pedagogy measures when there is situated triadic practice, serve-and-return and mediation of shared focus in the classroom.

The thesis of this article is restrictive. An eye-tracking system that marks “correct/incorrect” if the child follows the gaze, a social robot that runs RJA/IJA prompts, or a computer-vision model that scores “joint visual attention” from video do not constitute pedagogical joint attention in early childhood. At this stage, joint attention is situated triadic practice (adult–child–referent), serve-and-return and mediation of shared focus; not a gaze-matching score nor a robotic trial decoupled from teaching craft. So et al. (2023) compare robot-based intervention (RBI) versus human-based intervention (HBI) to improve joint attention in autistic children, with RJA/IJA measures (ESCS): a finding of clinical/experimental intervention, not of joint-attention craft in an ordinary early childhood classroom. Sani-Bozkurt and Bozkus-Genc (2023) systematically review social robots for joint-attention development in autism spectrum disorder: a synthesis of the robotic artefact in ASD, not an equivalence with everyday triadic mediation in kindergarten. de Belen et al. (2023) report eye-tracking correlates of response to joint attention in preschoolers with ASD: a finding of gaze as a correlate of RJA, not of pedagogical mediation of shared focus. Ozdemir, Akin-Bulbul and Yildiz (2025) compare visual attention in joint-attention bids between toddlers with ASD and typically developing toddlers: a finding of visual-attention patterns, not of a pedagogical dashboard declaring “joint attention completed.” That does not authorise translating “an eye-tracker is installed,” “there is a robot with RJA/IJA prompts,” or “there is CV that scores joint visual attention” into “pedagogical joint attention exists.” It authorises asking what was measured: often gaze following, RJA/IJA trials in laboratory or clinical rooms, or video labels; not the reiterated practice when an adult and a two-, three- or four-year-old share a referent, respond in serve-and-return and the educator mediates shared focus in classroom routines.

The problem is aggravated by four category confusions. First: joint attention is not generic adult–child “interaction quality” in the CLASS/turn-taking sense—a construct already treated in the series; here the object is the specific joint-attention construct (RJA, IJA, triadic gaze, serve-and-return oriented to a shared referent). Second: it is not SEL (social-emotional learning as a programme of emotions, empathy or self-regulation). Third: it is not inclusion/UDL. Fourth: it is not emergent literacy. This article does not recycle papers on literacy, music, motor skills, artistic creativity, generic adult–child interactions, executive functions, SEL, generic STEM/robotics, free play, documentation, numeracy, orality, formative assessment, UDL, participation, planning, teacher education, family–school, transition, leadership, inclusion, privacy, multilingualism, wellbeing, integrity, connectivity, AI literacy, parental mediation, UNESCO competence or genAI play/scaffolding. The question is one of joint attention: what counts as such when an early childhood centre “does AI and joint attention.”

There is a coordination economy: a log of correct gaze, a video of the robot running RJA/IJA prompts and a dashboard of “joint visual attention”; pedagogical joint attention—the adult who points to an object, waits for the child’s gaze, names the referent, responds to the infant’s serve and sustains shared focus at the table or in the yard—is slower to exhibit. The contributions are three: to reconstruct the state of the art that separates pedagogical joint attention from scoring eye-tracking, the robot with RJA/IJA prompts and labelling CV; to examine three families of empirical cases; and to offer four tests for deciding when a kindergarten may claim that pedagogical joint attention exists.

2. State of the art: from situated triadic joint-attention practice to the artefact on display

It is useful to separate four strata that the market of “AI for joint attention in early childhood” usually mixes. The first is the construct of joint attention at ages 0–6 as triadic adult–child–referent practice: RJA (following the other’s gaze or gesture toward an object), IJA (initiating shared focus by pointing, showing or alternating gaze) and shared gaze situated in routines (Bleiweiss & Sall, 2025; Ozdemir et al., 2025; NAEYC, 2022; OECD, 2021). The second is the pedagogical craft that cultivates it—serve-and-return, mediation of shared focus, intentional distribution of teacher attention and reflection on educator expertise—(Bleiweiss & Sall, 2025; Duthois et al., 2026; Isotalo et al., 2026; NAEYC, 2022; OECD, 2021, 2023). The third is evidence on social robots, eye-tracking and gaze correlates in joint attention, together with mappings of AI in ECE (So et al., 2023; Sani-Bozkurt & Bozkus-Genc, 2023; de Belen et al., 2023; Ozdemir et al., 2025; Mifsud et al., 2025; Chen, 2024; Su & Yang, 2022; Ljungcrantz, 2026; Nikolopoulou, 2025). The fourth is the rights, systems and developmentally appropriate practice framework that treats the 0–6 infant as a subject of situated triadic practice, not as a unit of gaze matching nor as a respondent to a robotic trial decoupled from the adult (UNESCO, 2021; Miao & Holmes, 2023; European Commission, 2022; U.S. Department of Education, 2023).

In the construct and craft stratum, Bleiweiss and Sall (2025) examine interventions to promote serve-and-return and joint attention as foundations of social communication: a finding and craft framework that anchors joint attention in mediated relational practice, not in a gaze classifier. Status: framework and intervention on serve-and-return and JA; pedagogical inference, marked as such: the craft a kindergarten may call pedagogical joint attention includes serve-and-return turns oriented to a shared referent and adult mediation of focus, not mere reproduction of RJA/IJA prompts by a robot. Duthois, Vanderlinde and Van Avermaet (2026) study with mobile eye-tracking the distribution of teachers’ attention in early childhood education: an empirical finding on which children do or do not receive the adult’s focus. Restrictive inference: eye-tracking here illuminates teacher attention; it does not authorise declaring pedagogical joint attention by a child gaze-following score. Isotalo, Ukkonen-Mikkola and Rutanen (2026) explore possibilities of eye-tracking for reflecting teachers’ expertise in ECEC: a framework of reflective artefact use, not an equivalence gaze-score = the child’s joint attention. NAEYC (2022) requires developmentally appropriate practice and sensitive interactions; OECD (2021, 2023) anchor quality in meaningful interactions and subordinated digitalisation. Inference: a robot does not, by itself, sustain the adult–child–object triad in routine. An adult who points, waits, names and responds to the child’s serve does.

In the artefact stratum, So et al. (2023) report RBI versus HBI for joint attention in autistic children with ESCS measures of RJA/IJA: the clinical ceiling of the robot as a trial agent. Sani-Bozkurt and Bozkus-Genc (2023) saturate the map of social robots for JA in ASD. de Belen et al. (2023) correlate eye-tracking and RJA in preschool with ASD. Ozdemir et al. (2025) describe visual attention in JA bids in toddlers. Mifsud, Bonello and Kucirkova (2025)—transfer marked: Maltese literacy classroom ages 5–6, not a pure JA study—show multifaceted roles of the social robot in early literacy lessons: social mediation, not only tutoring. Inference: robot in classroom ≠ automation of RJA/IJA as pedagogical joint attention. Chen (2024), Su and Yang (2022), Ljungcrantz (2026) and Nikolopoulou (2025) map AI in ECE without equating it to mediated triadic gaze.

In the systems stratum, UNESCO (2021) requires human oversight and AI ethics. Miao and Holmes (2023) set pedagogical validation and age thresholds for generative AI—a framework transferable to scoring pipelines and agents. The European Commission (2022) and the U.S. Department of Education (2023) subordinate AI to the educator’s professional judgement. Framework inference: a kindergarten may use tools under supervision; it may not declare pedagogical joint attention merely because eye-tracking, a robot with RJA/IJA prompts or joint-visual-attention CV was deployed.

A joint reading clarifies an asymmetry the market erases. On the craft side, Bleiweiss and Sall (2025), Duthois et al. (2026), Isotalo et al. (2026), NAEYC (2022) and OECD (2021, 2023) describe a slow practice of shared focus. On the artefact side, So et al. (2023), Sani-Bozkurt and Bozkus-Genc (2023), de Belen et al. (2023) and Ozdemir et al. (2025) deliver robotic trial, gaze correlate or visual pattern. Pedagogical inference, marked as such: to instrument is not to equate. Buying instrumentation buys visibility; it does not automatically buy pedagogical joint attention.

3. Review method

A critical narrative review was conducted, not a primary meta-analysis. The purpose was not to estimate a homogeneous effect size of the social robot or of eye-tracking on joint attention, but to articulate an argument of pedagogical category with verified sources. Inclusion criteria: (a) 2021–2026, with transfer explicitly marked when the sample, diagnosis or level does not equal an ordinary early childhood classroom ages 0–6; (b) joint attention, RJA, IJA, serve-and-return, eye-tracking, social robot, shared-gaze computer vision or AI in early childhood education; (c) relevance to kindergarten, daycare, CENDI, preschool or ages 0–6; (d) peer-reviewed journal, DOI or report from NAEYC, UNESCO, OECD, the European Commission or a department of education; (e) verifiable DOI or editorial page. Axes already used in this series were excluded as central objects—emergent literacy, music, motor/FMS, visual artistic creativity, generic CLASS/turn interaction quality, pedagogical documentation, SEL as substitute programme, executive functions, generic STEM/robotics, generic free play, numeracy, orality/narration, academic formative assessment, UDL/inclusion, participation, planning, teacher education, family–school, transition, leadership, privacy, multilingualism, wellbeing, integrity, connectivity, AI literacy, parental mediation, UNESCO competence and genAI play/scaffolding—although some appear as category boundaries.

The search was executed on 28 August 2026 (slot 21:02 America/Mexico_City) on DOI pages, Crossref, Elsevier, Springer, Taylor & Francis, Frontiers, BMC, Sage, JAIR, OECD iLibrary, UNESDOC, NAEYC and editorial sites. Each source was verified against at least one of those pages. Empirical finding, conceptual or normative framework, and pedagogical inference marked as such were distinguished. Priority was given to the distinction joint attention / generic interaction quality / SEL / inclusion-UDL / emergent literacy, and to the caution not to translate gaze correlates, RBI/HBI trials or joint-visual-attention scores into pedagogy of triadic shared gaze and classroom serve-and-return. No N, d, r, AUC or DOI was invented; when evidence comes from ASD, clinical or literacy samples (e.g. So et al., 2023; de Belen et al., 2023; Mifsud et al., 2025), transfer is marked.

4. Case 1. Scoring eye-tracking, a robot with RJA/IJA prompts or CV of joint visual attention do not constitute pedagogical joint attention

So, Cheng, Lam, Wong, Law, Huang, Ng, Chun and Wong (2023) publish in Frontiers in Psychiatry the comparison that best names the ceiling of the social robot when it is presented as pedagogical joint attention. They compare the effectiveness of a robot-based intervention (RBI) versus a human-based intervention (HBI) to improve joint attention in autistic children, with RJA and IJA assessed via ESCS. Status of the evidence. Empirical finding of robot versus human intervention on joint-attention measures in an ASD population; transfer marked: clinical/experimental autism context, not an ordinary early childhood classroom. Pedagogical inference, marked as such: this is the gesture a centre copies when it “does joint attention with a robot that runs RJA/IJA prompts.” The agent is switched on, gazes and points are directed, trials accumulate and achievement is declared. What exists is a technical ceiling of robot-mediated trial. Pedagogical joint attention, in Bleiweiss and Sall (2025), NAEYC (2022) and OECD (2021), requires situated triadic practice with an adult who mediates the referent in routine; not a log of robotic trials. A child may “improve on ESCS RJA/IJA measures” according to the study and, at the same time, never have sustained with their educator a reiterated shared focus on a play object in the everyday classroom.

Sani-Bozkurt and Bozkus-Genc (2023) saturate the artefact portrait in a systematic review of social robots for joint-attention development in ASD (International Journal of Disability, Development and Education). Status: synthesis of the robot–JA field in autism. Inference, marked as such: the existence of a corpus of robots for JA does not authorise declaring that the ordinary kindergarten “already does pedagogical joint attention” because it purchased a robot. de Belen, Pincham, Hodge, Silove, Sowmya, Bednarz and Eapen (2023) name the first artefact—eye-tracking as correlate—in BMC Psychiatry: eye-tracking correlates of response to joint attention in preschoolers with ASD. Status: empirical finding of gaze associated with RJA; ASD/clinical preschool transfer. Restrictive inference: eye-tracking correlate ≠ pedagogical mediation of shared focus. A dashboard that marks “correct/incorrect” if the child followed the gaze inherits that correlate’s grammar and turns it into a pedagogical verdict without craft.

Ozdemir, Akin-Bulbul and Yildiz (2025) reinforce the ceiling of gaze scoring from visual attention in joint-attention bids: comparison between toddlers with ASD and typically developing toddlers (Journal of Autism and Developmental Disorders). Status: empirical finding of visual-attention patterns to JA bids. Inference: describing or scoring visual attention in a bid does not, by itself, constitute classroom triadic practice. Mifsud, Bonello and Kucirkova (2025)—transfer marked: literacy in a Maltese classroom ages 5–6, not pure JA—show that the social robot in early childhood education assumes multifaceted roles of social mediation, not only tutoring. Inference: even when the robot enters the classroom, its presence does not sign pedagogical joint attention; it signs, at most, a social artefact that still requires adult craft. Chen (2024), Su and Yang (2022), Ljungcrantz (2026), Nikolopoulou (2025) and Miao and Holmes (2023) confirm the growth of AI in ECE and thresholds of pedagogical validation without evidence that eye-tracking-scorer, RJA/IJA robot or joint-visual-attention CV replaces triadic mediation.

The three artefacts share the same substitution grammar. Scoring eye-tracking replaces pedagogical observation of shared focus with a gaze-matching verdict. The robot with RJA/IJA prompts replaces the adult–child–referent triad with an agent trial. CV that scores joint visual attention replaces mediation of shared focus with a video label. Pedagogical inference, marked as such: pedagogical joint attention is not declared by the product catalogue; it is declared by reiterated practice of triadic shared gaze, serve-and-return and adult mediation in the classroom. So et al. (2023), Sani-Bozkurt and Bozkus-Genc (2023), de Belen et al. (2023), Ozdemir et al. (2025) and Mifsud et al. (2025) authorise saying that there was a robotic intervention, gaze correlate, visual-attention pattern or social mediation with a classroom robot; not signing situated RJA/IJA or reiterated serve-and-return as everyday kindergarten craft.

The risk of improper transfer should be underlined. Evidence in ASD (So et al., 2023; Sani-Bozkurt & Bozkus-Genc, 2023; de Belen et al., 2023; Ozdemir et al., 2025) is valuable in its clinical and experimental domain; it does not become, by commercial contiguity, a protocol of “JA with AI” for ordinary classrooms ages 0–6. Inference, marked as such: citing a robot that improves ESCS measures in autistic children does not authorise a CENDI to declare that its gaze-matching dashboard “already guarantees joint attention.” Transfer must be named; when it is not named, the artefact colonises the construct.

5. Case 2. What the kindergarten does when pedagogical joint attention exists: serve-and-return, situated RJA/IJA and triadic mediation

Bleiweiss and Sall (2025) situate serve-and-return and joint attention as foundations of social communication in TEACHING Exceptional Children: they examine interventions aimed at promoting those processes. Status: framework and intervention on social-communication craft. Pedagogical inference, marked as such: this is the craft an early childhood centre may call pedagogical joint attention when adult and child respond around a referent, the child responds to bids (RJA) or initiates shared focus (IJA) in routine, and the educator mediates—points, waits, names, expands—without replacing the triad by a robot or a score. Joint attention is not a percentage of gaze matching: it is growing mastery of sharing a focus with another on an object or event, in real time and with meaning. An eye-tracker may label “correct following” and, at the same time, never have captured whether the adult sustained serve-and-return when the child showed a block or pointed at the window. A robot may run an RJA prompt and empty the classroom of human mediation.

Duthois, Vanderlinde and Van Avermaet (2026) saturate the craft from the adult side: with mobile eye-tracking they ask which children fall outside the distribution of teacher attention in ECE. Status: empirical finding on distribution of the educator’s focus. Inference: when pedagogical joint attention exists in early childhood, there is an adult whose attention is distributed—with biases and opportunities for correction—toward shared referents with concrete girls and boys; the artefact serves, if at all, reflection on that distribution, not a declaration that the child “already has JA” because a model scored their gaze. Isotalo, Ukkonen-Mikkola and Rutanen (2026) explore eye-tracking as a possibility for reflecting teachers’ expertise in ECEC. Inference: legitimate use of eye-tracking at this stage leans toward the adult’s professional reflection, not toward autonomous scoring of the child as if joint attention were completed. NAEYC (2022) anchors sensitive interactions and developmentally appropriate practice; OECD (2021, 2023) anchor meaningful everyday interactions and subordinated digitalisation. Inference: triadic mediation means noticing whether there was a shared referent, whether the adult waited for the child’s gaze, whether the object was named and whether time for shared focus was protected against the dashboard’s haste. Pedagogical observation is professional reading of the triadic episode; not a log of “joint visual attention score.”

Situated RJA and IJA do not deny play: they orient it toward the shared referent. Mediation, in the sense of NAEYC (2022) and OECD (2021), is adult presence that organises, models, invites and expands: points to an insect in the yard and waits for the child to look; follows the child’s gaze toward a truck and names; responds with a return when the child shows a drawing; protects the episode against the score’s interruption. Serve-and-return completes the picture: the child “serves” (vocalises, points, shows) and the adult “returns” contingently, keeping focus on the same referent. Pedagogical inference, marked as such: a CV may label gaze overlap; it does not sign serve-and-return as a practice of meaning. Triadic shared gaze is the minimal gesture that distinguishes pedagogical joint attention from “having coincided in looking at a screen” (Bleiweiss & Sall, 2025; Duthois et al., 2026; Isotalo et al., 2026).

Duthois et al. (2026) also contribute an ethical warning internal to the craft: if teacher attention is distributed unequally, some children fall outside possible shared focus. Inference: a kindergarten that “does JA with AI” but does not examine whom the educator looks at reproduces inequity under a technological varnish. Isotalo et al. (2026) suggest that eye-tracking can serve reflection on that expertise; OECD (2023) requires that the digital not replace meaningful interaction. The contrast is sharp: the same type of sensor the market sells to score the child can, in teachers’ hands, illuminate adult practice. Only the second use approaches the framework of UNESCO (2021), the European Commission (2022) and the U.S. Department of Education (2023).

6. Case 3. Joint attention is not generic interaction quality, nor SEL, nor emergent literacy

The first category boundary is generic adult–child interaction quality (CLASS, turns, climate). That construct may coexist in a centre and was already treated in the series; it does not, by itself, constitute joint attention as RJA/IJA and adult–child–referent triadic gaze. Status: pedagogical inference marked as such, anchored in Bleiweiss and Sall (2025) and OECD (2021) as practice of shared focus and serve-and-return oriented to the referent. This article does not make generic interactional quality its central object: it names it to forbid the equivalence. A system may raise “turns” or “sensitivity” without having cultivated JA bids or mediation of the referent. Confusing “interaction quality improved” with “pedagogical joint attention exists” is the category error the market exploits when it sells gaze scoring as if it were the full construct.

The second boundary is SEL. Programmes of emotions, empathy or self-regulation are another construct; they are not joint attention. Inference: teaching to “name feelings” is not cultivating RJA/IJA or triadic gaze on an object. The third boundary is inclusion/UDL. Access adjustments and universal design may facilitate participation in triadic episodes; they do not, by themselves, turn an eye-tracking score into pedagogical joint attention, nor authorise reducing JA to an accessibility checklist. This article does not recycle the inclusion/UDL paper: it uses it only as a limit. The fourth boundary is emergent literacy. Shared reading and print awareness may empirically overlap with shared-focus episodes on the book; they do not turn joint attention into literacy. Mifsud et al. (2025), with transfer, illustrate robots in literacy lessons: classroom social mediation, not a signature of pedagogical JA as construct. Chen (2024), Su and Yang (2022), Ljungcrantz (2026) and Nikolopoulou (2025) map AI in ECE; Miao and Holmes (2023) set limits. Inference: the only AI use coherent with ages 0–6 remains on the side of the adult who organises serve-and-return and triadic mediation—as support for professional reflection on the distribution of teacher attention (Duthois et al., 2026; Isotalo et al., 2026) or for preparing bid repertoires—subject to pedagogical validation. It does not enter as autonomous eye-tracking that declares “correct/incorrect,” a robot that replaces the adult in RJA/IJA prompts, or CV that declares joint visual attention completed.

So et al. (2023) and Sani-Bozkurt and Bozkus-Genc (2023) reinforce the clinical boundary: ASD and robot evidence does not authorise declaring JA craft in an ordinary classroom by catalogue. de Belen et al. (2023) and Ozdemir et al. (2025) reinforce the correlate boundary: gaze and visual attention inform; they do not sign mediation. Inference: pedagogical joint attention includes the situated triadic bond; it is not exhausted by a matching score or a robotic trial.

7. Inferential framework: four tests to claim pedagogical joint attention, not an artefact

The framework that follows is pedagogical inference of this article, anchored in the cases and in verified instruments. It is not a new international standard. It distinguishes four tests. If a kindergarten, daycare, CENDI or early childhood school does not pass them, it may not declare that an eye-tracking system that marks “correct/incorrect” if the child follows the gaze, a social robot that runs RJA/IJA prompts, or a computer-vision model that scores “joint visual attention” from video constitute pedagogical joint attention.

7.1. Test of situated triadic practice (adult–child–referent), not of the gaze-matching score. Bleiweiss and Sall (2025), NAEYC (2022) and OECD (2021) define the craft as mediated practice around a shared focus. de Belen et al. (2023) and Ozdemir et al. (2025) estimate correlates or patterns of visual attention. Inference: evidence of pedagogical joint attention is verified in whether child and adult shared a referent with mediation in context. If the centre’s “evidence” is a dashboard of correct gaze or a “correct/incorrect” of following, the centre has done scoring, not pedagogical joint attention.

7.2. Test of serve-and-return and of situated classroom RJA/IJA, not of the robot with prompts. Bleiweiss and Sall (2025) situate serve-and-return and JA as social-communication craft. So et al. (2023) and Sani-Bozkurt and Bozkus-Genc (2023) document robots in JA/ASD. Mifsud et al. (2025)—transfer—show social mediation by the robot in literacy, not replacement of the adult. Inference: running RJA/IJA prompts with a robot does not demonstrate reiterated serve-and-return or triadic mediation in routine. A trial may be recorded; it does not sign the practice.

7.3. Test of adult mediation and pedagogical observation of shared focus, not of autonomous CV. OECD (2021, 2023) require meaningful interactions and subordinated digitalisation. Duthois et al. (2026) and Isotalo et al. (2026) orient eye-tracking toward distribution of teacher attention and reflection on expertise. Inference: a high “joint visual attention” score may raise a threshold without raising the quality of triadic mediation. Pedagogical observation reads whether there was a referent, waiting, naming and return; the dashboard counts a label.

7.4. Test of category distinction and professional judgement, not of the product catalogue. Joint attention ≠ generic interaction quality, SEL, inclusion/UDL or emergent literacy. UNESCO (2021), Miao and Holmes (2023), the European Commission (2022), the U.S. Department of Education (2023), NAEYC (2022) and OECD (2021, 2023) require human oversight, pedagogical validation and not replacing professional judgement. Inference: a centre may not treat the infant as a unit of gaze matching nor as a respondent to decoupled robotic trials. Pedagogical joint attention is not fulfilled by better scoring of gaze following. It is fulfilled by practising triadic shared gaze, serve-and-return, situated RJA/IJA and mediation of shared focus with adults who sustain the referent in the classroom.

The framework admits the digital when it is subordinated to professional reflection on teacher attention and to validated observation (Duthois et al., 2026; Isotalo et al., 2026; Miao & Holmes, 2023). It rejects declaring pedagogical joint attention by autonomous eye-tracking-scorer, robot with RJA/IJA prompts or joint-visual-attention CV (So et al., 2023; de Belen et al., 2023; Ozdemir et al., 2025; Sani-Bozkurt & Bozkus-Genc, 2023).

The four tests are read together. Passing only the first—there was a triadic episode—without mediation or category distinction is not enough. Passing only the fourth—the centre distinguishes constructs—without situated practice is not enough either.

8. Discussion

Three tensions organise the discussion. The first is between exhibiting a measurement or trial artefact of joint attention and exercising pedagogical joint attention. It is a finding that RBI and HBI are compared on JA in autistic children (So et al., 2023; ASD transfer); that a systematic corpus of social robots for JA in ASD exists (Sani-Bozkurt & Bozkus-Genc, 2023); that eye-tracking correlates with RJA in preschoolers with ASD (de Belen et al., 2023); that visual attention in JA bids differentiates patterns in toddlers (Ozdemir et al., 2025); and that social robots in ECE classrooms assume social-mediation roles in literacy (Mifsud et al., 2025; transfer). It is a framework that joint attention at ages 0–6 is played out in triadic practice, serve-and-return and mediation (Bleiweiss & Sall, 2025; Duthois et al., 2026; Isotalo et al., 2026; NAEYC, 2022). It is not a finding that eye-tracking-scorer, RJA/IJA robot or joint-visual-attention CV produce the craft the classroom requires. The three artefacts measure, trial or label what engineering knows how to deliver and declare what only mediated triadic practice would authorise.

The second is between automated gaze assessment and everyday pedagogy of shared focus. de Belen et al. (2023) and Ozdemir et al. (2025) objectify correlates or patterns of visual attention; dashboards of “correct following” objectify isolated items. Inference: insisting that the kindergarten “already has joint attention” because the model classifies gaze matching is inverted pedagogy. An evaluative proxy is taken and made to stand for triadic practice, serve-and-return and mediation. The infant becomes a matching unit; the adult, supervisor of the classifier.

The third is between robotic trial and integral joint-attention craft. So et al. (2023) and Sani-Bozkurt and Bozkus-Genc (2023) document the robot’s ceiling in JA/ASD; Mifsud et al. (2025) show social mediation without making the robot a substitute for the adult; Miao and Holmes (2023) and Nikolopoulou (2025) balance promises and challenges. Inference: fragmenting joint attention into a trial of RJA/IJA prompts and selling it as “JA with AI” confuses a technical product with situated practice. The only AI use coherent with ages 0–6 remains on the side of the adult who organises shared focus and observes the triad, subject to UNESCO, the European Commission and the U.S. Department of Education. Autonomous eye-tracking or a robot because the system needs a video is not that use.

Additional pedagogical inference, marked as such: the series of artefacts—scoring eye-tracking, RJA/IJA robot and joint-visual-attention CV—shares the same visibility economy. Each produces an output legible for coordination: a gaze log, a robotic trial, a number of visual overlap. OECD (2021) and NAEYC (2022) require, instead, that joint-attention development be verified in everyday classroom life. Chen (2024), Su and Yang (2022), Ljungcrantz (2026) and Nikolopoulou (2025) show growth of the AI field in early childhood education; that growth does not authorise the equivalence. Miao and Holmes (2023), UNESCO (2021), the European Commission (2022) and the U.S. Department of Education (2023) agree in subordinating AI to professional judgement.

A cross-reading of the cases reinforces the restrictive thesis without inventing effects the sources do not report. So et al. (2023) show that RBI and HBI can be compared on JA; they do not show craft of triadic mediation in an ordinary CENDI. de Belen et al. (2023) show eye-tracking correlates with RJA; they do not authorise autonomous pedagogical scoring. Ozdemir et al. (2025) show visual-attention patterns in bids; they do not sign situated serve-and-return. Duthois et al. (2026) and Isotalo et al. (2026) show usefulness of eye-tracking for looking at teacher attention and expertise; not for declaring the child’s JA by dashboard. Bleiweiss and Sall (2025) anchor the craft in serve-and-return and JA. The discussion does not deny that a robot or an eye-tracker may be associated with JA measures in specific contexts: it denies that that artefact is the pedagogical construct. If the centre exhibits a gaze log, a video of an RJA/IJA robot or a joint-visual-attention dashboard, it has documented an artefact; if it describes with observation who shared which referent with whom, what serve-and-return occurred and how RJA/IJA were mediated in routine, it may begin to speak of pedagogical joint attention.

The restrictive thesis is not technophobia. It is a category rule. Chen (2024), Su and Yang (2022), Ljungcrantz (2026), Nikolopoulou (2025) and Miao and Holmes (2023) show field growth and require child-centred pedagogical validation. None of that obliges accepting that a matching score, a robotic prompt or a video label are joint attention. It obliges asking who mediates the referent and what falls outside when coordination only watches the dashboard.

9. Limits

This review is narrative. It does not apply PRISMA nor estimate primary combined effects. So et al. (2023) validate RBI versus HBI in JA/ASD, not triadic mediation in an ordinary CENDI. Sani-Bozkurt and Bozkus-Genc (2023) review robots and JA in ASD, not the ordinary classroom as primary variable. de Belen et al. (2023) are eye-tracking correlates of RJA in preschool with ASD: transfer marked. Ozdemir et al. (2025) compare visual attention in bids in ASD versus typical toddlers, not institutional scoring. Duthois et al. (2026) and Isotalo et al. (2026) anchor eye-tracking in teacher attention and expertise, not in declaring child JA by model. Bleiweiss and Sall (2025) anchor serve-and-return and JA craft. Mifsud et al. (2025) are literacy/classroom ages 5–6 transfer. Chen (2024), Su and Yang (2022), Ljungcrantz (2026) and Nikolopoulou (2025) map AI in ECE. NAEYC, UNESCO and OECD are framework sources. No Latin American AI trials were located that compare situated triadic mediation against autonomous eye-tracking-scorer or RJA/IJA robot in CENDI. Inferences in section 7 are hypotheses of pedagogical category, not implementation evidence.

10. Conclusions

An eye-tracking system that marks “correct/incorrect” if the child follows the gaze, a social robot that runs RJA/IJA prompts, or a computer-vision model that scores “joint visual attention” from video do not constitute pedagogical joint attention in an early childhood education centre. Verified evidence does not authorise that declaration. So et al. (2023) compare RBI and HBI on joint attention in autistic children (ESCS RJA/IJA)—clinical/experimental intervention, not ordinary-classroom craft. Sani-Bozkurt and Bozkus-Genc (2023) synthesise social robots for JA in ASD—artefact corpus, not everyday pedagogical equivalence. de Belen et al. (2023) correlate eye-tracking and RJA in preschool with ASD—gaze correlate ≠ mediation. Ozdemir et al. (2025) describe visual attention in JA bids in toddlers—visual pattern ≠ dashboard of completed JA. Mifsud et al. (2025) show social-mediation roles of the robot in ECE literacy (transfer)—classroom robot ≠ declared pedagogical JA. By contrast, when pedagogical joint attention exists in early childhood, there is situated practice: serve-and-return and JA as social-communication craft (Bleiweiss & Sall, 2025); reflective distribution of teacher attention (Duthois et al., 2026); eye-tracking in service of educator expertise (Isotalo et al., 2026); sensitive interactions and subordinated digitalisation (NAEYC, 2022; OECD, 2021, 2023). Joint attention is distinguished from generic interaction quality, SEL, inclusion/UDL and emergent literacy. Current guidelines require human oversight, pedagogical validation and not replacing professional judgement (UNESCO, 2021; Miao & Holmes, 2023; European Commission, 2022; U.S. Department of Education, 2023; Chen, 2024; Su & Yang, 2022; Ljungcrantz, 2026; Nikolopoulou, 2025).

Where the sources do not measure an ordinary kindergarten, this article does not claim it. Where they measure robot, eye-tracking or visual attention in JA, it does not translate them into pedagogical joint attention. Accompanying girls and boys from birth to six in joint attention is to exercise situated triadic adult–child–referent practice, serve-and-return, RJA/IJA in routine and mediation of shared focus in the classroom. The rest is gaze-matching score, robotic prompt trial and algorithmic classification of joint visual attention. It is not pedagogical joint attention in early childhood education, and it must not be presented as what it is not.

Editorial Laboratory of NEXTECH.IA / Ingeniero Mitre.

References

  1. Bleiweiss, J. D., y Sall, N. (2025). Social communication foundations: Examining interventions to promote serve and return and joint attention. TEACHING Exceptional Children. https://doi.org/10.1177/00400599251346783
  2. Chen, J. J. (2024). A scoping study on AI affordances in early childhood education. Journal of Artificial Intelligence Research, 81, 701–740. https://doi.org/10.1613/jair.1.16882
  3. de Belen, R. A. J., Pincham, H., Hodge, A., Silove, N., Sowmya, A., Bednarz, T., y Eapen, V. (2023). Eye-tracking correlates of response to joint attention in preschool children with autism spectrum disorder. BMC Psychiatry, 23, 211. https://doi.org/10.1186/s12888-023-04585-3
  4. Duthois, T., Vanderlinde, R., y Van Avermaet, P. (2026). What children get overlooked? The distribution of teachers’ attention in early childhood education: a mobile eye-tracking study. Early Childhood Research Quarterly, 75, 1–14. https://doi.org/10.1016/j.ecresq.2025.12.004
  5. European Commission. (2022). Ethical guidelines on the use of artificial intelligence (AI) and data in teaching and learning for educators. Publications Office of the European Union. https://doi.org/10.2766/153756
  6. Isotalo, S., Ukkonen-Mikkola, T., y Rutanen, N. (2026). Possibilities of eye-tracking on reflecting teachers’ expertise in early childhood education and care. Journal of Early Childhood Education Research. https://doi.org/10.58955/jecer.176441
  7. Ljungcrantz, L. (2026). The interaction of AI and early childhood education. A state-of-the-art review 2020–2024. Early Childhood Education Journal, 54, 3565–3581. https://doi.org/10.1007/s10643-025-02079-3
  8. Miao, F., y Holmes, W. (2023). Guidance for generative AI in education and research. UNESCO. https://doi.org/10.54675/EWZM9535
  9. Mifsud, C. L., Bonello, C., y Kucirkova, N. I. (2025). Exploring the multifaceted roles of social robots in early childhood literacy lessons: Insights from a Maltese classroom. International Journal of Social Robotics, 17, 1235–1249. https://doi.org/10.1007/s12369-025-01290-x
  10. NAEYC. (2022). Developmentally appropriate practice in early childhood programs serving children from birth through age 8 (4.ª ed.). NAEYC. https://www.naeyc.org/resources/pubs/books/dap-fourth-edition
  11. Nikolopoulou, K. (2025). Child-centered integration of generative AI in early learning: Balancing promises and challenges. AI, Brain and Child. https://doi.org/10.1007/s44436-025-00023-1
  12. OECD. (2021). Starting Strong VI: Supporting meaningful interactions in early childhood education and care. OECD Publishing. https://doi.org/10.1787/f47a06ae-en
  13. OECD. (2023). Empowering young children in the digital age (Starting Strong). OECD Publishing. https://doi.org/10.1787/50967622-en
  14. Ozdemir, S., Akin-Bulbul, I., y Yildiz, G. I. (2025). Visual attention in joint attention bids: A comparison between toddlers with autism spectrum disorder and typically developing toddlers. Journal of Autism and Developmental Disorders, 55, 1081–1094. https://doi.org/10.1007/s10803-023-06224-y
  15. Sani-Bozkurt, S., y Bozkus-Genc, G. (2023). Social robots for joint attention development in autism spectrum disorder: A systematic review. International Journal of Disability, Development and Education, 70(5), 625–646. https://doi.org/10.1080/1034912X.2021.1905153
  16. So, W.-C., Cheng, C.-H., Lam, W.-Y., Wong, T., Law, W.-W., Huang, Y., Ng, K.-C., Chun, C.-L., y Wong, W. (2023). Comparing the effectiveness of robot-based to human-based intervention in improving joint attention in autistic children. Frontiers in Psychiatry, 14, 1114907. https://doi.org/10.3389/fpsyt.2023.1114907
  17. Su, J., y Yang, W. (2022). Artificial intelligence in early childhood education: A scoping review. Computers and Education: Artificial Intelligence, 3, 100049. https://doi.org/10.1016/j.caeai.2022.100049
  18. UNESCO. (2021). Recommendation on the ethics of artificial intelligence. UNESCO. https://unesdoc.unesco.org/ark:/48223/pf0000381137
  19. U.S. Department of Education, Office of Educational Technology. (2023). Artificial intelligence and the future of teaching and learning: Insights and recommendations. U.S. Department of Education. https://www.ed.gov/sites/ed/files/documents/ai-report/ai-report.pdf