1. Introduction and problem
In infant and preschool settings a promise circulates that is symmetrical to personalisation: artificial intelligence “observes” the child, “transcribes” the classroom, “completes” the portfolio, and “frees” the adult to be, at last, with childhood. That promise usually silences its condition of possibility. To transcribe, a microphone is needed. To tag, an ontology of categories. To “capture” childhood, a camera or a vest with a recorder. The resulting object is not abstract. It is a slice of a three-year-old’s day: a speaking turn with word error, a photo tagged “collaboration,” a clip classified as “free play.” Those who produce that slice are not defending a thesis: they are learning to be seen. If the school treats the slice as if it were already pedagogical documentation, the adult’s gaze —called listening in Reggio and written as a learning story in Aotearoa— risks being replaced by an algorithmic register.
The problem is not hostility to the tool or nostalgia for the anecdotal notebook. It is confusion among three objects. The first is an empirical finding: what is observed when a system classifies speakers and transcribes the classroom; what happens when a digital platform orients what the educator sees and tags; how learning stories are written and shared when the device mediates. The second is a still-current normative or classical framework: what counts as observation, documentation, and assessment in NAEYC’s (2020) developmentally appropriate practice statement, in the pedagogy of listening (Rinaldi, 2006), and in learning stories (Carr, 2001; Carr and Lee, 2012). The third is a pedagogical inference: what should not be sold as a solution —AI that “improves observation,” “frees the teacher to be with the child,” “makes richer portfolios”— if the source does not measure it. Mixing them produces a pedagogy of capture: the school “documents with AI” without being able to say whether what it archives is a situated interpretation or only a taggable trace.
The thesis of this article is restrictive. AI can transcribe and label; it does not document in the pedagogical sense unless there is situated adult interpretation. Where the source does not measure improved observation, it is not asserted. Where it does not measure time freed to be with the child, it is not promised. Where it does not measure portfolio richness, it is not invented. The gap is epistemic: what counts as documentation, not only what is extracted. This article does not treat consent/CNIL/ClassDojo (21 Aug, 1:02), teacher wellbeing (9:02), integrity/detectors (13:02), curricular AI literacy, UNESCO teacher competence, parental mediation, inclusion, play/PopBots, adaptive tutors, multilingualism, or “learning AI” gaps. The object is pedagogical documentation in early childhood education: observation, the portfolio, and the risk of replacing the adult’s gaze.
2. State of the art: to document is not to archive
Three strata should be kept apart. The first is conceptual: what pedagogical documentation means when the child is understood as a subject of a hundred languages and not as a data emitter. The second is the digital shift: what happens when paper yields to portfolio platforms and tagging systems. The third is the algorithmic shift: what automatic transcription, diarisation, and movement detection in preschool classrooms actually measure.
In the conceptual stratum, Rinaldi (2006) articulates the pedagogy of listening, documentation as research, and the relation between documenting and assessing. To document is not to pile up anonymous evidence: it is to make a learning process visible in order to interpret it with others —children, families, colleagues— and to return that interpretation to teaching. Edwards, Gandini, and Forman (2012) collect the Reggio Emilia experience in transformation, including the hundred languages poem attributed to Loris Malaguzzi: the child expresses herself in ways that are not reducible to adult speech, and the school that records only one of them impoverishes what it can see. Dahlberg, Moss, and Pence (2013) oppose to the “discourse of quality” —the reduction of questions of value to expert measurement— a language of evaluation as meaning-making. Carr (2001) shifts early childhood assessment from over-formal methods toward learning stories centred on dispositions —resilience, confidence to express ideas, collaborative problem-solving— because, if complex outcomes are not assessed, they drop out of teaching. Carr and Lee (2012, 2019) argue that those stories construct learner identities and offer a framework of narrative authorship and co-participation. Knauf (2020) summarises an empirical consensus: documentation constructs an image of the child and can be an instrument of participation. That is the criterion with which this article reads what AI does not do when it transcribes.
In the digital-shift stratum, Cowan and Flewitt (2021) document, in three London early years settings, the move from paper to e-portfolios and online learning journals. White et al. (2021) show that platforms are not neutral: they orient what the educator sees. Nuttall et al. (2023) identify two actions —tagging and monitoring— that connect technical operations with work motives. Blaisdell et al. (2021) interrogate children’s authorship of the stories. Sands and Lee (2024) show that sharing learning stories in an Aotearoa centre strengthens practice when there is dialogue among teachers. Restiglian et al. (2023) saturate, from Veneto, the professional dilemma of documenting in the platform era: they are not used here as a privacy axis, but as evidence that the very status of documentation becomes uncertain when the device leads.
In the algorithmic stratum, Elbaum, Perry, and Messinger (2024) survey automatic sensors in preschool: speaker classification, speech analysis, and movement detection across the day. Sun et al. (2024, 2025) measure transcription reliability and “who said what.” Chaparro-Moreno et al. (2024) measure the accuracy of a system (IDEAS) on therapist and child talk. Those sources quantify error, not pedagogical meaning. The gap is not the absence of technology: it is the scarcity of evidence that an automatic trace counts as documentation in the sense of Rinaldi, Carr, or NAEYC (2020).
In the normative stratum, NAEYC (2020) prescribes that observation, documentation, and assessment be ongoing, strategic, reflective, and purposeful; that a system exist to collect, make sense of, and use that information; that children, from preschool on, be encouraged to observe and document; and that digitally based assessments be limited, especially for young children who should have restricted screen exposure. Making sense is not an optional step: it is the criterion that distinguishes documenting from archiving. That framework does not measure AI. Its authority is normative.
3. Review method
A critical narrative review was conducted, not a meta-analysis. The purpose was not to estimate a homogeneous effect size —non-existent among a 110-minute transcription test, a six-educator pilot, and a professional practice statement— but to articulate what counts as pedagogical documentation when systems that transcribe and tag enter the scene. Inclusion: (a) pedagogical documentation, observation, portfolio, learning stories, transcription or automatic analytics in early childhood or preschool; (b) preferably 2021–2026, with still-current classics (Carr, 2001; Rinaldi, 2006; Edwards et al., 2012; Dahlberg et al., 2013; NAEYC, 2020; Carr and Lee, 2012, 2019; Knauf, 2020); (c) journal or proceedings with DOI, academic-press book, or professional-body statement; (d) verifiable DOI or official URL. Excluded as axis were ClassDojo privacy/datafication, teacher wellbeing, integrity and detectors, student curricular literacy, UNESCO teacher competence, parental mediation, disability inclusion, play/PopBots, adaptive tutors, multilingualism, and “learning AI” gaps.
The search was run on 21 August 2026 (around 17:03, America/Mexico_City) on DOI pages, IEEE Xplore, arXiv, ASHA, ScienceDirect, Taylor & Francis, SAGE, Springer, Routledge, NAEYC, EPAA, EECERJ, and university repositories. Each source was verified against at least one of those pages. Nuclear cases: (4) automatic transcription and sensors; (5) digital platforms and learning stories; (6) the epistemic status of the gaze. Restiglian et al. (2023) saturate platformisation; Manolev and the datafication of discipline are not used as a nuclear case —they already were on 21 August at 1:02.
The analysis distinguished three enunciative statuses. Empirical finding: what was observed in the sample or tool test. Normative or classical framework: what a body prescribes or what a pedagogical tradition holds as a criterion of what counts. Pedagogical inference: the translation to the early childhood classroom, marked as such. Limits are those of any narrative review (section 9).
4. Case 1. Transcribing the classroom is not documenting learning: Who Said What, WSW 2.0, and IDEAS
Sun, Londono, Elbaum, Estrada, Lazo, Vitale, Gonzalez Villasanti, Fusaroli, Perry, and Messinger (2024) presented at the IEEE International Conference on Development and Learning an automatic framework (ALICE to classify child/teacher speaker; Whisper large-v2 to transcribe) and compared it with a human expert on 110 minutes of preschool classroom recordings: 85 minutes from child microphones (n = 4 children, 3.8–4.6 years) and 25 minutes from teacher microphones (n = 2). Empirical finding: the proportion of agreement on utterance classification was .76, with error-corrected kappa of .50 and weighted F1 of .76. Word error rate was .15 for both teachers and children: 15% of words would need to be deleted, added, or changed to equate the automatic transcription with the expert’s. Speech features —mean length of utterance, proportion of questions, responses within 2.5 seconds— were similar when calculated separately from both transcriptions. Expert transcription of those 110 minutes required about 55 hours. The authors describe “substantial progress” in speech analysis that may support language-development research. They do not measure teacher observation, time with the child, or portfolio richness. The object is reliability of “who said what.”
Sun, Feng, Gutierrez, Londono, Xu, Elbaum, Narayanan, Perry, and Messinger (2025) update the framework (WSW 2.0: wav2vec2 classification and Whisper large-v2/v3) on 235 minutes (160 from 12 children and 75 from 5 teachers). Empirical finding: weighted F1 of .845, accuracy of .846, and kappa of .672 for child/teacher classification. Word error rate was .119 for teachers and .238 for children: child speech is, in this sample, harder. Intraclass correlations of features (utterance length, lexical diversity, questions, responses) ranged from .64 to .98. The framework was applied to more than 1,592 hours over two years. The authors claim potential for educational research and for guiding language interventions. Status: a finding of scalability and differential reliability (worse on the child). It is not a finding of pedagogical documentation. To convert “1,592 hours transcribed” into “childhood captured” or “improved observation” would be an illegitimate inference: the source does not measure the teacher’s gaze or portfolio quality.
Chaparro-Moreno, Gonzalez Villasanti, Justice, Sun, and Schmitt (2024) evaluated, in the Journal of Speech, Language, and Hearing Research, the IDEAS program (Interaction Detection in Early Childhood Settings) on 45 video-recorded speech-language therapy sessions with 27 therapists and 56 children. Empirical finding: diarisation is high for therapist talk and lower for child talk. Nine of ten therapist linguistic-unit estimates meet the authors’ accuracy criteria; none of the child linguistic-unit estimates do. The authors conclude that IDEAS is reliable mainly for measuring adult clinical talk. Pedagogical inference, marked as such: if the system that “listens” to the room fails systematically on child speech, the automatic trace is not a symmetrical witness of the child. Documenting the child with an instrument blind to her speech is, at best, documenting the adult around her.
Elbaum, Perry, and Messinger (2024), in Early Childhood Research Quarterly, review sensors that combine recorders and locators with algorithms: speaker classification, speech analysis, and movement detection. Review finding: those systems open a wider window onto variability of interactions and integration into classroom social networks, and they pose technical and ethical deployment challenges. The object is research on classroom processes that predict language and social development. It is not a portfolio trial. It does not claim that the sensor replaces pedagogical observation or frees it.
Accumulated status. Empirical finding: there is measurable progress in transcribing and classifying preschool speech; error is not zero; child speech is, in WSW 2.0 and IDEAS, the weak point; the product is a corpus of features (MLU, questions, movement), not an interpretation of dispositions or children’s theories. It is not a finding that AI documents. Pedagogical inference: a “who said what” file can be input to a learning story if an adult interprets it in context. Without that interpretation, it is a register. Confusing the two is the epistemic risk of this case.
5. Case 2. From paper to tag: digital platforms, learning stories, and what the educator comes to see
Cowan and Flewitt (2021) published in the International Journal of Early Years Education participatory ethnographic case studies in three multicultural London early years settings. Empirical finding: digital documentation —photo, video, audio, and text combined in e-portfolios and online journals— opens possibilities for capturing the dynamic, embodied character of learning and can make documentation more accessible to children and families. At the same time, many systems are designed primarily for adults, not for the child to access and contribute; that adult-oriented design risks marginalising the child’s voice as practice moves from paper to digital. The authors call for collaboration among researchers, educators, and designers. Status: a finding about digital documentation made by people, not about AI that transcribes or tags on its own. They do not measure that a model “makes richer portfolios.” The richness they describe is multimodality in the service of a pedagogy of listening, threatened precisely when the system cannot be used by the child.
White, Rooney, Gunn, and Nuttall (2021), in the Australasian Journal of Early Childhood, analysed video-stimulated recall interviews with six educators at four centres in Australia and New Zealand. Empirical finding: platforms orient how learning is seen and articulated. They propose four concepts: tag-ability, trackability, completeness, and co-constitution. Each is problematised against the ubiquity of the visual. What “counts” as learning tends to be what can be tagged, what can be tracked over time, what looks like a complete file, and what platform and user co-produce. It is not an RCT of children’s learning. It is a finding about the gaze: the device does not merely record what the adult already saw; it participates in what becomes visible.
Nuttall, Rooney, Gunn, and White (2023), in Technology, Pedagogy and Education, tested the usefulness of activity theory (Leontiev’s hierarchy: operations, actions, motives) on work-shadowing observation and interviews with seven teachers across four centres in Australia and New Zealand. Empirical finding: two actions —tagging and monitoring— connect basic technical operations with motives for using the platform. The authors flag, as a future agenda, datafication as a mode of governing educators’ work. That note is not developed here as a surveillance axis (already treated on 21 August at 1:02); it is retained as a labour-epistemic finding: documenting work shifts toward the tag and monitoring. They do not measure time “freed” to be with the child. Knauf (2020), in interviews with 24 teachers (12 in Germany and 12 in New Zealand) in settings that document intensively, found that time is the most cited requirement and that strategies —staff discussion, multiple use of the same text, sharing children, templates, delimited phases, mutual support, parallelisation with supervision, setting priorities, and digitalisation— serve mainly to make a task perceived as never-ending manageable. In New Zealand, collegial discussion and platforms are described as channels of communication with families; one teacher says the response cycle became “much deeper”: situated perception, not a measured effect. In Germany, formal cutting (phases, quotas) may document the chosen window, not what matters to the child. Knauf does not study AI. Her relevant finding is that quality, in teachers’ accounts, depends on interpretive discussion, not on the digital file.
Blaisdell, McNair, Addison, and Davis (2021), in the European Early Childhood Education Research Journal, report phase one of action research in a Scottish nursery on children’s participation rights in the authorship of learning stories. Empirical finding: putting Learning Stories into practice involves complex political and material considerations; children’s authorship needs to be accommodated through flexible, non- (or less) written methods. The title quotes a child’s question —“Why am I in all of these pictures?”— as a symptom of documentation that portrays without authorising. Sands and Lee (2024), in a case study of Greerton Early Learning Centre (Aotearoa), find that the dialogical process of sharing learning stories strengthens pedagogical practice: a zoom toward learning identities and a zoom toward the continuity of teaching and learning. Status: evidence that the learning story operates as inquiry among teachers. It is not evidence that an automatic tag produces the same effect. Carr (2001) and Carr and Lee (2012, 2019) define the genre: situated narrative, dispositions, learner identity, co-authorship. A “collaboration” tag is not that genre.
Accumulated status. Empirical finding: as documentation digitises, it gains potential multimodality and accessibility, and at the same time it lets itself be oriented by the taggable, the trackable, and the “complete”; the child may be left outside authorship; teachers’ work shifts toward tagging and monitoring; the pedagogical force that Sands and Lee (2024) and Knauf (2020) describe runs through interpretive dialogue among adults —and, in Blaisdell et al. (2021), through children’s authorship. It is not a finding that AI enriches portfolios. Pedagogical inference: automating the tag is to accelerate precisely the operation that White et al. (2021) and Nuttall et al. (2023) already mark as a mould of the gaze. Without an adult who interprets dispositions in context, the portfolio can become denser in files and poorer in documentation.
6. Case 3. What counts as documentation: the adult’s gaze, a hundred languages, and the algorithmic register
The third case is not an AI product. It is the epistemic criterion against which the first two are read. Rinaldi (2006) holds that documentation is, first of all, an educational and research instrument, and that recognising it as an assessment tool is an “antibody” against anonymous, decontextualised, only apparently objective instruments. Edwards, Gandini, and Forman (2012) collect that grammar: the competent child, listening, the atelier, documentation as democratic negotiation (Dahlberg’s chapter in that volume), and the use of digital media in Reggio as one of the languages, not as a substitute for the looking adult (Forman’s chapter). Classical framework, not a Whisper trial: the camera in Reggio does not document alone; the one who chooses the frame, discusses it, and returns it to the group documents.
Albin-Clark (2021), in Contemporary Issues in Early Childhood, shifts the question from what documentation means to what it is doing. On wall documentation of a child’s water exploration, she argues that the artefact invites creating and resisting actions and offers senses of belonging; she proposes that teachers shift between “matters of fact” and “matters of concern.” Qualitative classroom finding, with a new-materialist reading: documentation has agency in the space. It does not measure AI. Pedagogical inference, marked as such: if a photo panel already does something in the room —summons, excludes, belongs— an automatic stream of tagged clips will also do something. The question is not whether there is “more evidence,” but what form of childhood and of teacher that doing produces. Restiglian, Raffaghelli, Gottardo, and Zoroaster (2023), in Education Policy Analysis Archives, interviewed Veneto educators (14 individual and one group). Empirical finding: balancing technology-mediated documentation with other professional values is not linear; educators ask for policies and support for reflective use. Cited as saturation of a professional dilemma —what documentation is ceasing to be when the platform leads— not as a biometrics or ClassDojo case.
NAEYC (2020) requires that observations be used to plan, to communicate with the family, and to improve the program, and that the system not stop at collecting: it must make sense. It encourages the child, from preschool on, to observe and document. It limits digital assessments at ages when screen exposure should be restricted. Normative framework: a mass transcription that the child cannot read, contest, or co-author does not meet that criterion by being “complete.” Dahlberg, Moss, and Pence (2013) warn that the discourse of quality turns value into metric. Pedagogical inference: a dashboard of MLU, questions per minute, or minutes of “play” detected by vision is, if presented as documentation, a relapse into that discourse. It can be a research input (Elbaum et al., 2024; Sun et al., 2024). It is not, without interpretation, a learning story (Carr, 2001).
Accumulated status. Finding of tradition and classroom: to document is to listen, interpret, negotiate, and, where possible, co-author with the child; platforms already tension that status; NAEYC requires sense, not only archive. It is not a finding that computer vision “sees” the hundred languages. A network that classifies sitting, running, or jumping —even if accurate— would see a motor language, not the theory the child is trying out with water, a block, or silence. Pedagogical inference: replacing the adult’s gaze with the algorithmic register is not a technical shortcut. It is a change in what counts as knowledge of childhood.
7. Inferential framework: transcribing is not documenting
The framework that follows is this article’s pedagogical inference, anchored in the cases and the verified classics. It is not a new international standard. It distinguishes five tests. If a proposal of “documentation with AI” in early childhood education does not pass them, it is not deployed as a pedagogical solution.
7.1. Test of the transcription that does not interpret. Sun et al. (2024, 2025) and Chaparro-Moreno et al. (2024) measure word error, kappa, and linguistic units. Carr (2001) and Rinaldi (2006) require interpretation of dispositions, theories, and processes. Inference: an automatic file can feed a story; it does not write it. To say the system “documents” is an abuse of language. Where the source does not measure situated interpretation, this article does not invent it.
7.2. Test of the tag that moulds the gaze. White et al. (2021) and Nuttall et al. (2023) show that the taggable and the monitorable orient what is seen and done. Inference: automating the tag does not “free” the teacher for a freer gaze; it can harden the mould. Knauf (2020) found that discussion among colleagues is a resource of perceived quality. An auto-tag is not that discussion. Pedagogical time-saving is not promised: Knauf documents strategies for creating time, not an AI effect.
7.3. Test of completeness that is not richness. White et al. (2021) problematise “completeness” as a criterion of what platforms make visible. Cowan and Flewitt (2021) warn of adult-oriented design. Blaisdell et al. (2021) show the portrait without authorship. Inference: a portfolio with more clips is not, of itself, a richer portfolio. Richness, in Carr and Lee (2012, 2019) and in Sands and Lee (2024), is narrative, identitarian, and dialogical. No nuclear source measures that AI increases it.
7.4. Test that the adult’s gaze is not a bottleneck to eliminate. Capture marketing treats human observation as bias and scarcity. NAEYC (2020) treats it as professional practice: strategic, reflective, purposeful, sense-making. Edwards et al. (2012) and Rinaldi (2006) treat it as inquiring listening. Inference: replacing that gaze with a sensor does not correct a defect; it removes the organ of documentation. Elbaum et al. (2024) legitimate the sensor as an instrument of interaction research. Moving it into the everyday portfolio as if it were the same object is a genre jump. This article does not claim that the adult sees “better” than the machine on MLU or minutes of movement: on those features, Sun et al. (2025) report high agreement. It claims that those features are not pedagogical documentation.
7.5. Test of children’s authorship and the hundred languages. NAEYC (2020) asks that the child document. Blaisdell et al. (2021) ask for methods that are not only written. Edwards et al. (2012) recall that adult speech does not exhaust the child. Chaparro-Moreno et al. (2024) and Sun et al. (2025) show that child speech is the relative blind spot of transcribers. Inference: a system that hears the child worse and that does not read gesture, drawing, silence, or theory-in-act cannot, without an adult, document a childhood of a hundred languages. The “capture of childhood” is, strictly, the capture of what the model knows how to recognise.
Operationally, the framework admits: (a) automatic transcriptions and clips as input to a learning story written and interpreted by adults, with the child and family where possible; (b) digital platforms used for multimodality and dialogue, not as a completeness dashboard; (c) sensors in protocolled research, not as an everyday portfolio. It rejects: (d) calling the uninterpreted file documentation; (e) selling AI as improved observation or as liberation to “be with the child” without measurement; (f) auto-tagging dispositions; (g) treating the adult gaze as noise the algorithm corrects.
8. Discussion
Three tensions organise the discussion. The first is between trace and meaning. Sun et al. (2024, 2025) show that a preschool speech trace with known error can be produced at scale. Elbaum et al. (2024) add movement and networks. That is a measurement achievement. Rinaldi (2006), Carr (2001), and NAEYC (2020) require another operation: making sense in context, with others, in order to teach. Inference: the twenty-first century can transcribe the classroom and, at the same time, cease to document it. Abundance of traces does not resolve scarcity of interpretation; it can disguise it. Cowan and Flewitt (2021) already saw that ambivalence in pre-AI digital practice: more modes, and at the same time more design that excludes the child.
The second tension is between efficiency and gaze. Knauf (2020) documents that teachers perceive documentation as never-ending and invent strategies —including digitalisation— to make it fit. The market’s rhetorical leap is: if manual transcription cost 55 hours per 110 minutes (Sun et al., 2024), AI “frees.” The source does not measure that liberation in the room: it measures research-expert hours. Saving laboratory transcription is not returning presence to the adult. Nuttall et al. (2023) describe, on the contrary, new actions (tagging, monitoring). A teacher who spends the day tagging is not more with the child. Automating that mediation is a hypothesis of intensification, not of measured relief.
The third tension is between capture and listening. The metaphor of “capturing childhood” treats the child as the object of a sensor. The metaphor of listening (Rinaldi, 2006) treats the child as interlocutor of a shared inquiry. Albin-Clark (2021) recalls that documentation does things in space: it summons, resists, belongs. A stream of auto-tagged clips will also do things: normalise that what is not tagged did not happen; that what the model did not transcribe was not said; that a disposition is a form field. Dahlberg et al. (2013) called that, in another vocabulary, the discourse of quality. Restiglian et al. (2023) show educators who already perceive the dilemma. Sands and Lee (2024) show the other pole: the story shared among teachers as a motor of pedagogical change. That pole is human, slow, and situated. It has, in this corpus, no algorithmic equivalent.
If the object is early childhood education, the teacher does not need a sensor that replaces her. She needs the authority to say that a transcribed classroom is not, without more, a documented classroom, and that a complete portfolio is not, without more, a child understood. Replacing that judgement with a “who said what” model is to abdicate the pedagogy of listening. Keeping the trace as input and reserving the name of documentation for interpretation is, in this article, the only path coherent with the verified sources.
9. Limits
This review is narrative. It does not apply a full PRISMA protocol or estimate combined effects on “quality of observation” or “portfolio richness”: those variables, as AI effects, do not appear measured in the nuclear corpus. Sun et al. (2024, 2025) are IEEE proceedings on transcription reliability in U.S. preschool, not learning-story trials. Chaparro-Moreno et al. (2024) measure speech-language therapy, not the play corner. Elbaum et al. (2024) review research sensors, not portfolios. Cowan and Flewitt (2021) are three London cases of human digital documentation. White et al. (2021) and Nuttall et al. (2023) are pilots of six and seven teachers. Knauf (2020) selects settings that already document intensively and does not study AI. Blaisdell et al. (2021), Sands and Lee (2024), Albin-Clark (2021), and Restiglian et al. (2023) are situated cases or readings, not auto-tagging trials. Rinaldi (2006), Edwards et al. (2012), Carr (2001), Carr and Lee (2012, 2019), Dahlberg et al. (2013), and NAEYC (2020) are classics or frameworks, not 2026 evidence on Whisper. Geography is biased toward the United States, the United Kingdom, Aotearoa, Australia, Germany, Italy, and Scotland; no Latin American trial of automatic transcription or portfolio auto-tagging in early childhood education that met the criteria was located with the same degree of verification. That absence is a gap, not a proof of non-existence. Section 7 inferences are hypotheses of an epistemic-pedagogical threshold, not evidence of national implementation. This article does not claim that AI improves observation, frees the teacher to be with the child, makes richer portfolios, or documents: no nuclear source measures those promises.
10. Conclusions
In early childhood education, pedagogical documentation in the face of artificial intelligence is not resolved by accumulating traces. It is resolved by asking what counts as documentation when a system can transcribe the classroom and tag the portfolio. Three families of verified evidence support the argument. The 2024–2025 automatic frameworks transcribe and classify preschool speech with known error and relatively worse child performance; they are research instruments of “who said what,” not of learning stories (Sun et al., 2024, 2025; Chaparro-Moreno et al., 2024; Elbaum et al., 2024). Digital platforms already orient the gaze toward the taggable, the trackable, and the “complete,” shift work toward tagging and monitoring, and recover pedagogical force only when there is interpretive dialogue and, where possible, children’s authorship (Cowan and Flewitt, 2021; White et al., 2021; Nuttall et al., 2023; Knauf, 2020; Blaisdell et al., 2021; Sands and Lee, 2024). The traditions of Reggio, of learning stories, and of NAEYC require listening, sense, disposition, and learner identity, not an anonymous archive (Rinaldi, 2006; Edwards et al., 2012; Carr, 2001; Carr and Lee, 2012, 2019; Dahlberg et al., 2013; NAEYC, 2020; Albin-Clark, 2021).
The restrictive thesis holds. AI can transcribe and label. It does not document in the pedagogical sense unless there is situated adult interpretation. Where the sources do not measure that AI improves observation, frees the teacher to be with the child, or makes richer portfolios, this article does not invent it. The specific risk of early childhood education is not only extracting data —the axis of another text in this series. It is that what the adult no longer looks at ceases to count as documentation. Childhood is not “captured.” It is listened to. And listening, in the sources verified here, remains a human, collegial practice and, when the school does not abdicate, a children’s practice as well.
Ingeniero Mitre / Laboratorio Editorial de NEXTECH.IA
References
- Albin-Clark, J. (2021). What is documentation doing? Early childhood education teachers shifting from and between the meanings and actions of documentation practices. Contemporary Issues in Early Childhood, 22(2), 140–155. https://doi.org/10.1177/1463949120917157
- Blaisdell, C., McNair, L. J., Addison, L., and Davis, J. M. (2021). ‘Why am I in all of these pictures?’ From Learning Stories to Lived Stories: The politics of children’s participation rights in documentation practices. European Early Childhood Education Research Journal, 30(4), 572–585. https://doi.org/10.1080/1350293X.2021.2007970
- Carr, M. (2001). Assessment in early childhood settings: Learning stories. SAGE. https://au.sagepub.com/en-gb/anz/assessment-in-early-childhood-settings/book10420
- Carr, M., and Lee, W. (2012). Learning stories: Constructing learner identities in early education. SAGE. https://uk.sagepub.com/en-gb/eur/learning-stories/book235143
- Carr, M., and Lee, W. (2019). Learning stories in practice. SAGE. https://uk.sagepub.com/en-gb/eur/learning-stories-in-practice/book257126
- Chaparro-Moreno, L. J., Gonzalez Villasanti, H., Justice, L. M., Sun, J., and Schmitt, M. B. (2024). Accuracy of automatic processing of speech-language pathologist and child talk during school-based therapy sessions. Journal of Speech, Language, and Hearing Research, 67(8), 2669–2684. https://doi.org/10.1044/2024_JSLHR-23-00310
- Cowan, K., and Flewitt, R. (2021). Moving from paper-based to digital documentation in Early Childhood Education: Democratic potentials and challenges. International Journal of Early Years Education, 31(4), 888–906. https://doi.org/10.1080/09669760.2021.2013171
- Dahlberg, G., Moss, P., and Pence, A. (2013). Beyond quality in early childhood education and care: Languages of evaluation (3rd ed.). Routledge. https://doi.org/10.4324/9780203371114
- Edwards, C., Gandini, L., and Forman, G. (Eds.). (2012). The hundred languages of children: The Reggio Emilia experience in transformation (3rd ed.). Praeger. https://www.bloomsbury.com/us/hundred-languages-of-children-9780313359811/
- Elbaum, B., Perry, L. K., and Messinger, D. S. (2024). Investigating children’s interactions in preschool classrooms: An overview of research using automated sensing technologies. Early Childhood Research Quarterly, 66, 147–156. https://doi.org/10.1016/j.ecresq.2023.10.005
- Knauf, H. (2020). Documentation strategies: Pedagogical documentation from the perspective of early childhood teachers in New Zealand and Germany. Early Childhood Education Journal, 48(1), 11–19. https://doi.org/10.1007/s10643-019-00979-9
- National Association for the Education of Young Children. (2020). DAP: Observing, documenting, and assessing children’s development and learning. https://www.naeyc.org/resources/position-statements/dap/assessing-development
- Nuttall, J., Rooney, T., Gunn, A. C., and White, E. J. (2023). The impact of digital documentation platforms on early childhood educators’ work in Australia and New Zealand. Technology, Pedagogy and Education, 32(2), 257–273. https://doi.org/10.1080/1475939X.2023.2177720
- Restiglian, E., Raffaghelli, J. E., Gottardo, M., and Zoroaster, P. (2023). Pedagogical documentation in the era of digital platforms: Early childhood educators’ professionalism in a dilemma. Education Policy Analysis Archives, 31(137). https://doi.org/10.14507/epaa.31.7909
- Rinaldi, C. (2006). In dialogue with Reggio Emilia: Listening, researching and learning. Routledge. https://www.routledge.com/In-Dialogue-with-Reggio-Emilia-Listening-Researching-and-Learning/Rinaldi/p/book/9780367427047
- Sands, L., and Lee, W. (2024). Teacher inquiry and learning stories: A site for pedagogical change. European Early Childhood Education Research Journal. https://doi.org/10.1080/1350293X.2024.2381118
- Sun, A., Feng, T., Gutierrez, G., Londono, J. J., Xu, A., Elbaum, B., Narayanan, S., Perry, L. K., and Messinger, D. S. (2025). Who said what (WSW 2.0)? Enhanced automated analysis of preschool classroom speech. En 2025 IEEE International Conference on Development and Learning (ICDL). IEEE. https://doi.org/10.1109/icdl63968.2025.11204438
- Sun, A., Londono, J. J., Elbaum, B., Estrada, L., Lazo, R. J., Vitale, L., Gonzalez Villasanti, H., Fusaroli, R., Perry, L. K., and Messinger, D. S. (2024). Who said what? An automated approach to analyzing speech in preschool classrooms. En 2024 IEEE International Conference on Development and Learning (ICDL) (pp. 1–8). IEEE. https://doi.org/10.1109/icdl61372.2024.10644508
- White, E. J., Rooney, T., Gunn, A. C., and Nuttall, J. (2021). Understanding how early childhood educators ‘see’ learning through digitally cast eyes: Some preliminary concepts concerning the use of digital documentation platforms. Australasian Journal of Early Childhood, 46(1), 6–18. https://doi.org/10.1177/1836939120979066