1. Introduction and problem

In the 3-to-6 age band—kindergarten, preschool, CENDI, infant school—a package of three artefacts has been installed that claims to count as formative assessment. The first is a “progress” dashboard: a board that paints bars, traffic lights or percentiles and exhibits that the centre “follows” each child. The second is a continuous score: a number updated with clicks, response times or items and presented as evidence that there is “real-time” assessment. The third is a model that “diagnoses”: a classifier or a cognitively diagnostic system that assigns the infant a profile, a level or a set of attributes and declares, with that label, that it already knows what the child needs. All three are visible and cheap in coordination time. They allow a centre to exhibit that it “already does formative assessment with artificial intelligence”. The enunciative leap is huge: one moves from tabulating, scoring or diagnosing to asserting that formative assessment is taking place. That leap is authorised neither by the theory of the formative process nor by evidence from early childhood classrooms.

The thesis of this article is restrictive. A progress dashboard, a continuous score or a model that diagnoses the child does not constitute formative assessment. Formative assessment in early childhood is the adult’s situated judgement in the process: feedback the child can use, adjustment of the activity under way, and a view of the group, not a board that classifies. Heritage formulates the craft as a process in which teachers and students keep learning moving, organised around three questions that are not a report: where learners are going, where they are now, and how the gap between those states is closed (Heritage, 2022). Those questions are answered in the activity, not on a panel. NAEYC requires that observation, documentation and assessment be ongoing, strategic, reflective and purposeful, and warns that using assessments in ways that do not support the child’s education is not developmentally appropriate practice (NAEYC, 2022). An algorithm that first classifies and then delivers a colour is not assessing formatively. It is administering a difference the adult did not interpret.

The problem is aggravated by a craft reason that product sheets do not mention. Formative assessment is not a test, nor a maturity verdict, nor a portfolio that is filed, nor a tutor that doses items. The formative is not a property of the instrument, but of the use of evidence to modify the teaching and learning in which one is engaged (Heritage, 2022). Confusing the product—board, score, diagnosis—with the process is the category error this paper names. Yan, Li, Panadero, Yang, Yang, and Lao (2021), reviewing 52 studies, show that formative assessment depends on personal and contextual factors of the adult who assesses, not on the availability of a panel. Inference, marked as such: the market for “AI for formative assessment” inherits the error and automates it. It turns judgement—a fact of the craft—into a trait of the child that must be detected.

This paper does not recycle axes already treated in this series. The question is one of pedagogical category: what counts as formative assessment when an early childhood centre “does AI”. The contributions are three: to reconstruct the state of the art that separates situated judgement in the process from classification by board, score or diagnostic model; to examine three families of cases; and to offer four tests for deciding when a kindergarten may claim that it assesses formatively, and not only that it tabulates, scores or diagnoses.

2. State of the art: from the formative process to the board that classifies

It is useful to separate four strata that the market for “AI for formative assessment” usually mixes. The first is formative assessment as process: eliciting evidence while teaching, interpreting it and acting—feedback, adjustment, involvement of the learner (Heritage, 2022). The second is evidence of effects, still uneven and, decisively for this article, weakly anchored in the 3–6 band: a meta-analysis of 118 studies and 258 effect sizes reports a Hedges g of 0.25 in K–12 (Yao, Amos, Snider, & Brown, 2024); a meta-analysis in Turkey finds that computer-initiated feedback has a small effect (d = 0.42) against student-initiated feedback (d = 1.16) (Karaman, 2021); and an umbrella review of thirteen meta-analyses explicitly excludes preschool from its population (Sortwell, Trimble, Ferraz, Geelan, Hine, Ramirez-Campillo, Carter-Thuiller, Gkintoni, & Xuan, 2024). The third is AI in early childhood education: reviews that map prediction and classification of the child as an affordance distinct from others (Chen, 2024; Su & Yang, 2022; Ljungcrantz, 2026). The fourth is the rights and high-risk framework, which treats the 3–6-year-old as a subject of human oversight and, in European territory, of high-risk systems when AI evaluates learning outcomes or determines educational level.

In the process stratum, Heritage (2022) insists that formative assessment is not a type of test or an event. It is a process used by teachers and students to keep learning moving: goals and success criteria, evidence gathered during the course of the activity, actionable feedback and learner involvement. Yan and Pastore (2022) distinguish, on a scale validated with 449 teachers in Hong Kong and 309 in Italy, two poles of practice: teacher-directed and student-directed formative assessment. Status: psychometric finding in primary and secondary, not in kindergarten. Inference, marked as such: both poles presuppose an adult who interprets and a use of evidence in the activity. A board that paints a score is neither pole. It is a third object: a representation that may never enter judgement.

In the effects stratum, Yao et al. (2024) offer the most recent combined estimator: g = 0.25 for 118 studies, g = 0.22 in the United States, with larger effects in high school. The range includes K–12 but does not isolate 3–6. Karaman (2021), with 32 studies and 47 effect sizes, reports d = 0.72 overall and, in the subgroup that most matters to the thesis, d = 0.42 for computer-initiated feedback against d = 1.16 for student-initiated and d = 0.69 for adult-initiated. Sortwell et al. (2024) review thirteen meta-analyses: effects from trivial to large, GRADE certainty very low in nine of thirteen, and explicit exclusion of preschool. Inference: evidence that “formative assessment works” does not authorise treating any board, score or diagnosis as formative assessment in a kindergarten. It authorises, with caution, treating an adult’s use of evidence to adjust activity as an instructional hypothesis. It does not authorise automating a profile. And it leaves an empirical gap at ages 3–6 that the market fills with product.

In the early childhood AI stratum, Su and Yang (2022) review 17 studies from 1995 to 2021. Chen (2024) maps 18 articles from 11 countries and 16 journals (2005–2023; 14 of them in 2020–2023) covering 15,081 children aged 2 to 8, and extracts four affordances; the second is AI as technology for predicting or classifying children’s conditions. Ljungcrantz (2026) updates 2020–2024: 39 studies, 20 of them from 2024, with a predominance of ages 4 to 6. Status: field finding, not formative assessment. Inference: predicting or classifying the child is exactly what the market sells as “formative assessment”. Heritage’s framework does not authorise it as equivalent. Formative assessment does not start with a diagnosis. It starts with evidence used in the process.

In the rights stratum, the Recommendation on the ethics of AI requires human oversight and particular attention when children are involved (UNESCO, 2021). Miao and Holmes (2023) set a 13-year threshold for independent conversations with generative platforms and require pedagogical validation. UNICEF (2021) prioritises the best interests of the child. Regulation (EU) 2024/1689 treats as high-risk, in Annex III, systems intended to evaluate learning outcomes or to determine the appropriate educational level (European Union, 2024). The European Commission (2022) and the U.S. Department of Education (2023) agree on not replacing professional judgement. OECD (2021) anchors early childhood quality in process interactions; OECD (2023) documents digitalisation in ECEC across 30 countries and jurisdictions; TALIS Starting Strong 2024 situates joint observation as a weekly staff practice, not as the infant’s board (OECD, 2025). Inference: a five-year-old is not the user of a dashboard. The child is the subject of a judgement an adult exercises in the activity, with the group in view.

3. Review method

A critical narrative review was conducted, not a meta-analysis. The purpose was not to estimate a homogeneous effect size among an AI classifier, an Ontario kindergarten and a mathematics instrument used by educators, but to articulate an argument of pedagogical category with verified sources. Inclusion criteria: (a) 2021–2026, with Heritage (2022) and NAEYC (2022) when they update the craft; (b) formative assessment as a process of situated judgement, or AI in early childhood education ages 3–6; (c) relevance to kindergarten, preschool, CENDI or ages 3–6; (d) peer-reviewed journal, DOI, or report from NAEYC, UNESCO, OECD, UNICEF, the European Union, the European Commission or a department of education; (e) verifiable DOI or publisher page. Axes already used in this series were excluded as central objects, though some appear as limits. Adaptive testing is cited only as a limit, not as an axis of tutors.

The search was executed on 26 August 2026 on DOI pages, Springer, Elsevier, Wiley, SAGE, MDPI, Frontiers, OECD iLibrary, UNESDOC, UNICEF, EUR-Lex, ERIC, NAEYC and publisher sites. Each source was checked against at least one of those pages. The analysis distinguished three enunciative statuses. Empirical finding: what was observed or measured in the sample. Conceptual or normative framework: what a framework or a regulation prescribes. Pedagogical inference: the translation to kindergartens and CENDI, marked as such.

4. Case 1. Tabulation, scoring or diagnosis is not formative assessment

Chen (2024) publishes in the Journal of Artificial Intelligence Research the scoping study that best illustrates the leap the thesis rejects. With 18 articles from 11 countries and 16 journals, it covers 15,081 children aged 2 to 8 between 2005 and 2023; 14 of the 18 texts are from 2020–2023. Thematic analysis extracts four affordances of AI for use in ECE. The second is AI as technology for predicting or classifying children’s conditions. Status of the evidence. Empirical field finding: the “AI for early childhood” literature already distinguishes classifying the child as an affordance of its own. It is not a finding of formative assessment. Chen does not claim that predicting or classifying is formative assessment. Pedagogical inference, marked as such: the gesture a kindergarten copies when it “does formative assessment with a dashboard” is exactly this. A signal is taken—click, error, time, item, video—a profile is assigned and a colour, a score or a path is delivered. What there is is a verdict. Formative assessment, in Heritage (2022) and NAEYC (2022), asks that evidence be used to modify the activity under way. A classifier does not modify it. It condenses it.

Clements, Sarama, Tatsuoka, Banse, and Tatsuoka (2022) publish in the Journal of Research in Childhood Education the case that best names the third artefact when it is presented as formative support. They report CREMAT, a model of cognitively diagnostic adaptive assessments, illustrated with length measurement in first and second grade: identifying the level of thinking and cognitive components with a reasonable number of items. The authors conclude that the approach matters for assessing formatively and documenting progress. Status: finding of a diagnostic adaptive model in grades 1–2, not of a 3–6 kindergarten and not that a CAT constitutes formative assessment. Inference, marked as such: this is the limit the article admits citing. An adaptive test that diagnoses attributes may be a measurement instrument. It is not, by being adaptive or “diagnostic”, formative assessment. Heritage (2022) locates the formative in the use of evidence to close a gap in the activity. A profile of attributes does not close that gap. It describes it.

Liao, Zhang, Wang, and Luo (2024) saturate the portrait of the dashboard as an empirical object. In a ninth-grade biology classroom, with 125 students (63 treatment, 62 control), an AI-enabled visual report—NLP, cognitive diagnosis and visualisation—is fed by monthly exams. There is an interaction of intervention and time on achievement; the treatment group increases anxiety (d = 0.203) and self-efficacy (d = 1.793). Against teacher feedback, there are no significant differences except self-efficacy. Status: secondary finding, not 3–6. Inference, marked as such: the artefact the market sells to the kindergarten has been trialled where there are monthly exams and students who read a board. A four-year-old is not that user. Karademir, Di Mitri, Schneider, Jivet, Allmang, Gombert, Kubsch, Neumann, and Drachsler (2024) derive sixteen requirements for a learning-analytics cockpit in secondary school: the board should support the teacher’s feedback, not replace it; teachers report lack of time and difficulty making sense of what the dashboard shows. Inference: if in secondary school the board is not formative assessment but, at best, support for the adult, in kindergarten the leap is larger. A CENDI has no monthly exams. It has a group in view. Karaman (2021) measures the ceiling: machine-initiated feedback has the smallest effect. Yao et al. (2024) find more effect in high school. Sortwell et al. (2024) exclude preschool. Inference: a product sheet that says “dashboard + score + diagnosis = formative assessment in early childhood” does not pass that test.

Su and Yang (2022) and Ljungcrantz (2026) saturate the growth of the AI-in-ECE field without equating it to formative assessment. Seventeen studies to 2021 and thirty-nine in 2020–2024, with a peak in 2024 and a focus on ages 4–6, describe tools, robots, classification and personalisation. Status: finding of growth, not of equivalence. Inference: the 2024 surge does not authorise translating “more AI in kindergarten” as “more formative assessment”. It authorises asking what the AI does: whether it serves the adult’s judgement in the process or classifies the child to exhibit a board.

5. Case 2. What the kindergarten does when there is formative assessment: the adult judges in the process

Braund, DeLuca, Panadero, and Cheng (2021) publish in Frontiers in Education the study that best names, in this corpus, empirical formative assessment in kindergarten. In Ontario they observe eight classrooms. Eight kindergarten teachers and four early childhood educators complete semi-structured interviews at two points in 2019; 56 hours of observation per classroom are added, 448 hours in total. Thematic analysis yields four themes: authentic assessment practices; feedback as foundational; shared purposes between formative assessment and co-regulation; and the need to broaden conceptualisations of assessment in kindergarten. Participants describe classroom assessment as authentic and natural, and feedback—including peer feedback—as foundational. The last theme requires not reducing kindergarten assessment to a product. Status of the evidence. Empirical process finding in a densely observed N: formative assessment is feedback and adjustment in the activity, with the group in view. It is not an AI finding nor a finding that a board reproduces that craft. This article does not turn it into a SEL study: the object retained is feedback and adjustment, not a programme of socio-emotional competences. The context is public kindergarten in Ontario, with a teacher–educator pair; it is not generalised to a Latin American CENDI. Pedagogical inference, marked as such: this is the object an early childhood centre may properly call formative assessment. It is the adult’s situated judgement in the process. A dashboard does not give feedback to a four-year-old. A score does not adjust the group’s activity. A diagnostic model does not look at the group.

Pyle, DeLuca, Wickstrom, and Danniels (2022) saturate the portrait with eighteen kindergarten classrooms. Observation and interviews group teachers into three assessment profiles: informal, formal and blended; the blended group reflects contemporary policies. Status: finding of practice profiles, not a trial and not AI. Inference: the “formal” profile approaches the gesture the market automates—collecting evidence for a product or a score. The blended profile approaches the craft the thesis defends: assessment enters the activity and the adult decides what to do with what is seen. Apostolache (2024) adds twelve kindergarten teachers, with five to eighteen years of experience, in a focus group. They define formative assessment as a process that optimises instructional actions: constant observation to supply useful information and plan ongoing development. Status: finding of perceptions, N = 12, not classroom observation and not AI. Inference: when kindergarten teachers name formative assessment, they name observation, feedback and adjustment. They do not name a board. NAEYC (2022) requires continuity, strategy, reflection and purpose. OECD (2021, 2025) locates the engine in process interactions and joint observation as a weekly practice. Yan et al. (2021), in 52 studies, show that implementation depends on attitude, self-efficacy, training and internal support, not on a panel. Inference: a kindergarten does not implement formative assessment by buying a dashboard. Evidence is verified in use, not in the artefact.

6. Case 3. An instrument that serves the adult’s judgement is not a board that classifies the child

Grimmond, Neilsen-Hewett, and Howard (2022) publish in the Australasian Journal of Early Childhood the case that best illustrates the third gesture when it is placed on the right side of the craft. They examine the qualitative component of a mixed-methods study of the Numeracy and Mathematics Block-Based Assessment (NUMBBA), an instrument designed to embed formative assessment in naturalistic pedagogies. Sixteen educators used the validated tool. Thematic analysis of semi-structured interviews revealed shifts in knowledge of mathematics, in perception of children’s capacities, in attitudes towards mathematics and assessment, and in approaches to intentional pedagogy. Status of the evidence. Empirical finding of use by adults, not of a trial with a dashboard, not of AI, and not that the instrument “does” formative assessment by itself. The authors do not claim that NUMBBA classifies the child on a board. They claim that an instrument usable in the naturalistic setting can catalyse the educator’s judgement. Pedagogical inference, marked as such: this is the only place in which an artefact can approach the formative in early childhood without contradicting the framework. The adult is the one who judges. The tool, if it enters, enters as support for that judgement—what is seen, named, adjusted—not as a verdict. This article does not turn it into a study of play: the block is the medium; the object retained is the shift in adult judgement. A dashboard that paints a score without passing through that knowing is not the same object.

The distribution of roles is confirmed by contrast. Karademir et al. (2024) design a cockpit so that the teacher gives feedback, not so that the board assesses. Liao et al. (2024) find no clear superiority of the visual report over oral feedback and do find more anxiety. Clements et al. (2022) deliver a diagnosis of attributes. Chen (2024) maps classification of the child. Inference: AI and the instrument enter formative assessment in early childhood, if at all, on the side of the adult who observes, interprets and adjusts. They do not enter as the infant’s interlocutor or as a judge who labels. Miao and Holmes (2023) exclude those under 13. OECD (2023, 2025) situates digitalisation in staff and protection, not in an adaptive profile of the infant. UNESCO (2021) requires human oversight; UNICEF (2021), the best interests of the child; the European Commission (2022), not replacing professional judgement. Annex III of Regulation (EU) 2024/1689 includes automated evaluation and determination of educational level (European Union, 2024). Inference: a system that, in a CENDI, diagnoses the child and assigns a “progress level” is not a formative innovation. In that territory it is a high-risk practice if it determines level or assessment.

7. Inferential framework: four tests to claim formative assessment, not a board

The framework that follows is pedagogical inference of this article, anchored in the cases and in the verified instruments. It is not a new international standard. It distinguishes four tests. If a kindergarten, preschool, CENDI or infant school does not pass them, it cannot declare that a progress dashboard, a continuous score or a model that diagnoses the child constitutes formative assessment.

7.1. Test of process, not product. Heritage (2022) locates formative assessment in three questions answered during instruction: where one is going, where one is, how the gap is closed. NAEYC (2022) requires continuity, strategy, reflection and purpose. Braund et al. (2021) observe feedback and adjustment in 448 hours of kindergarten. Apostolache (2024) hears teachers who name constant observation to optimise action. Chen (2024) shows that the AI literature predicts and classifies. Clements et al. (2022) deliver a profile of attributes. Liao et al. (2024) paint a report from exams. Inference: evidence of formative assessment is verified in the use of evidence to modify the activity under way. If the “evidence” that the centre assesses formatively is a progress log, a traffic light or a diagnosis, the centre has done product, not process.

7.2. Test of the adult’s judgement, not classification of the child. Yan et al. (2021) show that implementation depends on the adult and the environment. Yan and Pastore (2022) measure practices directed by the teacher and by the student, not by a panel. Grimmond et al. (2022) show an instrument that shifts the knowledge and pedagogy of sixteen educators. Karademir et al. (2024) design the board so that the teacher gives feedback. Miao and Holmes (2023) exclude those under 13 as independent interlocutors. Inference: the 3–6-year-old is not the user of a model that “formatively assesses” them. The child is the subject of a judgement an adult exercises. If AI or the instrument enters, it enters as a workshop or as support for that judgement, subject to pedagogical validation. It does not enter as tutor, classifier or scorer of the infant.

7.3. Test of feedback that adjusts the activity and of the view of the group, not of the individual score. Braund et al. (2021) find feedback as foundational and an assessment that is not reduced to the isolated individual. Pyle et al. (2022) show that the blended profile integrates assessment into classroom activity. OECD (2021, 2025) locates quality in interactions and joint observation as a weekly practice. Karaman (2021) measures the smallest effect in computer-initiated feedback. Yao et al. (2024) find more effect in high school. Inference: a child’s continuous score is not a view of the group. A dashboard that ranks percentiles does not adjust circle time. At ages 3–6, the formative is played in the next move of the activity—changing the material, regrouping, returning a prompt, sustaining a shared observation—not in a number that updates.

7.4. Test of rights and high risk, not of the product board. UNESCO (2021) requires human oversight. UNICEF (2021) requires the best interests of the child. Annex III of Regulation (EU) 2024/1689 includes access, admission, assignment, evaluation of learning and educational level (European Union, 2024). The European Commission (2022) and the U.S. Department of Education (2023) require not replacing the teacher. Inference: an early childhood centre cannot treat the child as the object of an automated “progress” verdict. In European territory, a system that assigns level or assesses with AI is, save for the regulation’s safeguards, a high-risk practice. Outside that territory it remains a category error: formative assessment is not fulfilled by classifying better. It is fulfilled by judging in the process, with feedback, adjustment and a view of the group.

The framework admits the instrument that serves the adult’s judgement (Grimmond et al., 2022); the feedback and adjustment observed in kindergarten (Braund et al., 2021; Pyle et al., 2022; Apostolache, 2024); Heritage’s three questions (2022) and NAEYC’s position (2022); and Yao et al.’s g = 0.25 (2024) as a cautious ceiling of formative practices in K–12, not of boards. It rejects declaring formative assessment by a progress dashboard, a continuous score or a model that diagnoses the child (Chen, 2024; Clements et al., 2022; Liao et al., 2024; Karaman, 2021; Miao & Holmes, 2023).

8. Discussion

Three tensions organise the discussion. The first is between classifying and judging. It is a finding that the AI-in-ECE literature includes, as a distinct affordance, predicting or classifying children’s conditions (Chen, 2024); that the field grew to 39 studies in 2020–2024, with a peak in 2024 and a focus on ages 4–6 (Ljungcrantz, 2026); and that 17 earlier studies already reported tools and impacts (Su & Yang, 2022). It is a framework that formative assessment is a process of using evidence (Heritage, 2022; NAEYC, 2022). It is not a finding that a dashboard, a score or a diagnosis produces the craft Braund et al. (2021), Pyle et al. (2022) and Apostolache (2024) observe or hear. The policy of the three artefacts—board, score, model—measures what engineering knows how to measure and declares what only situated judgement would authorise.

The second is between the child as profile and the group as process. Clements et al. (2022) deliver cognitive attributes. Liao et al. (2024) deliver a visual report to students who sit monthly exams and, moreover, increase anxiety. Karademir et al. (2024) collect teachers who have no time and do not always know what to do with the board. Sortwell et al. (2024) leave preschool out of the synthesis of meta-analyses. Yao et al. (2024) find more effect in high school. Inference: insisting that the child “have a progress curve” while the adult stops looking at the group is inverted pedagogy. The profile blames the infant for what the process did not interpret. OECD (2021) had anchored quality in interactions. The diagnostic model does the inverse operation: it extracts the child from the group and returns the child as an itinerary.

The third is between the adult’s instrument and the child interlocutor. Grimmond et al. (2022) show educators whose judgement shifts when they use an instrument in the naturalistic setting. Karaman (2021) shows that machine-initiated feedback has the smallest effect. Miao and Holmes (2023) exclude those under 13. OECD (2023, 2025) situates digitalisation in staff and protection. Inference: the only use of AI or of a digital instrument that does not contradict formative assessment at ages 3–6 is the one that remains on the side of the adult who observes, interprets and adjusts, and is subject to pedagogical validation. An instrument that helps the educator see and decide may serve the formative. A model that diagnoses the child because the system needs a score does not.

9. Limits

This review is narrative. It does not apply PRISMA or estimate combined effects. Chen (2024) maps AI affordances in ECE ages 2–8, not implementations of formative assessment; transfer to the thesis is inference. Su and Yang (2022) cover 1995–2021; Ljungcrantz (2026), 2020–2024. Clements et al. (2022) are grades 1–2 and a diagnostic adaptive model; they are used as a limit, not as an axis of tutors. Liao et al. (2024) and Karademir et al. (2024) are secondary. Karaman (2021) does not isolate 3–6. Yao et al. (2024) cover K–12; Sortwell et al. (2024) exclude preschool. Yan et al. (2021) and Yan and Pastore (2022) cover mainly primary and secondary. Braund et al. (2021) are eight Ontario classrooms; they are not used as a SEL article. Pyle et al. (2022) are profiles, not a trial. Apostolache (2024) is an N = 12 of perceptions. Grimmond et al. (2022) are sixteen educators; they are not used as a play article. Heritage, NAEYC, UNESCO, UNICEF, the OECD, the European Union and the 2021–2023 guidelines are framework. No Latin American trials of formative assessment and AI in kindergarten measuring situated judgement against a dashboard were located. The inferences of section 7 are hypotheses of pedagogical category, not implementation evidence.

10. Conclusions

A progress dashboard, a continuous score or a model that diagnoses the child does not constitute formative assessment in an early childhood centre. The verified evidence does not authorise that declaration. Eighteen studies covering 15,081 children aged 2 to 8 distinguish, as an AI affordance, predicting or classifying the child; that is the field of AI in ECE, not formative assessment (Chen, 2024). Seventeen and thirty-nine reviews saturate the field’s growth without equating it to situated judgement (Su & Yang, 2022; Ljungcrantz, 2026). A cognitively diagnostic adaptive model informs attributes; it does not thereby assess formatively (Clements et al., 2022). An AI visual report in ninth grade increases anxiety and does not clearly outperform teacher feedback (Liao et al., 2024). Computer-initiated feedback has the smallest effect (Karaman, 2021). The K–12 g of 0.25 is not isolated to ages 3–6 and is larger in high school (Yao et al., 2024); the synthesis of meta-analyses leaves preschool out (Sortwell et al., 2024). By contrast, when there is formative assessment in early childhood, there is an adult who judges in the process: 448 hours of kindergarten with feedback as foundational (Braund et al., 2021), practice profiles in eighteen classrooms (Pyle et al., 2022), twelve teachers who name observation to optimise action (Apostolache, 2024), and sixteen educators whose judgement shifts with an instrument that does not classify the child on a board (Grimmond et al., 2022). Current law requires human oversight, the best interests of the child, an age threshold for independent conversations with generative platforms and, in the European Union, treating as high-risk the AI that determines access, admission, assignment, evaluation or educational level (UNESCO, 2021; UNICEF, 2021; Miao & Holmes, 2023; European Union, 2024; European Commission, 2022).

Where the sources do not measure a kindergarten, this article does not assert it. Where they measure a board, a score or a diagnosis, it does not translate them into formative assessment. Accompanying children aged three to six with formative assessment is to exercise situated judgement in the process—feedback, adjustment of the activity, a view of the group—with the adult as interpreter. The rest is a board that classifies. It is not formative assessment, and it should not be presented as what it is not.

Editorial Laboratory of NEXTECH.IA / Ingeniero Mitre.

References

  1. Apostolache, R. (2024). Kindergarten teachers’ perception on preschoolers’ formative assessment. Educatia 21, (27), Article 14. https://doi.org/10.24193/ed21.2024.27.14
  2. Braund, H., DeLuca, C., Panadero, E., & Cheng, L. (2021). Exploring formative assessment and co-regulation in kindergarten through interviews and direct observation. Frontiers in Education, 6, 732373. https://doi.org/10.3389/feduc.2021.732373
  3. Chen, J. J. (2024). A scoping study on AI affordances in early childhood education: Mapping the global landscape, identifying research gaps, and charting future research directions. Journal of Artificial Intelligence Research, 81, 701–740. https://doi.org/10.1613/jair.1.16882
  4. Clements, D. H., Sarama, J., Tatsuoka, C., Banse, H. W., & Tatsuoka, K. K. (2022). Evaluating a model for developing cognitively diagnostic adaptive assessments: The case of young children’s length measurement. Journal of Research in Childhood Education, 36(1), 143–158. https://doi.org/10.1080/02568543.2021.1895921
  5. European Commission. (2022). Ethical guidelines on the use of artificial intelligence (AI) and data in teaching and learning for educators. Publications Office of the European Union. https://doi.org/10.2766/153756
  6. Grimmond, J., Neilsen-Hewett, C., & Howard, S. J. (2022). The effectiveness of a formative play-based mathematics assessment in supporting the early childhood intentional pedagogue. Australasian Journal of Early Childhood, 47(4), 304–319. https://doi.org/10.1177/18369391221130787
  7. Heritage, M. (2022). Formative assessment: Making it happen in the classroom (2nd ed.). Corwin. https://doi.org/10.4135/9781071813706
  8. Karademir, O., Di Mitri, D., Schneider, J., Jivet, I., Allmang, J., Gombert, S., Kubsch, M., Neumann, K., & Drachsler, H. (2024). I don’t have time! But keep me in the loop: Co-designing requirements for a learning analytics cockpit with teachers. Journal of Computer Assisted Learning, 40(6), 2681–2699. https://doi.org/10.1111/jcal.12997
  9. Karaman, P. (2021). The effect of formative assessment practices on student learning: A meta-analysis study. International Journal of Assessment Tools in Education, 8(4), 801–817. https://doi.org/10.21449/ijate.870300
  10. Liao, X., Zhang, X., Wang, Z., & Luo, H. (2024). Design and implementation of an AI-enabled visual report tool as formative assessment to promote learning achievement and self-regulated learning: An experimental study. British Journal of Educational Technology, 55(3), 1253–1276. https://doi.org/10.1111/bjet.13424
  11. Ljungcrantz, L. (2026). The interaction of AI and early childhood education. A state-of-the-art review 2020–2024. Early Childhood Education Journal, 54, 3565–3581. https://doi.org/10.1007/s10643-025-02079-3
  12. Miao, F., & Holmes, W. (2023). Guidance for generative AI in education and research. UNESCO. https://doi.org/10.54675/EWZM9535
  13. NAEYC. (2022). Developmentally appropriate practice in early childhood programs serving children from birth through age 8 (4th ed.). NAEYC. https://www.naeyc.org/resources/pubs/books/dap-fourth-edition
  14. OECD. (2021). Starting Strong VI: Supporting meaningful interactions in early childhood education and care. OECD Publishing. https://doi.org/10.1787/f47a06ae-en
  15. OECD. (2023). Empowering young children in the digital age (Starting Strong). OECD Publishing. https://doi.org/10.1787/50967622-en
  16. OECD. (2025). Results from TALIS Starting Strong 2024: Strengthening early childhood education and care. OECD Publishing. https://doi.org/10.1787/20af08c0-en
  17. Pyle, A., DeLuca, C., Wickstrom, H., & Danniels, E. (2022). Connecting kindergarten teachers’ play-based learning profiles and their classroom assessment practices. Teaching and Teacher Education, 119, 103855. https://doi.org/10.1016/j.tate.2022.103855
  18. Sortwell, A., Trimble, K., Ferraz, R., Geelan, D. R., Hine, G., Ramirez-Campillo, R., Carter-Thuiller, B., Gkintoni, E., & Xuan, Q. (2024). A systematic review of meta-analyses on the impact of formative assessment on K-12 students’ learning: Toward sustainable quality education. Sustainability, 16(17), 7826. https://doi.org/10.3390/su16177826
  19. Su, J., & Yang, W. (2022). Artificial intelligence in early childhood education: A scoping review. Computers and Education: Artificial Intelligence, 3, 100049. https://doi.org/10.1016/j.caeai.2022.100049
  20. UNESCO. (2021). Recommendation on the ethics of artificial intelligence. UNESCO. https://unesdoc.unesco.org/ark:/48223/pf0000381137
  21. UNICEF. (2021). Policy guidance on AI for children 2.0. UNICEF Innocenti. https://www.unicef.org/innocenti/reports/policy-guidance-ai-children
  22. European Union. (2024). Regulation (EU) 2024/1689 of the European Parliament and of the Council of 13 June 2024 laying down harmonised rules on artificial intelligence. Official Journal of the European Union, L 1689. https://eur-lex.europa.eu/eli/reg/2024/1689/oj
  23. U.S. Department of Education, Office of Educational Technology. (2023). Artificial intelligence and the future of teaching and learning: Insights and recommendations. U.S. Department of Education. https://www.ed.gov/sites/ed/files/documents/ai-report/ai-report.pdf
  24. Yan, Z., Li, Z., Panadero, E., Yang, M., Yang, L., & Lao, H. (2021). A systematic review on factors influencing teachers’ intentions and implementations regarding formative assessment. Assessment in Education: Principles, Policy & Practice, 28(3), 228–260. https://doi.org/10.1080/0969594X.2021.1884042
  25. Yan, Z., & Pastore, S. (2022). Assessing teachers’ strategies in formative assessment: The Teacher Formative Assessment Practice Scale. Journal of Psychoeducational Assessment, 40(5), 592–604. https://doi.org/10.1177/07342829221075121
  26. Yao, Y., Amos, M., Snider, K., & Brown, T. (2024). The impact of formative assessment on K-12 learning: A meta-analysis. Educational Research and Evaluation, 29(7-8), 452–475. https://doi.org/10.1080/13803611.2024.2363831