1. Introduction and problem
In early childhood education, the word “inclusion” circulates with an ease that does not always match its legal content. The Convention on the Rights of Persons with Disabilities recognizes the right to an inclusive education system at all levels and to lifelong learning, with reasonable accommodation, individualized support, and augmentative and alternative modes of communication (United Nations, 2006, arts. 2, 7, 9, 21, and 24). General comment No. 4 of the Committee on the Rights of Persons with Disabilities specifies that inclusion is neither integration nor physical placement: it is a transformation of culture, policy, and practice to remove barriers and guarantee full participation (Committee on the Rights of Persons with Disabilities, 2016). That standard is not met by a device. Nor is it met by a model that “detects” or “personalizes.”
At the same time, AI enters kindergarten and early intervention with a promise of accessibility: speech recognizers for poorly intelligible speech, gaze communicators, and autism classifiers from video, audio, or electroencephalography. UNICEF estimates that nearly 240 million children—approximately one in ten—live with a disability and that, compared with their peers, they are 25% less likely to attend an early childhood education programme and 24% less likely to receive early stimulation and responsive care (UNICEF, 2021a). WHO and UNICEF further document that more than 2.5 billion people need assistive products and nearly one billion do not obtain them; in some low-income countries access may be as low as 3%, compared with around 90% in high-income settings (World Health Organization & UNICEF, 2022). Offering AI as “inclusion” without asking whom the model recognizes, who mediates it, and whose post is left unfilled is not a rights policy.
The thesis of this article is restrictive. AI does not include by itself. It can expand communicative access—the right to be understood, to choose, to play with an interlocutor—when configured as assistive technology, with human service, individualized goals, and usability evaluation. It can, at the same time, medicalize childhood by turning difference into a predictive label, systematically bias those who speak, look, or move in non-normative ways, and displace the speech-language pathologist, the occupational therapist, the support teacher, and the interpreter. Distinguishing those three destinations is not a rhetorical nuance: it is the difference between scaffolding and substitution.
The problem is compounded because the “AI for autism” literature often reports high areas under the curve without auditing equity or external validity (Valizadeh et al., 2024). Recognizers trained on adults of standard varieties fail with children, accents, and non-native speakers (Feng et al., 2024). Fairness roadmaps warn that atypical speech and AAC voice can lock people with disabilities out of systems sold as accessible (Guo et al., 2019; Trewin et al., 2019). None of this authorizes the claim that a kindergarten is “already inclusive” because it bought a licence.
This paper offers three contributions: to reconstruct the state of the art while separating norm, evidence, and inference; to examine three families of verified cases—communicative support, automated autism detection as risk evidence, and voice bias—without recycling previously published axes on tutors, PopBots, UNESCO teacher competence, or parental mediation; and to offer an inferential framework on when to support and when not to substitute, anchored in the Convention, UNICEF (2021b), and UNESCO (2021a), without turning them into a teacher workshop.
2. State of the art: inclusion, assistive technologies, and bias
Three strata should be kept apart. The first is normative: what treaties and general comments prescribe. The second is epidemiological and service-related: who remains outside early childhood education and assistive technologies. The third is technical-empirical: what AAC, speech-recognition, and diagnostic-classification systems measure, with which samples and which limits.
In the normative stratum, Article 7 of the Convention requires full enjoyment of rights, the best interests of the child, and the right to express views with assistance appropriate to age and disability; Article 9 requires accessibility, including ICTs; Article 21 protects freedom of expression through sign language, Braille, and augmentative and alternative modes; Article 24 links that right to an inclusive education system, teachers trained in those modes, and individualized support “in environments that maximize academic and social development, consistent with the goal of full inclusion” (United Nations, 2006). General comment No. 4 insists that inclusion requires transforming financing, administration, design, and monitoring, not adding a technological annex (Committee on the Rights of Persons with Disabilities, 2016). General comment No. 7 recalls the participation of children with disabilities through their representative organizations (Committee on the Rights of Persons with Disabilities, 2018). The Convention on the Rights of the Child and its General comment No. 25 extend those rights to the digital environment, including algorithms and artificial intelligence (United Nations, 1989; Committee on the Rights of the Child, 2021).
UNICEF’s policy guidance on AI for children, version 2.0, sets nine requirements; nuclear here are ensuring inclusion of and for children and prioritizing fairness and non-discrimination. The document warns that bias—unrepresentative data, context blindness, and use of outputs without human control—can produce long-lasting exclusion, and that children with disabilities must be prioritized (UNICEF, 2021b). UNESCO’s Recommendation on the ethics of AI, adopted by 193 States in 2021, anchors human rights, fairness, non-discrimination, and human oversight (UNESCO, 2021a). The generative-AI guidance adds ethical and pedagogical validation, data protection, and an age threshold for independent conversations with generative platforms (Miao & Holmes, 2023). None of these instruments asserts that a classifier “includes” a three-year-old. UNESCO’s report on inclusive early childhood care and education documents that disability, together with gender, ethnicity, language, and displacement, continues to deny access to services declared universal (UNESCO, 2021b).
In the service stratum, the Global report on assistive technology insists that support products—wheelchairs, hearing aids, communicators, cognition and communication apps—enable rights and have a return on investment, yet remain radically inequitable (World Health Organization & UNICEF, 2022). That gap matters for AI: a gaze- or voice-recognition model is not a fulfilled right if there is no device, maintenance, team training, and continuity after the trial ends. Hsieh et al. (2024) show this in Taiwan: eligibility, co-payment, and the paediatric novelty of gaze technology condition everyday use.
In the empirical stratum, high-tech AAC incorporates machine learning for prediction, non-normative speech recognition, and just-in-time boards. Costanzo et al. (2023) place Talkitt on that line. Hsieh et al. (2024) and Borgestig et al. (2021) place gaze as an access method. Holyfield et al. (2024) explore, with a Wizard-of-Oz prototype, automated augmented input. Those studies measure usability, participation, naming, or visual attention. They do not measure inclusion as systemic transformation.
In the bias stratum, Feng et al. (2024) quantify, in Dutch and Mandarin, bias by gender, age, and accent: native child speech (ages 7–11) is recognized worse than adolescent speech; non-native speech degrades word error rate by 15 to 24 points. Koenecke et al. (2020) documented, in five commercial systems, an error rate of 0.35 for Black speakers versus 0.19 for white speakers. Guo et al. (2019) warn of failure with dysarthria, disability-related accent, and AAC voice. Trewin et al. (2019) stress that disability is diverse, under-represented, and collides with the privacy of personalization data. Bias, here, is denial of communicative access.
In the detection stratum, Valizadeh et al. (2024) synthesize 344 studies of automated autism diagnosis across nine modalities. AUC lies between 0.71 and 0.90; EEG reaches 0.85–0.93. The authors conclude that the literature is “rife with bias and methodological/reporting flaws.” That statement is the finding this article retains: not the promise of earlier diagnosis, but evidence of a field that produces high figures with fragile clinical validity. Medicalizing early childhood with an algorithmic risk threshold, absent equity standards and clinical oversight, is precisely the scenario UNICEF describes as profiling that locks children into trajectories and restricts their development (UNICEF, 2021b).
3. Review method
A critical narrative review was conducted, not a meta-analysis. The purpose was not to estimate a comparable effect size—non-existent in homogeneous form across gaze AAC, speech recognizers, and diagnostic classifiers—but to articulate an argument of policy and practice with verified sources. Inclusion criteria were: (a) publication between 2021 and 2026, with justified exception for treaties and still-standing bias studies (United Nations, 2006, 1989; Koenecke et al., 2020; Guo et al., 2019; Trewin et al., 2019); (b) focus on early childhood (0–8 years) or on childhood disability when the study explicitly includes young children or child speech; (c) document type of peer-reviewed journal, DOI proceedings, or a United Nations, UNESCO, UNICEF, or WHO report; (d) access to a DOI page, institutional repository, or publisher site confirming authors, year, title, and findings. Sources whose DOI did not open to verifiable metadata, unreviewed opinion essays, and cases already used as the axis of previous articles in this series (PopBots, personalized tutors, UNESCO teacher competence, and parental mediation in the home) were excluded.
The search was executed on 20 August 2026 on DOI pages, PubMed, Frontiers, ScienceDirect, Taylor & Francis, PNAS, UNESDOC, UNICEF Innocenti, WHO, OHCHR, and university repositories (Delft, Jönköping, NSF PAR). Each cited source was verified against at least one of those pages. The corpus was organized as a gaze-and-speech support case, an automated-detection case treated as risk evidence, and a voice-bias case. For saturation, Borgestig et al. (2021) and Holyfield et al. (2024) were retained, together with the frameworks of UNICEF (2021a, 2021b), UNESCO (2021a, 2021b), Miao and Holmes (2023), and WHO and UNICEF (2022).
The analysis distinguished three enunciative statuses. Empirical finding: what was observed or measured in the sample. Normative framework: what a treaty, general comment, or policy guidance prescribes. Pedagogical inference: the translation this article proposes for kindergarten and support teams, marked as such. The limits are those of any narrative review: there is no full PRISMA protocol, there is language bias toward English and Italian, and Latin America and Africa are under-represented (section 9).
4. Case 1. Assistive technologies: gaze, atypical speech, and the ceiling of intelligent AAC
Hsieh, Granlund, Odom, Hwang, and Hemmingsson (2024) published in Disability and Rehabilitation: Assistive Technology (vol. 19, no. 2, pp. 492–505) a multiple-baseline design with four children aged three to six in Taiwan (mean 5.1 years), with dyskinetic cerebral palsy or a neurometabolic disorder, GMFCS, MACS, and CFCS levels IV–V, and low eye-control. The six-month intervention combined a gaze device, personalized software, two team meetings, and twelve individual supports at home or school. The use diary, checked against software logs, showed greater activity diversity (from one or two to three–six), higher frequency (from 0.3–2 to 2.2–3.2 days per week), and about 25 minutes per day of use (group effect 0.76). Six of eight play, communication, and learning goals were attained. Parents and teachers perceived clinically significant changes, with one exception on satisfaction. The authors stress that handing over the device is not enough: service, positioning, and everyday opportunities are part of the intervention. In maintenance, use declined without a personal device (Hsieh et al., 2024).
Status of the evidence. Empirical finding: in that minimal sample, with that service and duration, gaze technology was associated with more participation in computer play, choice, and, in one case, school tasks. It is not a finding that AI “includes” the child in the ordinary classroom, nor that it improves global cognitive development: the study does not measure that. Nor is it a randomized trial; it is a single-case design with four demonstrations, argued internal validity, and limited external validity. Frequency did not reach daily use. The authors interpret 20–30 minutes as optimal for young beginners, not as evidence of technological superiority.
Borgestig et al. (2021) add multicenter saturation: seventeen participants aged 3 to 26 in Sweden, Dubai, and the United States, seven months of an eye-gaze-controlled computer and assistive-technology-center services. Expressive communication and functional independence increased significantly; sixteen of seventeen expanded their repertoire; the average was 4.1 activities, 70 minutes per day, and use on 76% of days. The age range exceeds early childhood; the value is convergent: gaze can be an access method when there is service. It does not authorize withdrawing the therapist.
Costanzo et al. (2023) evaluated Talkitt, by Voiceitt, which translates poorly intelligible sounds into clear words in real time through a user-calibrated discrete recognizer. Of 34 enrolled participants with Down syndrome (5.54 to 28.9 years), 23 completed six months (mean age 9.44; mean IQ 59.78). The system was trained for at least twenty words or phrases per person. Usability, reported by caregivers, was high for instructions and interface and minimal for “use at school.” After Bonferroni correction, global adaptive behavior (ABAS-II: 55.6 to 62.4) and naming (BVL: 36.6 to 45.1) improved; articulation did not survive correction. Accuracy (60–95%) was not analyzed there. The authors declare a commercial conflict of two Voiceitt co-authors, no control of concomitant therapies, and a closed vocabulary (Costanzo et al., 2023).
Status of the evidence. Empirical finding: in a mixed sample of children, adolescents, and one young adult, with intensive calibration and clinical follow-up, personalized speech AAC was associated with improved naming and an adaptive composite, with low school uptake. It is not a finding of kindergarten inclusion: the study itself records the relative failure of use with classmates. It is not a finding of a “cure” of speech. Mean age (9.44) exceeds early childhood education; it is retained because the lower bound (5.54 years) touches the end of early childhood and because it illustrates the ceiling of a closed-vocabulary recognizer. Pedagogical inference, marked as such: a communicator that only “understands” twenty words does not replace the speech-language pathologist or the support teacher at recess or circle time.
Holyfield et al. (2024) describe a Wizard-of-Oz prototype—not a deployed model—that automates augmented input (photographs paired with the interlocutor’s speech) for three minimally verbal children on the autism spectrum. Results were variable; overall, visual attention and linguistic participation increased. The authors require further research on efficacy, reliability, and validity (Holyfield et al., 2024). Status: a concept test with a human in the loop, not a licence to dispense with the interlocutor.
Pedagogical inference of this article: mediated assistive technologies—gaze, calibrated speech, augmented input—can expand the right to communicate and to participate in play and tasks, which is exactly the object of Articles 21 and 24 of the Convention. That finding does not scale to the claim “AI includes.” Inclusion remains an attribute of the education system and of the human team that positions, models, interprets attempts, and sustains the device when the trial ends.
5. Case 2. Automated autism detection: high AUC, field bias, and medicalization
Valizadeh et al. (2024) published in Reviews in the Neurosciences (vol. 35, no. 2, pp. 141–163) a review with meta-analysis of automated autism diagnosis. In February 2022 they searched databases and grey literature, used adapted QUADAS-2, and synthesized by modality. They included 344 studies with 186,020 participants (51,129 estimated unique) across nine modalities; 232 entered the meta-analysis. AUC lay between 0.71 and 0.90; EEG reached 0.85–0.93. The explicit conclusion is that the literature is rife with bias and with methodological and reporting flaws (Valizadeh et al., 2024).
Status of the evidence. Empirical finding (of synthesis): the field produces high statistical-discrimination figures and, simultaneously, an unfavorable diagnosis of evidence quality. It is not a finding that AI diagnoses autism in a clinically valid way in a kindergarten. It is not a finding of equity across sexes, languages, camouflaged phenotypes, or minoritized populations: the review does not authorize that inference of success. Treating AUC as “early detection achieved” would be precisely the leap this article forbids.
The normative framework turns that leap into a rights problem. UNICEF warns that profiling can lock children into a profile, reinforce stereotypes, and limit trajectories, and that key decisions—diagnosis, benefits, school admission—must not be taken by AI alone, without a human in the loop (UNICEF, 2021b). UNESCO requires proportionality, human oversight, and non-discrimination (UNESCO, 2021a; Miao & Holmes, 2023). General comment No. 4 rejects reducing the child to a deficit that must be classified in order to be placed (Committee on the Rights of Persons with Disabilities, 2016). An “ASD-risk” classifier in a three-year-old classroom medicalizes difference, may bias against girls and minoritized languages, and may advance a label that conditions supports and separation.
Pedagogical inference, marked as such: kindergarten is not an algorithmic screening center. Observation, functional assessment and, where appropriate, interdisciplinary clinical evaluation belong to human specialists. A system that “prioritizes” whom to assess may be, at best, an auditable input for waiting lists; at worst, a barrier for those the model does not recognize. This article does not deny the clinical value of early evaluation by competent teams. It asserts that the diagnostic-AI literature, in the state documented by Valizadeh et al. (2024), does not support replacing those teams or adopting them uncritically in early childhood education.
6. Case 3. Bias of voice, accent, and non-normative speech
Feng, Halpern, Kudina, and Scharenborg (2024) published in Computer Speech & Language (vol. 84, art. 101567) a quantification of automatic speech-recognition bias by gender, age, and accent, in Dutch and Mandarin, with hybrid and end-to-end architectures. In Dutch (Jasmin-CGN corpus), native adolescents were recognized best; native children aged 7–11 were worst among natives (bias of 11.8 WER points on read speech with the hybrid system, p < 0.001). Non-native speech suffered 23.1 points of bias on read speech and 13.9 in human–machine interaction. Flemish speech was worse than speech from the Netherlands. Only a fraction of the bias was explained by pronunciation. In Mandarin they found no significant gender bias, but bias against accents from non-Mandarin regions, particularly Min (Feng et al., 2024).
Status of the evidence. Empirical finding: state-of-the-art systems do not recognize everyone equally. Child speech—even of children older than those in kindergarten, with a more stable vocal tract than a three-year-old—is already at a disadvantage. The finding does not measure disability; it measures age, accent, and non-nativeness. The inference toward kindergarten and atypical speech is therefore one of risk, not of equivalence: if the acoustic mismatch is already large at age seven, a generic recognizer in a four-year-old classroom, or facing uncalibrated dysarthria, cleft palate, or Down-syndrome speech, cannot be taken as accessible. Costanzo et al. (2023) illustrate the opposite path: intense personalization can make a closed vocabulary usable, at the cost of not generalizing to the classroom.
Koenecke et al. (2020) evaluated five commercial systems on 19.8 hours of interviews with 42 white speakers and 73 Black speakers in five U.S. cities. All showed racial disparities: average WER 0.35 versus 0.19, maintained on identical phrases, pointing to the acoustic models. The corpus is adult; it is cited as evidence that commercial recognition reproduces linguistic hierarchies. In a bilingual kindergarten or one of minoritized varieties, a voice assistant “for everyone” without error audit by group fails non-discrimination (UNICEF, 2021b).
Guo et al. (2019) map the risk that ASR will not work for atypical speech, disability-associated accents, and AAC-device voice, and that those who cannot speak will be excluded. Trewin et al. (2019) add that including disability in the data collides with privacy, and that habitual fairness techniques—designed for large groups and binary attributes—do not transfer as they stand. UNICEF further notes that facial recognition is less reliable on children’s faces and on groups defined by gender and ethnicity (UNICEF, 2021b). This article does not take that statement as a controlled trial of its own; it takes it as a risk framework for vision systems, convergent with voice bias.
Pedagogical inference: a system that does not recognize the child is not support; it is a new barrier. Before introducing voice or emotion recognition in kindergarten, the criterion is not product novelty but the disaggregated error rate for the group’s real voices, languages, and motilities, and the existence of a human repair channel. Without those data, the purchase is an act of faith, not of inclusion.
7. Inferential framework: when to support, when not to substitute
The framework that follows is a pedagogical inference of this article, anchored in the cases and in the verified instruments. It is not a new international standard. It distinguishes four tests. If a tool does not pass them, it is not deployed as “inclusion.”
7.1. Test of right, not of product. The first question is not “what can the model do?” but “which right is realized and which is put at risk?” Communication, accessibility, inclusive education, non-discrimination, privacy, and participation are enforceable rights (United Nations, 2006, 1989; Committee on the Rights of the Child, 2021; UNICEF, 2021b). A gaze communicator that allows choosing play, with GAS goals and service, can sit on the side of Articles 21 and 24. An autism classifier that labels without consent or a route of challenge sits on the side of the profiling UNICEF rejects. Assistive technology in the Global report is a means to rights, not an end; its scarcity in low-income countries forbids celebrating “inclusive AI” that exists only in trials of the North (World Health Organization & UNICEF, 2022; UNESCO, 2021b).
7.2. Test of irreplaceable human mediation. Hsieh et al. (2024) show that the device without service does not change the baseline of those who already had a tracker. Borgestig et al. (2021) integrate the team of an assistive-technology center. Costanzo et al. (2023) depend on calibration, track cutting, a speech-language pathologist, and a caregiver; school use fails. Holyfield et al. (2024) operate with a wizard of Oz: the “system” is a human. UNICEF requires a human in the loop for decisions that affect the child’s life (UNICEF, 2021b). UNESCO requires human oversight (UNESCO, 2021a). Inference: to support is to expand the capacity of the specialist and the support teacher—more opportunities for augmented input, more useful gaze time, less mechanical board load. To substitute is to dismiss, leave posts unfilled, or delegate diagnosis and conflict mediation. The second use is not authorized by the evidence reviewed here.
7.3. Test of disaggregated bias. Feng et al. (2024) and Koenecke et al. (2020) show that error is not uniform. Guo et al. (2019) and Trewin et al. (2019) predict disability-specific exclusion. A voice or vision system that does not publish error rates by age, language, accent, and speech condition cannot be declared support. Extreme personalization (Talkitt) mitigates bias at the cost of a closed vocabulary and sensitive voice data: that trade-off must be informed, reversible, and proportional (UNICEF, 2021b). Children’s voices are not uploaded to general-purpose generative models (Miao & Holmes, 2023).
7.4. Test of non-medicalization. Valizadeh et al. (2024) document a detection field with high metrics and fragile method. Turning that literature into a kindergarten protocol is to medicalize difference and advance labels with placement consequences. General comment No. 4 requires supports in the general system, not screening that separates (Committee on the Rights of Persons with Disabilities, 2016). Clinical detection, when it is due, is an act of a responsible team, with consent and a right of reply. Kindergarten observes, documents, and refers; it does not classify with a laboratory AUC.
Operationally, the framework admits: (a) personalized gaze or speech AAC, with service, goals, and usability review, for children who cannot access by other methods; (b) augmented-input prototypes with a specialist in the loop and continuous evaluation; (c) error audit of any recognizer before purchase. The framework rejects: (d) automated diagnosis or “ASD risk” as an educational decision; (e) generic voice assistants as the sole channel of participation; (f) replacement of speech-language pathologist, occupational therapist, interpreter, or support teacher by a licence; (g) independent conversation of a kindergarten child with a generative model (Miao & Holmes, 2023).
8. Discussion
Three tensions organize the discussion. The first is between communicative access and the myth of automatic inclusion. It is a finding that, with service, gaze technology can increase play, choice, and some tasks in three- to six-year-olds with complex disabilities (Hsieh et al., 2024), and that a calibrated recognizer may be associated with more naming in a Down-syndrome sample (Costanzo et al., 2023). It is a framework that Articles 9, 21, and 24 require facilitating augmentative modes and individualized supports (United Nations, 2006). It is not a finding that those devices accomplish inclusion. Costanzo et al. (2023) in fact measure the failure of school use. Hsieh et al. (2024) measure participation in computer activities, not belonging to the peer group. Confusing usability with inclusion is the categorical error that General comment No. 4 corrects (Committee on the Rights of Persons with Disabilities, 2016).
The second tension is between early detection and medicalization. The AI-for-autism field offers AUCs that, read uncritically, look like a solution to waiting lists. Valizadeh et al. (2024) read those same figures as a biased corpus. UNICEF describes profiling as a threat to development (UNICEF, 2021b). In early childhood, where play, language, and attention are still unstable, a false positive or false negative is not a laboratory error: it is a trajectory. This article’s inference is that the clinical urgency to evaluate does not translate into an urgency to automate.
The third tension is between declared accessibility and measured bias. The same systems sold to “give voice” recognize children, accents, and racialized speakers worse (Feng et al., 2024; Koenecke et al., 2020). Guo et al. (2019) anticipate failure with atypical speech. The inclusion by design that UNICEF (2021b) asks for requires diverse data and, at the same time, UNICEF and Trewin et al. (2019) recall the privacy cost of those data. There is no shortcut: either one personalizes with strict governance and a specialist, or one does not deploy. The third path—a general model “for all children”—is, on the bias evidence, a path of exclusion.
If the device requires calibration, positioning, augmented-language modeling, and error repair, the specialist is not a cost to be optimized: it is the condition of support. Substituting that person is incompatible with Article 24.4 (United Nations, 2006) and with the requirement not to leave vital decisions to a system without a human (UNICEF, 2021b). In the best documented case, AI expands that professional’s reach. It does not replace them.
9. Limits
This review is narrative. It does not apply a full PRISMA protocol or estimate combined effects. The nuclear support cases have small samples (n = 4 in Hsieh et al., 2024; n = 23 completers in Costanzo et al., 2023; n = 3 in Holyfield et al., 2024) and biased geography (Taiwan, Italy, the United States, Europe). Costanzo et al. (2023) mix ages and declare a commercial conflict. Feng et al. (2024) measure children aged 7–11, not 3–5; transfer to kindergarten is a risk inference. Koenecke et al. (2020) measure adults. Valizadeh et al. (2024) aggregate heterogeneous modalities and are not equivalent to a classroom trial. Latin American or African trials of AI AAC in early childhood were not located with the same degree of DOI openness; that absence is a gap, not proof of non-existence. The frameworks of the Convention, UNICEF, and UNESCO are prescriptive: their authority is normative, not empirical. The pedagogical inferences of section 7 are hypotheses of education policy, not implementation evidence.
10. Conclusions
Artificial intelligence does not include early childhood. It can, under mediated and auditable conditions, expand communicative access for girls and boys with complex disabilities: that is what gaze technology in four children aged three to six measures, with modesty and human service (Hsieh et al., 2024) and, with more caveats of age and conflict of interest, a calibrated speech communicator (Costanzo et al., 2023). It can, at the same time, medicalize when the automated-autism-detection literature—high in AUC, fragile in method (Valizadeh et al., 2024)—is taken as a licence to label in kindergarten. It can bias when speech recognition, already unequal for children, accents, and racialized speakers (Feng et al., 2024; Koenecke et al., 2020), is presented as a universal channel. It can displace the specialist when device purchase is offered as a substitute for the speech-language pathologist, the therapist, or the support teacher, against what the usability studies themselves show: without a team, school use collapses and the baseline does not move.
Standing law is not ambiguous. The Convention requires an inclusive system, reasonable accommodation, augmentative modes, and trained professionals, not a classifier (United Nations, 2006; Committee on the Rights of Persons with Disabilities, 2016). UNICEF requires inclusion, non-discrimination, privacy, and a human in the loop (UNICEF, 2021b). UNESCO requires ethics, human oversight and, in generative AI, age limits and pedagogical validation (UNESCO, 2021a; Miao & Holmes, 2023). Where the sources do not measure inclusion, development, cure, or valid clinical diagnosis, this article does not assert them. To support is to sustain communication and participation with assistive technology and with specialists. To substitute is another policy, and it is not justified.
Ingeniero Mitre / Laboratorio Editorial de NEXTECH.IA
References
- Borgestig, M., Al Khatib, I., Masayko, S., & Hemmingsson, H. (2021). The impact of eye-gaze controlled computer on communication and functional independence in children and young people with complex needs: A multicenter intervention study. Developmental Neurorehabilitation, 24(8), 511–524. https://doi.org/10.1080/17518423.2021.1903603
- Committee on the Rights of the Child. (2021). General comment No. 25 (2021) on children’s rights in relation to the digital environment (CRC/C/GC/25). United Nations. https://www.ohchr.org/en/documents/general-comments-and-recommendations/general-comment-no-25-2021-childrens-rights-relation
- Committee on the Rights of Persons with Disabilities. (2016). General comment No. 4 (2016) on the right to inclusive education (CRPD/C/GC/4). United Nations. https://www.ohchr.org/en/documents/general-comments-and-recommendations/general-comment-no-4-article-24-right-inclusive
- Committee on the Rights of Persons with Disabilities. (2018). General comment No. 7 (2018) on the participation of persons with disabilities, including children with disabilities, through their representative organizations, in the implementation and monitoring of the Convention (CRPD/C/GC/7). United Nations. https://digitallibrary.un.org/record/3899396
- Costanzo, F., Fucà, E., Caciolo, C., Ruà, D., Smolley, S., Weissberg, D., & Vicari, S. (2023). Talkitt: Toward a new instrument based on artificial intelligence for augmentative and alternative communication in children with Down syndrome. Frontiers in Psychology, 14, 1176683. https://doi.org/10.3389/fpsyg.2023.1176683
- Feng, S., Halpern, B. M., Kudina, O., & Scharenborg, O. (2024). Towards inclusive automatic speech recognition. Computer Speech & Language, 84, 101567. https://doi.org/10.1016/j.csl.2023.101567
- Guo, A., Kamar, E., Wortman Vaughan, J., Wallach, H., & Ringel Morris, M. (2019). Toward fairness in AI for people with disabilities: A research roadmap [Workshop paper]. ACM ASSETS 2019 Workshop on AI Fairness for People with Disabilities. https://doi.org/10.48550/arXiv.1907.02227
- Holyfield, C., MacNeil, S., Caldwell, N., Zimmerman, T. O., Lorah, E. R., Dragut, E., & Vucetic, S. (2024). Leveraging communication partner speech to automate augmented input for children on the autism spectrum who are minimally verbal: Prototype development and preliminary efficacy investigation. American Journal of Speech-Language Pathology, 33(3), 1174–1192. https://doi.org/10.1044/2023_AJSLP-23-00224
- Hsieh, Y.-H., Granlund, M., Odom, S. L., Hwang, A.-W., & Hemmingsson, H. (2024). Increasing participation in computer activities using eye-gaze assistive technology for children with complex needs. Disability and Rehabilitation: Assistive Technology, 19(2), 492–505. https://doi.org/10.1080/17483107.2022.2099988
- Koenecke, A., Nam, A., Lake, E., Nudell, J., Quartey, M., Mengesha, Z., Toups, C., Rickford, J. R., Jurafsky, D., & Goel, S. (2020). Racial disparities in automated speech recognition. Proceedings of the National Academy of Sciences, 117(14), 7684–7689. https://doi.org/10.1073/pnas.1915768117
- Miao, F., & Holmes, W. (2023). Guidance for generative AI in education and research. UNESCO. https://doi.org/10.54675/EWZM9535
- United Nations. (1989). Convention on the Rights of the Child. https://www.ohchr.org/en/instruments-mechanisms/instruments/convention-rights-child
- United Nations. (2006). Convention on the Rights of Persons with Disabilities. https://www.ohchr.org/en/instruments-mechanisms/instruments/convention-rights-persons-disabilities
- Trewin, S., Basson, S., Muller, M., Branham, S., Treviranus, J., Gruen, D., Hebert, D., Lyckowski, N., & Manser, E. (2019). Considerations for AI fairness for people with disabilities. AI Matters, 5(3), 40–63. https://doi.org/10.1145/3362077.3362086
- UNESCO. (2021a). Recommendation on the ethics of artificial intelligence. UNESCO. https://unesdoc.unesco.org/ark:/48223/pf0000381137
- UNESCO. (2021b). Inclusive early childhood care and education: From commitment to action. UNESCO. https://www.unesco.org/en/articles/inclusive-early-childhood-care-and-education-commitment-action
- UNICEF. (2021a). Seen, counted, included: Using data to shed light on the well-being of children with disabilities. UNICEF. https://data.unicef.org/resources/children-with-disabilities-report-2021/
- UNICEF. (2021b). Policy guidance on AI for children 2.0. UNICEF Office of Global Insight and Policy. https://www.unicef.org/innocenti/reports/policy-guidance-ai-children
- Valizadeh, A., Moassefi, M., Nakhostin-Ansari, A., Heidari Some’eh, S., Hosseini-Asl, H., Saghab Torbati, M., Aghajani, R., Maleki Ghorbani, Z., Menbari-Oskouie, I., Aghajani, F., Mirzamohamadi, A., Ghafouri, M., Faghani, S., & Memari, A. H. (2024). Automated diagnosis of autism with artificial intelligence: State of the art. Reviews in the Neurosciences, 35(2), 141–163. https://doi.org/10.1515/revneuro-2023-0050
- World Health Organization, & UNICEF. (2022). Global report on assistive technology. World Health Organization and UNICEF. https://www.who.int/publications/i/item/9789240049451