1. Introduction and problem

In early childhood and basic-education classrooms a promise of apparent technical innocence circulates: artificial intelligence “understands” the child, “speaks” their language, and “builds literacy” in twenty-first-century competencies. That promise usually treats language as a transparent channel. It is not. A language model or a speech recognizer does not “have a language”: it has a training distribution. If that distribution is dominated by English —or by a handful of high-resource languages— the system is not a neutral mediator. It is a filter that rewards certain varieties and degrades others. In early childhood and the first cycle, that filter coincides with the window in which a home language is consolidated, literacy begins, and it is decided which language “counts” as school language.

The problem is not hostility to English nor nostalgia for a pre-digital classroom. It is confusion among three objects. The first is an empirical finding: what is observed when a language model is evaluated on a parallel benchmark of 122 variants, when an ASR system is trained on a few hours of Quechua, Guarani, Bribri, Kotiria, or Wa’ikhana, or when NLP systems are asked to generate grammar exercises in Maya, Bribri, or Guarani. The second is a normative framework: what UNESCO’s Recommendation on the ethics of AI, UNESCO’s guidance on multilingual education, UNICEF, the Committee on the Rights of the Child, and the International Decade of Indigenous Languages require. The third is a pedagogical inference: what should not be deployed in the early childhood and basic-education classroom even if marketing says “multilingual,” “inclusive,” or “AI literacy.” Mixing them produces a pedagogy of English by omission: the school “uses AI” and, without measuring it, turns English —or a standard Spanish of web corpora— into the language of digital competence.

This article’s thesis is restrictive. AI is not linguistically neutral. In early childhood and basic education, a system trained mainly in English can widen language inequality. Joshi, Santy, Budhiraja, Bali, and Choudhury (2020) questioned the “language agnostic” status of NLP systems. Blasi, Anastasopoulos, and Neubig (2022) quantified systematic inequalities in technological utility. Bender, Gebru, McMillan-Major, and Shmitchell (2021) warned that ever-larger language models have been developed mainly for English. None of that equals a classroom trial. Nor does it equal the promise, frequent in education policy, that “AI closes gaps.” Where a source does not measure closure of a language gap, this article does not assert it. Where it does not measure revitalization of an Indigenous language, it does not promise it. The vacuum that organizes the work is that of an AI literacy that, in fact, only works —or only works well— in English.

This article does not address classroom privacy or datafication (21 August, 1:02), nor UNESCO/DAILy/PrimaryAI curricular literacy (20 August evening), nor inclusion by disability (20 August, 5:00), nor UNESCO teacher competence (19 August), nor parental mediation (19 extra), nor play or PopBots (18 August), nor tutors or adaptive assessment (16 August), nor gaps in “learning AI” (08-05). The object is language: model and ASR bias, minoritized languages —including Indigenous languages of Latin America when sources allow—, Spanish versus English, and the risk of an AI literacy that operates only in English.

2. State of the art: linguistic bias, resources, and the classroom as a language space

Three strata should be kept apart. The first is conceptual: what it means that an NLP or speech system is not linguistically neutral. The second is normative: which frameworks fix multilingualism, non-discrimination, and the right to learn in a language one understands. The third is empirical: what has been measured in models, recognizers, and educational-resource tasks, not in a trial of “AI that saves languages.”

In the conceptual stratum, Joshi et al. (2020) classified the world’s languages by NLP resource density and showed that a tiny fraction concentrates data, tools, and publications; the finding casts doubt on current models being language-independent. Bender et al. (2021) moved the argument to the scale of language models: English dominates development and data documentation is insufficient. Blasi et al. (2022) estimated the global utility of language technologies and concluded that the last decade’s progress is restricted to a minuscule subset of the world’s approximately 6,500 languages. Tonja, Balouchzahi, Butt, Kolesnikova, Ceballos, Gelbukh, and Solorio (2024) located that inequality in Latin America: Mexico recognizes 68 Indigenous languages, but only about 22 appear in NLP research; in Peru, of more than 70 languages, only 12. The rise in publications since 2021 coincides with the AmericasNLP workshop; it does not coincide with school coverage or with a classroom-measured closing of the gap.

In the normative stratum, UNESCO’s (2021) Recommendation on the ethics of AI requires that AI benefits be accessible to all, taking into account, among others, different language groups, and that States promote inclusive access to AI systems with locally relevant content and services, respecting multilingualism and cultural diversity. UNESCO (2003) had already recommended promoting multilingualism and universal access to cyberspace. UNESCO (2025), in Languages matter, holds that children should learn in a language they understand and that mother-tongue-based education, sustained for several years, is a condition of quality and inclusion; the associated GEM brief insists on home-language instruction for six to eight years (UNESCO, 2025b). Those claims are education policy, not evaluation of a commercial model. The global roadmap for multilingualism in the digital era (UNESCO, 2024) articulates language technologies with the International Decade of Indigenous Languages (2022–2032) and places data governance and technology deployment under community-driven decision-making. UNICEF (2021) requires inclusion by design. The Committee on the Rights of the Child (2021) affirms that Convention rights apply in the digital environment and that information must be accessible. Miao and Holmes (2023) warn that the speed of generative AI leaves users and education systems unprotected without regulation. None of these instruments measures an LLM’s performance in Spanish or Quechua; they prescribe a threshold of linguistic equity that products do not demonstrate by advertising themselves as “multilingual.”

In the empirical stratum, recent evidence concentrates on evaluation benchmarks and low-resource speech systems, not early-childhood classrooms. Bandarkar et al. (2024) built BELEBELE, a parallel reading-comprehension set in 122 variants. Li, Shi, Liu, Yang, Payani, Liu, and Du (2025) proposed Language Ranker and found a correlation between an LLM’s performance in a language and that language’s share of the pretraining corpus. Laurençon et al. (2022) documented ROOTS, BLOOM’s corpus: English accounts for 30.03%; Spanish, 10.85%; Indigenous languages of the Americas do not appear among the 46 natural languages. Romero et al. (2024) reported ASR for five Indigenous languages with hours of audio that, in several cases, do not reach a dozen. Chiruzzo et al. (2024) organized the first shared task on automatic educational materials for Bribri, Maya, and Guarani. The gap is not the absence of norms or of engineering papers: it is the distance between what is measured (character error, multiple-choice accuracy, generation of a grammatical variant) and what schools usually claim (equity, revitalization, AI literacy for every child).

3. Review method

A critical narrative review was conducted, not a meta-analysis. The purpose was not to estimate a homogeneous effect size of “AI and linguistic equity” —non-existent across a reading-comprehension benchmark, an ASR fine-tuning, and a shared exercise task— but to articulate the risk of English-only AI literacy with verified sources. Inclusion criteria were: (a) focus on linguistic bias of models or ASR, on minoritized or Indigenous languages, or on Spanish versus English, with pedagogical translation to early childhood and basic education; (b) preferred publication 2021–2026, with justified exception for still-binding frameworks (Joshi et al., 2020; UNESCO, 2003; Bender, 2011 cited through Bandarkar et al., 2024 and Joshi et al., 2020); (c) documentary type of journal or proceedings with DOI, agency report on UNESDOC or an official page, or recommendation text; (d) access to a DOI, publisher page, or official URL confirming authors, year, title, and scope. Excluded as nuclear cases were classroom privacy, curricular literacy as the axis, inclusion by disability, UNESCO teacher competence, parental mediation, PopBots/genAI play, tutors, and adaptive assessment. Koenecke, Nam, Lake, Nudell, Quartey, Mengesha, Toups, Rickford, Jurafsky, and Goel (2020) is used only as saturation of variety within English (African American English versus white English in commercial ASR), not as a duplicated nuclear case from a prior inclusion article.

The search was executed on 21 August 2026 (around 05:04, America/Mexico_City) in ACL Anthology, AAAI, JMLR, MDPI, PMLR, NeurIPS, UNESDOC, UNICEF Innocenti, OHCHR, PNAS, and DOI pages. Each cited source was verified against at least one of those pages. The corpus was organized into three nuclear cases: the LLM English/other-language gap; ASR for Indigenous languages of the Americas; generation of educational materials for Indigenous languages. The analysis distinguished three enunciative statuses. Empirical finding: what was observed in the sample or benchmark. Normative framework: what an agency prescribes. Pedagogical inference: the translation to the early childhood and basic-education classroom, marked as such. Limits are those of any narrative review (section 9).

4. Case 1. Language models: English as the baseline and Spanish as a “high” language that still loses

Bandarkar et al. (2024) published in ACL a parallel multiple-choice reading-comprehension benchmark —BELEBELE— with 900 questions aligned across 122 language variants, 27 families, and 29 scripts. The design allows accuracy to be compared on the same task, not an anecdote that “the model speaks Spanish.” Empirical finding: in five-shot in-context learning, Llama 2 with 70 billion parameters reaches 90.9% accuracy in English and 47.7% average in non-English languages; GPT-3.5-turbo, zero-shot with English instructions, reaches 87.7% in English and 50.7% non-English average. The model exceeds 50% accuracy in only 38.5% of languages (Llama 2) or 44.2% (GPT-3.5). Additional empirical finding, read from the same article’s tables: in Spanish (spa_Latn) GPT-3.5 scores 79.2, Llama-2-chat 70B 68.4, and Llama 2 70B 85.0, all below their respective English figures (87.7; 78.8; 90.9). In Guarani (grn_Latn) the figures fall to 34.2, 32.2, and 32.4, near chance (25%). GPT-3.5 performs best on the first twenty high-resource languages and, past the threshold of forty or fifty, falls behind smaller masked models trained on more balanced multilingual data (InfoXLM, XLM-V). Translate-Test —translate into English and then ask— outperforms, in 68 of 91 evaluated languages, Llama-2-chat 70B’s in situ comprehension. That is not a pedagogical success: it is evidence that the model “understands” better when the world is translated into English for it.

Li et al. (2025) published in AAAI Language Ranker, an intrinsic metric that compares internal language representations against an English baseline. Empirical finding: high-resource languages show higher cosine similarity with English; low-resource languages, lower. There is a strong correlation between that similarity and the language’s share of the pretraining corpus. In Llama 2, German (0.17% of the corpus, similarity 0.723) and French (0.16%, 0.737) outperform Kannada and Urdu (≤ 0.01%; similarities 0.236 and 0.275). The article does not evaluate classrooms; it evaluates representations. What the finding authorizes is a narrow claim: English is not one more language in the model’s internal space; it is the origin of the axis.

Laurençon et al. (2022) documented ROOTS, 1.6 TB in 59 languages (46 natural and 13 programming), used to train BLOOM. Empirical finding of composition: English 30.03%, Simplified Chinese 16.16%, French 12.9%, Spanish 10.85%, Portuguese 4.91%, Arabic 4.6%. BigScience Workshop (2024) describes BLOOM as an open 176-billion-parameter model trained on those 46 natural languages. Spanish is among the “privileged” languages of a project presented as an Anglocentric alternative; even so, English remains the largest block and Indigenous languages of the Americas do not enter the list. Bandarkar et al. (2024) further recall that Llama 2 reports 89.7% English data, 8.4% unidentified, and 1.9% in 26 other languages. The distance between “multilingual model” and “model with 2% rest of the world” is a finding of corpus documentation, not a metaphor.

Status of the evidence. Accumulated empirical finding: on a parallel benchmark, English outperforms Spanish; Spanish, a language of hundreds of millions of speakers, does not match English; an Indigenous language of the Americas present in the benchmark (Guarani) approaches chance in English-centric LLMs. Composition finding: even BLOOM, deliberately multilingual, assigns English triple the mass of Spanish and zero mass to Quechua, Nahuatl, or Maya. It is not a finding that a chatbot “closes the gap” if asked to “speak Spanish.” It is not a finding of children’s learning. Pedagogical inference, marked as such: if a child in early childhood or basic education interacts with a system whose reading comprehension, reasoning, and system instructions are calibrated in English, the “AI literacy” offered is not the same as that offered to an Anglophone child. Spanish does not save that asymmetry; it mitigates it halfway and leaves out whoever arrives in the classroom in a minoritized language.

5. Case 2. ASR and Indigenous languages: few hours, high error, and English as saturation of variety

Ebrahimi, Mager, Wiemerslage, Denisov, Oncevay, and colleagues (2023) described the second AmericasNLP competition (NeurIPS 2022) on speech-to-text recognition and translation for five Indigenous languages: Bribri, Guarani, Kotiria, Wa’ikhana, and Quechua. The call text is explicit: those languages have received very little attention from the machine-learning and NLP communities, and the lack of systems is articulated with inequalities affecting their speakers. That is a problem frame, not a measure of school inequality. Romero, Gómez, and Torre (2024) published in Applied Sciences the winning ASR-subtask approach: fine-tuning Wav2vec 2.0 XLS-R (300 M and 1 B parameters), speed augmentation, and Bayesian hyperparameter search, with about 36.65 h of transcribed speech from diverse sources. Empirical finding of error: Quechua, CER 12.14 and WER 48.98 (about 12 h of training); Guarani, CER 15.59 and WER 62.91 (under 1 h); Bribri, CER 34.70 and WER 69.03; Wa’ikhana, CER 35.23 and WER 68.42; Kotiria, CER 36.59 and WER 79.69, despite being the language with the most hours (near 30). The reported average CER is 26.85. The authors release models —the first open ASR models for Wa’ikhana and Kotiria— and warn that language complexity, not only data volume, conditions the result. They do not measure classroom use, a child speaker’s comprehension, or language revitalization.

Tonja et al. (2024) place that result on a wider map: machine translation concentrates about 40% of NLP publications for Latin American Indigenous languages; ASR and morphological analyzers barely exceed 5%. Quechua, Nahuatl, Shipibo-Konibo, Bribri, Mapudungun, Aymara, and Wixárika appear with more than one task; most languages recognized by States do not appear. The finding is bibliometric: unequal visibility in the research corpus, not pedagogical efficacy.

As saturation —not as a nuclear case—, Koenecke et al. (2020) showed that five commercial ASR systems (Amazon, Apple, Google, IBM, Microsoft) transcribed sociolinguistic interviews with average WER of 0.35 for Black speakers and 0.19 for white speakers, in U.S. English, with 19.8 matched hours. The gap was attributed to acoustic models and underrepresentation of African American English. It is cited only to anchor a variety thesis: bias does not begin when one leaves English; it already operates inside English when the variety does not match the training standard. Pedagogical inference, marked as such: a Spanish-speaking early-childhood child with code-switching, a child who speaks an Andean or Caribbean variety, or a child whose home language is not the standard Spanish of web corpora, cannot be counted as “covered” because a product lists “Spanish” in a menu. Dhawan, Rekesh, and Ginsburg (2023) document, for English–Spanish and English–Hindi, that code-switching ASR requires specific tokenization and language identification and that natural alternation data are scarce. That is not a classroom trial either; it is evidence that real bilingualism is not the sum of two catalogue monolingualisms.

Status of the evidence. Empirical finding: open ASR systems exist for five Indigenous languages of the Americas, with two-digit CER and WER that, in four of five languages, exceeds 60%. Saturation empirical finding: inside English, racialized variety already doubles word error. It is not a finding that ASR “preserves” or “rescues” a language. Romero et al. (2024) open a research path; they do not measure revitalization. Pedagogical inference: placing a classroom microphone or a “multilingual” voice assistant in front of an early-childhood child whose language is not on the menu —or is there with WER of 50 to 80%— is not linguistic inclusion. It is an intelligibility test the child did not choose and that the system, with high probability, fails. The teacher who interprets the system’s silence as the child’s silence commits a pedagogical error induced by the model.

6. Case 3. Automatic educational materials: Maya, Bribri, Guarani, and what the task does not measure

Chiruzzo, Denisov, Molina-Villegas, Fernandez-Sabido, Coto-Solano, and colleagues (2024) published the results of the first NLP shared task on creating educational materials for Indigenous languages of the Americas, at AmericasNLP 2024 (Mexico City). The task is not a classroom RCT nor a revitalization program. It consists in transforming a source sentence into a target sentence by changing a linguistic feature —usually associated with the verb: negation, aspect, tense— to produce pairs usable as grammar exercises. Languages: Bribri, Maya, and Guarani. Seven teams, 22 systems. The organizers describe the results as “very promising” in terms of the task metrics, and observe that large-scale models, often with some tuning, can work when they are not asked to generate a full sentence from scratch but to introduce localized changes. Empirical finding: it is possible, in a competition setting, to generate grammatical variations in three Indigenous languages with neural systems and LLMs. Design finding: the unit of success is linguistic transformation, not children’s comprehension, not teacher support, not mother-tongue time in class, not linguistic self-esteem, not the reduction of repetition or dropout that UNESCO (2025) associates —as a policy frame, not as an evaluation of these systems— with home-language education.

This case is retained precisely because it touches education and must not be inflated. A generated grammar exercise is not a class. An LLM that correctly changes aspect in Maya has not taught Maya to a six-year-old. The task shares the risk of turning an engineering achievement into a salvation narrative. Tonja et al. (2024) insist that progress must respect Indigenous community perspectives; UNESCO (2024) places community decision-making at the center of the long-term vision for language technologies. Those are prescriptions. Chiruzzo et al. (2024) do not measure them as outcomes. Nor do they measure harm: they do not report whether exercises contain errors a non-native teacher would miss, or whether the LLM displaces speakers as authors.

Status of the evidence. Empirical finding: there is a precedent of a shared task on automatic educational materials in Bribri, Maya, and Guarani, with international participation and generation metrics. It is not a finding of learning. It is not a finding that AI “supports intercultural bilingual education.” Pedagogical inference, marked as such: in early childhood education, the material that matters is not only the minimal sentence pair, but the voice, play, the family’s language, and the presence of an adult who speaks it. A generator can, at best, assist a teacher who already possesses the language. It does not replace that competence. Presenting it as AI literacy for children who speak minoritized languages inverts the burden: one literates the model in the child’s language, or one pretends to literate the child in a model that does not have it.

7. Inferential framework: five tests against English-only literacy

The framework that follows is this article’s pedagogical inference, anchored in the cases and in the verified instruments. It is not a new international standard. It distinguishes five admissibility tests. If a proposal of “AI in the classroom” of early childhood or basic education does not pass them, it is not deployed as literacy.

7.1. Home-language test. UNESCO (2025) holds that learning in a language one understands is a right and a condition of quality; UNICEF (2021) requires inclusion by design; the Committee on the Rights of the Child (2021) requires accessibility of digital information. Inference: a system that does not operate, or operates poorly, in the child’s language is not a literacy resource; it is a covert admissions filter. In early childhood, the home language is not “content” substitutable by English or by a corpus Spanish. Digital competence is not feigned where there is linguistic incomprehension by the system.

7.2. Distribution test, not menu test. Li et al. (2025), Laurençon et al. (2022), and Llama 2 documentation cited by Bandarkar et al. (2024) show that performance follows pretraining mass. A menu that lists “Spanish” or “Quechua” is not an evaluation. Inference: demand, before purchase or recommendation, error or accuracy figures in the classroom’s real variety —not in benchmark English nor in a Peninsular Spanish demo. If the vendor does not provide them, the system is not ready for children.

7.3. Variety and speech test. Romero et al. (2024) document WER of 49 to 80% on Indigenous languages with competition ASR; Koenecke et al. (2020) document, as saturation, a racial gap inside English; Dhawan et al. (2023) document the difficulty of Spanish–English code-switching. Inference: early-childhood speech —incomplete, dialectal, often alternating— is a harder case than an adult reading a corpus. A voice assistant that does not recognize the child does not “fail a little”; it teaches that their mouth is unreadable to the machine the school has authorized.

7.4. Non-salvation test. Chiruzzo et al. (2024) measure exercise generation, not revitalization. Tonja et al. (2024) measure publications, not vitality. UNESCO (2024) prescribes community governance; it does not certify products. Inference: it is forbidden, in this framework, to claim that AI closes language gaps or rescues Indigenous languages. That sentence, if it appears in a school project, is marketing, not evidence. What is admissible is a narrow use, with speakers, with a right to refuse the data, and without a salvation narrative.

7.5. Literacy in the classroom language test. Miao and Holmes (2023) and UNESCO (2021) require a human-centered approach and linguistic access. An “AI literacy” whose examples, interfaces, datasets, and success criteria are in English produces, by design, a second-order competence: the child must know English in order to learn to query the machine. Inference: that is not AI literacy; it is English mediated by a chatbot. In Latin American early childhood and basic education, that inversion is unacceptable as an equity policy. Operationally, the framework admits: (a) public evaluation of error in the classroom’s language and variety; (b) native materials and prompts, not on-the-fly translation into English; (c) teachers and speaker communities as authors, not as end users of an alien model; (d) the right not to use the system without losing the school activity. It rejects: (e) voice assistants unmeasured in the child’s language; (f) AI literacy available only in English; (g) the fiction that “Spanish is on the menu” equals English; (h) the narrative that an ASR with CER 12 or 36 “preserves” a language.

8. Discussion

Three tensions organize the discussion. The first is between Spanish as a “large” language and Spanish as a language subordinate to English in the model’s space. BELEBELE shows that Spanish scores below English on the same items (Bandarkar et al., 2024). ROOTS shows that, even in a multilingual project, English weighs three times as much (Laurençon et al., 2022). Inference: education policies that treat Spanish as sufficient against Anglocentric bias underestimate the gap. A Spanish-monolingual child in basic education does not receive the same system as an Anglophone child; they receive an approximation. Calling that approximation “AI in Spanish” without figures is opacity.

The second tension is between Indigenous languages as objects of papers and as classroom languages. AmericasNLP, Romero et al. (2024), and Chiruzzo et al. (2024) are advances of engineering and of a scientific community. Tonja et al. (2024) show that most recognized languages remain outside. UNESCO (2025) prescribes mother-tongue education; it does not evaluate those systems. The leap from a laboratory CER to a school language right is political, not technical. Whoever presents it as already solved invents a result.

The third tension is between AI literacy and English literacy. If the interface, the example corpus, and the criterion of a “good question to the model” are in English, the competence being taught is the machine’s English. Bandarkar et al. (2024) found that translating into English improves Llama-2-chat’s performance in most evaluated languages. That is an engineering recipe and, at the same time, a pedagogical diagnosis: the institutional shortcut will be to translate the child toward English so that AI “works.” In early childhood, that shortcut colonizes the home language in the name of a competence the child does not yet have.

Miao and Holmes (2023) warn of lack of protection given the speed of generative AI. UNESCO (2021) requires linguistic non-discrimination. The Committee on the Rights of the Child (2021) requires an accessible digital environment. Those frameworks are not met because a vendor offers ten languages in a dropdown. They are met, if they are met, when error in the child’s language is known, acceptable, and governed by the school and the community —not by the model.

9. Limits

This review is narrative. It does not apply a full PRISMA protocol nor estimate combined effects on learning or on linguistic vitality. BELEBELE is not a children’s test: its passages come from FLORES-200 (news, Wikijunior, WikiVoyage) and the questions were created in English and translated; Bandarkar et al. (2024) warn of translationese and Anglocentric design bias. Language Ranker measures representation similarity, not pedagogy. ROOTS and BLOOM document a corpus and a model, not a classroom. Romero et al. (2024) measure CER and WER under competition conditions, not intelligibility of spontaneous child speech. Chiruzzo et al. (2024) measure exercise generation, not student achievement. Koenecke et al. (2020) is saturation of variety in adult U.S. English, not early-childhood or Spanish evidence. Dhawan et al. (2023) is technical code-switching ASR, not school ethnography. The empirical geography of LLMs is biased toward global benchmarks; that of Indigenous languages, toward NLP workshops. No Latin American early-childhood classroom trial measuring AI literacy in Nahuatl, Maya, or Quechua was located with the same degree of verifiable openness. That absence is a gap, not proof of non-existence or of success. UNESCO, UNICEF, and Committee on the Rights of the Child frameworks are prescriptive. Section 7 inferences are hypotheses of an ethical-pedagogical threshold, not evidence of national implementation. This article does not claim that all AI in a non-English language is useless; it claims that, in the verified corpus, neither closure of language gaps nor rescue of Indigenous languages was demonstrated.

10. Conclusions

Artificial intelligence in early childhood and basic education is not linguistically neutral. A system trained mainly in English can widen language inequality in the window in which a mother tongue is consolidated. Three families of verified evidence support the argument. BELEBELE and Language Ranker show that performance follows English as the baseline and that Spanish, a high-resource language, still loses, while an Indigenous language present in the benchmark approaches chance (Bandarkar et al., 2024; Li et al., 2025). ROOTS and BLOOM show that a deliberately multilingual project still assigns English the largest mass and does not include Indigenous languages of the Americas (Laurençon et al., 2022; BigScience Workshop, 2024). AmericasNLP ASR shows word error of 49 to 80% with few hours of audio (Romero et al., 2024; Ebrahimi et al., 2023). The educational-materials task shows that exercises can be generated in Bribri, Maya, and Guarani, not that children learn better or that languages are revitalized (Chiruzzo et al., 2024).

Law and policy are not silent. UNESCO (2021, 2025) requires multilingualism, non-discrimination, and learning in a language one understands. UNICEF (2021) requires inclusion by design. The Committee on the Rights of the Child (2021) requires digital accessibility. Miao and Holmes (2023) mandate a human-centered approach to generative AI. UNESCO (2024) places communities in the governance of language technologies. Where sources do not measure gap closure or language rescue, this article does not invent it. The threshold is pedagogical before it is commercial: a competence that only works in English is not called AI literacy.

Ingeniero Mitre / Laboratorio Editorial de NEXTECH.IA

References

  1. Bandarkar, L., Liang, D., Muller, B., Artetxe, M., Shukla, S. N., Husa, D., Goyal, N., Krishnan, A., Zettlemoyer, L., and Khabsa, M. (2024). The Belebele benchmark: A parallel reading comprehension dataset in 122 language variants. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) (pp. 749–775). Association for Computational Linguistics. https://doi.org/10.18653/v1/2024.acl-long.44
  2. Bender, E. M., Gebru, T., McMillan-Major, A., and Shmitchell, S. (2021). On the dangers of stochastic parrots: Can language models be too big? In Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency (pp. 610–623). ACM. https://doi.org/10.1145/3442188.3445922
  3. BigScience Workshop. (2024). BLOOM: A 176B-parameter open-access multilingual language model. Journal of Machine Learning Research, 25(422), 1–74. https://jmlr.org/papers/v25/23-0581.html
  4. Blasi, D., Anastasopoulos, A., and Neubig, G. (2022). Systematic inequalities in language technology performance across the world’s languages. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) (pp. 5486–5505). Association for Computational Linguistics. https://doi.org/10.18653/v1/2022.acl-long.376
  5. Chiruzzo, L., Denisov, P., Molina-Villegas, A., Fernandez-Sabido, S., Coto-Solano, R., Agüero-Torales, M., Alvarez, A., Canul-Yah, S., Hau-Ucán, L., Ebrahimi, A., Pugh, R., Oncevay, A., Rijhwani, S., von der Wense, K., and Mager, M. (2024). Findings of the AmericasNLP 2024 Shared Task on the Creation of Educational Materials for Indigenous Languages. In Proceedings of the 4th Workshop on Natural Language Processing for Indigenous Languages of the Americas (AmericasNLP 2024) (pp. 224–235). Association for Computational Linguistics. https://doi.org/10.18653/v1/2024.americasnlp-1.27
  6. Committee on the Rights of the Child. (2021). General comment No. 25 (2021) on children’s rights in relation to the digital environment (CRC/C/GC/25). United Nations. https://www.ohchr.org/en/documents/general-comments-and-recommendations/general-comment-no-25-2021-childrens-rights-relation
  7. Dhawan, K., Rekesh, D., and Ginsburg, B. (2023). Unified model for code-switching speech recognition and language identification based on concatenated tokenizer. In Proceedings of the 6th Workshop on Computational Approaches to Linguistic Code-Switching (pp. 74–82). Association for Computational Linguistics. https://aclanthology.org/2023.calcs-1.7/
  8. Ebrahimi, A., Mager, M., Wiemerslage, A., Denisov, P., Oncevay, A., Liu, D., Koneru, S., Ugan, E. Y., Li, Z., Niehues, J., Romero, M., Torre, I. G., Alumäe, T., Kong, J., Polezhaev, S., Belousov, Y., Chen, W.-R., Sullivan, P., Adebara, I., … Kann, K. (2023). Findings of the Second AmericasNLP Competition on Speech-to-Text Translation. In Proceedings of the NeurIPS 2022 Competitions Track (PMLR 220, pp. 217–232). PMLR. https://proceedings.mlr.press/v220/ebrahimi23a.html
  9. Joshi, P., Santy, S., Budhiraja, A., Bali, K., and Choudhury, M. (2020). The state and fate of linguistic diversity and inclusion in the NLP world. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics (pp. 6282–6293). Association for Computational Linguistics. https://doi.org/10.18653/v1/2020.acl-main.560
  10. Koenecke, A., Nam, A., Lake, E., Nudell, J., Quartey, M., Mengesha, Z., Toups, C., Rickford, J. R., Jurafsky, D., and Goel, S. (2020). Racial disparities in automated speech recognition. Proceedings of the National Academy of Sciences, 117(14), 7684–7689. https://doi.org/10.1073/pnas.1915768117
  11. Laurençon, H., Saulnier, L., Wang, T., Akiki, C., Villanova del Moral, A., Le Scao, T., von Werra, L., Mou, C., González Ponferrada, E., Nguyen, H., Frohberg, J., Šaško, M., Lhoest, Q., McMillan-Major, A., Dupont, G., Biderman, S., Rogers, A., Ben Allal, L., De Toni, F., … Jernite, Y. (2022). The BigScience ROOTS corpus: A 1.6TB composite multilingual dataset. In Advances in Neural Information Processing Systems 35 (NeurIPS 2022) Datasets and Benchmarks Track. https://doi.org/10.52202/068431-2306
  12. Li, Z., Shi, Y., Liu, Z., Yang, F., Payani, A., Liu, N., and Du, M. (2025). Language Ranker: A metric for quantifying LLM performance across high and low-resource languages. Proceedings of the AAAI Conference on Artificial Intelligence, 39(27), 28186–28194. https://doi.org/10.1609/aaai.v39i27.35038
  13. Miao, F., and Holmes, W. (2023). Guidance for generative AI in education and research. UNESCO. https://doi.org/10.54675/EWZM9535
  14. Romero, M., Gómez, S., and Torre, I. G. (2024). Automatic speech recognition advancements for Indigenous languages of the Americas. Applied Sciences, 14(15), 6497. https://doi.org/10.3390/app14156497
  15. Tonja, A. L., Balouchzahi, F., Butt, S., Kolesnikova, O., Ceballos, H., Gelbukh, A., and Solorio, T. (2024). NLP progress in Indigenous Latin American languages. In Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers) (pp. 6972–6987). Association for Computational Linguistics. https://doi.org/10.18653/v1/2024.naacl-long.385
  16. UNESCO. (2003). Recommendation concerning the Promotion and Use of Multilingualism and Universal Access to Cyberspace. https://www.unesco.org/en/legal-affairs/recommendation-concerning-promotion-and-use-multilingualism-and-universal-access-cyberspace
  17. UNESCO. (2021). Recommendation on the ethics of artificial intelligence. https://unesdoc.unesco.org/ark:/48223/pf0000381137
  18. UNESCO. (2024). Global roadmap for multilingualism in the digital era: Advancing the role of language technologies. https://unesdoc.unesco.org/ark:/48223/pf0000393920
  19. UNESCO. (2025). Languages matter: Global guidance on multilingual education. https://unesdoc.unesco.org/ark:/48223/pf0000392477
  20. UNESCO. (2025b). The right multilingual education policies can unlock learning and inclusion (GEM Report advocacy brief). https://www.unesco.org/gem-report/en/publication/right-multilingual-education-policies-can-unlock-learning-and-inclusion
  21. UNICEF. (2021). Policy guidance on AI for children 2.0. UNICEF Office of Global Insight and Policy. https://www.unicef.org/innocenti/reports/policy-guidance-ai-children