Speech Recognition for Early Learning
New state-of-the-art open models for children's speech recognition, with roughly half the errors of leading open alternatives.
Voice is one of the most natural ways for kids to learn, explore, and show what they know. Yet today's automatic speech recognition (ASR) technology hardly understands them.
This work set out to close that gap. With support from the Gates Foundation, DrivenData assembled the largest dataset of transcribed English child speech for developing open ASR models, then ran On Top of Pasketti, an open challenge where experts from around the world competed to develop the best ASR models for kids' speech. The challenge resulted in models for two kinds of transcription: word-level, which captures which words a child said and supports use cases such as reading and reasoning assessments, and phonetic, which captures exactly how they said them and supports use cases such as screening for atypical speech development.
DrivenData then retrained the winning approaches on more data and released them as two open models, pasketti-word and pasketti-phonetic, whose performance by age is shown below.
This project represents a multi-year effort, bringing together contributions from education, data science, and research communities. This effort included:
-
Datasets, benchmarking, and competition design Oct 2024 – Jul 2025State-of-the-field research and benchmarking, a review of the child-speech data that already existed, acquiring and harmonizing what we could use, and designing the challenge to address the most pressing opportunities.
-
Phonetic and word transcription May 2025 – Jul 2026Based on the first phase, a supplemental investment funded annotation to address a critical lack of transcribed speech data, especially phonetic labels for child speech. Our transcription team produced 253 hours of newly annotated child speech, including 164 hours of phonetic labels.
-
Open competition Jan – Apr 2026Over the course of the live challenge, more than 800 participants made 2,000+ submissions across the modeling tracks and competed for $120,000 in prize money. The top-performing models reduced word error rates by half against the best existing open models for kids.
-
Public goods Apr – Sep 2026Our team streamlined and retrained the winning approaches on all the data, including audio too sensitive to share with solvers, and released the results: open models on Hugging Face, the new word-level and phonetic transcripts, and the winners' code.
Why children's speech is hard
Current automatic speech recognition (ASR) models transcribe adult speech well but struggle with children's speech.
Models trained on adult speech don't work for kids because child speech is fundamentally different. Kids have distinct vocal characteristics, different speech patterns, and make frequent speech shortcuts and pronunciation errors, all of which make understanding them difficult.
- Metathesis: elephant → ephelant
- Velar fronting: cup → tup
- Syllable deletion: banana → nana
- Errors combined: spaghetti → pasketti
This poor performance is a significant barrier to educational tools that support early learning. Addressing it requires assembling more representative data and developing better ASR models built for children's speech.
Models
Open, retrained speech recognition models for word- and phone-level transcription of children's speech.
The ASR models developed in this project were created in two phases:
- The competition surfaced the best modeling approaches. A global community with a wide variety of backgrounds explored the solution space to find the most effective open approaches, each evaluated against the same test set.
- Post-challenge retraining focused on the best final models for release. The winning approaches were streamlined to be easier to run, retrained on more data, and packaged with clear documentation and evaluation of results.
The result is two open-weight models on Hugging Face, one for each kind of transcription. Both are released under the BigCode OpenRAIL-M license, which allows commercial and noncommercial use subject to responsible-use restrictions, and each has a model card covering its training data, intended uses, evaluation, and known limitations:
- pasketti-word 🍝 Word-level transcription: which words the child said
- pasketti-phonetic 🍝 Phonetic (IPA) transcription: exactly how the child said it
Both final models build on the challenge's top models:
- pasketti-word is a full fine-tune of Qwen3-ASR-1.7B, trained on 372 hours of child speech from more than 3,600 children in clips of up to 20 seconds.
- pasketti-phonetic streamlines the challenge's large ensembles into a single WavLM Large model, trained on 116 hours of phonetically transcribed speech from about 1,700 children. A second, word-level output used only during training let it learn from an additional 371 hours of word-transcribed speech.
While the competition models were trained only on competition data, the retrained versions add a larger and more representative dataset, including sensitive data from realistic education populations that could not be shared with solvers, and data deliberately gathered from groups where current models perform worst, such as children with atypical speech development, non-white students, and students from low-income backgrounds. Each model was then tested on held-out data and checked for differences in performance by age, race, speech pathology, sex, setting, and source corpus.
Choosing the right model
The two models capture different outputs and therefore serve different purposes:
| How it sounded | Word transcript | Phonetic transcript |
|---|---|---|
| "Uh… eye doh-noh" | I don't know | ʔəː ʔɑi doʊ noʊ |
| "Eye-dohn-noh" | I don't know | ɑi doʊn noʊ |
| "Eye-dun-no" | I don't know | ɑi don no |
Word transcription tells you which words the child said. This supports educators and ed-tech developers with use cases such as:
- Reading fluency: needs an accurate word-for-word record to compare against the target text. The phonetic model can add mispronunciation detail.
- Language use assessment: measures vocabulary and syntax.
- Comprehension, problem solving, reasoning: scores the substance of an answer.
- Verbal interaction with tools: a voice interface needs to know what was said in order to respond.
Phonetic transcription tells you how the child said it. This further supports speech-language pathologists and clinicians with use cases such as:
- Speech screening: flags substitutions, omissions, and other production errors that word-level transcription smooths over. The word model helps confirm what was attempted.
- Early literacy screening: early indicators of reading difficulty often show up in phonological patterns, so phonetic transcripts may help flag children for follow-up.
- Supporting word-level transcription: as an interim output, a phonetic transcript records the sounds a child actually made, which may help downstream tools account for developmentally typical sound substitutions rather than mistaking them for errors.
How the final models perform
Both models improved on the leading open alternatives across every age and demographic group in their evaluation sets, demonstrating the importance of diverse training data.
- pasketti-word has a word error rate (WER) of 17%, about half that of KidWhisper (32%) and OpenAI Whisper (34%).
- pasketti-word also narrows the gap between white and non-white children, and improves on KidWhisper by 29% in classroom settings, which are especially difficult for models (and humans) because of background talk and noise.
- pasketti-phonetic has a phone error rate (PER) of 33%, compared with 59% for PhoneticXeus, cutting errors by more than a third in every age group.
- Children ages 3–4 remain the hardest to transcribe, but they are also where some of the biggest gains came compared with existing models.
Performance by age is shown at the top of this page. Full evaluation results for both models, including breakdowns by age, race, speech pathology, sex, and source corpus, are in the Pasketti model cards on Hugging Face.
As is, these models are usable, but should be tested and evaluated in the specific context and population they would be used in, given that performance will vary by age, dialect, setting, and task. They will be especially useful in contexts like literacy assessments, where a known target transcript is available to compare against, and can serve as strong starting points for fine-tuning to specific tasks, grade levels, or student populations. Where exact transcripts matter, outputs should be reviewed by a person.
For intended uses, out-of-scope uses, and known limitations, see the Pasketti model cards.
Data
The largest dataset of its kind for developing open ASR models for children's speech: 14 datasets identified, sourced, and augmented with new annotations to build the models above.
Data has been the biggest constraint to better ASR for children. Models that work well for adults learned from vast amounts of adult speech, and no comparable, representative dataset existed for kids. To assemble such a dataset, DrivenData reviewed the landscape of existing child-speech resources, then identified and sourced audio from public, research, and enterprise partners. Our transcription team added over 250 hours of new orthographic annotation and harmonized it with existing corpora into a common schema.
The result spans a wide range of children and settings: ages from preschoolers to young adults in clinical research data, typically developing kids alongside children with diagnosed or suspected speech and language differences, and dialects including African American English and General American English. The table below is a directory to each dataset — what it is, how much of it each model used, where to find the audio and the transcripts, and information on licensing outside the challenge.
Child speech data is sensitive and difficult to share. We are grateful for the work of the data providers who thoughtfully collected and provided the data used in this project under appropriate consents, and to the speakers represented whose voices have helped to advance this work.
| Dataset | Description | Usage | Audio | Transcripts |
|---|---|---|---|---|
| Arizona Child Acoustic Database Repository | Isolated words and prompts.Ages 2–7. | Word hours: 15.6 challenge, 13.8 final model Phonetic hours: 15.6 challenge, 15.7 final model |
Access via TalkBankNoncommercial use with attribution; commercial use requires permission. | Access via TalkBankOrthographic + phonetic labels. |
| Cameron | Story-telling recordings.Ages 3–5; speakers of African American English and General American English. | Word hours: 6.6 challenge, 7.2 final model Phonetic hours: 6.6 challenge, 7.3 final model |
Access via TalkBankNoncommercial use with attribution; commercial use requires permission. | New annotations forthcoming on TalkBankPhonetic + orthographic transcripts. |
| CMU Kids Corpus | Read-aloud recordings, drawn in part from a school chosen for its higher rate of reading difficulties.Ages 6–11. | Word hours: 3.7 challenge, 3.7 final model |
Access via LDCLicensed through the LDC under its terms. | Access via LDCOrthographic transcripts, included with the corpus. |
| CSLU: Kids' Speech Version 1.1 | Scripted and spontaneous speech recordings.Original corpus spans kindergarten through grade 10; per-child ages not on file. | Word hours: 92.1 challenge, 62.8 final model |
Access via LDCLicensed through the LDC under its terms. | Access via LDCOrthographic transcripts, included with the corpus. |
| Edmonton Narrative Norms Instrument | Narrative-elicitation recordings.Ages 4–10; roughly one-fifth with diagnosed language impairment. | Word hours: 26.5 challenge, 28.1 final model Phonetic hours: 26.5 challenge, 27.8 final model |
Access via TalkBankNoncommercial use with attribution; commercial use requires permission. | New annotations forthcoming on TalkBankPhonetic + orthographic transcripts. |
| Ellis Weismer Corpus | Clinical/research recordings of toddlers and preschoolers.Ages 2.5–5.5; late talkers alongside typically developing peers. | Word hours: 17.4 challenge, 23.2 final model Phonetic hours: 17.4 challenge, 16.4 final model |
Access via TalkBankNoncommercial use with attribution; commercial use requires permission. | New annotations forthcoming on TalkBankPhonetic + orthographic transcripts. |
| JIBO Kids | Child–social-robot interaction recordings (letter/digit identification, extended discourse).Pre-K through 1st grade. | Word hours: 3.8 challenge, 4.8 final model Phonetic hours: 5.0 challenge, 5.3 final model |
Access via ZenodoCommercial and noncommercial use with attribution to the JIBO Kids corpus. | New annotations forthcoming on the K-12 AI Infrastructure PlatformNew phonetic + orthographic transcripts. |
| My Science Tutor | Spoken-dialogue recordings.Grades 3–5. | Word hours: 169.0 challenge, 160.6 final model |
Access via LDCLicensed through the LDC under its terms. Commercial licenses through Boulder Learning Inc.; contact mystcorpus25@gmail.com. | Access via LDCOrthographic transcripts, included with the corpus. |
| Ohio Child Speech Corpus | Elicited-speech tasks in a science-museum lab; a social robot joined about 60% of sessions.Ages 4–9; mostly White; some children with parent-reported speech or language problems. | Word hours: 89.1 challenge, 86.7 final model |
Access via TalkBankNoncommercial use with attribution; commercial use requires permission. | Access via TalkBankOrthographic transcripts. |
| PERCEPT‑GFTA | Goldman-Fristoe articulation test recordings.Ages roughly 7–17, plus a few young adults; suspected speech sound disorders or childhood apraxia of speech alongside typically developing peers. | Word hours: 3.9 challenge, 4.2 final model Phonetic hours: 3.9 challenge, 4.2 final model |
Access via TalkBankNoncommercial use with attribution; commercial use requires permission. | New annotations forthcoming on TalkBankOrthographic + phonetic transcripts (part pre-existing, part new). |
| PERCEPT‑R | Speech pathology / perceptual research recordings.Ages 6–24; suspected or diagnosed speech sound disorders alongside typically developing peers. | Word hours: 29.6 challenge, 29.7 final model Phonetic hours: 29.6 challenge, 29.7 final model |
Access via TalkBankNoncommercial use with attribution; commercial use requires permission. | New annotations forthcoming on TalkBankPhonetic + orthographic transcripts. |
| ReadNet | Early-literacy assessment recordings (sentence and nonword repetition, sound blending and deletion).Kindergarten through 3rd grade; range of household incomes and racial backgrounds. | Word hours: 16.9 challenge, 16.6 final model Phonetic hours: 2.7 challenge, 25.1 final model |
ReadNet project pageNot publicly distributed. Contact the project PI, Yaacov Petscher, at ypetscher@fsu.edu. | ReadNet project pageOrthographic (from correct-answer prompts) + phonetic labels. Contact the project PI, Yaacov Petscher, at ypetscher@fsu.edu. |
| SPROUT | A growing repository of speech from ~300 children across diverse backgrounds.Age 4; some children flagged at risk for speech or language delays, alongside typically developing peers. | Word hours: 16.1 challenge, 29.9 final model Phonetic hours: 16.1 challenge, 28.0 final model |
SPROUT dataset pageNot publicly distributed. Contact Northwestern University's PedzSTAR lab. | SPROUT dataset pagePhonetic + orthographic transcripts. Contact Northwestern University's PedzSTAR lab. |
On Top of Pasketti Competition
An open online competition catalyzed development against the challenge data and elevated the best-performing approaches.
Children's speech recognition is a hard problem where the best approaches are not evident at the outset. An open challenge brings a large expert community to the problem, tests hundreds of models quickly and cost-effectively, and lets the best-performing solutions rise to the top of the leaderboard.
The On Top of Pasketti: Children's Speech Recognition Challenge featured two tracks: the Word Track, scored on word error rate, and the Phonetic Track, scored on IPA character error rate. There was also a bonus prize for solutions that performed best on noisy classroom audio. The challenge was set up to keep evaluation data private and test solutions by bringing the models to the data.
- TrainParticipants trained models on transcribed child speech (344 hours for word, 85 for phonetic), and could add outside data and models.
- Submit codeThey submitted trained models and inference code for running in a containerized test environment.
- Score privatelyEvery submission ran on never-seen test data, including audio too sensitive to share.
- Rank publiclyScores went to a public leaderboard, with the best performance displayed for each team.
Results
The challenge website drew more than 6,500 visitors from 179 countries. Over the course of the challenge, 828 participants submitted over 2,000 solutions for evaluation. Within the first weeks, the best submissions had passed the benchmarks, and they kept improving until the final days.
Final Word Track winners improved more than 50% over both OpenAI Whisper and KidWhisper. Phonetic Track winners improved 49% over a baseline that transcribes words with Parakeet and looks up their pronunciations in the CMU Pronouncing Dictionary. On noisy classroom audio, the Word Track bonus winner improved 39% over KidWhisper.
$120,000 in prizes went to eight winners across the two tracks and the noisy classroom bonus:
| Place | Word Track | Phonetic Track |
|---|---|---|
| 🥇 1st | $25,000Kotaro Watanabe | $25,000Cheng Huige |
| 🥈 2nd | $15,000Sunday | $15,000Team Epoch VI Rein Viegers, Maxim Cardenas Cruz, Willem Dieleman |
| 🥉 3rd | $10,000Tang Yongqwei | $10,000Tuan Dung Le |
| Noisy classroom bonus | $5,000 eachKotaro Watanabe, Sunday, Shiqi Li, and Mitchell DeHaven | |
Full write-ups of every winning approach are in the winners blog post, and all the code is in the winners' repository.
Resources
A consolidated list of shared models, data, and resources connected with this project.
Models
- pasketti-word 🍝 Hugging Face model card for word-level transcription
- pasketti-phonetic 🍝 Hugging Face model card for phonetic (IPA) transcription
- Winners' repository Code and write-ups for every prize-winning solution
Competition
- Competition overview Both tracks, results, and leaderboards
- Word Track Problem description, data, and final leaderboard
- Phonetic Track Problem description, data, and final leaderboard
- About the project Background, acknowledgements, and how to cite the competition and its data
Data
- Dataset directory Every dataset behind this project, with where to find its audio and transcripts
Knowledge
- Winners blog post Results and every winning approach
- Word Track reference implementation Tutorial: fine-tuning NVIDIA Parakeet
- Phonetic Track reference implementation Tutorial: fine-tuning Wav2Vec2
- SciPy 2026: Horton Hears a Word Talk by Katie Wetstone on building AI infrastructure for children's speech recognition
- Open-source packages for speech data in ML Background on working with speech data
- Further reading Curated papers and guides on children's ASR
Note on future opportunities
The open models, data, and resources from this project provide a foundation that also enables further development. Several opportunities identified from stakeholder interviews could make them easier to adopt, more trustworthy, and more useful in applied settings:
-
Engineering to facilitate fine-tuning: Develop a lightweight, open-source Python package for fine-tuning the models on an organization's own data, so teams can adapt them to the use cases and populations they serve.
What we've heard: “Assessing English speaking skills for children who are bilingual and whose first language is not English is extremely important for us.”
-
Trust signals and deeper evaluation: Add word- and phoneme-level confidence scores that flag predictions for human review, and evaluate the models further on young speakers, diverse dialects, and children at risk of speech disorders.
What we've heard: “It would be really helpful to have confidence scores to flag areas where manual validation may be needed.”
-
Moving between words and sounds: Convert between phonetic and word-level outputs for literacy assessment and research, and publish a citable dataset and paper that pair the new transcripts with their source audio.
What we've heard: “The interplay between words and phonemes is especially important in learning to read. There are patterns of speech that don't necessarily indicate a gap in literacy or a phonological deficiency, but are still useful to see and preserve.”
-
Extensions for real-world tools: Add timestamps, speaker diarization, and punctuation, and develop smaller variants that run efficiently or on-device.
What we've heard: Real-world deployment needs capabilities beyond transcription, such as diarization, timestamps, punctuation, and more efficient models.
While none of these is needed to use the models today, each would build directly on this work to support better ASR tooling for early learners. If you'd like to explore these opportunities, discuss other ideas, or share how you're using these models and data in your work, we'd love to hear from you.