Speech Recognition for Early Learning

New state-of-the-art open models for children's speech recognition, with roughly half the errors of leading open alternatives.

A teacher reads a picture book to a group of young children seated on the floor, overlaid with colored speech-waveform graphics.

Voice is one of the most natural ways for kids to learn, explore, and show what they know. Yet today's automatic speech recognition (ASR) technology hardly understands them.

This work set out to close that gap. With support from the Gates Foundation, DrivenData assembled the largest dataset of transcribed English child speech for developing open ASR models, then ran On Top of Pasketti, an open challenge where experts from around the world competed to develop the best ASR models for kids' speech. The challenge resulted in models for two kinds of transcription: word-level, which captures which words a child said and supports use cases such as reading and reasoning assessments, and phonetic, which captures exactly how they said them and supports use cases such as screening for atypical speech development.

DrivenData then retrained the winning approaches on more data and released them as two open models, pasketti-word and pasketti-phonetic, whose performance by age is shown below.

Word model: Word error rate Bar chart of word error rate by age group for pasketti-word, KidWhisper and OpenAI Whisper. pasketti-word has the lowest error in every group: 0.31 for ages 3-4, 0.11 for 5-7, 0.10 for 8-11, 0.05 for 12 and up, and 0.18 for unknown age.
Phonetic model: Phone error rate Bar chart of phone error rate by age group for pasketti-phonetic and PhoneticXeus. pasketti-phonetic has the lowest error in every group: 0.36 for ages 3-4, 0.32 for 5-7, 0.21 for 8-11, 0.16 for 12 and up, and 0.29 for unknown age, against 0.66, 0.52, 0.44, 0.42 and 0.49 for PhoneticXeus.
Error rate by age group on a held-out evaluation set: word error rate (roughly the share of words a model gets wrong) for the retrained word model (left), compared with OpenAI Whisper and KidWhisper (Whisper fine-tuned on a small amount of child speech), and phone error rate (PER, the share of speech sounds a model gets wrong) for the retrained phonetic model (right), compared with PhoneticXeus, a leading open phone recognizer. Lower is better. OpenAI Whisper performs poorly on the 12+ age group, which comes from particularly challenging recordings of individual speech sounds. Unknown reflects recordings without age information. For how the models were evaluated, and for performance by speech pathology, race, setting, sex, and source corpus, see the Pasketti model cards on Hugging Face.

This project represents a multi-year effort, bringing together contributions from education, data science, and research communities. This effort included:

  • Datasets, benchmarking, and competition design Oct 2024 – Jul 2025
    State-of-the-field research and benchmarking, a review of the child-speech data that already existed, acquiring and harmonizing what we could use, and designing the challenge to address the most pressing opportunities.
  • Phonetic and word transcription May 2025 – Jul 2026
    Based on the first phase, a supplemental investment funded annotation to address a critical lack of transcribed speech data, especially phonetic labels for child speech. Our transcription team produced 253 hours of newly annotated child speech, including 164 hours of phonetic labels.
  • Open competition Jan – Apr 2026
    Over the course of the live challenge, more than 800 participants made 2,000+ submissions across the modeling tracks and competed for $120,000 in prize money. The top-performing models reduced word error rates by half against the best existing open models for kids.
  • Public goods Apr – Sep 2026
    Our team streamlined and retrained the winning approaches on all the data, including audio too sensitive to share with solvers, and released the results: open models on Hugging Face, the new word-level and phonetic transcripts, and the winners' code.

Why children's speech is hard

Current automatic speech recognition (ASR) models transcribe adult speech well but struggle with children's speech.

Word error rate on child speech vs. adult benchmarks Bar chart of word error rate as of June 2025: leading ASR on adult speech 6 to 7 percent, professional human transcribers 8 to 10 percent, leading ASR on children under 6 40 to 76 percent.
Lower is better. Source: DrivenData benchmarking of OpenAI Whisper and KidWhisper on three preschool child-speech datasets (JIBO Kids, Ellis Weismer, Cameron), June 2025; adult ASR from the Hugging Face Open ASR Leaderboard, June 2025; human transcriber range from Radford et al. (2022).

Models trained on adult speech don't work for kids because child speech is fundamentally different. Kids have distinct vocal characteristics, different speech patterns, and make frequent speech shortcuts and pronunciation errors, all of which make understanding them difficult.

  • Metathesis: elephant → ephelant
  • Velar fronting: cup → tup
  • Syllable deletion: banana → nana
  • Errors combined: spaghetti → pasketti
Common patterns in how young children pronounce words, each shown as the intended word and how a child might say it.

This poor performance is a significant barrier to educational tools that support early learning. Addressing it requires assembling more representative data and developing better ASR models built for children's speech.

Models

Open, retrained speech recognition models for word- and phone-level transcription of children's speech.

The ASR models developed in this project were created in two phases:

  • The competition surfaced the best modeling approaches. A global community with a wide variety of backgrounds explored the solution space to find the most effective open approaches, each evaluated against the same test set.
  • Post-challenge retraining focused on the best final models for release. The winning approaches were streamlined to be easier to run, retrained on more data, and packaged with clear documentation and evaluation of results.

The result is two open-weight models on Hugging Face, one for each kind of transcription. Both are released under the BigCode OpenRAIL-M license, which allows commercial and noncommercial use subject to responsible-use restrictions, and each has a model card covering its training data, intended uses, evaluation, and known limitations:

Both final models build on the challenge's top models:

  • pasketti-word is a full fine-tune of Qwen3-ASR-1.7B, trained on 372 hours of child speech from more than 3,600 children in clips of up to 20 seconds.
  • pasketti-phonetic streamlines the challenge's large ensembles into a single WavLM Large model, trained on 116 hours of phonetically transcribed speech from about 1,700 children. A second, word-level output used only during training let it learn from an additional 371 hours of word-transcribed speech.

While the competition models were trained only on competition data, the retrained versions add a larger and more representative dataset, including sensitive data from realistic education populations that could not be shared with solvers, and data deliberately gathered from groups where current models perform worst, such as children with atypical speech development, non-white students, and students from low-income backgrounds. Each model was then tested on held-out data and checked for differences in performance by age, race, speech pathology, sex, setting, and source corpus.

Choosing the right model

The two models capture different outputs and therefore serve different purposes:

How it sounded Word transcript Phonetic transcript
"Uh… eye doh-noh"I don't knowʔəː ʔɑi doʊ noʊ
"Eye-dohn-noh"I don't knowɑi doʊn noʊ
"Eye-dun-no"I don't knowɑi don no
The same phrase as produced by three different children. The word transcript records which words each child said, in standard spelling; the phonetic transcript keeps every difference in how they were pronounced.

Word transcription tells you which words the child said. This supports educators and ed-tech developers with use cases such as:

  • Reading fluency: needs an accurate word-for-word record to compare against the target text. The phonetic model can add mispronunciation detail.
  • Language use assessment: measures vocabulary and syntax.
  • Comprehension, problem solving, reasoning: scores the substance of an answer.
  • Verbal interaction with tools: a voice interface needs to know what was said in order to respond.

Phonetic transcription tells you how the child said it. This further supports speech-language pathologists and clinicians with use cases such as:

  • Speech screening: flags substitutions, omissions, and other production errors that word-level transcription smooths over. The word model helps confirm what was attempted.
  • Early literacy screening: early indicators of reading difficulty often show up in phonological patterns, so phonetic transcripts may help flag children for follow-up.
  • Supporting word-level transcription: as an interim output, a phonetic transcript records the sounds a child actually made, which may help downstream tools account for developmentally typical sound substitutions rather than mistaking them for errors.

How the final models perform

Word model: Word error rate Bar chart of overall word error rate on the held-out evaluation set: pasketti-word 17%, KidWhisper 32%, OpenAI Whisper 34%. Lower is better.
Phonetic model: Phone error rate Bar chart of overall phone error rate on the held-out evaluation set: pasketti-phonetic 33%, PhoneticXeus 59%. Lower is better.
Error rate on held-out evaluation data; lower is better. The word model was evaluated on 148 hours of speech from more than 2,300 children as young as 3, against OpenAI Whisper and KidWhisper on the same clips. Every utterance is scored, and all three models run with protection against getting stuck repeating themselves. The phonetic model was evaluated on 44 hours of speech from 737 children, against PhoneticXeus, a leading open phone recognizer. Word boundaries are excluded from scoring; counting them, pasketti-phonetic's error rate is 30%.

Both models improved on the leading open alternatives across every age and demographic group in their evaluation sets, demonstrating the importance of diverse training data.

  • pasketti-word has a word error rate (WER) of 17%, about half that of KidWhisper (32%) and OpenAI Whisper (34%).
  • pasketti-word also narrows the gap between white and non-white children, and improves on KidWhisper by 29% in classroom settings, which are especially difficult for models (and humans) because of background talk and noise.
  • pasketti-phonetic has a phone error rate (PER) of 33%, compared with 59% for PhoneticXeus, cutting errors by more than a third in every age group.
  • Children ages 3–4 remain the hardest to transcribe, but they are also where some of the biggest gains came compared with existing models.
Word model: By setting Bar chart of word error rate by setting. In classrooms: pasketti-word 0.44, KidWhisper 0.62, Whisper 0.69. Outside classrooms: pasketti-word 0.15, KidWhisper 0.29, Whisper 0.31.
Word model: By race and corpus Bar chart of word error rate by race within the SPROUT and ReadNet corpora. In SPROUT, pasketti-word ranges from 0.28 to 0.33 across Black, Hispanic, multi-racial and white children, against 0.55 to 0.62 for KidWhisper and 0.60 to 0.66 for Whisper. In ReadNet, pasketti-word is 0.07 for Black, 0.08 for multi-racial and 0.05 for white children, against 0.56, 0.59 and 0.40 for Whisper.
Word error rate for pasketti-word compared with OpenAI Whisper and KidWhisper on the held-out evaluation set, by recording setting (left) and by race within each corpus (right). Lower is better. Race is compared within corpora because few corpora record it, and some groups are small.

Performance by age is shown at the top of this page. Full evaluation results for both models, including breakdowns by age, race, speech pathology, sex, and source corpus, are in the Pasketti model cards on Hugging Face.

As is, these models are usable, but should be tested and evaluated in the specific context and population they would be used in, given that performance will vary by age, dialect, setting, and task. They will be especially useful in contexts like literacy assessments, where a known target transcript is available to compare against, and can serve as strong starting points for fine-tuning to specific tasks, grade levels, or student populations. Where exact transcripts matter, outputs should be reviewed by a person.

For intended uses, out-of-scope uses, and known limitations, see the Pasketti model cards.

Data

The largest dataset of its kind for developing open ASR models for children's speech: 14 datasets identified, sourced, and augmented with new annotations to build the models above.

660K+ child utterances curated
520 hours of read, prompted & spontaneous child speech
14 contributing datasets, newly annotated or harmonized

Data has been the biggest constraint to better ASR for children. Models that work well for adults learned from vast amounts of adult speech, and no comparable, representative dataset existed for kids. To assemble such a dataset, DrivenData reviewed the landscape of existing child-speech resources, then identified and sourced audio from public, research, and enterprise partners. Our transcription team added over 250 hours of new orthographic annotation and harmonized it with existing corpora into a common schema.

The result spans a wide range of children and settings: ages from preschoolers to young adults in clinical research data, typically developing kids alongside children with diagnosed or suspected speech and language differences, and dialects including African American English and General American English. The table below is a directory to each dataset — what it is, how much of it each model used, where to find the audio and the transcripts, and information on licensing outside the challenge.

Child speech data is sensitive and difficult to share. We are grateful for the work of the data providers who thoughtfully collected and provided the data used in this project under appropriate consents, and to the speakers represented whose voices have helped to advance this work.

Dataset Description Usage Audio Transcripts
Arizona Child Acoustic Database Repository Isolated words and prompts.Ages 2–7.
Word hours: 15.6 challenge, 13.8 final model
Phonetic hours: 15.6 challenge, 15.7 final model
Access via TalkBankNoncommercial use with attribution; commercial use requires permission. Access via TalkBankOrthographic + phonetic labels.
Cameron Story-telling recordings.Ages 3–5; speakers of African American English and General American English.
Word hours: 6.6 challenge, 7.2 final model
Phonetic hours: 6.6 challenge, 7.3 final model
Access via TalkBankNoncommercial use with attribution; commercial use requires permission. New annotations forthcoming on TalkBankPhonetic + orthographic transcripts.
CMU Kids Corpus Read-aloud recordings, drawn in part from a school chosen for its higher rate of reading difficulties.Ages 6–11.
Word hours: 3.7 challenge, 3.7 final model
Access via LDCLicensed through the LDC under its terms. Access via LDCOrthographic transcripts, included with the corpus.
CSLU: Kids' Speech Version 1.1 Scripted and spontaneous speech recordings.Original corpus spans kindergarten through grade 10; per-child ages not on file.
Word hours: 92.1 challenge, 62.8 final model
Access via LDCLicensed through the LDC under its terms. Access via LDCOrthographic transcripts, included with the corpus.
Edmonton Narrative Norms Instrument Narrative-elicitation recordings.Ages 4–10; roughly one-fifth with diagnosed language impairment.
Word hours: 26.5 challenge, 28.1 final model
Phonetic hours: 26.5 challenge, 27.8 final model
Access via TalkBankNoncommercial use with attribution; commercial use requires permission. New annotations forthcoming on TalkBankPhonetic + orthographic transcripts.
Ellis Weismer Corpus Clinical/research recordings of toddlers and preschoolers.Ages 2.5–5.5; late talkers alongside typically developing peers.
Word hours: 17.4 challenge, 23.2 final model
Phonetic hours: 17.4 challenge, 16.4 final model
Access via TalkBankNoncommercial use with attribution; commercial use requires permission. New annotations forthcoming on TalkBankPhonetic + orthographic transcripts.
JIBO Kids Child–social-robot interaction recordings (letter/digit identification, extended discourse).Pre-K through 1st grade.
Word hours: 3.8 challenge, 4.8 final model
Phonetic hours: 5.0 challenge, 5.3 final model
Access via ZenodoCommercial and noncommercial use with attribution to the JIBO Kids corpus. New annotations forthcoming on the K-12 AI Infrastructure PlatformNew phonetic + orthographic transcripts.
My Science Tutor Spoken-dialogue recordings.Grades 3–5.
Word hours: 169.0 challenge, 160.6 final model
Access via LDCLicensed through the LDC under its terms. Commercial licenses through Boulder Learning Inc.; contact mystcorpus25@gmail.com. Access via LDCOrthographic transcripts, included with the corpus.
Ohio Child Speech Corpus Elicited-speech tasks in a science-museum lab; a social robot joined about 60% of sessions.Ages 4–9; mostly White; some children with parent-reported speech or language problems.
Word hours: 89.1 challenge, 86.7 final model
Access via TalkBankNoncommercial use with attribution; commercial use requires permission. Access via TalkBankOrthographic transcripts.
PERCEPT‑GFTA Goldman-Fristoe articulation test recordings.Ages roughly 7–17, plus a few young adults; suspected speech sound disorders or childhood apraxia of speech alongside typically developing peers.
Word hours: 3.9 challenge, 4.2 final model
Phonetic hours: 3.9 challenge, 4.2 final model
Access via TalkBankNoncommercial use with attribution; commercial use requires permission. New annotations forthcoming on TalkBankOrthographic + phonetic transcripts (part pre-existing, part new).
PERCEPT‑R Speech pathology / perceptual research recordings.Ages 6–24; suspected or diagnosed speech sound disorders alongside typically developing peers.
Word hours: 29.6 challenge, 29.7 final model
Phonetic hours: 29.6 challenge, 29.7 final model
Access via TalkBankNoncommercial use with attribution; commercial use requires permission. New annotations forthcoming on TalkBankPhonetic + orthographic transcripts.
ReadNet Early-literacy assessment recordings (sentence and nonword repetition, sound blending and deletion).Kindergarten through 3rd grade; range of household incomes and racial backgrounds.
Word hours: 16.9 challenge, 16.6 final model
Phonetic hours: 2.7 challenge, 25.1 final model
ReadNet project pageNot publicly distributed. Contact the project PI, Yaacov Petscher, at ypetscher@fsu.edu. ReadNet project pageOrthographic (from correct-answer prompts) + phonetic labels. Contact the project PI, Yaacov Petscher, at ypetscher@fsu.edu.
SPROUT A growing repository of speech from ~300 children across diverse backgrounds.Age 4; some children flagged at risk for speech or language delays, alongside typically developing peers.
Word hours: 16.1 challenge, 29.9 final model
Phonetic hours: 16.1 challenge, 28.0 final model
SPROUT dataset pageNot publicly distributed. Contact Northwestern University's PedzSTAR lab. SPROUT dataset pagePhonetic + orthographic transcripts. Contact Northwestern University's PedzSTAR lab.
The Usage column gives the hours of child speech each model used at two stages: challenge, the On Top of Pasketti challenge's Word or Phonetic Track (training and test combined), and final model, the released pasketti-word or pasketti-phonetic model (training and evaluation combined). Phonetic hours count only speech with phonetic labels. The final models drew on an expanded, re-cleaned dataset, so their hours can differ from the challenge's. A 14th dataset (real classroom audio) is withheld from public naming here at its data partner's request. Word Track participants also received classroom background-noise recordings from the University of Maryland's RealClass project to augment training audio.

On Top of Pasketti Competition

An open online competition catalyzed development against the challenge data and elevated the best-performing approaches.

Children's speech recognition is a hard problem where the best approaches are not evident at the outset. An open challenge brings a large expert community to the problem, tests hundreds of models quickly and cost-effectively, and lets the best-performing solutions rise to the top of the leaderboard.

The On Top of Pasketti: Children's Speech Recognition Challenge featured two tracks: the Word Track, scored on word error rate, and the Phonetic Track, scored on IPA character error rate. There was also a bonus prize for solutions that performed best on noisy classroom audio. The challenge was set up to keep evaluation data private and test solutions by bringing the models to the data.

  1. TrainParticipants trained models on transcribed child speech (344 hours for word, 85 for phonetic), and could add outside data and models.
  2. Submit codeThey submitted trained models and inference code for running in a containerized test environment.
  3. Score privatelyEvery submission ran on never-seen test data, including audio too sensitive to share.
  4. Rank publiclyScores went to a public leaderboard, with the best performance displayed for each team.
Running submissions as code on DrivenData's platform confirmed each solution actually works, and meant the test audio never had to leave the platform.

Results

The challenge website drew more than 6,500 visitors from 179 countries. Over the course of the challenge, 828 participants submitted over 2,000 solutions for evaluation. Within the first weeks, the best submissions had passed the benchmarks, and they kept improving until the final days.

Word Track: 1,542 submissions evaluated Line chart of the best Word Track leaderboard score over time, falling from about 0.49 word error rate in early February 2026 to below the benchmark (about 0.32) within days, and to about 0.19 by early April.
Phonetic Track: 684 submissions evaluated Line chart of the best Phonetic Track leaderboard score over time, falling from about 0.51 character error rate in early February 2026 to below the benchmark (about 0.35) within days, and to about 0.26 by late March.
Best leaderboard score over time; lower is better. Dotted lines mark out-of-the-box Parakeet (Word Track) and the Wav2Vec2 reference implementation (Phonetic Track). The top three finished at 0.19–0.20 WER (vs. 0.41 for KidWhisper and 0.49 for OpenAI Whisper) and 0.26 CER (vs. 0.52 for Parakeet with the CMU Pronouncing Dictionary). See the Word Track and Phonetic Track reference implementations.

Final Word Track winners improved more than 50% over both OpenAI Whisper and KidWhisper. Phonetic Track winners improved 49% over a baseline that transcribes words with Parakeet and looks up their pronunciations in the CMU Pronouncing Dictionary. On noisy classroom audio, the Word Track bonus winner improved 39% over KidWhisper.

$120,000 in prizes went to eight winners across the two tracks and the noisy classroom bonus:

Place Word Track Phonetic Track
🥇 1st $25,000Kotaro Watanabe $25,000Cheng Huige
🥈 2nd $15,000Sunday $15,000Team Epoch VI Rein Viegers, Maxim Cardenas Cruz, Willem Dieleman
🥉 3rd $10,000Tang Yongqwei $10,000Tuan Dung Le
Noisy classroom bonus $5,000 eachKotaro Watanabe, Sunday, Shiqi Li, and Mitchell DeHaven

Full write-ups of every winning approach are in the winners blog post, and all the code is in the winners' repository.

Resources

A consolidated list of shared models, data, and resources connected with this project.

Models

Competition

Data

Knowledge

Note on future opportunities

The open models, data, and resources from this project provide a foundation that also enables further development. Several opportunities identified from stakeholder interviews could make them easier to adopt, more trustworthy, and more useful in applied settings:

  • Engineering to facilitate fine-tuning: Develop a lightweight, open-source Python package for fine-tuning the models on an organization's own data, so teams can adapt them to the use cases and populations they serve.

    What we've heard: “Assessing English speaking skills for children who are bilingual and whose first language is not English is extremely important for us.”

  • Trust signals and deeper evaluation: Add word- and phoneme-level confidence scores that flag predictions for human review, and evaluate the models further on young speakers, diverse dialects, and children at risk of speech disorders.

    What we've heard: “It would be really helpful to have confidence scores to flag areas where manual validation may be needed.”

  • Moving between words and sounds: Convert between phonetic and word-level outputs for literacy assessment and research, and publish a citable dataset and paper that pair the new transcripts with their source audio.

    What we've heard: “The interplay between words and phonemes is especially important in learning to read. There are patterns of speech that don't necessarily indicate a gap in literacy or a phonological deficiency, but are still useful to see and preserve.”

  • Extensions for real-world tools: Add timestamps, speaker diarization, and punctuation, and develop smaller variants that run efficiently or on-device.

    What we've heard: Real-world deployment needs capabilities beyond transcription, such as diarization, timestamps, punctuation, and more efficient models.

Quotes from stakeholder interviews conducted for this project.

While none of these is needed to use the models today, each would build directly on this work to support better ASR tooling for early learners. If you'd like to explore these opportunities, discuss other ideas, or share how you're using these models and data in your work, we'd love to hear from you.