Google Research Names WAXAL Speech Challenge Winners After 8,837 Submissions

The contest tested approaches to transcribing Lingala and Shona with open speech data. The next question is whether those results can translate into useful voice services.

By 2 min read
Google Research Names WAXAL Speech Challenge Winners After 8,837 Submissions
Google Research Names WAXAL Speech Challenge Winners After 8,837 Submissions

Listen to this story

The audio brief

About 0:44
0:000:44
Read transcript
Team Pasketti has won Google Research’s WAXAL speech challenge, narrowly beating two other finalists in a contest built to improve transcription for African languages. The competition drew 1,462 participants from 100 countries, including 42 African countries, and produced 8,837 submitted solutions after more than 24,000 hours of work. The challenge used open speech data from WAXAL, a resource that Google says has grown from 27 to 32 Sub-Saharan African languages. Its initial release included about 1,846 hours of transcribed speech for automatic speech recognition, plus more than 565 hours intended for speech generation. The data is available under a Creative Commons attribution license. But the contest focused on just two languages: Lingala and Shona. The leading teams took notably different routes. Pasketti combined eight models, using their different transcription errors to improve the final result. The second-place system processed each audio clip several ways before choosing its most trusted transcription. Third place went to an approach tuned to language structure, including words that run together in Lingala and surrounding context in Shona. The larger opportunity is access. Google says speakers may use these languages fluently while finding written interfaces impractical, and many current AI systems still do not understand them. WAXAL could eventually support services for crop prices, health advice, or education by voice—but those are possibilities, not announced products. The key question is whether gains on this focused benchmark can become reliable voice tools as more languages are added.

Story brief

3 key points

The WAXAL ASR Challenge shows how open data and global competition can push African-language speech recognition beyond a research release. Google says 1,462 entrants submitted 8,837 solutions over more than 24,000 hours, with Team Pasketti narrowly leading competitors from Burkina Faso and Niger. The underlying WAXAL resource has expanded from 27 to 32 languages and includes speech data for recognition and...

  1. 01

    WAXAL’s initial release included roughly 1,846 hours of transcribed speech and 565 hours for speech generation under CC-BY-4.0.

  2. 02

    The competition focused on Lingala and Shona, rather than testing all 32 languages in the expanded resource.

  3. 03

    Pasketti used eight-model ensembling; other finalists used repeated audio processing and language-specific context.

Google Research has named the winners of its WAXAL ASR Challenge with Zindi, a competition to improve AI transcription for African languages. Team Pasketti took first place, followed by Alban Nyantudre and Abdourahamane Ide Salifou, after a contest that drew 1,462 participants from 100 countries.

From open speech data to a competition

The challenge builds on WAXAL, an open-source speech collection Google now describes as covering 32 African languages. Its initial March release covered 27 Sub-Saharan African languages and included about 1,846 hours of transcribed natural speech for automatic speech recognition, plus more than 565 hours of recordings for speech generation. The resources were released under a Creative Commons CC-BY-4.0 license.

Google Research and Zindi asked entrants to build systems that understand and transcribe spoken language, starting with Lingala and Shona. The earlier dataset release was designed around both recognition and speech generation; the contest narrowed the immediate task to recognition. Google frames the effort around an access problem: people may speak these languages fluently while finding written interfaces less practical, and many current AI systems do not understand them.

The challenge at a glance
1,462Participants

Entrants came from 100 countries, including 42 African countries.

8,837Submitted solutions

Participants spent more than 24,000 hours on the problem, according to Google.

Different routes to a transcription result

Google said the top three scores were remarkably close. But the entries approached the same task differently, illustrating the range of techniques competitors used with the challenge data rather than a single prescribed solution.

  • First-place Team Pasketti, represented by Roman Solovyev and enes3774 from Kazakhstan and Canada, combined eight models so their different transcription errors could be weighed against one another.
  • Second-place Nyantudre of Burkina Faso had his system process the same audio clip in several ways before selecting the transcription it trusted most.
  • Third-place Salifou of Niger adapted his approach to language patterns, accounting for words running together in Lingala and using surrounding context for Shona.

The work now points beyond the leaderboard

Nyantudre, a machine-learning engineer who builds speech tools for his mother tongue, Mooré, described why voice access matters to him:

Plenty of people I grew up around speak their language perfectly but cannot read or write it. That is why speech is the only realistic way for those people to use technology at all.

Google says WAXAL is intended to keep expanding to additional languages. It points to future possibilities such as checking crop prices, finding health advice, and learning by speaking to a phone in a familiar language. Those examples remain future possibilities, not announced services.

Sources

  1. research.googleWAXAL: A large-scale open resource for African language speech technology
  2. blog.googleMeet the winners of the WAXAL speech recognition challenge

Loading discussion...

YOUR READING SPACE

Notifications