Skip to the card

Card 232 of 9992026-08-31 issue

Observed arrival · 2026-08-31

Dataset.ET Is Teaching AI the Languages of Ethiopia

dataset.africa Observed source
Editorial interest 79/100 Selection signal · not a rating of the site

An open speech, text, and translation corpus project for more than 80 Ethiopian languages.

Landing page captured for the 2026-08-31 issue.

Field notes

The project describes a four-stage pipeline: contributors submit sentences, paragraphs, documents, or language material; community reviewers validate it; the team cleans and annotates accepted data; researchers and developers use the corpus for model training. The directory gives Amharic 500k+ items, Afaan Oromoo 450k+, and Somali and Tigrinya 200k+ each, while Sidamo and Wolaytta are marked “Coming soon.”

Observed signals

Read the marks

Editorial observations of this landing page, not a rating.

OpenPublic substance visible
LoginAccess appeared gated
ProPolished or operationally mature
NicheUnusually specific use
HumanPersonal, local, civic, or handmade
ƒJavaScriptBrowser-side code central

One card from the complete issue

Habitat Reduced to One Settee

219,482 arrived 1,000 judged 999 catalogued Enter the complete issue
dataset.africa

Landing page observed 2026-08-31. The live site may have changed.