Observed arrival · 2026-08-31
Dataset.ET Is Teaching AI the Languages of Ethiopia
An open speech, text, and translation corpus project for more than 80 Ethiopian languages.
○Open
⊠Login
$Paid
†Ads
✦Pretty
●Pro
◎Niche
◉Human
⚑Risk
ƒJS
Field notes
The project describes a four-stage pipeline: contributors submit sentences, paragraphs, documents, or language material; community reviewers validate it; the team cleans and annotates accepted data; researchers and developers use the corpus for model training. The directory gives Amharic 500k+ items, Afaan Oromoo 450k+, and Somali and Tigrinya 200k+ each, while Sidamo and Wolaytta are marked “Coming soon.”
Observed signals
Read the marks
Editorial observations of this landing page, not a rating.
○OpenPublic substance visible
⊠LoginAccess appeared gated
●ProPolished or operationally mature
◎NicheUnusually specific use
◉HumanPersonal, local, civic, or handmade
ƒJavaScriptBrowser-side code central
One card from the complete issue