Open speech data for Kashmiri.
Zoon means moon. It is also the largest open speech corpus ever attempted for Kóshur — built so the language can live inside the voice technology the rest of the world takes for granted.
A corpus the size of a language.
Two datasets, recorded with volunteers from across the valley and released free for anyone to use.
700
hours of expressive text-to-speech (target)
500
hours of speech-to-text (target)
8
dedicated voices, full of feeling
500
speakers from across the valley
One voice the machines can learn.
700 hours, full of feeling
Around eight dedicated speakers record not flat studio reading, but speech carrying real emotion — laughter, lament, lullaby. Enough for a voice that actually sounds Kashmiri.
500 hours, 500 voices
Gathered from everyday speakers — every district, dialect, age and accent — so recognition works for the whole valley, not just newsreaders.
Cleaned, labelled, aligned
Every clip is transcribed, timestamped and quality-checked — turning raw recordings into data a model can actually learn from.
Free for everyone
The finished corpus is released openly for researchers, developers and anyone who wants to build Kashmiri into their tools.
Four steps, about ten minutes.
01Sign up
Make an account at record.yimberzol.org. It takes a minute and runs in your browser — no app to install.
02Read aloud
We show you short Kashmiri sentences, each with the script and a Roman transliteration. You read them out and record.
03We check
A person reviews every clip before it's accepted, so the dataset stays clean and usable.
04Released openly
Accepted recordings go into an open dataset that anyone can use to build Kashmiri speech tools.
Your accent belongs in the corpus.
No studio, no signup fuss — just your phone and your voice. Every dialect you add makes Kashmiri speech tech work for more people.