زوٗن · Project Zoon

Open speech data for Kashmiri.

Zoon means moon. It is also the largest open speech corpus ever attempted for Kóshur — built so the language can live inside the voice technology the rest of the world takes for granted.

By the numbers

A corpus the size of a language.

Two datasets, recorded with volunteers from across the valley and released free for anyone to use.

700

hours of expressive text-to-speech (target)

500

hours of speech-to-text (target)

8

dedicated voices, full of feeling

500

speakers from across the valley

Two halves

One voice the machines can learn.

Expressive TTS

700 hours, full of feeling

Around eight dedicated speakers record not flat studio reading, but speech carrying real emotion — laughter, lament, lullaby. Enough for a voice that actually sounds Kashmiri.

Speech-to-text

500 hours, 500 voices

Gathered from everyday speakers — every district, dialect, age and accent — so recognition works for the whole valley, not just newsreaders.

Curate & align

Cleaned, labelled, aligned

Every clip is transcribed, timestamped and quality-checked — turning raw recordings into data a model can actually learn from.

Release openly

Free for everyone

The finished corpus is released openly for researchers, developers and anyone who wants to build Kashmiri into their tools.

How recording works

Four steps, about ten minutes.

01Sign up

Make an account at record.yimberzol.org. It takes a minute and runs in your browser — no app to install.

02Read aloud

We show you short Kashmiri sentences, each with the script and a Roman transliteration. You read them out and record.

03We check

A person reviews every clip before it's accepted, so the dataset stays clean and usable.

04Released openly

Accepted recordings go into an open dataset that anyone can use to build Kashmiri speech tools.

Your part

Your accent belongs in the corpus.

No studio, no signup fuss — just your phone and your voice. Every dialect you add makes Kashmiri speech tech work for more people.