Datasets ready to license.
12 datasets, 680 hours. Play a sample from any card. Every figure was measured from the audio, and the badges say what each dataset is fit for.
2 of 12 datasets · 13 h of speech
Two-speaker general conversation in Marathi, recorded on one channel with speakers labelled. 9 hours. Transcripts are time-aligned and written in Devanagari script, with English words kept as spoken.
9 hOne channelTranscribed0% English37.8 dB · clean- ASR
- Full duplex
- Turn-taking
- Voice agents
- Diarisation
- TTS
Call-centre conversations in Marathi, banking and insurance, recorded on one channel with speakers labelled. 4 hours. Transcripts are time-aligned and written in Devanagari script, with English words kept as spoken.
4 hOne channelTranscribed0% English35.5 dB · clean- ASR
- Full duplex
- Turn-taking
- Voice agents
- Diarisation
- TTS