Kenpath Labs
Get started

Datasets ready to license.

12 datasets, 680 hours. Play a sample from any card. Every figure was measured from the audio, and the badges say what each dataset is fit for.

3 of 12 datasets · 381 h of speech

  • Scripted call-centre conversations in Hindi, insurance, recorded with each speaker on a separate channel. 356 hours. Transcripts are time-aligned and written in Devanagari script, with English words kept as spoken. Layouts across the set: 1,135 two-channel, 710 one side of a call.

    356 h
    Two channels
    Transcribed
    17% English
    33.7 dB · clean
    • ASR
    • Full duplex
    • Turn-taking
    • Voice agents
    • Diarisation
    • TTS
    490 speakersGet a sample
  • Hindi and Tamil full-duplex conversation data: 1,265 two-channel call-centre calls with each speaker on a separate channel, overlaps, backchannels and turn timing preserved, transcripts time-aligned per channel. 264 hours.

    264 h
    Two channels
    Transcribed
    16% English
    29.2 dB · some background
    • ASR
    • Full duplex
    • Turn-taking
    • Voice agents
    • Diarisation
    • TTS
    529 speakersGet a sample
  • Scripted call-centre conversations in Tamil, telecom, delivery, e-commerce and banking, recorded with each speaker on a separate channel. 25 hours. Transcripts are time-aligned and written in Tamil script, with English words kept as spoken. Layouts across the set: 2 one side of a call, 132 two-channel.

    25 h
    Two channels
    Transcribed
    29% English
    31.2 dB · clean
    • ASR
    • Full duplex
    • Turn-taking
    • Voice agents
    • Diarisation
    • TTS
    93 speakersGet a sample