Kenpath Labs
Get started

Sample bundles

Hear the data before you license it.

Each bundle is 10 to 30 minutes of a dataset, cut the way the full set is delivered. Choose one, and we email you a download link that works for 24 hours.

What is in a bundle

  • A README with the dataset card excerpt for that dataset.
  • The audio, in the same layout and format as the full dataset.
  • Transcripts and metadata in the same files and fields you would receive with the full set.
  • LICENSE-SAMPLE.txt: the bundle is for evaluation only.

How to read it

Numbers that identify a person appear as [pii] in the transcripts and as a tone in the audio. Other tags in transcripts are in square brackets, like [filler] and [overlap].

Delivery layouts for the full sets are on formats; licences are quoted per use, as described on licensing.

Whole conversations

Complete recordings with speaker-labelled transcripts.

  • Hindi call-centre conversations, insurance

  • Hindi general conversation

  • Tamil call-centre conversations, consumer surveys

  • Tamil two-channel customer-service calls

  • Tamil general conversation

  • Marathi call-centre conversations, banking and insurance

  • Marathi general conversation

Full-duplex speech to speech

Each speaker on a separate channel, timing preserved.

Speech recognition

Audio with time-aligned transcripts in native script.

  • Hindi speech recognition utterances

  • Tamil speech recognition utterances

Diarisation

Single-channel recordings with speaker turns marked.

  • Hindi speaker diarization

  • Tamil speaker diarization

Your bundle

Choose a bundle.

The link works for 24 hours. We keep a record of each request.