[Lokah](https://kenpathlabs.com/lokah) / [Datasets](https://kenpathlabs.com/lokah/datasets) / Samples

Sample bundles

# Hear the data before you license it.

Each bundle is 10 to 30 minutes of a dataset, cut the way the full set is delivered. Choose one, and we email you a download link that works for 24 hours.

## What is in a bundle

- A README with the dataset card excerpt for that dataset.
- The audio, in the same layout and format as the full dataset.
- Transcripts and metadata in the same files and fields you would receive with the full set.
- LICENSE-SAMPLE.txt: the bundle is for evaluation only.

## How to read it

Numbers that identify a person appear as \[pii\] in the transcripts and as a tone in the audio. Other tags in transcripts are in square brackets, like \[filler\] and \[overlap\].

Delivery layouts for the full sets are on [formats](https://kenpathlabs.com/lokah/formats); licences are quoted per use, as described on [licensing](https://kenpathlabs.com/lokah/licensing).

## Whole conversations

Complete recordings with speaker-labelled transcripts.

- [Hindi call-centre conversations, insurance](https://kenpathlabs.com/lokah/datasets/hindi-call-centre-insurance)
    
- [Hindi general conversation](https://kenpathlabs.com/lokah/datasets/hindi-general-conversation)
    
- [Tamil call-centre conversations, consumer surveys](https://kenpathlabs.com/lokah/datasets/tamil-call-centre-customer-service)
    
- [Tamil two-channel customer-service calls](https://kenpathlabs.com/lokah/datasets/tamil-call-centre-banking-and-retail)
    
- [Tamil general conversation](https://kenpathlabs.com/lokah/datasets/tamil-general-conversation)
    
- [Marathi call-centre conversations, banking and insurance](https://kenpathlabs.com/lokah/datasets/marathi-call-centre-banking-and-insurance)
    
- [Marathi general conversation](https://kenpathlabs.com/lokah/datasets/marathi-general-conversation)
    

## Full-duplex speech to speech

Each speaker on a separate channel, timing preserved.

- [Hindi and Tamil full-duplex call-centre conversations](https://kenpathlabs.com/lokah/datasets/hindi-tamil-full-duplex-call-centre-conversations)
    

## Speech recognition

Audio with time-aligned transcripts in native script.

- [Hindi speech recognition utterances](https://kenpathlabs.com/lokah/datasets/hindi-speech-recognition-utterances)
    
- [Tamil speech recognition utterances](https://kenpathlabs.com/lokah/datasets/tamil-speech-recognition-utterances)
    

## Diarisation

Single-channel recordings with speaker turns marked.

- [Hindi speaker diarization](https://kenpathlabs.com/lokah/datasets/hindi-speaker-diarization)
    
- [Tamil speaker diarization](https://kenpathlabs.com/lokah/datasets/tamil-speaker-diarization)

---

Source: https://kenpathlabs.com/lokah/samples
Site index for agents: https://kenpathlabs.com/llms.txt
