---
pretty_name: "Hindi speaker diarization"
language:
- hi
language_details: hi-IN
license: other
license_name: lokah-custom
license_link: https://kenpathlabs.com/lokah/licensing
task_categories:
- automatic-speech-recognition
size_categories:
- 1K<n<10K
tags:
- Hindi speech dataset
- conversational speech
- speaker diarization dataset
- call centre and spontaneous conversation
- speech recognition training data
- code-switching
---

# Hindi speaker diarization

> Hindi speaker diarization data: whole call-centre and everyday two-speaker conversations with every speaker turn marked, 432,929 turns in RTTM, for training and scoring who spoke when. 321 hours across 1,394 recordings.

`LK-SP-HIN-004` · [Get a quote](https://kenpathlabs.com/lokah/datasets/hindi-speaker-diarization#contact) · [Record as JSON](https://kenpathlabs.com/api/lokah/datasets/hindi-speaker-diarization) · [Croissant](https://kenpathlabs.com/lokah/datasets/hindi-speaker-diarization/croissant.json)

Hindi speaker diarization data: whole call-centre and everyday two-speaker conversations with every speaker turn marked, 432,929 turns in RTTM, for training and scoring who spoke when. 321 hours across 1,394 recordings. Measured from 6 sample conversations (24 minutes): 16 kHz, 16-bit PCM WAV, one mono file per conversation. Across them 9% of transcript words are English written in Latin script, 17% of the time is silence, and there are 320 turns. Every recording has a complete, segment-level transcript. Personal data: redacted. Licence: custom, quoted per use.

## Specification

| Field | Value | Note |
| --- | --- | --- |
| id | LK-SP-HIN-004 |  |
| type | speech · speaker diarization · call centre and general |  |
| language | हिन्दी · Hindi · hi-IN |  |
| hours | 321 h |  |
| channels | One channel, speakers labelled in the transcript |  |
| files | 1,394 files · 432,929 speaker turns | counted across the full set |
| layouts | 1,394 mono conversation | counted across the full set |
| audio | FLAC · 16 kHz · 16-bit | measured across the full set |
| bandwidth | wideband (8 kHz) |  |
| snr | 30.3 dB median | across the full set |
| release | v1.0 |  |
| transcript | time-aligned by segment · Devanagari script · delivered as JSON |  |
| speakers | 619 across the full set · id and gender per speaker |  |
| pii | Numbers of 4 or more digits, digit strings spoken as words, and email addresses are masked as [pii] in text and replaced by a tone in audio. First names are not masked. |  |
| source | Recorded for the dataset |  |
| review | Every recording has a complete, segment-level transcript |  |
| personal data | Redacted |  |
| licence | Custom | quoted per use |


## What it is good for

| Task | Fit | Why |
| --- | --- | --- |
| Speech recognition | partly | Whole recordings with time-aligned transcripts; built for who spoke when, not for utterance-level training. |
| Full-duplex speech to speech | no | Single-channel recordings. |
| Turn-taking | yes | Every change of speaker is marked: 432,929 turns. |
| Voice agents | partly | Teaches an agent who is speaking, and when. |
| Diarisation | yes | Built for it: RTTM turn files for 1,394 recordings. |
| Text to speech | no | Conversational call audio, not studio voice. |

## Conversation profile

Measured from 6 full conversations (24 minutes, 320 turns, 3402 words).

| Measure | Value |
| --- | --- |
| Talk time, speaker 1 / speaker 2 | 43% / 57% |
| Silence | 17% |
| Overlapping speech | 9% |
| English words, written in Latin script | 9% |
| Speaking rate | 168 words a minute |

One mixed channel: segment edges were placed by an annotator, so turn timing is approximate and overlap is an event label.

## Audio quality

Frame RMS at 20 ms on the channels mixed to mono. Noise floor is the 10th percentile, speech level the 90th; SNR is their difference. Bandwidth is the highest frequency at which speech still rises 6 dB above the recording's own noise spectrum.

| Measure | Value | Reading |
| --- | --- | --- |
| Speech above noise floor (SNR) | 54.4 dB (28.3 to 81.7) | clean |
| Noise floor | -70.2 dBFS |  |
| Effective bandwidth | 8.0 kHz | wideband (8 kHz) across the full set |
| Clipping | 0.009% of samples |  |
| DNSMOS P.835 (1 to 5) | background 3.42, speech 3.04, overall 2.56 | listener-rated quality, estimated |

## Samples

4 public excerpts, 44.855 seconds each, from different conversations in the dataset.

- Sample 1, Open conversation, from a 4-minute conversation: [audio](https://kenpathlabs.com/lokah-samples/hindi-speaker-diarization--1.m4a) · [segments](https://kenpathlabs.com/lokah-samples/hindi-speaker-diarization--1.segments.json)
- Sample 2, Open conversation, from a 4-minute conversation: [audio](https://kenpathlabs.com/lokah-samples/hindi-speaker-diarization--2.m4a) · [segments](https://kenpathlabs.com/lokah-samples/hindi-speaker-diarization--2.segments.json)
- Sample 3, Open conversation, from a 4-minute conversation: [audio](https://kenpathlabs.com/lokah-samples/hindi-speaker-diarization--3.m4a) · [segments](https://kenpathlabs.com/lokah-samples/hindi-speaker-diarization--3.segments.json)
- Sample 4, Open conversation, from a 4-minute conversation (Speaker 1: male; Speaker 2: male): [audio](https://kenpathlabs.com/lokah-samples/hindi-speaker-diarization--4.m4a) · [segments](https://kenpathlabs.com/lokah-samples/hindi-speaker-diarization--4.segments.json)

### Sample 1 transcript

Open conversation, starting at 0:31. Audio sha256 `ce52248430a77ce3dac29efedbc75ee4ae26ec76c4c40b32530ed82d80fe535a`.

| Start | Speaker | Text |
| --- | --- | --- |
| 0:00 | Speaker 1 | hello |
| 0:01 | Speaker 2 | हाँ madam [filler] |
| 0:03 | Speaker 1 | hello |
| 0:05 | Speaker 2 | hello |
| 0:06 | Speaker 1 | हाँ राजू |
| 0:08 | Speaker 2 | हाँ |
| 0:10 | Speaker 1 | बोलो |
| 0:12 | Speaker 2 | [filler] हम तो यहाँ में है घर में तुम्हार |
| 0:18 | Speaker 1 | hello |
| 0:19 | Speaker 2 | hello [overlap] हाँ क्या कर रहा है? |
| 0:22 | Speaker 1 | [overlap] हाँ |
| 0:22 | Speaker 1 | hello |
| 0:24 | Speaker 2 | हाँ क्या कर रहा है? |
| 0:25 | Speaker 1 | अरे तुम बात मत करो ना. record कर रहा था तो hello |
| 0:30 | Speaker 2 | hello |
| 0:31 | Speaker 1 | हाँ राजू |
| 0:33 | Speaker 2 | ओय |
| 0:33 | Speaker 1 | हाँ बोलो |
| 0:35 | Speaker 2 | क्या करना है? बोलके आप जब |
| 0:37 | Speaker 1 | हं हम तो ऐसे बैठा रहा था घर पे |
| 0:41 | Speaker 2 | हं अच्छा अच्छा कल का loanding में है ना. बस्ती में है. |

## Licence

Licence: custom, quoted per use (training, evaluation or both; internal or commercial; exclusive or not). Delivered in the layout your training code reads: https://kenpathlabs.com/lokah/formats.

---

Machine-readable: [llms.txt](https://kenpathlabs.com/llms.txt) · [JSON API](https://kenpathlabs.com/api/lokah/datasets) · [OpenAPI](https://kenpathlabs.com/lokah/openapi.json) · [MCP](https://kenpathlabs.com/api/lokah/mcp) · [for agents](https://kenpathlabs.com/lokah/for-agents)
