---
pretty_name: "Hindi call-centre conversations, insurance"
language:
- hi
language_details: hi-IN
license: other
license_name: lokah-custom
license_link: https://kenpathlabs.com/lokah/licensing
task_categories:
- audio-to-audio
- automatic-speech-recognition
size_categories:
- 1K<n<10K
tags:
- Hindi speech dataset
- conversational speech
- call centre audio
- dual-channel audio
- full-duplex speech-to-speech training data
- speech recognition training data
---

# Hindi call-centre conversations, insurance

> Scripted call-centre conversations in Hindi, insurance, recorded with each speaker on a separate channel. 356 hours. Transcripts are time-aligned and written in Devanagari script, with English words kept as spoken. Layouts across the set: 1,135 two-channel, 710 one side of a call.

`LK-SP-HIN-001` · [Get a quote](https://kenpathlabs.com/lokah/datasets/hindi-call-centre-insurance#contact) · [Record as JSON](https://kenpathlabs.com/api/lokah/datasets/hindi-call-centre-insurance) · [Croissant](https://kenpathlabs.com/lokah/datasets/hindi-call-centre-insurance/croissant.json)

Scripted call-centre conversations in Hindi, insurance, recorded with each speaker on a separate channel. 356 hours. Transcripts are time-aligned and written in Devanagari script, with English words kept as spoken. Layouts across the set: 1,135 two-channel, 710 one side of a call. Measured from 6 sample conversations (25 minutes): 16 kHz, 16-bit PCM WAV, two mono files per conversation. Across them 17% of transcript words are English written in Latin script, 12% of the time is silence, and there are 236 turns. 12 distinct voices in the sample (3 male, 9 female). Every recording has a complete, segment-level transcript. Personal data: redacted. Licence: custom, quoted per use.

## Specification

| Field | Value | Note |
| --- | --- | --- |
| id | LK-SP-HIN-001 |  |
| type | speech · conversational · call centre |  |
| language | हिन्दी · Hindi · hi-IN |  |
| hours | 356 h |  |
| channels | Two channels, one per speaker |  |
| files | 1,845 files · 1,845 conversations | counted across the full set |
| layouts | 1,135 two-channel, 710 one side of a call | counted across the full set |
| audio | FLAC · 16 kHz · 16-bit | measured across the full set |
| bandwidth | wideband (8 kHz) | no telephone-band files |
| snr | 33.7 dB median | across the full set |
| release | v1.0 |  |
| transcript | time-aligned by segment · Devanagari script · delivered as JSON |  |
| speakers | 490 across the full set · id and gender per speaker |  |
| pii | Numbers of 4 or more digits, digit strings spoken as words, and email addresses are masked as [pii] in text and replaced by a tone in audio. First names are not masked. |  |
| source | Recorded for the dataset |  |
| review | Every recording has a complete, segment-level transcript |  |
| personal data | Redacted |  |
| licence | Custom | quoted per use |


## What it is good for

| Task | Fit | Why |
| --- | --- | --- |
| Speech recognition | yes | Time-aligned transcripts in native script, English kept as spoken. 33.7 dB median SNR across the full set, wideband (53.3 dB in the sample conversations). |
| Full-duplex speech to speech | yes | Each speaker on a separate channel, so overlap, backchannels and turn timing survive. Moshi and PersonaPlex train on exactly this layout. |
| Turn-taking and endpointing | yes | Gaps and overlaps at every change of speaker are measured from the two channels; see the profile. |
| Voice agents for support | yes | Agent and customer turns in a real support flow. |
| Speaker diarisation | yes | Speaker-attributed segments across full conversations. |
| Text to speech | partly | 33.7 dB median SNR across the full set, wideband (53.3 dB in the sample conversations): clean and wideband enough for conversational prosody data, though not a studio voice. |

## Conversation profile

Measured from 6 full conversations (25 minutes, 236 turns, 4074 words).

| Measure | Value |
| --- | --- |
| Talk time, speaker 1 / speaker 2 | 56% / 44% |
| Silence | 18% |
| Overlapping speech | 6% |
| English words, written in Latin script | 17% |
| Speaking rate | 188 words a minute |

### Turn-taking, measured from the two channels

Voice activity detected on each channel at 10 ms; IPUs bounded by more than 200 ms of silence; talkspurts under 90 ms dropped.

| Per minute | This dataset | Fisher corpus |
| --- | --- | --- |
| Inter-pausal units | 23 | 21.6 |
| Pauses | 10.5 | 7 |
| Gaps | 5 | 7.5 |
| Overlaps | 7.3 | 6.5 |
| Backchannels | 4.3 | not reported |

Floor-transfer offset: median +0.16 s, 10th to 90th percentile -0.41 s to +1.49 s. 33% of 186 changes of speaker were overlapped. Channel isolation: -40.7 dB.

## Audio quality

Frame RMS at 20 ms on the channels mixed to mono. Noise floor is the 10th percentile, speech level the 90th; SNR is their difference. Bandwidth is the highest frequency at which speech still rises 6 dB above the recording's own noise spectrum.

| Measure | Value | Reading |
| --- | --- | --- |
| Speech above noise floor (SNR) | 53.3 dB (30 to 70.1) | clean |
| Noise floor | -77.5 dBFS |  |
| Effective bandwidth | 8.0 kHz | wideband (8 kHz) across the full set |
| Clipping | 0.000% of samples |  |
| DNSMOS P.835 (1 to 5) | background 3.56, speech 3.08, overall 2.59 | listener-rated quality, estimated |

## Samples

4 public excerpts, 32.52 seconds each, from different conversations in the dataset.

- Sample 1, Call-centre conversation, from a 4-minute conversation (Speaker 1: female; Speaker 2: female): [audio](https://kenpathlabs.com/lokah-samples/hindi-call-centre-insurance--1.m4a) · [segments](https://kenpathlabs.com/lokah-samples/hindi-call-centre-insurance--1.segments.json)
- Sample 2, Call-centre conversation, from a 4-minute conversation (Speaker 1: male; Speaker 2: female): [audio](https://kenpathlabs.com/lokah-samples/hindi-call-centre-insurance--2.m4a) · [segments](https://kenpathlabs.com/lokah-samples/hindi-call-centre-insurance--2.segments.json)
- Sample 3, Call-centre conversation, from a 4-minute conversation (Speaker 1: female; Speaker 2: female): [audio](https://kenpathlabs.com/lokah-samples/hindi-call-centre-insurance--3.m4a) · [segments](https://kenpathlabs.com/lokah-samples/hindi-call-centre-insurance--3.segments.json)
- Sample 4, Call-centre conversation, from a 4-minute conversation (Speaker 1: female; Speaker 2: female): [audio](https://kenpathlabs.com/lokah-samples/hindi-call-centre-insurance--4.m4a) · [segments](https://kenpathlabs.com/lokah-samples/hindi-call-centre-insurance--4.segments.json)

### Sample 1 transcript

Call-centre conversation, starting at 0:00. Stereo: left channel is speaker 1, right channel is speaker 2. Audio sha256 `5afb5f5e853e5673b2aab82375b865d9e96a62cf060dd765cd536010f71669e1`.

| Start | Speaker | Text |
| --- | --- | --- |
| 0:00 | Speaker 1 | hello |
| 0:01 | Speaker 2 | hello |
| 0:02 | Speaker 1 | OK ma'am आप कविता बात कर रहे हो |
| 0:05 | Speaker 2 | yes kavita spiking आप कोन बात कर रहे हो |
| 0:09 | Speaker 1 | OK ma'am में अर्पित बात कर रही हूँ |
| 0:12 | Speaker 2 | हाँजी बोलिए |
| 0:13 | Speaker 1 | मेने आपको न car insurance के बारे मे details देने के लिए call किया है |
| 0:19 | Speaker 2 | OK आ बोलिए |
| 0:22 | Speaker 1 | OK ma'am तो आपके पास car insurance है या फिर आपको renew करवन है है या फीर आपको new लेना है |
| 0:29 | Speaker 2 | आ actually मुझे renew करवना है |

## Licence

Licence: custom, quoted per use (training, evaluation or both; internal or commercial; exclusive or not). Delivered in the layout your training code reads: https://kenpathlabs.com/lokah/formats.

---

Machine-readable: [llms.txt](https://kenpathlabs.com/llms.txt) · [JSON API](https://kenpathlabs.com/api/lokah/datasets) · [OpenAPI](https://kenpathlabs.com/lokah/openapi.json) · [MCP](https://kenpathlabs.com/api/lokah/mcp) · [for agents](https://kenpathlabs.com/lokah/for-agents)
