---
pretty_name: "Hindi and Tamil full-duplex call-centre conversations"
language:
- hi
language_details: hi-IN
license: other
license_name: lokah-custom
license_link: https://kenpathlabs.com/lokah/licensing
task_categories:
- audio-to-audio
- automatic-speech-recognition
size_categories:
- 1K<n<10K
tags:
- Hindi speech dataset
- Tamil speech dataset
- conversational speech
- full-duplex dialogue data
- call centre audio
- dual-channel audio
---

# Hindi and Tamil full-duplex call-centre conversations

> Hindi and Tamil full-duplex conversation data: 1,265 two-channel call-centre calls with each speaker on a separate channel, overlaps, backchannels and turn timing preserved, transcripts time-aligned per channel. 264 hours.

`LK-SP-MUL-001` · [Get a quote](https://kenpathlabs.com/lokah/datasets/hindi-tamil-full-duplex-call-centre-conversations#contact) · [Record as JSON](https://kenpathlabs.com/api/lokah/datasets/hindi-tamil-full-duplex-call-centre-conversations) · [Croissant](https://kenpathlabs.com/lokah/datasets/hindi-tamil-full-duplex-call-centre-conversations/croissant.json)

Hindi and Tamil full-duplex conversation data: 1,265 two-channel call-centre calls with each speaker on a separate channel, overlaps, backchannels and turn timing preserved, transcripts time-aligned per channel. 264 hours. Measured from 6 sample conversations (25 minutes): 16 kHz, 16-bit PCM WAV, two mono files per conversation. Across them 16% of transcript words are English written in Latin script, 11% of the time is silence, and there are 278 turns. 11 distinct voices in the sample (6 male, 5 female). Every recording has a complete, segment-level transcript. Personal data: redacted. Licence: custom, quoted per use.

## Specification

| Field | Value | Note |
| --- | --- | --- |
| id | LK-SP-MUL-001 |  |
| type | speech · full-duplex calls · call centre |  |
| language | हिन्दी · Hindi · hi-IN / தமிழ் · Tamil · ta-IN |  |
| hours | 264 h |  |
| channels | Two channels, one per speaker |  |
| files | 1,265 files · 1,265 calls | counted across the full set |
| layouts | 1,265 mono conversation | counted across the full set |
| audio | FLAC · 16 kHz · 16-bit | measured across the full set |
| bandwidth | wideband (8 kHz) |  |
| snr | 29.2 dB median | across the full set |
| release | v1.0 |  |
| transcript | time-aligned by segment · Devanagari and Tamil script · delivered as JSON |  |
| speakers | 529 across the full set · id and gender per speaker |  |
| pii | Numbers of 4 or more digits, digit strings spoken as words, and email addresses are masked as [pii] in text and replaced by a tone in audio. First names are not masked. |  |
| source | Recorded for the dataset |  |
| review | Every recording has a complete, segment-level transcript |  |
| personal data | Redacted |  |
| licence | Custom | quoted per use |


## What it is good for

| Task | Fit | Why |
| --- | --- | --- |
| Speech recognition | partly | Per-channel transcripts are included; built for dialogue timing, not for utterance-level training. |
| Full-duplex speech to speech | yes | Built for it: 1,265 two-channel calls, each speaker on a separate channel, overlaps and timing intact. |
| Turn-taking | yes | Floor transfers, overlaps and backchannels are measurable from the two channels. |
| Voice agents | yes | Real agent and customer turns in a support flow, the condition a deployed agent hears. |
| Diarisation | partly | Two speakers, already separated by channel; useful as ground truth, not as a hard case. |
| Text to speech | no | Conversational call audio, not studio voice. |

## Conversation profile

Measured from 6 full conversations (25 minutes, 278 turns, 3540 words).

| Measure | Value |
| --- | --- |
| Talk time, speaker 1 / speaker 2 | 54% / 46% |
| Silence | 23% |
| Overlapping speech | 8% |
| English words, written in Latin script | 16% |
| Speaking rate | 161.6 words a minute |

### Turn-taking, measured from the two channels

Voice activity detected on each channel at 10 ms; IPUs bounded by more than 200 ms of silence; talkspurts under 90 ms dropped.

| Per minute | This dataset | Fisher corpus |
| --- | --- | --- |
| Inter-pausal units | 34.3 | 21.6 |
| Pauses | 11.7 | 7 |
| Gaps | 8.6 | 7.5 |
| Overlaps | 13.7 | 6.5 |
| Backchannels | 8.6 | not reported |

Floor-transfer offset: median +0.20 s, 10th to 90th percentile -0.38 s to +1.21 s. 35% of 331 changes of speaker were overlapped. Channel isolation: -39.4 dB.

## Audio quality

Frame RMS at 20 ms on the channels mixed to mono. Noise floor is the 10th percentile, speech level the 90th; SNR is their difference. Bandwidth is the highest frequency at which speech still rises 6 dB above the recording's own noise spectrum.

| Measure | Value | Reading |
| --- | --- | --- |
| Speech above noise floor (SNR) | 53.2 dB (31.5 to 61.1) | clean |
| Noise floor | -84.1 dBFS |  |
| Effective bandwidth | 7.9 kHz | wideband (8 kHz) across the full set |
| Clipping | 0.000% of samples |  |
| DNSMOS P.835 (1 to 5) | background 3.55, speech 3.17, overall 2.65 | listener-rated quality, estimated |

## Samples

4 public excerpts, 42.898 seconds each, from different conversations in the dataset.

- Sample 1, Two-channel call, from a 4-minute conversation (Speaker 1: male; Speaker 2: male): [audio](https://kenpathlabs.com/lokah-samples/hindi-tamil-full-duplex-call-centre-conversations--1.m4a) · [segments](https://kenpathlabs.com/lokah-samples/hindi-tamil-full-duplex-call-centre-conversations--1.segments.json)
- Sample 2, Two-channel call, from a 4-minute conversation (Speaker 1: male; Speaker 2: male): [audio](https://kenpathlabs.com/lokah-samples/hindi-tamil-full-duplex-call-centre-conversations--2.m4a) · [segments](https://kenpathlabs.com/lokah-samples/hindi-tamil-full-duplex-call-centre-conversations--2.segments.json)
- Sample 3, Two-channel call, from a 4-minute conversation (Speaker 1: male; Speaker 2: female): [audio](https://kenpathlabs.com/lokah-samples/hindi-tamil-full-duplex-call-centre-conversations--3.m4a) · [segments](https://kenpathlabs.com/lokah-samples/hindi-tamil-full-duplex-call-centre-conversations--3.segments.json)
- Sample 4, Two-channel call, from a 4-minute conversation (Speaker 1: female; Speaker 2: male): [audio](https://kenpathlabs.com/lokah-samples/hindi-tamil-full-duplex-call-centre-conversations--4.m4a) · [segments](https://kenpathlabs.com/lokah-samples/hindi-tamil-full-duplex-call-centre-conversations--4.segments.json)

### Sample 1 transcript

Two-channel call, starting at 0:03. Stereo: left channel is speaker 1, right channel is speaker 2. Audio sha256 `0be133fcc76d86e7c37e671583aba4c7af90ce1d9f031790dc89f4e4209f37c4`.

| Start | Speaker | Text |
| --- | --- | --- |
| 0:00 | Speaker 1 | hello |
| 0:01 | Speaker 2 | hello |
| 0:03 | Speaker 1 | नमस्कार sir |
| 0:04 | Speaker 2 | नमश्कार. |
| 0:06 | Speaker 1 | sir आप company |
| 0:08 | Speaker 2 | जी बताइए आप कौन? |
| 0:11 | Speaker 1 | sir मैं अरविंदर सिंह बात कर रहा हूँ. |
| 0:15 | Speaker 2 | जी |
| 0:17 | Speaker 1 | OK |
| 0:20 | Speaker 2 | जी जरूर आप की ही सेवा में बैठे हैं. |
| 0:23 | Speaker 1 | actually मैंने अभी तक कोई policy ली नहीं है और कोई बीमा भी नहीं करवाया है. तो मुझे उसके बारे में जानकारी नहीं है तो sir मुझे थोड़ा आप विस्तार से उसके बारे में बता सके बीमा के. |
| 0:27 | Speaker 2 | जी |
| 0:34 | Speaker 2 | अच्छा जी जी जी. |
| 0:36 | Speaker 1 | जी sir actually क्या है कि बीमा मुझे तो दो तीन बीमे करवाने है. |
| 0:41 | Speaker 2 | अच्छा. |

## Licence

Licence: custom, quoted per use (training, evaluation or both; internal or commercial; exclusive or not). Delivered in the layout your training code reads: https://kenpathlabs.com/lokah/formats.

---

Machine-readable: [llms.txt](https://kenpathlabs.com/llms.txt) · [JSON API](https://kenpathlabs.com/api/lokah/datasets) · [OpenAPI](https://kenpathlabs.com/lokah/openapi.json) · [MCP](https://kenpathlabs.com/api/lokah/mcp) · [for agents](https://kenpathlabs.com/lokah/for-agents)
