---
pretty_name: "Tamil two-channel customer-service calls"
language:
- ta
language_details: ta-IN
license: other
license_name: lokah-custom
license_link: https://kenpathlabs.com/lokah/licensing
task_categories:
- audio-to-audio
- automatic-speech-recognition
size_categories:
- n<1K
tags:
- Tamil speech dataset
- conversational speech
- call centre audio
- dual-channel audio
- full-duplex speech-to-speech training data
- speech recognition training data
---

# Tamil two-channel customer-service calls

> Scripted call-centre conversations in Tamil, telecom, delivery, e-commerce and banking, recorded with each speaker on a separate channel. 25 hours. Transcripts are time-aligned and written in Tamil script, with English words kept as spoken. Layouts across the set: 2 one side of a call, 132 two-channel.

`LK-SP-TAM-001` · [Get a quote](https://kenpathlabs.com/lokah/datasets/tamil-call-centre-banking-and-retail#contact) · [Record as JSON](https://kenpathlabs.com/api/lokah/datasets/tamil-call-centre-banking-and-retail) · [Croissant](https://kenpathlabs.com/lokah/datasets/tamil-call-centre-banking-and-retail/croissant.json)

Scripted call-centre conversations in Tamil, telecom, delivery, e-commerce and banking, recorded with each speaker on a separate channel. 25 hours. Transcripts are time-aligned and written in Tamil script, with English words kept as spoken. Layouts across the set: 2 one side of a call, 132 two-channel. Measured from 6 sample conversations (24 minutes): 16 kHz, 16-bit PCM WAV, two mono files per conversation. Across them 29% of transcript words are English written in Latin script, 20% of the time is silence, and there are 269 turns. 10 distinct voices in the sample (1 male, 7 female, 2 unknown). Every recording has a complete, segment-level transcript. Personal data: redacted. Licence: custom, quoted per use.

## Specification

| Field | Value | Note |
| --- | --- | --- |
| id | LK-SP-TAM-001 |  |
| type | speech · conversational · call centre |  |
| language | தமிழ் · Tamil · ta-IN |  |
| hours | 25 h |  |
| channels | Two channels, one per speaker |  |
| files | 134 files · 134 conversations | counted across the full set |
| layouts | 2 one side of a call, 132 two-channel | counted across the full set |
| audio | FLAC · 16 kHz · 16-bit | measured across the full set |
| bandwidth | wideband (8 kHz) | no telephone-band files |
| snr | 31.2 dB median | across the full set |
| release | v1.0 |  |
| transcript | time-aligned by segment · Tamil script · delivered as JSON |  |
| speakers | 93 across the full set · id and gender per speaker |  |
| pii | Numbers of 4 or more digits, digit strings spoken as words, and email addresses are masked as [pii] in text and replaced by a tone in audio. First names are not masked. |  |
| source | Recorded for the dataset |  |
| review | Every recording has a complete, segment-level transcript |  |
| personal data | Redacted |  |
| licence | Custom | quoted per use |


## What it is good for

| Task | Fit | Why |
| --- | --- | --- |
| Speech recognition | yes | Time-aligned transcripts in native script, English kept as spoken. 31.2 dB median SNR across the full set, wideband (42.7 dB in the sample conversations). |
| Full-duplex speech to speech | yes | Each speaker on a separate channel, so overlap, backchannels and turn timing survive. Moshi and PersonaPlex train on exactly this layout. |
| Turn-taking and endpointing | yes | Gaps and overlaps at every change of speaker are measured from the two channels; see the profile. |
| Voice agents for support | yes | Agent and customer turns in a real support flow. |
| Speaker diarisation | yes | Speaker-attributed segments across full conversations. |
| Text to speech | partly | 31.2 dB median SNR across the full set, wideband (42.7 dB in the sample conversations): clean and wideband enough for conversational prosody data, though not a studio voice. |

## Conversation profile

Measured from 6 full conversations (24 minutes, 269 turns, 3067 words).

| Measure | Value |
| --- | --- |
| Talk time, speaker 1 / speaker 2 | 47% / 53% |
| Silence | 25% |
| Overlapping speech | 4% |
| English words, written in Latin script | 29% |
| Speaking rate | 156.1 words a minute |

### Turn-taking, measured from the two channels

Voice activity detected on each channel at 10 ms; IPUs bounded by more than 200 ms of silence; talkspurts under 90 ms dropped.

| Per minute | This dataset | Fisher corpus |
| --- | --- | --- |
| Inter-pausal units | 21.2 | 21.6 |
| Pauses | 6.5 | 7 |
| Gaps | 8.9 | 7.5 |
| Overlaps | 5.5 | 6.5 |
| Backchannels | 4.4 | not reported |

Floor-transfer offset: median +1.05 s, 10th to 90th percentile +0.07 s to +1.92 s. 8% of 240 changes of speaker were overlapped. Channel isolation: -36.3 dB.

## Audio quality

Frame RMS at 20 ms on the channels mixed to mono. Noise floor is the 10th percentile, speech level the 90th; SNR is their difference. Bandwidth is the highest frequency at which speech still rises 6 dB above the recording's own noise spectrum.

| Measure | Value | Reading |
| --- | --- | --- |
| Speech above noise floor (SNR) | 42.7 dB (28.2 to 60.7) | clean |
| Noise floor | -70.5 dBFS |  |
| Effective bandwidth | 7.8 kHz | wideband (8 kHz) across the full set |
| Clipping | 0.000% of samples |  |
| DNSMOS P.835 (1 to 5) | background 3.78, speech 3.3, overall 2.84 | listener-rated quality, estimated |

## Samples

4 public excerpts, 43.177 seconds each, from different conversations in the dataset.

- Sample 1, Call-centre conversation, from a 4-minute conversation (Speaker 1: female; Speaker 2: female): [audio](https://kenpathlabs.com/lokah-samples/tamil-call-centre-banking-and-retail--1.m4a) · [segments](https://kenpathlabs.com/lokah-samples/tamil-call-centre-banking-and-retail--1.segments.json)
- Sample 2, Call-centre conversation, from a 4-minute conversation (Speaker 1: male; Speaker 2: female): [audio](https://kenpathlabs.com/lokah-samples/tamil-call-centre-banking-and-retail--2.m4a) · [segments](https://kenpathlabs.com/lokah-samples/tamil-call-centre-banking-and-retail--2.segments.json)
- Sample 3, Call-centre conversation, from a 4-minute conversation (Speaker 1: female; Speaker 2: female): [audio](https://kenpathlabs.com/lokah-samples/tamil-call-centre-banking-and-retail--3.m4a) · [segments](https://kenpathlabs.com/lokah-samples/tamil-call-centre-banking-and-retail--3.segments.json)
- Sample 4, Call-centre conversation, from a 4-minute conversation: [audio](https://kenpathlabs.com/lokah-samples/tamil-call-centre-banking-and-retail--4.m4a) · [segments](https://kenpathlabs.com/lokah-samples/tamil-call-centre-banking-and-retail--4.segments.json)

### Sample 1 transcript

Call-centre conversation, starting at 0:00. Stereo: left channel is speaker 1, right channel is speaker 2. Audio sha256 `2ed5586e4e3d98739f12ee23beaced07cdf0a1b6d1fadb88534bc9de532a8cb5`.

| Start | Speaker | Text |
| --- | --- | --- |
| 0:00 | Speaker 2 | hello |
| 0:00 | Speaker 1 | shree five |
| 0:03 | Speaker 1 | shree five care center வாடிக்கையாளர் மையத்துக்கு உங்களை வரவேற்கிறம் நான் திவ்யா என்ன சேவை தேவை உங்களுக்கு |
| 0:13 | Speaker 2 | madam வணக்கம் madam |
| 0:16 | Speaker 1 | வணக்கம் வணக்கம் சொல்லுங்க உங்க பேரு |
| 0:20 | Speaker 2 | madam என் பேர் காயத்திரி |
| 0:22 | Speaker 1 | ஆ சொல்லுங்க காயத்திரி |
| 0:25 | Speaker 2 | நாங்க சேலத்துல இருந்து call பண்ணிருக்கொங்க |
| 0:28 | Speaker 1 | ஆ சொல்லுங்க சொல்லுங்க காயத்திரி |
| 0:31 | Speaker 2 | madam எங்க அம்மாக்கு ஒடம்பு சரி இல்ல ஒரு ஒரு வாரத்துக்கு முன்னாடிதான் operation பண்ணிருகோம் இப்ப அவங்க விட்ல இருக்காங்க அவங்கள விட்ல வெச்சு பாத்துகோங்க சொல்லி discharge பண்ணிடாங்க |
| 0:38 | Speaker 1 | ம்ம் |

## Licence

Licence: custom, quoted per use (training, evaluation or both; internal or commercial; exclusive or not). Delivered in the layout your training code reads: https://kenpathlabs.com/lokah/formats.

---

Machine-readable: [llms.txt](https://kenpathlabs.com/llms.txt) · [JSON API](https://kenpathlabs.com/api/lokah/datasets) · [OpenAPI](https://kenpathlabs.com/lokah/openapi.json) · [MCP](https://kenpathlabs.com/api/lokah/mcp) · [for agents](https://kenpathlabs.com/lokah/for-agents)
