---
pretty_name: "Tamil speaker diarization"
language:
- ta
language_details: ta-IN
license: other
license_name: lokah-custom
license_link: https://kenpathlabs.com/lokah/licensing
task_categories:
- automatic-speech-recognition
size_categories:
- 1K<n<10K
tags:
- Tamil speech dataset
- conversational speech
- speaker diarization dataset
- call centre and spontaneous conversation
- speech recognition training data
- code-switching
---

# Tamil speaker diarization

> Tamil speaker diarization data: whole call-centre and everyday two-speaker conversations with every speaker turn marked, 345,825 turns in RTTM, for training and scoring who spoke when. 229 hours across 1,307 recordings.

`LK-SP-TAM-005` · [Get a quote](https://kenpathlabs.com/lokah/datasets/tamil-speaker-diarization#contact) · [Record as JSON](https://kenpathlabs.com/api/lokah/datasets/tamil-speaker-diarization) · [Croissant](https://kenpathlabs.com/lokah/datasets/tamil-speaker-diarization/croissant.json)

Tamil speaker diarization data: whole call-centre and everyday two-speaker conversations with every speaker turn marked, 345,825 turns in RTTM, for training and scoring who spoke when. 229 hours across 1,307 recordings. Measured from 6 sample conversations (25 minutes): 16 kHz, 16-bit PCM WAV, one mono file per conversation. Across them 39% of transcript words are English written in Latin script, 5% of the time is silence, and there are 385 turns. Every recording has a complete, segment-level transcript. Personal data: redacted. Licence: custom, quoted per use.

## Specification

| Field | Value | Note |
| --- | --- | --- |
| id | LK-SP-TAM-005 |  |
| type | speech · speaker diarization · call centre and general |  |
| language | தமிழ் · Tamil · ta-IN |  |
| hours | 229 h |  |
| channels | One channel, speakers labelled in the transcript |  |
| files | 1,307 files · 345,825 speaker turns | counted across the full set |
| layouts | 1,307 mono conversation | counted across the full set |
| audio | FLAC · 16 kHz · 16-bit | measured across the full set |
| bandwidth | wideband (8 kHz) |  |
| snr | 25.7 dB median | across the full set |
| release | v1.0 |  |
| transcript | time-aligned by segment · Tamil script · delivered as JSON |  |
| speakers | 1,238 across the full set · id and gender per speaker |  |
| pii | Numbers of 4 or more digits, digit strings spoken as words, and email addresses are masked as [pii] in text and replaced by a tone in audio. First names are not masked. |  |
| source | Recorded for the dataset |  |
| review | Every recording has a complete, segment-level transcript |  |
| personal data | Redacted |  |
| licence | Custom | quoted per use |


## What it is good for

| Task | Fit | Why |
| --- | --- | --- |
| Speech recognition | partly | Whole recordings with time-aligned transcripts; built for who spoke when, not for utterance-level training. |
| Full-duplex speech to speech | no | Single-channel recordings. |
| Turn-taking | yes | Every change of speaker is marked: 345,825 turns. |
| Voice agents | partly | Teaches an agent who is speaking, and when. |
| Diarisation | yes | Built for it: RTTM turn files for 1,307 recordings. |
| Text to speech | no | Conversational call audio, not studio voice. |

## Conversation profile

Measured from 6 full conversations (25 minutes, 385 turns, 3164 words).

| Measure | Value |
| --- | --- |
| Talk time, speaker 1 / speaker 2 | 67% / 33% |
| Silence | 5% |
| Overlapping speech | 0% |
| English words, written in Latin script | 39% |
| Speaking rate | 136.2 words a minute |

One mixed channel: segment edges were placed by an annotator, so turn timing is approximate and overlap is an event label.

## Audio quality

Frame RMS at 20 ms on the channels mixed to mono. Noise floor is the 10th percentile, speech level the 90th; SNR is their difference. Bandwidth is the highest frequency at which speech still rises 6 dB above the recording's own noise spectrum.

| Measure | Value | Reading |
| --- | --- | --- |
| Speech above noise floor (SNR) | 47.9 dB (27.3 to 98.7) | clean |
| Noise floor | -74 dBFS |  |
| Effective bandwidth | 7.9 kHz | wideband (8 kHz) across the full set |
| Clipping | 0.000% of samples |  |
| DNSMOS P.835 (1 to 5) | background 3.06, speech 3.04, overall 2.37 | listener-rated quality, estimated |

## Samples

4 public excerpts, 44.131 seconds each, from different conversations in the dataset.

- Sample 1, Open conversation, from a 4-minute conversation (Speaker 1: male; Speaker 2: female): [audio](https://kenpathlabs.com/lokah-samples/tamil-speaker-diarization--1.m4a) · [segments](https://kenpathlabs.com/lokah-samples/tamil-speaker-diarization--1.segments.json)
- Sample 2, Open conversation, from a 4-minute conversation (Speaker 1: male; Speaker 2: female): [audio](https://kenpathlabs.com/lokah-samples/tamil-speaker-diarization--2.m4a) · [segments](https://kenpathlabs.com/lokah-samples/tamil-speaker-diarization--2.segments.json)
- Sample 3, Open conversation, from a 4-minute conversation (Speaker 1: male; Speaker 2: female): [audio](https://kenpathlabs.com/lokah-samples/tamil-speaker-diarization--3.m4a) · [segments](https://kenpathlabs.com/lokah-samples/tamil-speaker-diarization--3.segments.json)
- Sample 4, Open conversation, from a 4-minute conversation (Speaker 1: female; Speaker 2: male): [audio](https://kenpathlabs.com/lokah-samples/tamil-speaker-diarization--4.m4a) · [segments](https://kenpathlabs.com/lokah-samples/tamil-speaker-diarization--4.segments.json)

### Sample 1 transcript

Open conversation, starting at 1:05. Audio sha256 `6ad726466610b5ae14cb72ca4e618af7cedcd2e9198a1deb0594e98821391226`.

| Start | Speaker | Text |
| --- | --- | --- |
| 0:00 | Speaker 1 | hair ஐ தொட்டு பாக்கும்போது soft ஆ இருக்கறதுக்கு madam, |
| 0:02 | Speaker 2 | four. |
| 0:03 | Speaker 1 | hair வந்து சிக்கு இல்லாம இருக்கா easy யா சீவுறதுக்கு வந்து பாக்கும்போது, |
| 0:07 | Speaker 2 | four. |
| 0:08 | Speaker 1 | hair shining க்கு madam, |
| 0:09 | Speaker 2 | five. |
| 0:10 | Speaker 1 | hair வந்து dry ஆகாம வச்சுருக்குறதுக்கு madam, |
| 0:13 | Speaker 2 | four. |
| 0:15 | Speaker 1 | [filler] conditioning effect க்கு madam, [silence] |
| 0:18 | Speaker 2 | [filler] five. |
| 0:20 | Speaker 1 | [unintelligible] துகள்களா hair ல தங்காம easy யா remove ஆகுறதுக்கு madam, |
| 0:23 | Speaker 2 | four. [silence] |
| 0:24 | Speaker 1 | body heat reduce பண்றதுக்கு madam, [silence] |
| 0:27 | Speaker 2 | four. [silence] |
| 0:29 | Speaker 1 | நுரையோட அளவுக்கு, [silence] |
| 0:31 | Speaker 2 | three. [silence] |
| 0:33 | Speaker 1 | smooth ஆ இருக்குறதுக்கு madam, hair smooth ஆ soft ஆ இருக்குறதுக்கு madam. |
| 0:35 | Speaker 2 | four. |
| 0:36 | Speaker 1 | dandruff control பண்றதுக்கு madam [silence], |
| 0:39 | Speaker 2 | four. |
| 0:40 | Speaker 1 | [unintelligible] எல்லாம் remove பண்றதுக்கு madam, [silence] |
| 0:43 | Speaker 2 | five. |

## Licence

Licence: custom, quoted per use (training, evaluation or both; internal or commercial; exclusive or not). Delivered in the layout your training code reads: https://kenpathlabs.com/lokah/formats.

---

Machine-readable: [llms.txt](https://kenpathlabs.com/llms.txt) · [JSON API](https://kenpathlabs.com/api/lokah/datasets) · [OpenAPI](https://kenpathlabs.com/lokah/openapi.json) · [MCP](https://kenpathlabs.com/api/lokah/mcp) · [for agents](https://kenpathlabs.com/lokah/for-agents)
