---
pretty_name: "Tamil general conversation"
language:
- ta
language_details: ta-IN
license: other
license_name: lokah-custom
license_link: https://kenpathlabs.com/lokah/licensing
task_categories:
- automatic-speech-recognition
size_categories:
- 1K<n<10K
tags:
- Tamil speech dataset
- conversational speech
- spontaneous conversation
- speech recognition training data
- code-switching
- tam
---

# Tamil general conversation

> Two-speaker general conversation in Tamil, recorded on one channel with speakers labelled. 105 hours. Transcripts are time-aligned and written in Tamil script, with English words kept as spoken.

`LK-SP-TAM-003` · [Get a quote](https://kenpathlabs.com/lokah/datasets/tamil-general-conversation#contact) · [Record as JSON](https://kenpathlabs.com/api/lokah/datasets/tamil-general-conversation) · [Croissant](https://kenpathlabs.com/lokah/datasets/tamil-general-conversation/croissant.json)

Two-speaker general conversation in Tamil, recorded on one channel with speakers labelled. 105 hours. Transcripts are time-aligned and written in Tamil script, with English words kept as spoken. Measured from 6 sample conversations (25 minutes): 16 kHz, 16-bit PCM WAV, one mono file per conversation. Across them 36% of transcript words are English written in Latin script, 6% of the time is silence, and there are 384 turns. 12 distinct voices in the sample (9 female, 3 male). Every recording has a complete, segment-level transcript. Personal data: redacted. Licence: custom, quoted per use.

## Specification

| Field | Value | Note |
| --- | --- | --- |
| id | LK-SP-TAM-003 |  |
| type | speech · conversational · general |  |
| language | தமிழ் · Tamil · ta-IN |  |
| hours | 105 h |  |
| channels | One channel, speakers labelled in the transcript |  |
| files | 464 files · 464 conversations | counted across the full set |
| layouts | 464 mono conversation | counted across the full set |
| audio | FLAC · 16 kHz · 16-bit | measured across the full set |
| bandwidth | wideband (8 kHz) | no telephone-band files |
| snr | 25.7 dB median | across the full set |
| release | v1.0 |  |
| transcript | time-aligned by segment · Tamil script · delivered as JSON |  |
| speakers | 279 across the full set · id and gender per speaker |  |
| pii | Numbers of 4 or more digits, digit strings spoken as words, and email addresses are masked as [pii] in text and replaced by a tone in audio. First names are not masked. |  |
| source | Recorded for the dataset |  |
| review | Every recording has a complete, segment-level transcript |  |
| personal data | Redacted |  |
| licence | Custom | quoted per use |


## What it is good for

| Task | Fit | Why |
| --- | --- | --- |
| Speech recognition | yes | Time-aligned transcripts in native script, English kept as spoken. 25.7 dB median SNR across the full set, wideband (42.1 dB in the sample conversations). |
| Full-duplex speech to speech | partly | One mixed channel. Turns are labelled, but overlapping speech cannot be separated, which these models need. |
| Turn-taking and endpointing | partly | Turn boundaries come from the transcript, so gaps are approximate and overlap is marked, not separated. |
| Voice agents for support | partly | Open conversation, not a support flow. Useful for language and prosody, not for task structure. |
| Speaker diarisation | yes | Speaker-attributed segments across full conversations. |
| Text to speech | no | 25.7 dB median SNR across the full set, wideband (42.1 dB in the sample conversations): too much background for a voice model. |

## Conversation profile

Measured from 6 full conversations (25 minutes, 384 turns, 3333 words).

| Measure | Value |
| --- | --- |
| Talk time, speaker 1 / speaker 2 | 59% / 41% |
| Silence | 6% |
| Overlapping speech | 0% |
| English words, written in Latin script | 36% |
| Speaking rate | 143.7 words a minute |

One mixed channel: segment edges were placed by an annotator, so turn timing is approximate and overlap is an event label.

## Audio quality

Frame RMS at 20 ms on the channels mixed to mono. Noise floor is the 10th percentile, speech level the 90th; SNR is their difference. Bandwidth is the highest frequency at which speech still rises 6 dB above the recording's own noise spectrum.

| Measure | Value | Reading |
| --- | --- | --- |
| Speech above noise floor (SNR) | 42.1 dB (29.4 to 56.4) | clean |
| Noise floor | -60.2 dBFS |  |
| Effective bandwidth | 8.0 kHz | wideband (8 kHz) across the full set |
| Clipping | 0.000% of samples |  |
| DNSMOS P.835 (1 to 5) | background 3, speech 3, overall 2.33 | listener-rated quality, estimated |

## Samples

4 public excerpts, 44.686 seconds each, from different conversations in the dataset.

- Sample 1, Open conversation, from a 4-minute conversation (Speaker 1: female; Speaker 2: female): [audio](https://kenpathlabs.com/lokah-samples/tamil-general-conversation--1.m4a) · [segments](https://kenpathlabs.com/lokah-samples/tamil-general-conversation--1.segments.json)
- Sample 2, Open conversation, from a 4-minute conversation (Speaker 1: male; Speaker 2: female): [audio](https://kenpathlabs.com/lokah-samples/tamil-general-conversation--2.m4a) · [segments](https://kenpathlabs.com/lokah-samples/tamil-general-conversation--2.segments.json)
- Sample 3, Open conversation, from a 4-minute conversation (Speaker 1: male; Speaker 2: female): [audio](https://kenpathlabs.com/lokah-samples/tamil-general-conversation--3.m4a) · [segments](https://kenpathlabs.com/lokah-samples/tamil-general-conversation--3.segments.json)
- Sample 4, Open conversation, from a 4-minute conversation (Speaker 1: female; Speaker 2: male): [audio](https://kenpathlabs.com/lokah-samples/tamil-general-conversation--4.m4a) · [segments](https://kenpathlabs.com/lokah-samples/tamil-general-conversation--4.segments.json)

### Sample 1 transcript

Open conversation, starting at 0:00. Audio sha256 `5f6ca60c1d7ed0e2558383d4cbe91fab14a5be6c3f80729d49c058744a8fbba7`.

| Start | Speaker | Text |
| --- | --- | --- |
| 0:00 | Speaker 1 | hello. |
| 0:01 | Speaker 2 | hello. |
| 0:02 | Speaker 1 | [overlap] [filler] யாரு பேசுறீங்க [overlap] hello, good morning mam. |
| 0:04 | Speaker 2 | நான் தியா, நாங்க அபிராமி finance ல இருந்து பேசுறோம் mam. |
| 0:07 | Speaker 1 | yeah, சொல்லுங்க mam. |
| 0:09 | Speaker 2 | [filler], mam உங்களோட name அ நாங்க verify பண்ணிக்கிறேன் mam. |
| 0:13 | Speaker 1 | [overlap] yeah, [unintelligible] [overlap] உங்களோட name வந்து |
| 0:14 | Speaker 2 | பிரீத்தி சிவக்குமார் தான mam? |
| 0:16 | Speaker 1 | அ yes mam. |
| 0:17 | Speaker 2 | OK mam, உங்களோட place வந்து மதுரை தான mam? |
| 0:21 | Speaker 1 | [filler] ஆமா. |
| 0:21 | Speaker 2 | [filler] OK mam, உங்களோட address நான் சொல்றேன் mam correct டானு பாத்துக்கோங்க mam. |
| 0:26 | Speaker 1 | [filler] OK mam சொல்லுங்க. |
| 0:27 | Speaker 2 | [suppressed] twenty nine six bar |
| 0:29 | Speaker 1 | [filler] |
| 0:30 | Speaker 2 | [suppressed] அம்மன் கோவில், north street மதுரை correct அ mam. |
| 0:33 | Speaker 1 | ம் correct மா. |
| 0:34 | Speaker 2 | [filler] OK mam, mam உங்களுக்கு loan எடுக்குறதுக்கு ஏதாவது idea இருக்கா mam? |
| 0:40 | Speaker 1 | ஆமா, அது அந்த மாதிரி ஒரு idea இருந்தனால தான் உங்க website பாத்தேன் நானு. |

## Licence

Licence: custom, quoted per use (training, evaluation or both; internal or commercial; exclusive or not). Delivered in the layout your training code reads: https://kenpathlabs.com/lokah/formats.

---

Machine-readable: [llms.txt](https://kenpathlabs.com/llms.txt) · [JSON API](https://kenpathlabs.com/api/lokah/datasets) · [OpenAPI](https://kenpathlabs.com/lokah/openapi.json) · [MCP](https://kenpathlabs.com/api/lokah/mcp) · [for agents](https://kenpathlabs.com/lokah/for-agents)
