---
pretty_name: "Tamil speech recognition utterances"
language:
- ta
language_details: ta-IN
license: other
license_name: lokah-custom
license_link: https://kenpathlabs.com/lokah/licensing
task_categories:
- automatic-speech-recognition
size_categories:
- 1K<n<10K
tags:
- Tamil speech dataset
- conversational speech
- ASR training data
- call centre and spontaneous conversation
- speech recognition training data
- code-switching
---

# Tamil speech recognition utterances

> Tamil speech recognition data: 157,231 single-speaker utterance clips of conversational Tamil, call-centre and everyday, each with its own time-aligned transcript in Tamil script, English words kept as spoken. 201 hours of audio, 1,748,023 transcribed words.

`LK-SP-TAM-004` · [Get a quote](https://kenpathlabs.com/lokah/datasets/tamil-speech-recognition-utterances#contact) · [Record as JSON](https://kenpathlabs.com/api/lokah/datasets/tamil-speech-recognition-utterances) · [Croissant](https://kenpathlabs.com/lokah/datasets/tamil-speech-recognition-utterances/croissant.json)

Tamil speech recognition data: 157,231 single-speaker utterance clips of conversational Tamil, call-centre and everyday, each with its own time-aligned transcript in Tamil script, English words kept as spoken. 201 hours of audio, 1,748,023 transcribed words. Every recording has a complete, segment-level transcript. Personal data: redacted. Licence: custom, quoted per use.

## Specification

| Field | Value | Note |
| --- | --- | --- |
| id | LK-SP-TAM-004 |  |
| type | speech · utterances for ASR · call centre and general |  |
| language | தமிழ் · Tamil · ta-IN |  |
| hours | 201 h |  |
| channels | One channel, one speaker per clip |  |
| files | 157,231 files · 157,231 utterances | counted across the full set |
| audio | FLAC · 16 kHz · 16-bit | measured across the full set |
| release | v1.0 |  |
| transcript | time-aligned by segment · Tamil script |  |
| speakers | 1,236 across the full set · id and gender per speaker |  |
| pii | Numbers of 4 or more digits, digit strings spoken as words, and email addresses are masked as [pii] in text and replaced by a tone in audio. First names are not masked. |  |
| source | Recorded for the dataset |  |
| review | Every recording has a complete, segment-level transcript |  |
| personal data | Redacted |  |
| licence | Custom | quoted per use |


## What it is good for

| Task | Fit | Why |
| --- | --- | --- |
| Speech recognition | yes | Built for it: 157,231 utterance clips, each with its own time-aligned transcript, noise left as recorded. |
| Full-duplex speech to speech | no | Clips are single utterances; the conversation timing is not in this dataset. |
| Turn-taking | no | No turn structure: each clip is one utterance. |
| Voice agents | partly | Good for the recogniser in an agent; the dialogue itself is not here. |
| Diarisation | no | Each clip holds one speaker. |
| Text to speech | no | Conversational call audio, not studio voice. |

## Conversation profile

This dataset is utterance clips, not whole conversations, so there is no conversation profile.


## Sample

No public excerpt. A sample is sent on request.

## Licence

Licence: custom, quoted per use (training, evaluation or both; internal or commercial; exclusive or not). Delivered in the layout your training code reads: https://kenpathlabs.com/lokah/formats.

---

Machine-readable: [llms.txt](https://kenpathlabs.com/llms.txt) · [JSON API](https://kenpathlabs.com/api/lokah/datasets) · [OpenAPI](https://kenpathlabs.com/lokah/openapi.json) · [MCP](https://kenpathlabs.com/api/lokah/mcp) · [for agents](https://kenpathlabs.com/lokah/for-agents)
