தமிழ்
Tamil speech recognition utterances
Tamil speech recognition data: 157,231 single-speaker utterance clips of conversational Tamil, call-centre and everyday, each with its own time-aligned transcript in Tamil script, English words kept as spoken. 201 hours of audio, 1,748,023 transcribed words.
- ASR
- Full duplex
- Turn-taking
- Voice agents
- Diarisation
- TTS
The full set, measured.
Figures measured from the delivered files.
- of audio
- 201 h
- 157,231 utterances
- audio files
- 157,231
- FLAC · 16 kHz · 16-bit
- transcribed words
- 1.7 M
- 157,231 turns
Personal data. Numbers of 4 or more digits, digit strings spoken as words, and email addresses are masked as [pii] in text and replaced by a tone in audio. First names are not masked.
Sample · 6 of 151 clips
Utterance clips, as delivered
One speaker, one utterance, its transcript beneath. Press play on any line.
ஆ, சொல்லுங்க, அங்கேருந்து தான் call பண்றீங்களா?
Female3.3 s70 dB SNR
OK hair வந்து சிக்கு இல்லாம easy ஆ சீவரதுக்கு அப்படி பாக்கும்போது
Male4.8 s63 dB SNR
OK mam. வேற எதாச்சும் facilities இருக்குதா?
Female3.9 s61 dB SNR
OK OK குடுப்பாங்க நா சொல்ற நா சொல்ற நீங்க கேட்டுகோங்கோ அவங்க குடுப்பாங்க.
Male3.7 s58 dB SNR
that means, for example, some people stay OK, we give you ten to peace discount but that is like ten days rupees discount on the next product device. So that is the corruption
Male11.5 s57 dB SNR
ஆ, அந்த மாதிரி model ன்னா sudden mam. actual லா sudden ஆ வந்துட்டு நீங்க வந்து இந்த long days க்கு போகும்போது இந்த luggage எல்லாம் கொண்டு போறதுக்கு நிறைய lossage இருக்குமா? இல்லை உங்களுக்கு? இல்லை இப்படி.
Male10.7 s56 dB SNR
Speakers
Across the whole dataset
- 1,236
- distinct voices
- 316 M · 784 F · 136 unknown
- by gender
Voices are grouped from the recordings themselves, so the count is an estimate. Gender is labelled, with voice-model checks.
Voices in the sample
In the transcripts
1,748,023 words across 157,231 utterances
Tamil script, English as spoken
Words are written in the script the speaker would use; English words stay in Latin script where they were said, so code-mixing is preserved as spoken.
One transcript per clip
Each clip carries its own transcript with its start and end time in the source recording, its speaker id and gender. No word-level timestamps.
One tag vocabulary, in square brackets
Anything that is not a spoken word is a bracket tag from a single list: [pii] for masked personal data, [filler] for hesitations, [overlap] where both speak at once. Plain-text fields carry no tags.
Personal data masked, in text and audio
Numbers of 4 or more digits, digit strings spoken as words, and email addresses are masked as [pii] in text and replaced by a tone in audio. First names are not masked.
Delivery layouts are on formats.
Start with a sample. License the full set when it fits.
The full set
Get a quote
Say what you will train and how many hours you need. We reply by email with the licence terms and a price for this set.
- Licensed per use, priced per dataset
- Delivered in the layout your training stack reads
- Every figure on this page comes with the delivery
Sample first
Get a sample by email
Ten to thirty minutes of this dataset, with transcripts, in the same files and fields as the full delivery. The link works for 24 hours.
- Real recordings from this dataset
- Same layout, naming and fields as the full set
- For evaluation only
What it is good for.
- Speech recognitionFitsBuilt for it: 157,231 utterance clips, each with its own time-aligned transcript, noise left as recorded.
- Full-duplex speech to speechNoClips are single utterances; the conversation timing is not in this dataset.
- Turn-takingNoNo turn structure: each clip is one utterance.
- Voice agentsPartlyGood for the recogniser in an agent; the dialogue itself is not here.
- DiarisationNoEach clip holds one speaker.
- Text to speechNoConversational call audio, not studio voice.
Specification
Figures marked measured were read from the audio. Anything we cannot confirm is listed under Ask us about.
Machine-readable
- id
- LK-SP-TAM-004
- type
- speech · utterances for ASR · call centre and general
- language
- தமிழ் · Tamil · ta-IN
- hours
- 201 h
- channels
- One channel, one speaker per clip
- files
- 157,231 files · 157,231 utterancescounted across the full set
- audio
- FLAC · 16 kHz · 16-bitmeasured across the full set
- release
- v1.0
- transcript
- time-aligned by segment · Tamil script
- speakers
- 1,236 across the full set · id and gender per speaker
- pii
- Numbers of 4 or more digits, digit strings spoken as words, and email addresses are masked as [pii] in text and replaced by a tone in audio. First names are not masked.
- source
- Recorded for the dataset
- review
- Every recording has a complete, segment-level transcript
- personal data
- Redacted
- licence
- Customquoted per use
Not exactly what you need?
A different domain, more hours, another channel layout or speaker mix. Tell us, and it becomes a collection built to the same specification.
Scope a collectionRelated datasets
Start a collection.
Tell us what you want to collect and from whom. We reply with how we would run it.
- Speech, images, documents, feedback or annotation
- A new collection, or a dataset from the catalogue
- Licensing and data handling
Prefer to reach out directly? Write to hello@kenpathlabs.com.