हिन्दीதமிழ்
Hindi and Tamil full-duplex call-centre conversations
Hindi and Tamil full-duplex conversation data: 1,265 two-channel call-centre calls with each speaker on a separate channel, overlaps, backchannels and turn timing preserved, transcripts time-aligned per channel. 264 hours.
- ASR
- Full duplex
- Turn-taking
- Voice agents
- Diarisation
- TTS
The full set, measured.
Figures measured from the delivered files.
- of audio
- 264 h
- 1,265 calls
- audio files
- 1,265
- FLAC · 16 kHz · 16-bit
- transcribed words
- 2.4 M
- 173,888 turns
- one channel
- 1,265
- bandwidth
- wideband
- every file measured
- median SNR
- 29.2 dB
- some background
Personal data. Numbers of 4 or more digits, digit strings spoken as words, and email addresses are masked as [pii] in text and replaced by a tone in audio. First names are not masked.
Sample 1 of 4
Two-channel call
0:42 from a 4-minute conversation, at 0:03
Speaker 1
Male
Speaker 2
Male
Speakers
Across the whole dataset
- 529
- distinct voices
- 281 F · 231 M · 17 unknown
- by gender
Voices are grouped from the recordings themselves, so the count is an estimate. Gender is labelled, with voice-model checks.
Voices in the sample
In the transcripts
2,386,257 words in 173,888 turns
Devanagari and Tamil script, English as spoken
Words are written in the script the speaker would use; English words stay in Latin script where they were said, so code-mixing is preserved as spoken.
Aligned per segment, speakers labelled
One segment per turn with its start and end time and the speaker, no word-level timestamps, on the channel that speaker was recorded to. Overlaps are kept as overlapping segments.
One tag vocabulary, in square brackets
Anything that is not a spoken word is a bracket tag from a single list: [pii] for masked personal data, [filler] for hesitations, [overlap] where both speak at once. Plain-text fields carry no tags.
Personal data masked, in text and audio
Numbers of 4 or more digits, digit strings spoken as words, and email addresses are masked as [pii] in text and replaced by a tone in audio. First names are not masked.
Delivery layouts are on formats.
Conversation profile
Across every conversation in the dataset
- Speaker 1 72%
- Speaker 2 28%
- Speech 73%
- Overlap 8%
- Silence 20%
- Latin script 20%
- Native script 80%
Across 1,265 two-channel calls. The middle 80% of changes fall between -0.66 s and +1.38 s; a negative gap means the next speaker came in early.
Turn-taking events a minute across the set, beside the Fisher corpus
| Event | This dataset | Fisher |
|---|---|---|
| Inter-pausal units | 27.4 | 21.6 |
| Pauses | 10.2 | 7.0 |
| Gaps | 7.1 | 7.5 |
| Overlaps | 3.8 | 6.5 |
| Backchannels | 5.9 | not reported |
- This dataset
- Fisher corpus
Channel isolation: 45 dB. How much of the other speaker bleeds into each channel. Lower is cleaner.
- 12.4 min
- median conversation
- 1,265
- conversations
- 7.4
- turns a minute
- 151
- words a minute
Audio quality
Across the whole dataset, then in the sample conversations.
Some background across the set. Bandwidth wideband (8 kHz).
Quiet rooms and close microphones. Fine for any speech task, including voice models. Noise floor -84.1 dBFS, speech at -28.1 dBFS.
Wideband speech. The full range a 16 kHz recording can hold is present.
- Background3.55
- Speech3.17
- Overall2.65
No clipping.
Start with a sample. License the full set when it fits.
The full set
Get a quote
Say what you will train and how many hours you need. We reply by email with the licence terms and a price for this set.
- Licensed per use, priced per dataset
- Delivered in the layout your training stack reads
- Every figure on this page comes with the delivery
Sample first
Get a sample by email
Ten to thirty minutes of this dataset, with transcripts, in the same files and fields as the full delivery. The link works for 24 hours.
- Real recordings from this dataset
- Same layout, naming and fields as the full set
- For evaluation only
What it is good for.
- Speech recognitionPartlyPer-channel transcripts are included; built for dialogue timing, not for utterance-level training.
- Full-duplex speech to speechFitsBuilt for it: 1,265 two-channel calls, each speaker on a separate channel, overlaps and timing intact.
- Turn-takingFitsFloor transfers, overlaps and backchannels are measurable from the two channels.
- Voice agentsFitsReal agent and customer turns in a support flow, the condition a deployed agent hears.
- DiarisationPartlyTwo speakers, already separated by channel; useful as ground truth, not as a hard case.
- Text to speechNoConversational call audio, not studio voice.
Specification
Figures marked measured were read from the audio. Anything we cannot confirm is listed under Ask us about.
Machine-readable
- id
- LK-SP-MUL-001
- type
- speech · full-duplex calls · call centre
- language
- हिन्दी · Hindi · hi-IN / தமிழ் · Tamil · ta-IN
- hours
- 264 h
- channels
- Two channels, one per speaker
- files
- 1,265 files · 1,265 callscounted across the full set
- layouts
- 1,265 mono conversationcounted across the full set
- audio
- FLAC · 16 kHz · 16-bitmeasured across the full set
- bandwidth
- wideband (8 kHz)
- snr
- 29.2 dB medianacross the full set
- release
- v1.0
- transcript
- time-aligned by segment · Devanagari and Tamil script · delivered as JSON
- speakers
- 529 across the full set · id and gender per speaker
- pii
- Numbers of 4 or more digits, digit strings spoken as words, and email addresses are masked as [pii] in text and replaced by a tone in audio. First names are not masked.
- source
- Recorded for the dataset
- review
- Every recording has a complete, segment-level transcript
- personal data
- Redacted
- licence
- Customquoted per use
Not exactly what you need?
A different domain, more hours, another channel layout or speaker mix. Tell us, and it becomes a collection built to the same specification.
Scope a collectionRelated datasets
Start a collection.
Tell us what you want to collect and from whom. We reply with how we would run it.
- Speech, images, documents, feedback or annotation
- A new collection, or a dataset from the catalogue
- Licensing and data handling
Prefer to reach out directly? Write to hello@kenpathlabs.com.