हिन्दी
Hindi call-centre conversations, insurance
Scripted call-centre conversations in Hindi, insurance, recorded with each speaker on a separate channel. 356 hours. Transcripts are time-aligned and written in Devanagari script, with English words kept as spoken. Layouts across the set: 1,135 two-channel, 710 one side of a call.
- ASR
- Full duplex
- Turn-taking
- Voice agents
- Diarisation
- TTS
The full set, measured.
Figures measured from the delivered files.
- of audio
- 356 h
- 1,845 conversations
- audio files
- 1,845
- FLAC · 16 kHz · 16-bit
- transcribed words
- 2.7 M
- 195,671 turns
- two channels
- 1,135
- 710 one side of a call
- bandwidth
- wideband
- every file measured
- median SNR
- 33.7 dB
- clean
Personal data. Numbers of 4 or more digits, digit strings spoken as words, and email addresses are masked as [pii] in text and replaced by a tone in audio. First names are not masked.
Sample 1 of 4
Call-centre conversation
0:32 from a 4-minute conversation, at 0:00
Speaker 1
Female
Speaker 2
Female
Speakers
Across the whole dataset
- 490
- distinct voices
- 237 F · 252 M · 1 unknown
- by gender
Voices are grouped from the recordings themselves, so the count is an estimate. Gender is labelled, with voice-model checks.
Voices in the sample
In the transcripts
2,703,381 words in 195,671 turns
Devanagari script, English as spoken
Words are written in the script the speaker would use; English words stay in Latin script where they were said, so code-mixing is preserved as spoken.
Aligned per segment, speakers labelled
One segment per turn with its start and end time and the speaker, no word-level timestamps, on the channel that speaker was recorded to. Overlaps are kept as overlapping segments.
One tag vocabulary, in square brackets
Anything that is not a spoken word is a bracket tag from a single list: [pii] for masked personal data, [filler] for hesitations, [overlap] where both speak at once. Plain-text fields carry no tags.
Personal data masked, in text and audio
Numbers of 4 or more digits, digit strings spoken as words, and email addresses are masked as [pii] in text and replaced by a tone in audio. First names are not masked.
Delivery layouts are on formats.
Conversation profile
Across every conversation in the dataset
- Speaker 1 78%
- Speaker 2 22%
- Speech 60%
- Overlap 8%
- Silence 32%
- Latin script 17%
- Native script 83%
Across 1,135 two-channel calls. The middle 80% of changes fall between -0.67 s and +1.33 s; a negative gap means the next speaker came in early.
Turn-taking events a minute across the set, beside the Fisher corpus
| Event | This dataset | Fisher |
|---|---|---|
| Inter-pausal units | 27.4 | 21.6 |
| Pauses | 10.4 | 7.0 |
| Gaps | 6.8 | 7.5 |
| Overlaps | 3.9 | 6.5 |
| Backchannels | 5.9 | not reported |
- This dataset
- Fisher corpus
Channel isolation: 46 dB. How much of the other speaker bleeds into each channel. Lower is cleaner.
- 12.1 min
- median conversation
- 1,845
- conversations
- 4.8
- turns a minute
- 127
- words a minute
Audio quality
Across the whole dataset, then in the sample conversations.
Clean across the set. Bandwidth wideband (8 kHz), no telephone-band files.
Quiet rooms and close microphones. Fine for any speech task, including voice models. Noise floor -77.5 dBFS, speech at -26.5 dBFS.
Wideband speech. The full range a 16 kHz recording can hold is present.
- Background3.56
- Speech3.08
- Overall2.59
No clipping.
Start with a sample. License the full set when it fits.
The full set
Get a quote
Say what you will train and how many hours you need. We reply by email with the licence terms and a price for this set.
- Licensed per use, priced per dataset
- Delivered in the layout your training stack reads
- Every figure on this page comes with the delivery
Sample first
Get a sample by email
Ten to thirty minutes of this dataset, with transcripts, in the same files and fields as the full delivery. The link works for 24 hours.
- Real recordings from this dataset
- Same layout, naming and fields as the full set
- For evaluation only
What it is good for.
- Speech recognitionFitsTime-aligned transcripts in native script, English kept as spoken. 33.7 dB median SNR across the full set, wideband (53.3 dB in the sample conversations).
- Full-duplex speech to speechFitsEach speaker on a separate channel, so overlap, backchannels and turn timing survive. Moshi and PersonaPlex train on exactly this layout.
- Turn-taking and endpointingFitsGaps and overlaps at every change of speaker are measured from the two channels; see the profile.
- Voice agents for supportFitsAgent and customer turns in a real support flow.
- Speaker diarisationFitsSpeaker-attributed segments across full conversations.
- Text to speechPartly33.7 dB median SNR across the full set, wideband (53.3 dB in the sample conversations): clean and wideband enough for conversational prosody data, though not a studio voice.
Specification
Figures marked measured were read from the audio. Anything we cannot confirm is listed under Ask us about.
Machine-readable
- id
- LK-SP-HIN-001
- type
- speech · conversational · call centre
- language
- हिन्दी · Hindi · hi-IN
- hours
- 356 h
- channels
- Two channels, one per speaker
- files
- 1,845 files · 1,845 conversationscounted across the full set
- layouts
- 1,135 two-channel, 710 one side of a callcounted across the full set
- audio
- FLAC · 16 kHz · 16-bitmeasured across the full set
- bandwidth
- wideband (8 kHz)no telephone-band files
- snr
- 33.7 dB medianacross the full set
- release
- v1.0
- transcript
- time-aligned by segment · Devanagari script · delivered as JSON
- speakers
- 490 across the full set · id and gender per speaker
- pii
- Numbers of 4 or more digits, digit strings spoken as words, and email addresses are masked as [pii] in text and replaced by a tone in audio. First names are not masked.
- source
- Recorded for the dataset
- review
- Every recording has a complete, segment-level transcript
- personal data
- Redacted
- licence
- Customquoted per use
Not exactly what you need?
A different domain, more hours, another channel layout or speaker mix. Tell us, and it becomes a collection built to the same specification.
Scope a collectionRelated datasets
Start a collection.
Tell us what you want to collect and from whom. We reply with how we would run it.
- Speech, images, documents, feedback or annotation
- A new collection, or a dataset from the catalogue
- Licensing and data handling
Prefer to reach out directly? Write to hello@kenpathlabs.com.