Kenpath Labs
Get started

हिन्दी

Hindi call-centre conversations, insurance

Scripted call-centre conversations in Hindi, insurance, recorded with each speaker on a separate channel. 356 hours. Transcripts are time-aligned and written in Devanagari script, with English words kept as spoken. Layouts across the set: 1,135 two-channel, 710 one side of a call.

356 h
Two channels
Transcribed
16 kHz · 16-bit FLAC
Transcribed in full
Fit for
  • ASR
  • Full duplex
  • Turn-taking
  • Voice agents
  • Diarisation
  • TTS

The full set, measured.

Figures measured from the delivered files.

of audio
356 h
1,845 conversations
audio files
1,845
FLAC · 16 kHz · 16-bit
transcribed words
2.7 M
195,671 turns
two channels
1,135
710 one side of a call
bandwidth
wideband
every file measured
median SNR
33.7 dB
clean

Personal data. Numbers of 4 or more digits, digit strings spoken as words, and email addresses are masked as [pii] in text and replaced by a tone in audio. First names are not masked.

Sample 1 of 4

Call-centre conversation

0:32 from a 4-minute conversation, at 0:00

Speaker 1

Female

Speaker 2

Female

0:00 / 0:32
Channel 1 · Speaker 1Channel 2 · Speaker 2

Speakers

Across the whole dataset

490
distinct voices
237 F · 252 M · 1 unknown
by gender

Voices are grouped from the recordings themselves, so the count is an estimate. Gender is labelled, with voice-model checks.

Voices in the sample

In the transcripts

2,703,381 words in 195,671 turns

  • Devanagari script, English as spoken

    Words are written in the script the speaker would use; English words stay in Latin script where they were said, so code-mixing is preserved as spoken.

  • Aligned per segment, speakers labelled

    One segment per turn with its start and end time and the speaker, no word-level timestamps, on the channel that speaker was recorded to. Overlaps are kept as overlapping segments.

  • One tag vocabulary, in square brackets

    Anything that is not a spoken word is a bracket tag from a single list: [pii] for masked personal data, [filler] for hesitations, [overlap] where both speak at once. Plain-text fields carry no tags.

  • Personal data masked, in text and audio

    Numbers of 4 or more digits, digit strings spoken as words, and email addresses are masked as [pii] in text and replaced by a tone in audio. First names are not masked.

Delivery layouts are on formats.

Conversation profile

Across every conversation in the dataset

Talk time78% / 22%
  • Speaker 1 78%
  • Speaker 2 22%
Speech, overlap and silence60% / 8% / 32%
  • Speech 60%
  • Overlap 8%
  • Silence 32%
Transcript words in Latin script17%
  • Latin script 17%
  • Native script 83%
Gap at each change of speakermedian +0.17 s

Across 1,135 two-channel calls. The middle 80% of changes fall between -0.67 s and +1.33 s; a negative gap means the next speaker came in early.

Turn-taking events a minute across the set, beside the Fisher corpus

Turn-taking events per minute, this dataset beside the Fisher corpus
EventThis datasetFisher
Inter-pausal units27.421.6
Pauses10.47.0
Gaps6.87.5
Overlaps3.96.5
Backchannels5.9not reported
  • This dataset
  • Fisher corpus

Channel isolation: 46 dB. How much of the other speaker bleeds into each channel. Lower is cleaner.

12.1 min
median conversation
1,845
conversations
4.8
turns a minute
127
words a minute

Audio quality

Across the whole dataset, then in the sample conversations.

Across all 1,845 files33.7 dB median SNR
noisysome backgroundclean

Clean across the set. Bandwidth wideband (8 kHz), no telephone-band files.

In the sample conversations53.3 dB (30 to 70.1)
noisysome backgroundclean

Quiet rooms and close microphones. Fine for any speech task, including voice models. Noise floor -77.5 dBFS, speech at -26.5 dBFS.

Effective bandwidth8.0 kHz
0telephone band8 kHz

Wideband speech. The full range a 16 kHz recording can hold is present.

Listener-rated quality, estimatedDNSMOS P.835
  • Background3.56
  • Speech3.08
  • Overall2.59

No clipping.

Start with a sample. License the full set when it fits.

The full set

Get a quote

Say what you will train and how many hours you need. We reply by email with the licence terms and a price for this set.

  • Licensed per use, priced per dataset
  • Delivered in the layout your training stack reads
  • Every figure on this page comes with the delivery

Sample first

Get a sample by email

Ten to thirty minutes of this dataset, with transcripts, in the same files and fields as the full delivery. The link works for 24 hours.

  • Real recordings from this dataset
  • Same layout, naming and fields as the full set
  • For evaluation only

What it is good for.

  • Speech recognitionFitsTime-aligned transcripts in native script, English kept as spoken. 33.7 dB median SNR across the full set, wideband (53.3 dB in the sample conversations).
  • Full-duplex speech to speechFitsEach speaker on a separate channel, so overlap, backchannels and turn timing survive. Moshi and PersonaPlex train on exactly this layout.
  • Turn-taking and endpointingFitsGaps and overlaps at every change of speaker are measured from the two channels; see the profile.
  • Voice agents for supportFitsAgent and customer turns in a real support flow.
  • Speaker diarisationFitsSpeaker-attributed segments across full conversations.
  • Text to speechPartly33.7 dB median SNR across the full set, wideband (53.3 dB in the sample conversations): clean and wideband enough for conversational prosody data, though not a studio voice.

Specification

Figures marked measured were read from the audio. Anything we cannot confirm is listed under Ask us about.

id
LK-SP-HIN-001
type
speech · conversational · call centre
language
हिन्दी · Hindi · hi-IN
hours
356 h
channels
Two channels, one per speaker
files
1,845 files · 1,845 conversationscounted across the full set
layouts
1,135 two-channel, 710 one side of a callcounted across the full set
audio
FLAC · 16 kHz · 16-bitmeasured across the full set
bandwidth
wideband (8 kHz)no telephone-band files
snr
33.7 dB medianacross the full set
release
v1.0
transcript
time-aligned by segment · Devanagari script · delivered as JSON
speakers
490 across the full set · id and gender per speaker
pii
Numbers of 4 or more digits, digit strings spoken as words, and email addresses are masked as [pii] in text and replaced by a tone in audio. First names are not masked.
source
Recorded for the dataset
review
Every recording has a complete, segment-level transcript
personal data
Redacted
licence
Customquoted per use

Not exactly what you need?

A different domain, more hours, another channel layout or speaker mix. Tell us, and it becomes a collection built to the same specification.

Scope a collection

Related datasets

Talk to us

Start a collection.

Tell us what you want to collect and from whom. We reply with how we would run it.

  • Speech, images, documents, feedback or annotation
  • A new collection, or a dataset from the catalogue
  • Licensing and data handling

Prefer to reach out directly? Write to hello@kenpathlabs.com.