Kenpath Labs
Get started

অসমীয়াEnglishગુજરાતીहिन्दीಕನ್ನಡമലയാളംमराठीଓଡ଼ିଆਪੰਜਾਬੀతెలుగుاردو

NewSample release · recorded to order

Indian-language calls between people and an AI voice agent

Phone calls between a person and an AI voice agent in eleven Indian languages, on recruitment, telecom, healthcare, e-commerce and insurance, each voice on its own channel at 48 kHz. A sample release: 22 calls, 83 minutes, with verbatim transcripts of both sides. Recorded to order in your languages, domains and scenarios, at any volume.

22 calls · 83 min
11 languages
Two channels
Transcribed
48 kHz · 16-bit FLAC
Transcribed in full
Fit for
  • ASR
  • Full duplex
  • Turn-taking
  • Voice agents
  • Diarisation
  • TTS

The sample, measured.

Figures measured from every call in the sample release. A collection recorded to order is measured the same way before delivery.

of audio
83 min
22 conversations
audio files
22
FLAC · 48 kHz · 16-bit
transcribed words
10,219
544 turns
two channels
22
bandwidth
super-wideband
11 of 22 files measured
median SNR
34.4 dB
clean

Personal data. Numbers of 4 or more digits, digit strings spoken as words, and email addresses are masked as [pii] in text and replaced by a tone in audio. First names are not masked.

Built to order.

A sample of a collection recorded to order. The full set is built to your brief. What you hear on this page is the spec: the same channels, sample rate and transcripts, in the languages, domains and scenarios you name, at the volume you need.

In the sample
11 languages: Assamese, Indian English, Gujarati, Hindi, Kannada, Malayalam, Marathi, Odia, Punjabi, Telugu, Urdu
Languages next
Bengali, Tamil, then yours on request
Layout
Each speaker on a separate channel
Audio
48 kHz · 16-bit
  • Recorded to order in the languages, domains and scenarios you need, at the same spec, and scales to large volumes
  • Each call is recorded on two channels, the person on channel 1 and the AI voice agent on channel 2
  • 48 kHz, 16-bit FLAC master
  • Verbatim transcripts of both sides, in each language's own script, with English words as spoken

Sample 1 of 22 · Marathi

Recruitment: phone screening

The whole conversation, 4:40

Caller

Female

AI voice agent

Female · synthetic voice

0:00 / 4:40
Channel 1 · CallerChannel 2 · AI voice agent

Speakers

Across the whole dataset

17 · 5 M · 12 F
distinct callers
1 · F
AI agent voice

Callers are people from our contributor network, counted by contributor as the collection team recorded them. The agent is one synthetic voice on every call.

Voices in the sample

In the transcripts

10,219 words in 544 turns

  • Bengali-Assamese and Latin and Gujarati and Devanagari and Kannada and Malayalam and Odia and Gurmukhi and Telugu and Perso-Arabic script, English as spoken

    Words are written in the script the speaker would use; English words stay in Latin script where they were said, so code-mixing is preserved as spoken.

  • Aligned per segment, speakers labelled

    One segment per turn with its start and end time and the speaker, no word-level timestamps, on the channel that speaker was recorded to. Overlaps are kept as overlapping segments.

  • One tag vocabulary, in square brackets

    Anything that is not a spoken word is a bracket tag from a single list: [pii] for masked personal data, [filler] for hesitations, [overlap] where both speak at once. Plain-text fields carry no tags.

  • Personal data masked, in text and audio

    Numbers of 4 or more digits, digit strings spoken as words, and email addresses are masked as [pii] in text and replaced by a tone in audio. First names are not masked.

Delivery layouts are on formats.

Conversation profile

Across every conversation in the dataset

Talk time59% / 41%
  • Caller 59%
  • AI voice agent 41%
Speech, overlap and silence75% / 2% / 23%
  • Speech 75%
  • Overlap 2%
  • Silence 23%
Transcript words in Latin script18%
  • Latin script 18%
  • Native script 82%
Gap at each change of speakermedian +1.35 s

Across 22 two-channel calls. The middle 80% of changes fall between -0.29 s and +2.26 s; a negative gap means the next speaker came in early.

Turn-taking events a minute across the set, beside the Fisher corpus

Turn-taking events per minute, this dataset beside the Fisher corpus
EventThis datasetFisher
Inter-pausal units18.121.6
Pauses8.97.0
Gaps5.97.5
Overlaps0.96.5
Backchannels1.9not reported
  • This dataset
  • Fisher corpus

Channel isolation: 60 dB. How much of the other speaker bleeds into each channel. Lower is cleaner.

3.8 min
median conversation
22
conversations
6.1
turns a minute
123
words a minute

Audio quality

Across every call in the sample release.

Across all 22 files34.4 dB median SNR
noisysome backgroundclean

Clean across the set. Measured on the caller's channel; the agent's channel is synthetic and silent between turns. Bandwidth super-wideband (16 kHz).

Effective bandwidth8.0 kHz
0telephone band8 kHz

Wideband or better. This reading is taken on a 16 kHz copy, so it stops at 8 kHz; measured file by file, the set holds 10 wideband, 11 super-wideband, 1 fullband.

Listener-rated quality, estimatedDNSMOS P.835
  • Background3.92
  • Speech3.50
  • Overall3.16

Clipping in 0.00% of the audio.

Start with a sample. License the full set when it fits.

Recorded to order

Request a collection

Name the languages, domains and scenarios, and the hours you need. The collection is recorded to the spec you hear in the sample, at any volume, and delivered in stages.

  • Same spec as the sample: two channels, 48 kHz, verbatim transcripts of both sides
  • Your languages, domains and call scenarios
  • Licensed per use, priced per collection, delivered in stages

Sample first

Get a sample by email

Ten to thirty minutes of this dataset, with transcripts, in the same files and fields as the full delivery. The link works for 24 hours.

  • Real recordings from this dataset
  • Same layout, naming and fields as the full set
  • For evaluation only

What it is good for.

  • Speech recognitionFitsTime-aligned transcripts in native script, English kept as spoken. 34.4 dB median SNR across the full set, wideband (39.5 dB in the sample conversations).
  • Full-duplex speech to speechFitsEach speaker on a separate channel, so overlap, backchannels and turn timing survive. Moshi and PersonaPlex train on exactly this layout.
  • Turn-taking and endpointingFitsGaps and overlaps at every change of speaker are measured from the two channels; see the profile.
  • Voice agents for supportFitsCaller and agent turns across support and recruitment scenarios, with the agent on its own channel.
  • Speaker diarisationFitsSpeaker-attributed segments across full conversations.
  • Text to speechPartly34.4 dB median SNR across the full set, wideband (39.5 dB in the sample conversations): clean and wideband enough for conversational prosody data, though not a studio voice.

Specification

Figures marked measured were read from the audio. Anything we cannot confirm is listed under Ask us about.

id
LK-SP-MUL-001
type
speech · conversational · call centre
language
অসমীয়া · Assamese · as-IN / English · Indian English · en-IN / ગુજરાતી · Gujarati · gu-IN / हिन्दी · Hindi · hi-IN / ಕನ್ನಡ · Kannada · kn-IN / മലയാളം · Malayalam · ml-IN / मराठी · Marathi · mr-IN / ଓଡ଼ିଆ · Odia · or-IN / ਪੰਜਾਬੀ · Punjabi · pa-IN / తెలుగు · Telugu · te-IN / اردو · Urdu · ur-IN
hours
1 h
channels
Two channels, one per speaker
files
22 files · 22 conversationscounted across the sample release
layouts
22 two-channelcounted across the sample release
audio
FLAC · 48 kHz · 16-bitmeasured across the sample release
bandwidth
super-wideband (16 kHz)
snr
34.4 dB medianacross the sample release, on the caller's channel
release
sample releasethe full collection is recorded to order
transcript
time-aligned by segment · Bengali-Assamese and Latin and Gujarati and Devanagari and Kannada and Malayalam and Odia and Gurmukhi and Telugu and Perso-Arabic script · delivered as JSON
speakers
17 callers and one AI voice agent (female synthetic voice)
pii
Numbers of 4 or more digits, digit strings spoken as words, and email addresses are masked as [pii] in text and replaced by a tone in audio. First names are not masked.
source
Recorded to order with an AI voice agent
review
Every recording has a complete, segment-level transcript
personal data
Redacted
licence
Customquoted per use

Not exactly what you need?

A different domain, more hours, another channel layout or speaker mix. Tell us, and it becomes a collection built to the same specification.

Scope a collection

Related datasets

Talk to us

Start a collection.

Tell us what you want to collect and from whom. We reply with how we would run it.

  • Speech, images, documents, feedback or annotation
  • A new collection, or a dataset from the catalogue
  • Licensing and data handling

Prefer to reach out directly? Write to hello@kenpathlabs.com.