অসমীয়াEnglishગુજરાતીहिन्दीಕನ್ನಡമലയാളംमराठीଓଡ଼ିଆਪੰਜਾਬੀతెలుగుاردو
NewSample release · recorded to order
Indian-language calls between people and an AI voice agent
Phone calls between a person and an AI voice agent in eleven Indian languages, on recruitment, telecom, healthcare, e-commerce and insurance, each voice on its own channel at 48 kHz. A sample release: 22 calls, 83 minutes, with verbatim transcripts of both sides. Recorded to order in your languages, domains and scenarios, at any volume.
- ASR
- Full duplex
- Turn-taking
- Voice agents
- Diarisation
- TTS
The sample, measured.
Figures measured from every call in the sample release. A collection recorded to order is measured the same way before delivery.
- of audio
- 83 min
- 22 conversations
- audio files
- 22
- FLAC · 48 kHz · 16-bit
- transcribed words
- 10,219
- 544 turns
- two channels
- 22
- bandwidth
- super-wideband
- 11 of 22 files measured
- median SNR
- 34.4 dB
- clean
Personal data. Numbers of 4 or more digits, digit strings spoken as words, and email addresses are masked as [pii] in text and replaced by a tone in audio. First names are not masked.
Built to order.
A sample of a collection recorded to order. The full set is built to your brief. What you hear on this page is the spec: the same channels, sample rate and transcripts, in the languages, domains and scenarios you name, at the volume you need.
- In the sample
- 11 languages: Assamese, Indian English, Gujarati, Hindi, Kannada, Malayalam, Marathi, Odia, Punjabi, Telugu, Urdu
- Languages next
- Bengali, Tamil, then yours on request
- Layout
- Each speaker on a separate channel
- Audio
- 48 kHz · 16-bit
- Recorded to order in the languages, domains and scenarios you need, at the same spec, and scales to large volumes
- Each call is recorded on two channels, the person on channel 1 and the AI voice agent on channel 2
- 48 kHz, 16-bit FLAC master
- Verbatim transcripts of both sides, in each language's own script, with English words as spoken
Sample 1 of 22 · Marathi
Recruitment: phone screening
The whole conversation, 4:40
Caller
Female
AI voice agent
Female · synthetic voice
Speakers
Across the whole dataset
- 17 · 5 M · 12 F
- distinct callers
- 1 · F
- AI agent voice
Callers are people from our contributor network, counted by contributor as the collection team recorded them. The agent is one synthetic voice on every call.
Voices in the sample
In the transcripts
10,219 words in 544 turns
Bengali-Assamese and Latin and Gujarati and Devanagari and Kannada and Malayalam and Odia and Gurmukhi and Telugu and Perso-Arabic script, English as spoken
Words are written in the script the speaker would use; English words stay in Latin script where they were said, so code-mixing is preserved as spoken.
Aligned per segment, speakers labelled
One segment per turn with its start and end time and the speaker, no word-level timestamps, on the channel that speaker was recorded to. Overlaps are kept as overlapping segments.
One tag vocabulary, in square brackets
Anything that is not a spoken word is a bracket tag from a single list: [pii] for masked personal data, [filler] for hesitations, [overlap] where both speak at once. Plain-text fields carry no tags.
Personal data masked, in text and audio
Numbers of 4 or more digits, digit strings spoken as words, and email addresses are masked as [pii] in text and replaced by a tone in audio. First names are not masked.
Delivery layouts are on formats.
Conversation profile
Across every conversation in the dataset
- Caller 59%
- AI voice agent 41%
- Speech 75%
- Overlap 2%
- Silence 23%
- Latin script 18%
- Native script 82%
Across 22 two-channel calls. The middle 80% of changes fall between -0.29 s and +2.26 s; a negative gap means the next speaker came in early.
Turn-taking events a minute across the set, beside the Fisher corpus
| Event | This dataset | Fisher |
|---|---|---|
| Inter-pausal units | 18.1 | 21.6 |
| Pauses | 8.9 | 7.0 |
| Gaps | 5.9 | 7.5 |
| Overlaps | 0.9 | 6.5 |
| Backchannels | 1.9 | not reported |
- This dataset
- Fisher corpus
Channel isolation: 60 dB. How much of the other speaker bleeds into each channel. Lower is cleaner.
- 3.8 min
- median conversation
- 22
- conversations
- 6.1
- turns a minute
- 123
- words a minute
Audio quality
Across every call in the sample release.
Clean across the set. Measured on the caller's channel; the agent's channel is synthetic and silent between turns. Bandwidth super-wideband (16 kHz).
Wideband or better. This reading is taken on a 16 kHz copy, so it stops at 8 kHz; measured file by file, the set holds 10 wideband, 11 super-wideband, 1 fullband.
- Background3.92
- Speech3.50
- Overall3.16
Clipping in 0.00% of the audio.
Start with a sample. License the full set when it fits.
Recorded to order
Request a collection
Name the languages, domains and scenarios, and the hours you need. The collection is recorded to the spec you hear in the sample, at any volume, and delivered in stages.
- Same spec as the sample: two channels, 48 kHz, verbatim transcripts of both sides
- Your languages, domains and call scenarios
- Licensed per use, priced per collection, delivered in stages
Sample first
Get a sample by email
Ten to thirty minutes of this dataset, with transcripts, in the same files and fields as the full delivery. The link works for 24 hours.
- Real recordings from this dataset
- Same layout, naming and fields as the full set
- For evaluation only
What it is good for.
- Speech recognitionFitsTime-aligned transcripts in native script, English kept as spoken. 34.4 dB median SNR across the full set, wideband (39.5 dB in the sample conversations).
- Full-duplex speech to speechFitsEach speaker on a separate channel, so overlap, backchannels and turn timing survive. Moshi and PersonaPlex train on exactly this layout.
- Turn-taking and endpointingFitsGaps and overlaps at every change of speaker are measured from the two channels; see the profile.
- Voice agents for supportFitsCaller and agent turns across support and recruitment scenarios, with the agent on its own channel.
- Speaker diarisationFitsSpeaker-attributed segments across full conversations.
- Text to speechPartly34.4 dB median SNR across the full set, wideband (39.5 dB in the sample conversations): clean and wideband enough for conversational prosody data, though not a studio voice.
Specification
Figures marked measured were read from the audio. Anything we cannot confirm is listed under Ask us about.
Machine-readable
- id
- LK-SP-MUL-001
- type
- speech · conversational · call centre
- language
- অসমীয়া · Assamese · as-IN / English · Indian English · en-IN / ગુજરાતી · Gujarati · gu-IN / हिन्दी · Hindi · hi-IN / ಕನ್ನಡ · Kannada · kn-IN / മലയാളം · Malayalam · ml-IN / मराठी · Marathi · mr-IN / ଓଡ଼ିଆ · Odia · or-IN / ਪੰਜਾਬੀ · Punjabi · pa-IN / తెలుగు · Telugu · te-IN / اردو · Urdu · ur-IN
- hours
- 1 h
- channels
- Two channels, one per speaker
- files
- 22 files · 22 conversationscounted across the sample release
- layouts
- 22 two-channelcounted across the sample release
- audio
- FLAC · 48 kHz · 16-bitmeasured across the sample release
- bandwidth
- super-wideband (16 kHz)
- snr
- 34.4 dB medianacross the sample release, on the caller's channel
- release
- sample releasethe full collection is recorded to order
- transcript
- time-aligned by segment · Bengali-Assamese and Latin and Gujarati and Devanagari and Kannada and Malayalam and Odia and Gurmukhi and Telugu and Perso-Arabic script · delivered as JSON
- speakers
- 17 callers and one AI voice agent (female synthetic voice)
- pii
- Numbers of 4 or more digits, digit strings spoken as words, and email addresses are masked as [pii] in text and replaced by a tone in audio. First names are not masked.
- source
- Recorded to order with an AI voice agent
- review
- Every recording has a complete, segment-level transcript
- personal data
- Redacted
- licence
- Customquoted per use
Not exactly what you need?
A different domain, more hours, another channel layout or speaker mix. Tell us, and it becomes a collection built to the same specification.
Scope a collectionRelated datasets
Start a collection.
Tell us what you want to collect and from whom. We reply with how we would run it.
- Speech, images, documents, feedback or annotation
- A new collection, or a dataset from the catalogue
- Licensing and data handling
Prefer to reach out directly? Write to hello@kenpathlabs.com.