Kenpath Labs
Get started

हिन्दी

Hindi speech recognition utterances

Hindi speech recognition data: 218,428 single-speaker utterance clips of conversational Hindi, call-centre and everyday, each with its own time-aligned transcript in Devanagari script, English words kept as spoken. 350 hours of audio, 3,377,105 transcribed words.

350 h
One channel
Transcribed
16 kHz · 16-bit FLAC
Transcribed in full
Fit for
  • ASR
  • Full duplex
  • Turn-taking
  • Voice agents
  • Diarisation
  • TTS

The full set, measured.

Figures measured from the delivered files.

of audio
350 h
218,428 utterances
audio files
218,428
FLAC · 16 kHz · 16-bit
transcribed words
3.4 M
218,428 turns

Personal data. Numbers of 4 or more digits, digit strings spoken as words, and email addresses are masked as [pii] in text and replaced by a tone in audio. First names are not masked.

Sample · 6 of 142 clips

Utterance clips, as delivered

One speaker, one utterance, its transcript beneath. Press play on any line.

  1. thank you so much ma'am phone करने के लिए धन्यवाद

    Female3.1 s92 dB SNR

  2. हाँ! हाँ-हाँ

    Female2.8 s86 dB SNR

  3. तो आप्ता के वजह से तो इतनी सारी आफत पड़ती हैं की सब जनता परेशान है.

    Male7.0 s73 dB SNR

  4. आपकी गाड़ी जो हैं उसको amount मिलेगा आपका जो भी गाड़ी पे नुकसान आ रहा हैं ठीक हैं उसके ऊपर आपको insurance मिलेगा आपको कुछ होता हैं तो आपका insurance नहीं मिलेगा.

    Male11.6 s72 dB SNR

  5. आपको दस अंकों वाला बताना है ma'am

    Male3.1 s68 dB SNR

  6. ma'am उनका नाम अजित होंगा. नंबर मे आपको नहीं दे सकता ma'am.

    Male3.3 s67 dB SNR

Speakers

Across the whole dataset

667
distinct voices
325 F · 330 M · 12 unknown
by gender

Voices are grouped from the recordings themselves, so the count is an estimate. Gender is labelled, with voice-model checks.

Voices in the sample

In the transcripts

3,377,105 words across 218,428 utterances

  • Devanagari script, English as spoken

    Words are written in the script the speaker would use; English words stay in Latin script where they were said, so code-mixing is preserved as spoken.

  • One transcript per clip

    Each clip carries its own transcript with its start and end time in the source recording, its speaker id and gender. No word-level timestamps.

  • One tag vocabulary, in square brackets

    Anything that is not a spoken word is a bracket tag from a single list: [pii] for masked personal data, [filler] for hesitations, [overlap] where both speak at once. Plain-text fields carry no tags.

  • Personal data masked, in text and audio

    Numbers of 4 or more digits, digit strings spoken as words, and email addresses are masked as [pii] in text and replaced by a tone in audio. First names are not masked.

Delivery layouts are on formats.

Start with a sample. License the full set when it fits.

The full set

Get a quote

Say what you will train and how many hours you need. We reply by email with the licence terms and a price for this set.

  • Licensed per use, priced per dataset
  • Delivered in the layout your training stack reads
  • Every figure on this page comes with the delivery

Sample first

Get a sample by email

Ten to thirty minutes of this dataset, with transcripts, in the same files and fields as the full delivery. The link works for 24 hours.

  • Real recordings from this dataset
  • Same layout, naming and fields as the full set
  • For evaluation only

What it is good for.

  • Speech recognitionFitsBuilt for it: 218,428 utterance clips, each with its own time-aligned transcript, noise left as recorded.
  • Full-duplex speech to speechNoClips are single utterances; the conversation timing is not in this dataset.
  • Turn-takingNoNo turn structure: each clip is one utterance.
  • Voice agentsPartlyGood for the recogniser in an agent; the dialogue itself is not here.
  • DiarisationNoEach clip holds one speaker.
  • Text to speechNoConversational call audio, not studio voice.

Specification

Figures marked measured were read from the audio. Anything we cannot confirm is listed under Ask us about.

id
LK-SP-HIN-003
type
speech · utterances for ASR · call centre and general
language
हिन्दी · Hindi · hi-IN
hours
350 h
channels
One channel, one speaker per clip
files
218,428 files · 218,428 utterancescounted across the full set
audio
FLAC · 16 kHz · 16-bitmeasured across the full set
release
v1.0
transcript
time-aligned by segment · Devanagari script
speakers
667 across the full set · id and gender per speaker
pii
Numbers of 4 or more digits, digit strings spoken as words, and email addresses are masked as [pii] in text and replaced by a tone in audio. First names are not masked.
source
Recorded for the dataset
review
Every recording has a complete, segment-level transcript
personal data
Redacted
licence
Customquoted per use

Not exactly what you need?

A different domain, more hours, another channel layout or speaker mix. Tell us, and it becomes a collection built to the same specification.

Scope a collection

Related datasets

Talk to us

Start a collection.

Tell us what you want to collect and from whom. We reply with how we would run it.

  • Speech, images, documents, feedback or annotation
  • A new collection, or a dataset from the catalogue
  • Licensing and data handling

Prefer to reach out directly? Write to hello@kenpathlabs.com.