---
pretty_name: "Hindi general conversation"
language:
- hi
language_details: hi-IN
license: other
license_name: lokah-custom
license_link: https://kenpathlabs.com/lokah/licensing
task_categories:
- automatic-speech-recognition
size_categories:
- n<1K
tags:
- Hindi speech dataset
- conversational speech
- spontaneous conversation
- speech recognition training data
- code-switching
- hin
---

# Hindi general conversation

> Two-speaker general conversation in Hindi, recorded on one channel with speakers labelled. 81 hours. Transcripts are time-aligned and written in Devanagari script, with English words kept as spoken.

`LK-SP-HIN-002` · [Get a quote](https://kenpathlabs.com/lokah/datasets/hindi-general-conversation#contact) · [Record as JSON](https://kenpathlabs.com/api/lokah/datasets/hindi-general-conversation) · [Croissant](https://kenpathlabs.com/lokah/datasets/hindi-general-conversation/croissant.json)

Two-speaker general conversation in Hindi, recorded on one channel with speakers labelled. 81 hours. Transcripts are time-aligned and written in Devanagari script, with English words kept as spoken. Measured from 6 sample conversations (25 minutes): 16 kHz, 16-bit PCM WAV, one mono file per conversation. Across them 14% of transcript words are English written in Latin script, 17% of the time is silence, and there are 359 turns. 11 distinct voices in the sample (4 male, 7 female). Every recording has a complete, segment-level transcript. Personal data: redacted. Licence: custom, quoted per use.

## Specification

| Field | Value | Note |
| --- | --- | --- |
| id | LK-SP-HIN-002 |  |
| type | speech · conversational · general |  |
| language | हिन्दी · Hindi · hi-IN |  |
| hours | 81 h |  |
| channels | One channel, speakers labelled in the transcript |  |
| files | 259 files · 259 conversations | counted across the full set |
| layouts | 259 mono conversation | counted across the full set |
| audio | FLAC · 16 kHz · 16-bit | measured across the full set |
| bandwidth | wideband (8 kHz) | no telephone-band files |
| snr | 36.8 dB median | across the full set |
| release | v1.0 |  |
| transcript | time-aligned by segment · Devanagari script · delivered as JSON |  |
| speakers | 177 across the full set · id and gender per speaker |  |
| pii | Numbers of 4 or more digits, digit strings spoken as words, and email addresses are masked as [pii] in text and replaced by a tone in audio. First names are not masked. |  |
| source | Recorded for the dataset |  |
| review | Every recording has a complete, segment-level transcript |  |
| personal data | Redacted |  |
| licence | Custom | quoted per use |


## What it is good for

| Task | Fit | Why |
| --- | --- | --- |
| Speech recognition | yes | Time-aligned transcripts in native script, English kept as spoken. 36.8 dB median SNR across the full set, wideband (53 dB in the sample conversations). |
| Full-duplex speech to speech | partly | One mixed channel. Turns are labelled, but overlapping speech cannot be separated, which these models need. |
| Turn-taking and endpointing | partly | Turn boundaries come from the transcript, so gaps are approximate and overlap is marked, not separated. |
| Voice agents for support | partly | Open conversation, not a support flow. Useful for language and prosody, not for task structure. |
| Speaker diarisation | yes | Speaker-attributed segments across full conversations. |
| Text to speech | partly | 36.8 dB median SNR across the full set, wideband (53 dB in the sample conversations): clean and wideband enough for conversational prosody data, though not a studio voice. |

## Conversation profile

Measured from 6 full conversations (25 minutes, 359 turns, 3631 words).

| Measure | Value |
| --- | --- |
| Talk time, speaker 1 / speaker 2 | 48% / 52% |
| Silence | 17% |
| Overlapping speech | 1% |
| English words, written in Latin script | 14% |
| Speaking rate | 176.8 words a minute |

One mixed channel: segment edges were placed by an annotator, so turn timing is approximate and overlap is an event label (1 in this conversation).

## Audio quality

Frame RMS at 20 ms on the channels mixed to mono. Noise floor is the 10th percentile, speech level the 90th; SNR is their difference. Bandwidth is the highest frequency at which speech still rises 6 dB above the recording's own noise spectrum.

| Measure | Value | Reading |
| --- | --- | --- |
| Speech above noise floor (SNR) | 53 dB (36.2 to 107.2) | some background |
| Noise floor | -75.7 dBFS |  |
| Effective bandwidth | 8.0 kHz | wideband (8 kHz) across the full set |
| Clipping | 0.015% of samples |  |
| DNSMOS P.835 (1 to 5) | background 3.75, speech 3.29, overall 2.88 | listener-rated quality, estimated |

## Samples

4 public excerpts, 44.161 seconds each, from different conversations in the dataset.

- Sample 1, Open conversation, from a 4-minute conversation (Speaker 1: male; Speaker 2: male): [audio](https://kenpathlabs.com/lokah-samples/hindi-general-conversation--1.m4a) · [segments](https://kenpathlabs.com/lokah-samples/hindi-general-conversation--1.segments.json)
- Sample 2, Open conversation, from a 4-minute conversation (Speaker 1: female; Speaker 2: female): [audio](https://kenpathlabs.com/lokah-samples/hindi-general-conversation--2.m4a) · [segments](https://kenpathlabs.com/lokah-samples/hindi-general-conversation--2.segments.json)
- Sample 3, Open conversation, from a 4-minute conversation (Speaker 1: male; Speaker 2: male): [audio](https://kenpathlabs.com/lokah-samples/hindi-general-conversation--3.m4a) · [segments](https://kenpathlabs.com/lokah-samples/hindi-general-conversation--3.segments.json)
- Sample 4, Open conversation, from a 4-minute conversation (Speaker 1: female; Speaker 2: female): [audio](https://kenpathlabs.com/lokah-samples/hindi-general-conversation--4.m4a) · [segments](https://kenpathlabs.com/lokah-samples/hindi-general-conversation--4.segments.json)

### Sample 1 transcript

Open conversation, starting at 0:20. Audio sha256 `0909ab3c1e2cf776a2610409d611e09665f7ab34304210efbf3aa6b05b5b8542`.

| Start | Speaker | Text |
| --- | --- | --- |
| 0:00 | Speaker 2 | हम्म शुरू हो गया कल से शुरू हुआ है। |
| 0:02 | Speaker 1 | ओह I C C ना? |
| 0:04 | Speaker 2 | हम्म |
| 0:05 | Speaker 1 | अच्छा ठीक है। |
| 0:09 | Speaker 2 | क्या कल तो opening था ना। |
| 0:11 | Speaker 1 | हम्म |
| 0:12 | Speaker 2 | england और new zeeland का बीच में था। |
| 0:14 | Speaker 1 | हाँ हाँ। |
| 0:16 | Speaker 2 | नया एक stadium बना है इंडिया में नरेंद्र मोदी stadium ना। |
| 0:19 | Speaker 1 | हाँ हाँ हाँ। |
| 0:19 | Speaker 2 | हैदराबाद में। |
| 0:20 | Speaker 1 | ह्म्म्म |
| 0:21 | Speaker 2 | [filler] वो हमें वो था इनके साथ [unintelligible] [silence] क्या था ये उसको क्या बोलते हो। |
| 0:27 | Speaker 1 | [filler] fans लोग जो? |
| 0:29 | Speaker 2 | fans लोग [overlap] जो |
| 0:29 | Speaker 1 | [overlap] attendance |
| 0:30 | Speaker 1 | नहीं था ना ज्यादा। |
| 0:31 | Speaker 2 | attendance नहीं था forty three thousand कुछ ही था। |
| 0:33 | Speaker 1 | [filler] वाह forty three thousand मामूली है लेकिन बड़ा है नरेंद्र मोदी stadium जो? |
| 0:37 | Speaker 2 | बहुत बड़ा है उसमें एक लाख से ज्यादा capacity है। |
| 0:40 | Speaker 1 | हम्म |
| 0:41 | Speaker 2 | तो उसमें इतना कम आया है। |
| 0:43 | Speaker 1 | हाँ। |

## Licence

Licence: custom, quoted per use (training, evaluation or both; internal or commercial; exclusive or not). Delivered in the layout your training code reads: https://kenpathlabs.com/lokah/formats.

---

Machine-readable: [llms.txt](https://kenpathlabs.com/llms.txt) · [JSON API](https://kenpathlabs.com/api/lokah/datasets) · [OpenAPI](https://kenpathlabs.com/lokah/openapi.json) · [MCP](https://kenpathlabs.com/api/lokah/mcp) · [for agents](https://kenpathlabs.com/lokah/for-agents)
