---
pretty_name: "Marathi general conversation"
language:
- mr
language_details: mr-IN
license: other
license_name: lokah-custom
license_link: https://kenpathlabs.com/lokah/licensing
task_categories:
- automatic-speech-recognition
size_categories:
- n<1K
tags:
- Marathi speech dataset
- conversational speech
- spontaneous conversation
- speech recognition training data
- code-switching
- mar
---

# Marathi general conversation

> Two-speaker general conversation in Marathi, recorded on one channel with speakers labelled. 9 hours. Transcripts are time-aligned and written in Devanagari script, with English words kept as spoken.

`LK-SP-MAR-001` · [Get a quote](https://kenpathlabs.com/lokah/datasets/marathi-general-conversation#contact) · [Record as JSON](https://kenpathlabs.com/api/lokah/datasets/marathi-general-conversation) · [Croissant](https://kenpathlabs.com/lokah/datasets/marathi-general-conversation/croissant.json)

Two-speaker general conversation in Marathi, recorded on one channel with speakers labelled. 9 hours. Transcripts are time-aligned and written in Devanagari script, with English words kept as spoken. Measured from 6 sample conversations (25 minutes): 16 kHz, 16-bit PCM WAV, one mono file per conversation. Across them 0% of transcript words are English written in Latin script, 17% of the time is silence, and there are 428 turns. 12 distinct voices in the sample (6 male, 5 female, 1 unknown). Every recording has a complete, segment-level transcript. Personal data: redacted. Licence: custom, quoted per use.

## Specification

| Field | Value | Note |
| --- | --- | --- |
| id | LK-SP-MAR-001 |  |
| type | speech · conversational · general |  |
| language | मराठी · Marathi · mr-IN |  |
| hours | 9 h |  |
| channels | One channel, speakers labelled in the transcript |  |
| files | 36 files · 36 conversations | counted across the full set |
| layouts | 36 mono conversation | counted across the full set |
| audio | FLAC · 16 kHz · 16-bit | measured across the full set |
| bandwidth | telephone (4 kHz) | 9 telephone-band files |
| snr | 37.8 dB median | across the full set |
| release | v1.0 |  |
| transcript | time-aligned by segment · Devanagari script · delivered as JSON |  |
| speakers | 40 across the full set · id and gender per speaker |  |
| pii | Numbers of 4 or more digits, digit strings spoken as words, and email addresses are masked as [pii] in text and replaced by a tone in audio. First names are not masked. |  |
| source | Recorded for the dataset |  |
| review | Every recording has a complete, segment-level transcript |  |
| personal data | Redacted |  |
| licence | Custom | quoted per use |


## What it is good for

| Task | Fit | Why |
| --- | --- | --- |
| Speech recognition | yes | Time-aligned transcripts in native script, English kept as spoken. 37.8 dB median SNR across the full set (60 dB in the sample conversations). |
| Full-duplex speech to speech | partly | One mixed channel. Turns are labelled, but overlapping speech cannot be separated, which these models need. |
| Turn-taking and endpointing | partly | Turn boundaries come from the transcript, so gaps are approximate and overlap is marked, not separated. |
| Voice agents for support | partly | Open conversation, not a support flow. Useful for language and prosody, not for task structure. |
| Speaker diarisation | yes | Speaker-attributed segments across full conversations. |
| Text to speech | partly | 37.8 dB median SNR across the full set (60 dB in the sample conversations): clean and wideband enough for conversational prosody data, though not a studio voice. |

## Conversation profile

Measured from 6 full conversations (25 minutes, 428 turns, 2877 words).

| Measure | Value |
| --- | --- |
| Talk time, speaker 1 / speaker 2 | 49% / 51% |
| Silence | 17% |
| Overlapping speech | 2% |
| English words, written in Latin script | 0% |
| Speaking rate | 138.3 words a minute |

One mixed channel: segment edges were placed by an annotator, so turn timing is approximate and overlap is an event label.

## Audio quality

Frame RMS at 20 ms on the channels mixed to mono. Noise floor is the 10th percentile, speech level the 90th; SNR is their difference. Bandwidth is the highest frequency at which speech still rises 6 dB above the recording's own noise spectrum.

| Measure | Value | Reading |
| --- | --- | --- |
| Speech above noise floor (SNR) | 60 dB (48.5 to 102.1) | some background |
| Noise floor | -74.9 dBFS |  |
| Effective bandwidth | 8.0 kHz | telephone (4 kHz) across the full set |
| Clipping | 3.004% of samples |  |
| DNSMOS P.835 (1 to 5) | background 3.48, speech 3, overall 2.55 | listener-rated quality, estimated |

## Samples

4 public excerpts, 44.767 seconds each, from different conversations in the dataset.

- Sample 1, Open conversation, from a 4-minute conversation (Speaker 1: male; Speaker 2: male): [audio](https://kenpathlabs.com/lokah-samples/marathi-general-conversation--1.m4a) · [segments](https://kenpathlabs.com/lokah-samples/marathi-general-conversation--1.segments.json)
- Sample 2, Open conversation, from a 4-minute conversation (Speaker 1: female; Speaker 2: male): [audio](https://kenpathlabs.com/lokah-samples/marathi-general-conversation--2.m4a) · [segments](https://kenpathlabs.com/lokah-samples/marathi-general-conversation--2.segments.json)
- Sample 3, Open conversation, from a 4-minute conversation (Speaker 1: male; Speaker 2: female): [audio](https://kenpathlabs.com/lokah-samples/marathi-general-conversation--3.m4a) · [segments](https://kenpathlabs.com/lokah-samples/marathi-general-conversation--3.segments.json)
- Sample 4, Open conversation, from a 4-minute conversation (Speaker 1: female; Speaker 2: gender unknown): [audio](https://kenpathlabs.com/lokah-samples/marathi-general-conversation--4.m4a) · [segments](https://kenpathlabs.com/lokah-samples/marathi-general-conversation--4.segments.json)

### Sample 1 transcript

Open conversation, starting at 3:11. Audio sha256 `afdd448433a6ee118b59e364eea170bac31298b30a9bad1d0c3bd4234233361d`.

| Start | Speaker | Text |
| --- | --- | --- |
| 0:00 | Speaker 1 | हा |
| 0:00 | Speaker 2 | मोठं असत थोडं |
| 0:01 | Speaker 1 | हा |
| 0:02 | Speaker 2 | त्याच्यामध्ये पण बऱ्याच [unclear] येतात. हि दोन्ही एम प्रकार चाललेत. दोन्ही जाती उत्तमेत. |
| 0:08 | Speaker 1 | हा |
| 0:08 | Speaker 2 | काही प्रॉब्लेम नाही. |
| 0:10 | Speaker 1 | ह्ह |
| 0:10 | Speaker 2 | पण हि कशी थोडी मम जी चार नंबर जो असतो ना विषय |
| 0:15 | Speaker 1 | हा |
| 0:16 | Speaker 2 | ते थोडी आकाराला थोडी कमी असते. |
| 0:19 | Speaker 1 | हा म्हणजे त्याची साईझ कमी असते. |
| 0:21 | Speaker 2 | साईझ कमी असते आणि पण ती भरगोस लागते |
| 0:25 | Speaker 1 | भरगोस लागते हा |
| 0:27 | Speaker 2 | [unclear] एका घडला जवळ जवळ दहा दहा पंधरा पंधरा बिया असतात काजू मध्ये [unclear] असतात. |
| 0:31 | Speaker 1 | अच्छा अच्छा हा हा |
| 0:33 | Speaker 2 | आणि अअ वेंगुर्ला सातचा तर काही प्रश्नच नाही ती एकदम जाडी ब |
| 0:38 | Speaker 2 | म मजबूत बी असते ती |
| 0:40 | Speaker 1 | हा मोठी असते |
| 0:41 | Speaker 1 | हा फल ते एकदम मोठं असतं. |
| 0:43 | Speaker 2 | हा बरोबर |

## Licence

Licence: custom, quoted per use (training, evaluation or both; internal or commercial; exclusive or not). Delivered in the layout your training code reads: https://kenpathlabs.com/lokah/formats.

---

Machine-readable: [llms.txt](https://kenpathlabs.com/llms.txt) · [JSON API](https://kenpathlabs.com/api/lokah/datasets) · [OpenAPI](https://kenpathlabs.com/lokah/openapi.json) · [MCP](https://kenpathlabs.com/api/lokah/mcp) · [for agents](https://kenpathlabs.com/lokah/for-agents)
