# Kenpath Labs — full context > Kenpath Labs is a frontier AI and data company in Bengaluru, India, with two > products. Svara TTS Turbo (model id svara-tts-turbo) is a > text-to-speech model available via API. The lab also ships Svara TTS v1, an > open-source TTS foundation model (Apache 2.0, 1M+ downloads). > Its second product is Lokah, a human data platform for AI: speech, images, documents, human feedback and expert annotation. > Contact: hello@kenpathlabs.com | https://kenpathlabs.com This is the extended overview of Kenpath Labs for LLM and agent context. For a shorter summary, see https://kenpathlabs.com/llms.txt. --- ## About Kenpath Labs Kenpath Labs originated from Kenpath and now operates as its own lab: open speech research, original models, and a live inference platform serving Svara TTS Turbo behind one API. The lab is small on purpose and based in Bengaluru, India. Its research starts from Indic languages, Indian voices, and Indian deployment constraints, then generalises outward. Positioning: people trust what speaks their language, and they buy from it — true across markets worldwide, not only India. Speech that speaks the customer's own language is the value proposition; latency and streaming are what make it hold a real conversation. Website: https://kenpathlabs.com Platform/console: https://platform.kenpathlabs.com (live, self-serve) API base URL: https://api.kenpathlabs.com Docs: https://docs.kenpathlabs.com Contact: hello@kenpathlabs.com Tagline: "Frontier AI and data. Built in India, for the world." Site structure (old URLs redirect permanently): - https://kenpathlabs.com/ — the company page: Svara and Lokah side by side - https://kenpathlabs.com/svara — the Svara TTS Turbo page (was / and /svara-tts-turbo) - https://kenpathlabs.com/developers — API overview (was /api) - https://kenpathlabs.com/open-source — Svara TTS v1 + tooling (was /svara-tts-v1) - https://kenpathlabs.com/pricing — Svara TTS Turbo's plans and the comparison table. Lokah datasets are not priced here; a licence is quoted per use (https://kenpathlabs.com/lokah/licensing) - https://kenpathlabs.com/lokah — Lokah, the lab's second product - https://kenpathlabs.com/lokah/datasets — the catalogue of licensable conversational speech datasets, each with a measured specification - https://kenpathlabs.com/lokah/licensing — datasets are licensed per use; what a licence is quoted on and how personal data is handled - https://kenpathlabs.com/lokah/formats — the source deliverable and the layouts it converts to - https://kenpathlabs.com/lokah/for-agents — when to use the catalogue and every machine interface to it (JSON API, MCP, Croissant, Markdown) - https://kenpathlabs.com/compare — comparisons with other TTS APIs (/compare/elevenlabs, /compare/sarvam, /compare/elevenlabs-alternatives): published rates dated and linked to each vendor's own pricing page. ElevenLabs lists $0.10 per 1,000 characters (Eleven v3 API) where Svara TTS Turbo is $0.01 on the same unit; Sarvam's Bulbul v3 lists ₹3 per 1,000 characters across 11 languages where Svara TTS Turbo is ₹1 across 80+ languages. The pages also state plainly where each competitor is the better choice. - https://kenpathlabs.com/contact-sales — Svara TTS Turbo at enterprise scale, licensing a Lokah dataset, and scoping a data collection - https://kenpathlabs.com/careers — working at the lab - https://kenpathlabs.com/careers/open-roles — the list of open roles - https://kenpathlabs.com/trust — trust centre (security, residency, certifications) - https://kenpathlabs.com/brand — brand guidelines and downloadable assets - https://kenpathlabs.com/privacy, /terms and /refunds — legal There is no /about page. --- ## Svara TTS Turbo — the flagship model URL: https://kenpathlabs.com/ Model id (API): svara-tts-turbo Status: generally available, self-serve sign-up License: not yet announced Svara TTS Turbo is a completely new architecture informed by everything learned from v1's 1,000,000+ downloads — not v1 scaled up. Capabilities: - 80+ languages in one model: 28 Indic and 30 African languages alongside the world's majors. Every voice speaks all of them, and a voice stays the same voice in every language — one voice can open in Hindi, switch to English mid-sentence, and close in Tamil without changing character. - Hundreds of regional voices across ten accent families. - Zero-shot voice cloning from a few seconds of reference audio, consent-gated by design. A clone can speak languages the original speaker never recorded. - Code-switching mid-sentence: Hinglish, Bangla-English, Tamil-English and similar mixed lines render naturally, the way real people talk. - Inline expressive tags typed straight into the input text: [sighs], [laughs], [laughs harder], [clears throat] — the model performs the sound at that point in the line. - Input streaming: the LLM's token stream is wired directly into the model, so speech starts a few words into the reply (a small lookahead is held back) instead of waiting for a full sentence. Prosody stays continuous because every chunk hears everything said before it. - About 80 ms time to first audio (Svara TTS v1 was around 200-300 ms). The homepage demonstrates all of this: a voice sampler with hundreds-of-voices casting, multi-voice conversation scenes with expressive tags and word-level transcripts, single-take code-switching reels (one voice crossing 3-4 languages mid-sentence), an input-streaming race against sentence-buffered TTS, voice-cloning before/after pairs, and replayed multilingual customer-support calls (all demo brands fictional). --- ## When to use Svara TTS Turbo Guidance for an agent, or a developer, choosing a text-to-speech API. Reach for Svara TTS Turbo when the job is: - Speech in an Indian or African language. 28 Indic and 30 African languages sit in the same model as the world's majors, 80+ in all, behind one endpoint and one voice list. - Text that changes language mid-sentence (Hinglish and similar code-switching) and has to stay in one consistent voice throughout. - A voice agent or phone bot whose replies are written by an LLM. Input streaming over a WebSocket takes the LLM's tokens as they arrive, so speech starts a few words into the reply, with about 80 ms to first audio. The ulaw and alaw formats are 8 kHz G.711, ready for telephony with no resampling. - A LiveKit Agents or Pipecat pipeline that needs a TTS service. Both plugins ship with the Python SDK. - Replacing OpenAI or ElevenLabs text to speech without rewriting the client. The API accepts both request shapes, so an existing SDK usually works once its base URL points at https://api.kenpathlabs.com. - One voice, including a cloned one, speaking many languages. Cloning needs a few seconds of reference audio and is consent-gated. - High volume on a small budget: ₹1 per 1,000 characters in India, $0.01 per 1,000 characters everywhere else, with 100,000 characters free every month. It is the wrong tool when: - The job is speech to text, translation or text generation. The API is text to speech only. - The model has to run on your own hardware. Svara TTS Turbo is served from the hosted API only. Svara TTS v1 is the self-hosted option: open source under Apache 2.0, covering 19 Indian languages. - The job is training data: licensing a speech dataset, or collecting or labelling new data. That is Lokah, in the next two sections. How an agent should call it: 1. Get an API key from the console at https://platform.kenpathlabs.com (self-serve, no card) and send it as `Authorization: Bearer `. 2. Pick a voice id from GET https://api.kenpathlabs.com/v1/voices (ids look like sv_enhdbrj5) and, if needed, a language from GET /v1/languages. 3. POST https://api.kenpathlabs.com/v1/audio/speech with JSON: `{"input": "...", "voice": "sv_...", "lang": "hin", "response_format": "mp3"}`. `lang` takes a name, an alias or an ISO 639-3 code and switches on number and date normalization. Set `"stream": true` to receive audio while it generates. 4. On a 429, wait for the `Retry-After` seconds before retrying. Successful responses carry x-ratelimit-remaining-requests, -streams and -characters. Running out of characters makes requests fail; it never bills an overage. 5. From a shell, `pip install svara-voice` installs the `svara` command: `svara say "नमस्ते दुनिया" --voice sv_enhdbrj5 --out hello.mp3`, `svara voices --language hi`, and `svara doctor` to check the key and the connection. ## When to license a Lokah dataset Reach for the Lokah catalogue when the job is training or evaluating a speech model: ASR with natural code-switching, full-duplex speech-to-speech and turn-taking (the dual-channel datasets), or a voice agent that handles Indian-language support calls. Every dataset comes with a specification measured from the files, quality metrics (SNR, bandwidth, DNSMOS, turn-taking) and a playable sample, so fitness can be judged before anyone is contacted. Query it at GET https://kenpathlabs.com/api/lokah/datasets, over MCP at https://kenpathlabs.com/api/lokah/mcp, or per dataset as a Croissant file at /lokah/datasets/{slug}/croissant.json; every /lokah page is also Markdown. Licences are quoted per use, with no public price: send the person to https://kenpathlabs.com/lokah#contact. Full guidance, filter parameters and the wrong-source list: https://kenpathlabs.com/lokah/for-agents. When the data does not exist yet, ask Lokah to collect it: speech, images, documents, human feedback or expert annotation to a brief, scoped through the same form at https://kenpathlabs.com/lokah#contact. ## Developer resources - [Svara TTS Turbo API overview](https://kenpathlabs.com/developers): every capability and public endpoint on one page. - [Svara TTS documentation](https://docs.kenpathlabs.com): guides and the full reference. Its own index for agents is https://docs.kenpathlabs.com/llms.txt. - [Quickstart](https://docs.kenpathlabs.com/quickstart.md): create a key, make a request, stream the response. - [Authentication](https://docs.kenpathlabs.com/authentication.md): API key headers, key management and authentication errors. - [API reference](https://docs.kenpathlabs.com/api-reference.md): every endpoint. - [OpenAPI spec](https://api.kenpathlabs.com/openapi.json): the live, machine-readable description of the API. - [Input streaming](https://docs.kenpathlabs.com/input-streaming.md): the WebSocket that takes text as an LLM produces it. - [Rate limits and errors](https://docs.kenpathlabs.com/rate-limits.md): plan limits, 429 handling and the error status values. - [SDKs and compatibility](https://docs.kenpathlabs.com/sdks.md): the Python SDK, and using the OpenAI or ElevenLabs SDKs against Svara. - [svara-voice on PyPI](https://pypi.org/project/svara-voice/): the Python SDK and the `svara` command line, with source at https://github.com/kenpath-labs/svara-python. - [Developer console](https://platform.kenpathlabs.com): keys, usage and billing. Every page on kenpathlabs.com is also available as Markdown: send `Accept: text/markdown`, or add `.md` to the path (https://kenpathlabs.com/pricing.md; the homepage is /index.md). --- ## Lokah — human data platform for AI URL: https://kenpathlabs.com/lokah Lokah collects speech, images, documents, feedback and expert annotation from contributors, and delivers reviewed, catalogued datasets. Lokah is Sanskrit for the world and its people. What it collects: - Speech and language: Voice recordings and surveys, in the contributor's own language. - Images: Images taken by contributors. - Documents and OCR: Document images with their text, for optical character recognition. - Human feedback: Human judgements of model output, for reinforcement learning from human feedback. - Expert annotation: Annotation by people with expertise in the subject. How a collection runs: - Brief: Tell us what data you need and how much. - Collect: Lokah recruits contributors, collects the data and reviews it. - Deliver: You receive a catalogued dataset with its consent terms. Why Lokah: - Reviewed: Every submission is reviewed before it enters a dataset. - Multilingual: Lokah reaches contributors across languages and regions, and each works in their own language. - Consented: Every contributor agrees to how their data is used. The terms come with the dataset. - Paid: Contributors are paid for their work. The catalogue: Each collection is listed with what it contains and what it can be used for. - Listed by data type, language and region - Consent terms attached to every dataset Who it is for: - Model teams: Training, evaluation and preference data for speech, vision, document and language models. - Researchers and public programmes: Surveys and field data collection, run on contributors' own phones. - Companies with a field workforce: A survey of which languages your workforce speaks, and where. Lokah is sales-led: there is no self-serve sign-up and no published price. Enquiries go through the form at https://kenpathlabs.com/lokah#contact or hello@kenpathlabs.com. The console at platform.kenpathlabs.com and the pricing page are for Svara TTS Turbo only. --- ## Svara TTS v1 — open source URL: https://kenpathlabs.com/open-source License: Apache 2.0 (weights and methods open, forever) Status: v1 in production use worldwide Svara TTS v1 is an open-source multilingual TTS foundation model for Indian languages. It speaks 19 languages in native scripts (Hindi, Bengali, Marathi, Telugu, Kannada, Tamil, Malayalam, Gujarati, Punjabi, Assamese, Odia, Nepali, Bhojpuri, Sanskrit, Manipuri, Bodo, Rajasthani, Chhattisgarhi, and Maithili), with emotion conditioning via expressive tags (, , , and more). Built like a language model for speech — easy to fine-tune with a few hours of audio. Proof points: - 1,000,000+ downloads on Hugging Face - Peaked at #7 among TTS models on Hugging Face (Feb 2026) Links: - Model: https://huggingface.co/kenpath/svara-tts-v1 - Demo space: https://huggingface.co/spaces/kenpath/svara-tts - Inference toolkit: https://github.com/Kenpath/svara-tts-inference - Colab: https://colab.research.google.com/drive/15YxFo1DzdQNbFUIZ1HJA4AN4oHqKxGtg - Release blog: https://huggingface.co/blog/kenpath/svara-tts-open-multilingual-speech-for-india Other open-source tooling: - GitHub org: https://github.com/kenpath-labs - Hugging Face org: https://huggingface.co/kenpath - svara-tts-inference (Python): inference & deployment toolkit for Svara TTS - indic-text-normalization (Python): text-normalization utilities for Indic TTS pipelines --- ## API surface URL: https://kenpathlabs.com/developers Base URL: https://api.kenpathlabs.com Reference and guides: https://docs.kenpathlabs.com The API exposes TWO request shapes over the same model, which is its strongest practical feature — an existing client usually works unchanged: - OpenAI-compatible: POST /v1/audio/speech (set stream to start audio early), plus /v1/audio/encode and /v1/audio/decode for cloning codes. - ElevenLabs-compatible: POST /v1/text-to-speech/{voice_id}, with /stream and /with-timestamps variants. Other groups: voices (/v1/voices, /v2/voices for search and filter, per-voice detail, preview, add, settings), languages (/v1/languages), models (/v1/models), account (/v1/user, /v1/user/subscription) and /health. Response formats: mp3 (default, 128 kbps), opus, pcm (raw 24 kHz s16le, lowest latency), ulaw and alaw (8 kHz G.711, straight out of the endpoint for telephony, so nothing has to be resampled), and more. Parameters worth knowing: voice, lang (name, alias or ISO-3 — enables normalization), response_format, speed (0.7 to 1.5, pitch preserved), stream, mode (tts or voice_clone with a reference), normalize (numbers, dates and symbols read aloud properly; needs lang), pronunciation_dictionary_id. SDK: svara-voice for Python, with sync and async clients, streaming, eager input streaming, and [livekit] and [pipecat] extras that provide TTS plugins for those pipelines. Install it with `pip install svara-voice` (https://pypi.org/project/svara-voice/); source is at https://github.com/kenpath-labs/svara-python. CLI: the same package installs a `svara` command, so a script or an agent can use the API without writing a client: - `svara say "नमस्ते दुनिया" --voice sv_enhdbrj5 --out hello.mp3` - `svara say "Your call is important" -v sv_enhdbrj5 -f ulaw -r 8000 -o prompt.ulaw` - `svara voices --language hi` - `svara doctor` checks the key and the connection, with timings. Rate limits: limits attach to the workspace, not to a key. A 429 carries a machine-readable status and, where it applies, a Retry-After header in seconds. Successful responses carry x-ratelimit-remaining-requests, x-ratelimit-remaining-streams and x-ratelimit-remaining-characters. Details: https://docs.kenpathlabs.com/rate-limits.md Machine-readable entry points: OpenAPI at https://api.kenpathlabs.com/openapi.json, the docs index at https://docs.kenpathlabs.com/llms.txt, and every page of kenpathlabs.com as Markdown (send Accept: text/markdown, or add .md to the path). --- ## Pricing URL: https://kenpathlabs.com/pricing One rate at every tier, with no volume discount: ₹1 per 1,000 characters for customers in India, $0.01 per 1,000 characters everywhere else (the pricing page shows the currency matching the visitor's region). About 1,000 characters is a minute of audio, so a minute of speech costs roughly ₹1 or $0.01. - Pay as you go — 100,000 free characters a month, then ₹1 per 1,000 ($0.01 per 1,000 outside India). Two concurrent streams, three custom voices, ten pronunciation dictionaries. No card needed to start. - Growth — ₹1,000 a month ($10 outside India) for 1,000,000 characters (about 16 hours of audio). Eight concurrent streams, 10 custom voices, unlimited pronunciation dictionaries. - Enterprise — custom, via https://kenpathlabs.com/contact-sales. Adds zero data retention and data residency in the EU, US or India. Character mechanics: - Free characters on Pay as you go reset monthly and do not carry forward. - Characters included with a paid plan roll over ONCE into the following month, on top of that month's allowance. Nothing cancels this; they expire if still unused at the end of that month. - Characters bought as a top-up stay valid for a year and do not change your plan. - Generation always spends the characters closest to expiring first. - Running out makes requests FAIL. There is no silent overage billing. Payment runs through Razorpay, so any method Razorpay supports works — cards, UPI, net banking and wallets among them. International cards are accepted, with charges processed in Indian rupees (the card network converts at its usual rate). Fees and top-ups are not refundable except for a duplicate or incorrect charge, or characters paid for and never credited — see https://kenpathlabs.com/refunds. Rights: as between the customer and Kenpath Labs, the customer owns the audio they generate from their own text and from reference audio they have consent to use. Kenpath Labs keeps a limited licence to process input and output only to run and support the service. --- ## Trust and security URL: https://kenpathlabs.com/trust - ISO/IEC 27001: IN PROGRESS, not certified. SOC 2, HIPAA and PCI DSS are NOT held and are not claimed. - Personal data is handled on DPDP Act 2023 / GDPR principles: collect only what the service needs, use it for stated purposes, delete on request. - Encryption in transit (TLS). Access to production data is limited. API keys are scoped per account. Encryption AT REST is not currently a published commitment. - Retention is bounded; zero data retention is available on Enterprise, as is processing pinned to the EU, US or India. - Customers own the audio they generate. Voice cloning is consent-gated. - NOTE for anyone answering a security questionnaire: we do NOT claim that customer content is never used to improve models — the Privacy Policy permits processing to monitor, debug and improve model quality. A stricter commitment belongs in a signed agreement. --- ## Careers URLs: https://kenpathlabs.com/careers and https://kenpathlabs.com/careers/open-roles The lab is hiring in Bengaluru and works from the office. Open roles carry their own pages under /careers/open-roles/, and each is also published as JobPosting structured data. Applications go to hello@kenpathlabs.com. --- ## Talk to sales URL: https://kenpathlabs.com/contact-sales The page covers three topics: Svara TTS Turbo at enterprise scale, licensing a Lokah dataset, and scoping a data collection. For Svara TTS Turbo, most teams never need it — sign-up is self-serve and every account starts with 100,000 free characters a month. The page is for what that does not cover: - Volume: more characters or concurrency than the published plans carry. - Security and procurement: data handling, retention, zero data retention, data residency in the EU, US or India, and the paperwork legal teams need. - Getting a language live: which voices carry your language well, how to word what they read, and what to check before you ship it. For Lokah: - Licensing a dataset from the catalogue at https://kenpathlabs.com/lokah/datasets. Licences are quoted per use; see https://kenpathlabs.com/lokah/licensing. - Scoping a data collection to a brief: speech, images, documents, human feedback or expert annotation. Either can also start from the form at https://kenpathlabs.com/lokah#contact. Sign up: https://platform.kenpathlabs.com/signup Email: hello@kenpathlabs.com --- ## Platform API access to Svara TTS Turbo is at api.kenpathlabs.com, with a console for keys and usage at platform.kenpathlabs.com. Sign-up is self-serve: every account starts with 100,000 free characters a month, and paid usage is one rate on every plan — ₹1 per 1,000 characters in India, $0.01 per 1,000 characters everywhere else. The platform serves Svara TTS Turbo only — Svara TTS v1 remains open source for self-hosting. The API follows the OpenAI-style audio/speech shape, with WebSocket input streaming for agent pipelines. --- ## Lokah: the catalogue in full --- # Lokah for agents > Lokah is the data platform of Kenpath Labs, a frontier AI and data company in Bengaluru. Its catalogue holds licensable datasets for training and evaluating AI, conversational speech in Indian languages and English first, each with a playable sample, a specification measured from the files and a conversation profile. 12 datasets, 680 hours, 3 languages. ## When to use this catalogue - You need two-speaker conversational speech in Hindi or Tamil today; other Indian languages are collected to a brief. - You are training or evaluating a full-duplex speech-to-speech model (Moshi, PersonaPlex and similar) and need dual-channel audio, where each speaker is on a separate channel. Filter with channels=dual. - You need speech recognition data with natural code-switching: transcripts are in native script with English words kept as spoken, and each dataset reports the share of Latin-script words across the whole set. - You need call-centre conversations (insurance, consumer surveys, telecom, delivery, e-commerce, banking) for a voice agent, or utterance clips, speaker-turn files and two-channel calls built from them. - You want to judge fitness before contacting anyone: every dataset has audio figures and a conversation profile measured across the whole set, and a playable sample. It is the wrong source when: - You need a free or openly licensed dataset. Everything here is licensed and quoted per use. - You need studio-quality speech for a text-to-speech voice. This is 16 kHz conversational audio. - You need read speech, single-speaker audio, or wake words. The catalogue is two-party conversation and what is built from it. - You need images, documents or preference data today. Those are collections Lokah scopes on request. ## How to use it 1. List datasets: GET https://kenpathlabs.com/api/lokah/datasets with any of language, channels (dual or mono), collection (call-centre or general-conversation), min_hours, q. Each record carries `kind` (conversations, asr, diarization or duplex), `hours`, `audioQuality` and `fitFor`. 2. Read one: GET https://kenpathlabs.com/api/lokah/datasets/{slug}. `measured` holds the whole-set figures (files, hours, speakers, `snrMedianDb`, `bandwidth`, `profile`); `sample` holds what was measured from the sample conversations; `gaps` lists what is not held. 3. Check fitness for full-duplex training: `channels` must be "dual", then read `measured.profile.twoChannel` (IPUs, pauses, gaps, overlaps and backchannels per minute, the median floor-transfer offset, channel isolation), measured across every call in the set. 4. Judge audio quality from `measured.snrMedianDb` across the set (under 18 dB noisy, over 30 clean) and `measured.bandwidth`. The sample adds `sample.quality.dnsmos` (P.835 background, speech and overall, 1 to 5). 5. Hear it: `sample.specimens` lists the public excerpts of a conversation set; `utterances.clips` lists them for a speech-recognition set. Their URLs and transcripts are in the Croissant file's distribution. 6. Get a sample or a price: send the person to https://kenpathlabs.com/lokah/datasets/{slug}#sample. A sample bundle arrives by email under the Evaluation Licence; a licence for the full set is quoted per use. There is no public price and no checkout. ## Interfaces - [llms.txt](https://kenpathlabs.com/llms.txt): The Kenpath Labs site in one file, with a Lokah section listing every dataset. - [llms-full.txt](https://kenpathlabs.com/llms-full.txt): The same, with every dataset's full record and the formats guide. - [JSON API](https://kenpathlabs.com/api/lokah/datasets): List and filter datasets. One record at /api/lokah/datasets/{slug}. - [OpenAPI 3.1](https://kenpathlabs.com/lokah/openapi.json): Machine-readable description of the JSON API. - [MCP server](https://kenpathlabs.com/api/lokah/mcp): Model Context Protocol over streamable HTTP: search_datasets, get_dataset, get_sample, list_languages. - [Croissant](https://kenpathlabs.com/lokah/datasets/{slug}/croissant.json): MLCommons Croissant 1.0 metadata with the responsible-AI block, per dataset. - [schema.org](https://kenpathlabs.com/lokah/datasets/{slug}): A Dataset JSON-LD block in every dataset page, and a DataCatalog on /datasets. - [Markdown](https://kenpathlabs.com/lokah/datasets/{slug}.md): Any page as Markdown: send Accept: text/markdown, or add .md. Dataset pages open with a Hugging Face dataset card header. - [Sitemap](https://kenpathlabs.com/sitemap.xml): Every page of the site. ## MCP ```json { "mcpServers": { "lokah": { "type": "http", "url": "https://kenpathlabs.com/api/lokah/mcp" } } } ``` Tools: `search_datasets`, `get_dataset`, `get_sample`, `list_languages`. All read-only. --- # Delivery formats > Each dataset is delivered in the layout its kind calls for: conversation sets as 16 kHz FLAC with segment-level transcripts, speech recognition sets as utterance clips in shards, diarization sets with RTTM turn files, two-channel sets as stereo files. Every layout below is a conversion of that source. ## Stereo conversations for Moshi and PersonaPlex For: Full-duplex speech to speech. Needs a dual-channel dataset. Kyutai's moshi-finetune, which NVIDIA PersonaPlex builds on, trains from stereo files: the left channel is the voice the model learns to produce, the right channel is the user it listens to. A dual-channel call maps straight onto it: agent left, customer right. ``` data/ lokah-hi-insurance.jsonl {"path": "stereo/HN_0001.wav", "duration": 907.2} stereo/ HN_0001.wav 2 channels: left = agent, right = customer HN_0001.json {"alignments": [["hello", [1.01, 1.53], "SPEAKER_MAIN"], ...]} ``` ```bash # two mono files per call -> one stereo file, agent on the left ffmpeg -i HN_0001_1.wav -i HN_0001_2.wav \ -filter_complex "[0:a][1:a]amerge=inputs=2" -ac 2 stereo/HN_0001.wav # word-level alignments for the text stream (moshi-finetune's own script) python annotate.py data/lokah-hi-insurance.jsonl --lang hi ``` - Audio is resampled to 24 kHz for the Mimi codec by the loader. The source is 16 kHz, so nothing above 8 kHz is present. - The text stream needs word-level timestamps. Transcripts here are aligned by segment, so annotate.py, or a forced aligner run against the human transcript, produces the word times. - PersonaPlex publishes no fine-tuning format of its own. It was trained on Fisher, which is also two-channel telephone speech. Source: [kyutai-labs/moshi-finetune](https://github.com/kyutai-labs/moshi-finetune) ## Hugging Face datasets, Parquet For: Any training stack that reads the Hub. One row per conversation, audio bytes embedded in Parquet shards, with the transcript segments and speaker metadata as columns. ``` data/train-00000-of-00012.parquet conversation_id string audio Audio(sampling_rate=16000, num_channels=2) segments list> speaker_a_id string speaker_a_gender speaker_b_id string speaker_b_gender language string BCP-47, e.g. hi-IN domain string ``` ```python from datasets import load_dataset, Audio ds = load_dataset("your-org/lokah-hi-insurance", split="train") # keep both channels: older versions of datasets downmix stereo to mono ds = ds.cast_column("audio", Audio(sampling_rate=24000, num_channels=2)) row = ds[0] row["audio"]["array"].shape # (2, num_samples): agent, customer ``` - num_channels arrived in datasets 4.4. Before that the Audio feature downmixed stereo to mono, which silently destroys a dual-channel dataset. Pin 4.4 or later. - The card carries task_categories (audio-to-audio for dual channel, automatic-speech-recognition where transcribed), language, and the gating fields. Source: [Hugging Face: audio datasets](https://huggingface.co/docs/datasets/audio_dataset) ## WebDataset shards For: Streaming at scale. Tar shards of about 1 GB. Files that share a prefix are one example, so each conversation is its audio plus one JSON with segments and speakers. ``` train/00000.tar HN_0001.wav HN_0001.json HN_0002.wav HN_0002.json ``` ```python from datasets import load_dataset ds = load_dataset("webdataset", data_dir="train", split="train", streaming=True) next(iter(ds)).keys() # dict_keys(['__key__', 'wav', 'json']) ``` - One column per file suffix. A dual-channel conversation can also ship as HN_0001.agent.wav and HN_0001.customer.wav. Source: [Hugging Face: WebDataset](https://huggingface.co/docs/hub/datasets-webdataset) ## NeMo manifests For: Speech recognition, and NeMo's duplex speech-to-speech. For recognition, a JSON-lines manifest of segments. For NeMo's speechlm2 duplex models, Lhotse cuts with the user's audio as the recording, the agent's as the target, and supervisions that carry a user or assistant role. ``` manifest.jsonl {"audio_filepath": "seg/HN_0001_0007.wav", "duration": 6.59, "text": "OK actually ma'am ...", "lang": "hi"} ``` ```python # one manifest line per transcript segment import json for seg in segments: print(json.dumps({"audio_filepath": seg["path"], "duration": seg["end"] - seg["start"], "text": seg["text"], "lang": "hi"}, ensure_ascii=False)) ``` - Transcripts keep English words in Latin script inside native-script text. Decide on a normalisation before training a recogniser on them. Source: [NVIDIA NeMo: ASR datasets](https://github.com/NVIDIA-NeMo/NeMo/blob/main/docs/source/asr/datasets.rst) ## Turn-taking terms used on this site - **Inter-pausal unit (IPU)**: Continuous speech on one channel, bounded by more than 200 ms of silence on both sides. - **Pause**: Silence between two IPUs of the same speaker. - **Gap**: Silence between IPUs of different speakers. - **Overlap**: Both channels in speech at once. - **Backchannel**: A short IPU that sits entirely inside an IPU of the other speaker, like a murmured yes. - **Floor-transfer offset**: At a change of speaker, the next start minus the previous end. Negative is an overlapped transfer, positive a gap. Published per-minute reference for the Fisher corpus: 21.6 IPUs, 7 pauses, 7.5 gaps, 6.5 overlaps. - [Nguyen et al. 2022, Generative Spoken Dialogue Language Modeling (dGSLM)](https://arxiv.org/abs/2203.16502) - [Heldner and Edlund 2010, Pauses, gaps and overlaps in conversations](https://doi.org/10.1016/j.wocn.2010.08.002) - [Roy et al. 2026, PersonaPlex: Voice and Role Control for Full Duplex Conversational Speech Models](https://arxiv.org/abs/2602.06053) --- # Hindi call-centre conversations, insurance > Scripted call-centre conversations in Hindi, insurance, recorded with each speaker on a separate channel. 356 hours. Transcripts are time-aligned and written in Devanagari script, with English words kept as spoken. Layouts across the set: 1,135 two-channel, 710 one side of a call. `LK-SP-HIN-001` · [Get a quote](https://kenpathlabs.com/lokah/datasets/hindi-call-centre-insurance#contact) · [Record as JSON](https://kenpathlabs.com/api/lokah/datasets/hindi-call-centre-insurance) · [Croissant](https://kenpathlabs.com/lokah/datasets/hindi-call-centre-insurance/croissant.json) Scripted call-centre conversations in Hindi, insurance, recorded with each speaker on a separate channel. 356 hours. Transcripts are time-aligned and written in Devanagari script, with English words kept as spoken. Layouts across the set: 1,135 two-channel, 710 one side of a call. Measured from 6 sample conversations (25 minutes): 16 kHz, 16-bit PCM WAV, two mono files per conversation. Across them 17% of transcript words are English written in Latin script, 12% of the time is silence, and there are 236 turns. 12 distinct voices in the sample (3 male, 9 female). Every recording has a complete, segment-level transcript. Personal data: redacted. Licence: custom, quoted per use. ## Specification | Field | Value | Note | | --- | --- | --- | | id | LK-SP-HIN-001 | | | type | speech · conversational · call centre | | | language | हिन्दी · Hindi · hi-IN | | | hours | 356 h | | | channels | Two channels, one per speaker | | | files | 1,845 files · 1,845 conversations | counted across the full set | | layouts | 1,135 two-channel, 710 one side of a call | counted across the full set | | audio | FLAC · 16 kHz · 16-bit | measured across the full set | | bandwidth | wideband (8 kHz) | no telephone-band files | | snr | 33.7 dB median | across the full set | | release | v1.0 | | | transcript | time-aligned by segment · Devanagari script · delivered as JSON | | | speakers | 490 across the full set · id and gender per speaker | | | pii | Numbers of 4 or more digits, digit strings spoken as words, and email addresses are masked as [pii] in text and replaced by a tone in audio. First names are not masked. | | | source | Recorded for the dataset | | | review | Every recording has a complete, segment-level transcript | | | personal data | Redacted | | | licence | Custom | quoted per use | ## What it is good for | Task | Fit | Why | | --- | --- | --- | | Speech recognition | yes | Time-aligned transcripts in native script, English kept as spoken. 33.7 dB median SNR across the full set, wideband (53.3 dB in the sample conversations). | | Full-duplex speech to speech | yes | Each speaker on a separate channel, so overlap, backchannels and turn timing survive. Moshi and PersonaPlex train on exactly this layout. | | Turn-taking and endpointing | yes | Gaps and overlaps at every change of speaker are measured from the two channels; see the profile. | | Voice agents for support | yes | Agent and customer turns in a real support flow. | | Speaker diarisation | yes | Speaker-attributed segments across full conversations. | | Text to speech | partly | 33.7 dB median SNR across the full set, wideband (53.3 dB in the sample conversations): clean and wideband enough for conversational prosody data, though not a studio voice. | ## Conversation profile Measured from 6 full conversations (25 minutes, 236 turns, 4074 words). | Measure | Value | | --- | --- | | Talk time, speaker 1 / speaker 2 | 56% / 44% | | Silence | 18% | | Overlapping speech | 6% | | English words, written in Latin script | 17% | | Speaking rate | 188 words a minute | ### Turn-taking, measured from the two channels Voice activity detected on each channel at 10 ms; IPUs bounded by more than 200 ms of silence; talkspurts under 90 ms dropped. | Per minute | This dataset | Fisher corpus | | --- | --- | --- | | Inter-pausal units | 23 | 21.6 | | Pauses | 10.5 | 7 | | Gaps | 5 | 7.5 | | Overlaps | 7.3 | 6.5 | | Backchannels | 4.3 | not reported | Floor-transfer offset: median +0.16 s, 10th to 90th percentile -0.41 s to +1.49 s. 33% of 186 changes of speaker were overlapped. Channel isolation: -40.7 dB. ## Audio quality Frame RMS at 20 ms on the channels mixed to mono. Noise floor is the 10th percentile, speech level the 90th; SNR is their difference. Bandwidth is the highest frequency at which speech still rises 6 dB above the recording's own noise spectrum. | Measure | Value | Reading | | --- | --- | --- | | Speech above noise floor (SNR) | 53.3 dB (30 to 70.1) | clean | | Noise floor | -77.5 dBFS | | | Effective bandwidth | 8.0 kHz | wideband (8 kHz) across the full set | | Clipping | 0.000% of samples | | | DNSMOS P.835 (1 to 5) | background 3.56, speech 3.08, overall 2.59 | listener-rated quality, estimated | ## Samples 4 public excerpts, 32.52 seconds each, from different conversations in the dataset. - Sample 1, Call-centre conversation, from a 4-minute conversation (Speaker 1: female; Speaker 2: female): [audio](https://kenpathlabs.com/lokah-samples/hindi-call-centre-insurance--1.m4a) · [segments](https://kenpathlabs.com/lokah-samples/hindi-call-centre-insurance--1.segments.json) - Sample 2, Call-centre conversation, from a 4-minute conversation (Speaker 1: male; Speaker 2: female): [audio](https://kenpathlabs.com/lokah-samples/hindi-call-centre-insurance--2.m4a) · [segments](https://kenpathlabs.com/lokah-samples/hindi-call-centre-insurance--2.segments.json) - Sample 3, Call-centre conversation, from a 4-minute conversation (Speaker 1: female; Speaker 2: female): [audio](https://kenpathlabs.com/lokah-samples/hindi-call-centre-insurance--3.m4a) · [segments](https://kenpathlabs.com/lokah-samples/hindi-call-centre-insurance--3.segments.json) - Sample 4, Call-centre conversation, from a 4-minute conversation (Speaker 1: female; Speaker 2: female): [audio](https://kenpathlabs.com/lokah-samples/hindi-call-centre-insurance--4.m4a) · [segments](https://kenpathlabs.com/lokah-samples/hindi-call-centre-insurance--4.segments.json) ### Sample 1 transcript Call-centre conversation, starting at 0:00. Stereo: left channel is speaker 1, right channel is speaker 2. Audio sha256 `5afb5f5e853e5673b2aab82375b865d9e96a62cf060dd765cd536010f71669e1`. | Start | Speaker | Text | | --- | --- | --- | | 0:00 | Speaker 1 | hello | | 0:01 | Speaker 2 | hello | | 0:02 | Speaker 1 | OK ma'am आप कविता बात कर रहे हो | | 0:05 | Speaker 2 | yes kavita spiking आप कोन बात कर रहे हो | | 0:09 | Speaker 1 | OK ma'am में अर्पित बात कर रही हूँ | | 0:12 | Speaker 2 | हाँजी बोलिए | | 0:13 | Speaker 1 | मेने आपको न car insurance के बारे मे details देने के लिए call किया है | | 0:19 | Speaker 2 | OK आ बोलिए | | 0:22 | Speaker 1 | OK ma'am तो आपके पास car insurance है या फिर आपको renew करवन है है या फीर आपको new लेना है | | 0:29 | Speaker 2 | आ actually मुझे renew करवना है | ## Licence Licence: custom, quoted per use (training, evaluation or both; internal or commercial; exclusive or not). Delivered in the layout your training code reads: https://kenpathlabs.com/lokah/formats. --- # Hindi speech recognition utterances > Hindi speech recognition data: 218,428 single-speaker utterance clips of conversational Hindi, call-centre and everyday, each with its own time-aligned transcript in Devanagari script, English words kept as spoken. 350 hours of audio, 3,377,105 transcribed words. `LK-SP-HIN-003` · [Get a quote](https://kenpathlabs.com/lokah/datasets/hindi-speech-recognition-utterances#contact) · [Record as JSON](https://kenpathlabs.com/api/lokah/datasets/hindi-speech-recognition-utterances) · [Croissant](https://kenpathlabs.com/lokah/datasets/hindi-speech-recognition-utterances/croissant.json) Hindi speech recognition data: 218,428 single-speaker utterance clips of conversational Hindi, call-centre and everyday, each with its own time-aligned transcript in Devanagari script, English words kept as spoken. 350 hours of audio, 3,377,105 transcribed words. Every recording has a complete, segment-level transcript. Personal data: redacted. Licence: custom, quoted per use. ## Specification | Field | Value | Note | | --- | --- | --- | | id | LK-SP-HIN-003 | | | type | speech · utterances for ASR · call centre and general | | | language | हिन्दी · Hindi · hi-IN | | | hours | 350 h | | | channels | One channel, one speaker per clip | | | files | 218,428 files · 218,428 utterances | counted across the full set | | audio | FLAC · 16 kHz · 16-bit | measured across the full set | | release | v1.0 | | | transcript | time-aligned by segment · Devanagari script | | | speakers | 667 across the full set · id and gender per speaker | | | pii | Numbers of 4 or more digits, digit strings spoken as words, and email addresses are masked as [pii] in text and replaced by a tone in audio. First names are not masked. | | | source | Recorded for the dataset | | | review | Every recording has a complete, segment-level transcript | | | personal data | Redacted | | | licence | Custom | quoted per use | ## What it is good for | Task | Fit | Why | | --- | --- | --- | | Speech recognition | yes | Built for it: 218,428 utterance clips, each with its own time-aligned transcript, noise left as recorded. | | Full-duplex speech to speech | no | Clips are single utterances; the conversation timing is not in this dataset. | | Turn-taking | no | No turn structure: each clip is one utterance. | | Voice agents | partly | Good for the recogniser in an agent; the dialogue itself is not here. | | Diarisation | no | Each clip holds one speaker. | | Text to speech | no | Conversational call audio, not studio voice. | ## Conversation profile This dataset is utterance clips, not whole conversations, so there is no conversation profile. ## Sample No public excerpt. A sample is sent on request. ## Licence Licence: custom, quoted per use (training, evaluation or both; internal or commercial; exclusive or not). Delivered in the layout your training code reads: https://kenpathlabs.com/lokah/formats. --- # Hindi speaker diarization > Hindi speaker diarization data: whole call-centre and everyday two-speaker conversations with every speaker turn marked, 432,929 turns in RTTM, for training and scoring who spoke when. 321 hours across 1,394 recordings. `LK-SP-HIN-004` · [Get a quote](https://kenpathlabs.com/lokah/datasets/hindi-speaker-diarization#contact) · [Record as JSON](https://kenpathlabs.com/api/lokah/datasets/hindi-speaker-diarization) · [Croissant](https://kenpathlabs.com/lokah/datasets/hindi-speaker-diarization/croissant.json) Hindi speaker diarization data: whole call-centre and everyday two-speaker conversations with every speaker turn marked, 432,929 turns in RTTM, for training and scoring who spoke when. 321 hours across 1,394 recordings. Measured from 6 sample conversations (24 minutes): 16 kHz, 16-bit PCM WAV, one mono file per conversation. Across them 9% of transcript words are English written in Latin script, 17% of the time is silence, and there are 320 turns. Every recording has a complete, segment-level transcript. Personal data: redacted. Licence: custom, quoted per use. ## Specification | Field | Value | Note | | --- | --- | --- | | id | LK-SP-HIN-004 | | | type | speech · speaker diarization · call centre and general | | | language | हिन्दी · Hindi · hi-IN | | | hours | 321 h | | | channels | One channel, speakers labelled in the transcript | | | files | 1,394 files · 432,929 speaker turns | counted across the full set | | layouts | 1,394 mono conversation | counted across the full set | | audio | FLAC · 16 kHz · 16-bit | measured across the full set | | bandwidth | wideband (8 kHz) | | | snr | 30.3 dB median | across the full set | | release | v1.0 | | | transcript | time-aligned by segment · Devanagari script · delivered as JSON | | | speakers | 619 across the full set · id and gender per speaker | | | pii | Numbers of 4 or more digits, digit strings spoken as words, and email addresses are masked as [pii] in text and replaced by a tone in audio. First names are not masked. | | | source | Recorded for the dataset | | | review | Every recording has a complete, segment-level transcript | | | personal data | Redacted | | | licence | Custom | quoted per use | ## What it is good for | Task | Fit | Why | | --- | --- | --- | | Speech recognition | partly | Whole recordings with time-aligned transcripts; built for who spoke when, not for utterance-level training. | | Full-duplex speech to speech | no | Single-channel recordings. | | Turn-taking | yes | Every change of speaker is marked: 432,929 turns. | | Voice agents | partly | Teaches an agent who is speaking, and when. | | Diarisation | yes | Built for it: RTTM turn files for 1,394 recordings. | | Text to speech | no | Conversational call audio, not studio voice. | ## Conversation profile Measured from 6 full conversations (24 minutes, 320 turns, 3402 words). | Measure | Value | | --- | --- | | Talk time, speaker 1 / speaker 2 | 43% / 57% | | Silence | 17% | | Overlapping speech | 9% | | English words, written in Latin script | 9% | | Speaking rate | 168 words a minute | One mixed channel: segment edges were placed by an annotator, so turn timing is approximate and overlap is an event label. ## Audio quality Frame RMS at 20 ms on the channels mixed to mono. Noise floor is the 10th percentile, speech level the 90th; SNR is their difference. Bandwidth is the highest frequency at which speech still rises 6 dB above the recording's own noise spectrum. | Measure | Value | Reading | | --- | --- | --- | | Speech above noise floor (SNR) | 54.4 dB (28.3 to 81.7) | clean | | Noise floor | -70.2 dBFS | | | Effective bandwidth | 8.0 kHz | wideband (8 kHz) across the full set | | Clipping | 0.009% of samples | | | DNSMOS P.835 (1 to 5) | background 3.42, speech 3.04, overall 2.56 | listener-rated quality, estimated | ## Samples 4 public excerpts, 44.855 seconds each, from different conversations in the dataset. - Sample 1, Open conversation, from a 4-minute conversation: [audio](https://kenpathlabs.com/lokah-samples/hindi-speaker-diarization--1.m4a) · [segments](https://kenpathlabs.com/lokah-samples/hindi-speaker-diarization--1.segments.json) - Sample 2, Open conversation, from a 4-minute conversation: [audio](https://kenpathlabs.com/lokah-samples/hindi-speaker-diarization--2.m4a) · [segments](https://kenpathlabs.com/lokah-samples/hindi-speaker-diarization--2.segments.json) - Sample 3, Open conversation, from a 4-minute conversation: [audio](https://kenpathlabs.com/lokah-samples/hindi-speaker-diarization--3.m4a) · [segments](https://kenpathlabs.com/lokah-samples/hindi-speaker-diarization--3.segments.json) - Sample 4, Open conversation, from a 4-minute conversation (Speaker 1: male; Speaker 2: male): [audio](https://kenpathlabs.com/lokah-samples/hindi-speaker-diarization--4.m4a) · [segments](https://kenpathlabs.com/lokah-samples/hindi-speaker-diarization--4.segments.json) ### Sample 1 transcript Open conversation, starting at 0:31. Audio sha256 `ce52248430a77ce3dac29efedbc75ee4ae26ec76c4c40b32530ed82d80fe535a`. | Start | Speaker | Text | | --- | --- | --- | | 0:00 | Speaker 1 | hello | | 0:01 | Speaker 2 | हाँ madam [filler] | | 0:03 | Speaker 1 | hello | | 0:05 | Speaker 2 | hello | | 0:06 | Speaker 1 | हाँ राजू | | 0:08 | Speaker 2 | हाँ | | 0:10 | Speaker 1 | बोलो | | 0:12 | Speaker 2 | [filler] हम तो यहाँ में है घर में तुम्हार | | 0:18 | Speaker 1 | hello | | 0:19 | Speaker 2 | hello [overlap] हाँ क्या कर रहा है? | | 0:22 | Speaker 1 | [overlap] हाँ | | 0:22 | Speaker 1 | hello | | 0:24 | Speaker 2 | हाँ क्या कर रहा है? | | 0:25 | Speaker 1 | अरे तुम बात मत करो ना. record कर रहा था तो hello | | 0:30 | Speaker 2 | hello | | 0:31 | Speaker 1 | हाँ राजू | | 0:33 | Speaker 2 | ओय | | 0:33 | Speaker 1 | हाँ बोलो | | 0:35 | Speaker 2 | क्या करना है? बोलके आप जब | | 0:37 | Speaker 1 | हं हम तो ऐसे बैठा रहा था घर पे | | 0:41 | Speaker 2 | हं अच्छा अच्छा कल का loanding में है ना. बस्ती में है. | ## Licence Licence: custom, quoted per use (training, evaluation or both; internal or commercial; exclusive or not). Delivered in the layout your training code reads: https://kenpathlabs.com/lokah/formats. --- # Hindi and Tamil full-duplex call-centre conversations > Hindi and Tamil full-duplex conversation data: 1,265 two-channel call-centre calls with each speaker on a separate channel, overlaps, backchannels and turn timing preserved, transcripts time-aligned per channel. 264 hours. `LK-SP-MUL-001` · [Get a quote](https://kenpathlabs.com/lokah/datasets/hindi-tamil-full-duplex-call-centre-conversations#contact) · [Record as JSON](https://kenpathlabs.com/api/lokah/datasets/hindi-tamil-full-duplex-call-centre-conversations) · [Croissant](https://kenpathlabs.com/lokah/datasets/hindi-tamil-full-duplex-call-centre-conversations/croissant.json) Hindi and Tamil full-duplex conversation data: 1,265 two-channel call-centre calls with each speaker on a separate channel, overlaps, backchannels and turn timing preserved, transcripts time-aligned per channel. 264 hours. Measured from 6 sample conversations (25 minutes): 16 kHz, 16-bit PCM WAV, two mono files per conversation. Across them 16% of transcript words are English written in Latin script, 11% of the time is silence, and there are 278 turns. 11 distinct voices in the sample (6 male, 5 female). Every recording has a complete, segment-level transcript. Personal data: redacted. Licence: custom, quoted per use. ## Specification | Field | Value | Note | | --- | --- | --- | | id | LK-SP-MUL-001 | | | type | speech · full-duplex calls · call centre | | | language | हिन्दी · Hindi · hi-IN / தமிழ் · Tamil · ta-IN | | | hours | 264 h | | | channels | Two channels, one per speaker | | | files | 1,265 files · 1,265 calls | counted across the full set | | layouts | 1,265 mono conversation | counted across the full set | | audio | FLAC · 16 kHz · 16-bit | measured across the full set | | bandwidth | wideband (8 kHz) | | | snr | 29.2 dB median | across the full set | | release | v1.0 | | | transcript | time-aligned by segment · Devanagari and Tamil script · delivered as JSON | | | speakers | 529 across the full set · id and gender per speaker | | | pii | Numbers of 4 or more digits, digit strings spoken as words, and email addresses are masked as [pii] in text and replaced by a tone in audio. First names are not masked. | | | source | Recorded for the dataset | | | review | Every recording has a complete, segment-level transcript | | | personal data | Redacted | | | licence | Custom | quoted per use | ## What it is good for | Task | Fit | Why | | --- | --- | --- | | Speech recognition | partly | Per-channel transcripts are included; built for dialogue timing, not for utterance-level training. | | Full-duplex speech to speech | yes | Built for it: 1,265 two-channel calls, each speaker on a separate channel, overlaps and timing intact. | | Turn-taking | yes | Floor transfers, overlaps and backchannels are measurable from the two channels. | | Voice agents | yes | Real agent and customer turns in a support flow, the condition a deployed agent hears. | | Diarisation | partly | Two speakers, already separated by channel; useful as ground truth, not as a hard case. | | Text to speech | no | Conversational call audio, not studio voice. | ## Conversation profile Measured from 6 full conversations (25 minutes, 278 turns, 3540 words). | Measure | Value | | --- | --- | | Talk time, speaker 1 / speaker 2 | 54% / 46% | | Silence | 23% | | Overlapping speech | 8% | | English words, written in Latin script | 16% | | Speaking rate | 161.6 words a minute | ### Turn-taking, measured from the two channels Voice activity detected on each channel at 10 ms; IPUs bounded by more than 200 ms of silence; talkspurts under 90 ms dropped. | Per minute | This dataset | Fisher corpus | | --- | --- | --- | | Inter-pausal units | 34.3 | 21.6 | | Pauses | 11.7 | 7 | | Gaps | 8.6 | 7.5 | | Overlaps | 13.7 | 6.5 | | Backchannels | 8.6 | not reported | Floor-transfer offset: median +0.20 s, 10th to 90th percentile -0.38 s to +1.21 s. 35% of 331 changes of speaker were overlapped. Channel isolation: -39.4 dB. ## Audio quality Frame RMS at 20 ms on the channels mixed to mono. Noise floor is the 10th percentile, speech level the 90th; SNR is their difference. Bandwidth is the highest frequency at which speech still rises 6 dB above the recording's own noise spectrum. | Measure | Value | Reading | | --- | --- | --- | | Speech above noise floor (SNR) | 53.2 dB (31.5 to 61.1) | clean | | Noise floor | -84.1 dBFS | | | Effective bandwidth | 7.9 kHz | wideband (8 kHz) across the full set | | Clipping | 0.000% of samples | | | DNSMOS P.835 (1 to 5) | background 3.55, speech 3.17, overall 2.65 | listener-rated quality, estimated | ## Samples 4 public excerpts, 42.898 seconds each, from different conversations in the dataset. - Sample 1, Two-channel call, from a 4-minute conversation (Speaker 1: male; Speaker 2: male): [audio](https://kenpathlabs.com/lokah-samples/hindi-tamil-full-duplex-call-centre-conversations--1.m4a) · [segments](https://kenpathlabs.com/lokah-samples/hindi-tamil-full-duplex-call-centre-conversations--1.segments.json) - Sample 2, Two-channel call, from a 4-minute conversation (Speaker 1: male; Speaker 2: male): [audio](https://kenpathlabs.com/lokah-samples/hindi-tamil-full-duplex-call-centre-conversations--2.m4a) · [segments](https://kenpathlabs.com/lokah-samples/hindi-tamil-full-duplex-call-centre-conversations--2.segments.json) - Sample 3, Two-channel call, from a 4-minute conversation (Speaker 1: male; Speaker 2: female): [audio](https://kenpathlabs.com/lokah-samples/hindi-tamil-full-duplex-call-centre-conversations--3.m4a) · [segments](https://kenpathlabs.com/lokah-samples/hindi-tamil-full-duplex-call-centre-conversations--3.segments.json) - Sample 4, Two-channel call, from a 4-minute conversation (Speaker 1: female; Speaker 2: male): [audio](https://kenpathlabs.com/lokah-samples/hindi-tamil-full-duplex-call-centre-conversations--4.m4a) · [segments](https://kenpathlabs.com/lokah-samples/hindi-tamil-full-duplex-call-centre-conversations--4.segments.json) ### Sample 1 transcript Two-channel call, starting at 0:03. Stereo: left channel is speaker 1, right channel is speaker 2. Audio sha256 `0be133fcc76d86e7c37e671583aba4c7af90ce1d9f031790dc89f4e4209f37c4`. | Start | Speaker | Text | | --- | --- | --- | | 0:00 | Speaker 1 | hello | | 0:01 | Speaker 2 | hello | | 0:03 | Speaker 1 | नमस्कार sir | | 0:04 | Speaker 2 | नमश्कार. | | 0:06 | Speaker 1 | sir आप company | | 0:08 | Speaker 2 | जी बताइए आप कौन? | | 0:11 | Speaker 1 | sir मैं अरविंदर सिंह बात कर रहा हूँ. | | 0:15 | Speaker 2 | जी | | 0:17 | Speaker 1 | OK | | 0:20 | Speaker 2 | जी जरूर आप की ही सेवा में बैठे हैं. | | 0:23 | Speaker 1 | actually मैंने अभी तक कोई policy ली नहीं है और कोई बीमा भी नहीं करवाया है. तो मुझे उसके बारे में जानकारी नहीं है तो sir मुझे थोड़ा आप विस्तार से उसके बारे में बता सके बीमा के. | | 0:27 | Speaker 2 | जी | | 0:34 | Speaker 2 | अच्छा जी जी जी. | | 0:36 | Speaker 1 | जी sir actually क्या है कि बीमा मुझे तो दो तीन बीमे करवाने है. | | 0:41 | Speaker 2 | अच्छा. | ## Licence Licence: custom, quoted per use (training, evaluation or both; internal or commercial; exclusive or not). Delivered in the layout your training code reads: https://kenpathlabs.com/lokah/formats. --- # Tamil speaker diarization > Tamil speaker diarization data: whole call-centre and everyday two-speaker conversations with every speaker turn marked, 345,825 turns in RTTM, for training and scoring who spoke when. 229 hours across 1,307 recordings. `LK-SP-TAM-005` · [Get a quote](https://kenpathlabs.com/lokah/datasets/tamil-speaker-diarization#contact) · [Record as JSON](https://kenpathlabs.com/api/lokah/datasets/tamil-speaker-diarization) · [Croissant](https://kenpathlabs.com/lokah/datasets/tamil-speaker-diarization/croissant.json) Tamil speaker diarization data: whole call-centre and everyday two-speaker conversations with every speaker turn marked, 345,825 turns in RTTM, for training and scoring who spoke when. 229 hours across 1,307 recordings. Measured from 6 sample conversations (25 minutes): 16 kHz, 16-bit PCM WAV, one mono file per conversation. Across them 39% of transcript words are English written in Latin script, 5% of the time is silence, and there are 385 turns. Every recording has a complete, segment-level transcript. Personal data: redacted. Licence: custom, quoted per use. ## Specification | Field | Value | Note | | --- | --- | --- | | id | LK-SP-TAM-005 | | | type | speech · speaker diarization · call centre and general | | | language | தமிழ் · Tamil · ta-IN | | | hours | 229 h | | | channels | One channel, speakers labelled in the transcript | | | files | 1,307 files · 345,825 speaker turns | counted across the full set | | layouts | 1,307 mono conversation | counted across the full set | | audio | FLAC · 16 kHz · 16-bit | measured across the full set | | bandwidth | wideband (8 kHz) | | | snr | 25.7 dB median | across the full set | | release | v1.0 | | | transcript | time-aligned by segment · Tamil script · delivered as JSON | | | speakers | 1,238 across the full set · id and gender per speaker | | | pii | Numbers of 4 or more digits, digit strings spoken as words, and email addresses are masked as [pii] in text and replaced by a tone in audio. First names are not masked. | | | source | Recorded for the dataset | | | review | Every recording has a complete, segment-level transcript | | | personal data | Redacted | | | licence | Custom | quoted per use | ## What it is good for | Task | Fit | Why | | --- | --- | --- | | Speech recognition | partly | Whole recordings with time-aligned transcripts; built for who spoke when, not for utterance-level training. | | Full-duplex speech to speech | no | Single-channel recordings. | | Turn-taking | yes | Every change of speaker is marked: 345,825 turns. | | Voice agents | partly | Teaches an agent who is speaking, and when. | | Diarisation | yes | Built for it: RTTM turn files for 1,307 recordings. | | Text to speech | no | Conversational call audio, not studio voice. | ## Conversation profile Measured from 6 full conversations (25 minutes, 385 turns, 3164 words). | Measure | Value | | --- | --- | | Talk time, speaker 1 / speaker 2 | 67% / 33% | | Silence | 5% | | Overlapping speech | 0% | | English words, written in Latin script | 39% | | Speaking rate | 136.2 words a minute | One mixed channel: segment edges were placed by an annotator, so turn timing is approximate and overlap is an event label. ## Audio quality Frame RMS at 20 ms on the channels mixed to mono. Noise floor is the 10th percentile, speech level the 90th; SNR is their difference. Bandwidth is the highest frequency at which speech still rises 6 dB above the recording's own noise spectrum. | Measure | Value | Reading | | --- | --- | --- | | Speech above noise floor (SNR) | 47.9 dB (27.3 to 98.7) | clean | | Noise floor | -74 dBFS | | | Effective bandwidth | 7.9 kHz | wideband (8 kHz) across the full set | | Clipping | 0.000% of samples | | | DNSMOS P.835 (1 to 5) | background 3.06, speech 3.04, overall 2.37 | listener-rated quality, estimated | ## Samples 4 public excerpts, 44.131 seconds each, from different conversations in the dataset. - Sample 1, Open conversation, from a 4-minute conversation (Speaker 1: male; Speaker 2: female): [audio](https://kenpathlabs.com/lokah-samples/tamil-speaker-diarization--1.m4a) · [segments](https://kenpathlabs.com/lokah-samples/tamil-speaker-diarization--1.segments.json) - Sample 2, Open conversation, from a 4-minute conversation (Speaker 1: male; Speaker 2: female): [audio](https://kenpathlabs.com/lokah-samples/tamil-speaker-diarization--2.m4a) · [segments](https://kenpathlabs.com/lokah-samples/tamil-speaker-diarization--2.segments.json) - Sample 3, Open conversation, from a 4-minute conversation (Speaker 1: male; Speaker 2: female): [audio](https://kenpathlabs.com/lokah-samples/tamil-speaker-diarization--3.m4a) · [segments](https://kenpathlabs.com/lokah-samples/tamil-speaker-diarization--3.segments.json) - Sample 4, Open conversation, from a 4-minute conversation (Speaker 1: female; Speaker 2: male): [audio](https://kenpathlabs.com/lokah-samples/tamil-speaker-diarization--4.m4a) · [segments](https://kenpathlabs.com/lokah-samples/tamil-speaker-diarization--4.segments.json) ### Sample 1 transcript Open conversation, starting at 1:05. Audio sha256 `6ad726466610b5ae14cb72ca4e618af7cedcd2e9198a1deb0594e98821391226`. | Start | Speaker | Text | | --- | --- | --- | | 0:00 | Speaker 1 | hair ஐ தொட்டு பாக்கும்போது soft ஆ இருக்கறதுக்கு madam, | | 0:02 | Speaker 2 | four. | | 0:03 | Speaker 1 | hair வந்து சிக்கு இல்லாம இருக்கா easy யா சீவுறதுக்கு வந்து பாக்கும்போது, | | 0:07 | Speaker 2 | four. | | 0:08 | Speaker 1 | hair shining க்கு madam, | | 0:09 | Speaker 2 | five. | | 0:10 | Speaker 1 | hair வந்து dry ஆகாம வச்சுருக்குறதுக்கு madam, | | 0:13 | Speaker 2 | four. | | 0:15 | Speaker 1 | [filler] conditioning effect க்கு madam, [silence] | | 0:18 | Speaker 2 | [filler] five. | | 0:20 | Speaker 1 | [unintelligible] துகள்களா hair ல தங்காம easy யா remove ஆகுறதுக்கு madam, | | 0:23 | Speaker 2 | four. [silence] | | 0:24 | Speaker 1 | body heat reduce பண்றதுக்கு madam, [silence] | | 0:27 | Speaker 2 | four. [silence] | | 0:29 | Speaker 1 | நுரையோட அளவுக்கு, [silence] | | 0:31 | Speaker 2 | three. [silence] | | 0:33 | Speaker 1 | smooth ஆ இருக்குறதுக்கு madam, hair smooth ஆ soft ஆ இருக்குறதுக்கு madam. | | 0:35 | Speaker 2 | four. | | 0:36 | Speaker 1 | dandruff control பண்றதுக்கு madam [silence], | | 0:39 | Speaker 2 | four. | | 0:40 | Speaker 1 | [unintelligible] எல்லாம் remove பண்றதுக்கு madam, [silence] | | 0:43 | Speaker 2 | five. | ## Licence Licence: custom, quoted per use (training, evaluation or both; internal or commercial; exclusive or not). Delivered in the layout your training code reads: https://kenpathlabs.com/lokah/formats. --- # Tamil speech recognition utterances > Tamil speech recognition data: 157,231 single-speaker utterance clips of conversational Tamil, call-centre and everyday, each with its own time-aligned transcript in Tamil script, English words kept as spoken. 201 hours of audio, 1,748,023 transcribed words. `LK-SP-TAM-004` · [Get a quote](https://kenpathlabs.com/lokah/datasets/tamil-speech-recognition-utterances#contact) · [Record as JSON](https://kenpathlabs.com/api/lokah/datasets/tamil-speech-recognition-utterances) · [Croissant](https://kenpathlabs.com/lokah/datasets/tamil-speech-recognition-utterances/croissant.json) Tamil speech recognition data: 157,231 single-speaker utterance clips of conversational Tamil, call-centre and everyday, each with its own time-aligned transcript in Tamil script, English words kept as spoken. 201 hours of audio, 1,748,023 transcribed words. Every recording has a complete, segment-level transcript. Personal data: redacted. Licence: custom, quoted per use. ## Specification | Field | Value | Note | | --- | --- | --- | | id | LK-SP-TAM-004 | | | type | speech · utterances for ASR · call centre and general | | | language | தமிழ் · Tamil · ta-IN | | | hours | 201 h | | | channels | One channel, one speaker per clip | | | files | 157,231 files · 157,231 utterances | counted across the full set | | audio | FLAC · 16 kHz · 16-bit | measured across the full set | | release | v1.0 | | | transcript | time-aligned by segment · Tamil script | | | speakers | 1,236 across the full set · id and gender per speaker | | | pii | Numbers of 4 or more digits, digit strings spoken as words, and email addresses are masked as [pii] in text and replaced by a tone in audio. First names are not masked. | | | source | Recorded for the dataset | | | review | Every recording has a complete, segment-level transcript | | | personal data | Redacted | | | licence | Custom | quoted per use | ## What it is good for | Task | Fit | Why | | --- | --- | --- | | Speech recognition | yes | Built for it: 157,231 utterance clips, each with its own time-aligned transcript, noise left as recorded. | | Full-duplex speech to speech | no | Clips are single utterances; the conversation timing is not in this dataset. | | Turn-taking | no | No turn structure: each clip is one utterance. | | Voice agents | partly | Good for the recogniser in an agent; the dialogue itself is not here. | | Diarisation | no | Each clip holds one speaker. | | Text to speech | no | Conversational call audio, not studio voice. | ## Conversation profile This dataset is utterance clips, not whole conversations, so there is no conversation profile. ## Sample No public excerpt. A sample is sent on request. ## Licence Licence: custom, quoted per use (training, evaluation or both; internal or commercial; exclusive or not). Delivered in the layout your training code reads: https://kenpathlabs.com/lokah/formats. --- # Tamil general conversation > Two-speaker general conversation in Tamil, recorded on one channel with speakers labelled. 105 hours. Transcripts are time-aligned and written in Tamil script, with English words kept as spoken. `LK-SP-TAM-003` · [Get a quote](https://kenpathlabs.com/lokah/datasets/tamil-general-conversation#contact) · [Record as JSON](https://kenpathlabs.com/api/lokah/datasets/tamil-general-conversation) · [Croissant](https://kenpathlabs.com/lokah/datasets/tamil-general-conversation/croissant.json) Two-speaker general conversation in Tamil, recorded on one channel with speakers labelled. 105 hours. Transcripts are time-aligned and written in Tamil script, with English words kept as spoken. Measured from 6 sample conversations (25 minutes): 16 kHz, 16-bit PCM WAV, one mono file per conversation. Across them 36% of transcript words are English written in Latin script, 6% of the time is silence, and there are 384 turns. 12 distinct voices in the sample (9 female, 3 male). Every recording has a complete, segment-level transcript. Personal data: redacted. Licence: custom, quoted per use. ## Specification | Field | Value | Note | | --- | --- | --- | | id | LK-SP-TAM-003 | | | type | speech · conversational · general | | | language | தமிழ் · Tamil · ta-IN | | | hours | 105 h | | | channels | One channel, speakers labelled in the transcript | | | files | 464 files · 464 conversations | counted across the full set | | layouts | 464 mono conversation | counted across the full set | | audio | FLAC · 16 kHz · 16-bit | measured across the full set | | bandwidth | wideband (8 kHz) | no telephone-band files | | snr | 25.7 dB median | across the full set | | release | v1.0 | | | transcript | time-aligned by segment · Tamil script · delivered as JSON | | | speakers | 279 across the full set · id and gender per speaker | | | pii | Numbers of 4 or more digits, digit strings spoken as words, and email addresses are masked as [pii] in text and replaced by a tone in audio. First names are not masked. | | | source | Recorded for the dataset | | | review | Every recording has a complete, segment-level transcript | | | personal data | Redacted | | | licence | Custom | quoted per use | ## What it is good for | Task | Fit | Why | | --- | --- | --- | | Speech recognition | yes | Time-aligned transcripts in native script, English kept as spoken. 25.7 dB median SNR across the full set, wideband (42.1 dB in the sample conversations). | | Full-duplex speech to speech | partly | One mixed channel. Turns are labelled, but overlapping speech cannot be separated, which these models need. | | Turn-taking and endpointing | partly | Turn boundaries come from the transcript, so gaps are approximate and overlap is marked, not separated. | | Voice agents for support | partly | Open conversation, not a support flow. Useful for language and prosody, not for task structure. | | Speaker diarisation | yes | Speaker-attributed segments across full conversations. | | Text to speech | no | 25.7 dB median SNR across the full set, wideband (42.1 dB in the sample conversations): too much background for a voice model. | ## Conversation profile Measured from 6 full conversations (25 minutes, 384 turns, 3333 words). | Measure | Value | | --- | --- | | Talk time, speaker 1 / speaker 2 | 59% / 41% | | Silence | 6% | | Overlapping speech | 0% | | English words, written in Latin script | 36% | | Speaking rate | 143.7 words a minute | One mixed channel: segment edges were placed by an annotator, so turn timing is approximate and overlap is an event label. ## Audio quality Frame RMS at 20 ms on the channels mixed to mono. Noise floor is the 10th percentile, speech level the 90th; SNR is their difference. Bandwidth is the highest frequency at which speech still rises 6 dB above the recording's own noise spectrum. | Measure | Value | Reading | | --- | --- | --- | | Speech above noise floor (SNR) | 42.1 dB (29.4 to 56.4) | clean | | Noise floor | -60.2 dBFS | | | Effective bandwidth | 8.0 kHz | wideband (8 kHz) across the full set | | Clipping | 0.000% of samples | | | DNSMOS P.835 (1 to 5) | background 3, speech 3, overall 2.33 | listener-rated quality, estimated | ## Samples 4 public excerpts, 44.686 seconds each, from different conversations in the dataset. - Sample 1, Open conversation, from a 4-minute conversation (Speaker 1: female; Speaker 2: female): [audio](https://kenpathlabs.com/lokah-samples/tamil-general-conversation--1.m4a) · [segments](https://kenpathlabs.com/lokah-samples/tamil-general-conversation--1.segments.json) - Sample 2, Open conversation, from a 4-minute conversation (Speaker 1: male; Speaker 2: female): [audio](https://kenpathlabs.com/lokah-samples/tamil-general-conversation--2.m4a) · [segments](https://kenpathlabs.com/lokah-samples/tamil-general-conversation--2.segments.json) - Sample 3, Open conversation, from a 4-minute conversation (Speaker 1: male; Speaker 2: female): [audio](https://kenpathlabs.com/lokah-samples/tamil-general-conversation--3.m4a) · [segments](https://kenpathlabs.com/lokah-samples/tamil-general-conversation--3.segments.json) - Sample 4, Open conversation, from a 4-minute conversation (Speaker 1: female; Speaker 2: male): [audio](https://kenpathlabs.com/lokah-samples/tamil-general-conversation--4.m4a) · [segments](https://kenpathlabs.com/lokah-samples/tamil-general-conversation--4.segments.json) ### Sample 1 transcript Open conversation, starting at 0:00. Audio sha256 `5f6ca60c1d7ed0e2558383d4cbe91fab14a5be6c3f80729d49c058744a8fbba7`. | Start | Speaker | Text | | --- | --- | --- | | 0:00 | Speaker 1 | hello. | | 0:01 | Speaker 2 | hello. | | 0:02 | Speaker 1 | [overlap] [filler] யாரு பேசுறீங்க [overlap] hello, good morning mam. | | 0:04 | Speaker 2 | நான் தியா, நாங்க அபிராமி finance ல இருந்து பேசுறோம் mam. | | 0:07 | Speaker 1 | yeah, சொல்லுங்க mam. | | 0:09 | Speaker 2 | [filler], mam உங்களோட name அ நாங்க verify பண்ணிக்கிறேன் mam. | | 0:13 | Speaker 1 | [overlap] yeah, [unintelligible] [overlap] உங்களோட name வந்து | | 0:14 | Speaker 2 | பிரீத்தி சிவக்குமார் தான mam? | | 0:16 | Speaker 1 | அ yes mam. | | 0:17 | Speaker 2 | OK mam, உங்களோட place வந்து மதுரை தான mam? | | 0:21 | Speaker 1 | [filler] ஆமா. | | 0:21 | Speaker 2 | [filler] OK mam, உங்களோட address நான் சொல்றேன் mam correct டானு பாத்துக்கோங்க mam. | | 0:26 | Speaker 1 | [filler] OK mam சொல்லுங்க. | | 0:27 | Speaker 2 | [suppressed] twenty nine six bar | | 0:29 | Speaker 1 | [filler] | | 0:30 | Speaker 2 | [suppressed] அம்மன் கோவில், north street மதுரை correct அ mam. | | 0:33 | Speaker 1 | ம் correct மா. | | 0:34 | Speaker 2 | [filler] OK mam, mam உங்களுக்கு loan எடுக்குறதுக்கு ஏதாவது idea இருக்கா mam? | | 0:40 | Speaker 1 | ஆமா, அது அந்த மாதிரி ஒரு idea இருந்தனால தான் உங்க website பாத்தேன் நானு. | ## Licence Licence: custom, quoted per use (training, evaluation or both; internal or commercial; exclusive or not). Delivered in the layout your training code reads: https://kenpathlabs.com/lokah/formats. --- # Tamil call-centre conversations, consumer surveys > Call-centre conversations in Tamil, consumer surveys, recorded on one channel with speakers labelled. 99 hours. Transcripts are time-aligned and written in Tamil script, with English words kept as spoken. `LK-SP-TAM-002` · [Get a quote](https://kenpathlabs.com/lokah/datasets/tamil-call-centre-customer-service#contact) · [Record as JSON](https://kenpathlabs.com/api/lokah/datasets/tamil-call-centre-customer-service) · [Croissant](https://kenpathlabs.com/lokah/datasets/tamil-call-centre-customer-service/croissant.json) Call-centre conversations in Tamil, consumer surveys, recorded on one channel with speakers labelled. 99 hours. Transcripts are time-aligned and written in Tamil script, with English words kept as spoken. Measured from 6 sample conversations (25 minutes): 16 kHz, 16-bit PCM WAV, one mono file per conversation. Across them 39% of transcript words are English written in Latin script, 9% of the time is silence, and there are 452 turns. 11 distinct voices in the sample (6 female, 5 male). Every recording has a complete, segment-level transcript. Personal data: redacted. Licence: custom, quoted per use. ## Specification | Field | Value | Note | | --- | --- | --- | | id | LK-SP-TAM-002 | | | type | speech · conversational · call centre | | | language | தமிழ் · Tamil · ta-IN | | | hours | 99 h | | | channels | One channel, speakers labelled in the transcript | | | files | 715 files · 715 conversations | counted across the full set | | layouts | 715 mono conversation | counted across the full set | | audio | FLAC · 16 kHz · 16-bit | measured across the full set | | bandwidth | wideband (8 kHz) | no telephone-band files | | snr | 22.9 dB median | across the full set | | release | v1.0 | | | transcript | time-aligned by segment · Tamil script · delivered as JSON | | | speakers | 883 across the full set · id and gender per speaker | | | pii | Numbers of 4 or more digits, digit strings spoken as words, and email addresses are masked as [pii] in text and replaced by a tone in audio. First names are not masked. | | | source | Recorded for the dataset | | | review | Every recording has a complete, segment-level transcript | | | personal data | Redacted | | | licence | Custom | quoted per use | ## What it is good for | Task | Fit | Why | | --- | --- | --- | | Speech recognition | yes | Time-aligned transcripts in native script, English kept as spoken. 22.9 dB median SNR across the full set, wideband (47 dB in the sample conversations). | | Full-duplex speech to speech | partly | One mixed channel. Turns are labelled, but overlapping speech cannot be separated, which these models need. | | Turn-taking and endpointing | partly | Turn boundaries come from the transcript, so gaps are approximate and overlap is marked, not separated. | | Voice agents for support | yes | Agent and customer turns in a real support flow. | | Speaker diarisation | yes | Speaker-attributed segments across full conversations. | | Text to speech | no | 22.9 dB median SNR across the full set, wideband (47 dB in the sample conversations): too much background for a voice model. | ## Conversation profile Measured from 6 full conversations (25 minutes, 452 turns, 2800 words). | Measure | Value | | --- | --- | | Talk time, speaker 1 / speaker 2 | 58% / 42% | | Silence | 9% | | Overlapping speech | 0% | | English words, written in Latin script | 39% | | Speaking rate | 124.7 words a minute | One mixed channel: segment edges were placed by an annotator, so turn timing is approximate and overlap is an event label. ## Audio quality Frame RMS at 20 ms on the channels mixed to mono. Noise floor is the 10th percentile, speech level the 90th; SNR is their difference. Bandwidth is the highest frequency at which speech still rises 6 dB above the recording's own noise spectrum. | Measure | Value | Reading | | --- | --- | --- | | Speech above noise floor (SNR) | 47 dB (38.9 to 97.9) | clean | | Noise floor | -71 dBFS | | | Effective bandwidth | 7.9 kHz | wideband (8 kHz) across the full set | | Clipping | 0.000% of samples | | | DNSMOS P.835 (1 to 5) | background 3.45, speech 3.03, overall 2.54 | listener-rated quality, estimated | ## Samples 4 public excerpts, 40.965 seconds each, from different conversations in the dataset. - Sample 1, Call-centre conversation, from a 4-minute conversation (Speaker 1: male; Speaker 2: female): [audio](https://kenpathlabs.com/lokah-samples/tamil-call-centre-customer-service--1.m4a) · [segments](https://kenpathlabs.com/lokah-samples/tamil-call-centre-customer-service--1.segments.json) - Sample 2, Call-centre conversation, from a 4-minute conversation (Speaker 1: female; Speaker 2: male): [audio](https://kenpathlabs.com/lokah-samples/tamil-call-centre-customer-service--2.m4a) · [segments](https://kenpathlabs.com/lokah-samples/tamil-call-centre-customer-service--2.segments.json) - Sample 3, Call-centre conversation, from a 4-minute conversation (Speaker 1: female; Speaker 2: male): [audio](https://kenpathlabs.com/lokah-samples/tamil-call-centre-customer-service--3.m4a) · [segments](https://kenpathlabs.com/lokah-samples/tamil-call-centre-customer-service--3.segments.json) - Sample 4, Call-centre conversation, from a 4-minute conversation (Speaker 1: male; Speaker 2: female): [audio](https://kenpathlabs.com/lokah-samples/tamil-call-centre-customer-service--4.m4a) · [segments](https://kenpathlabs.com/lokah-samples/tamil-call-centre-customer-service--4.segments.json) ### Sample 1 transcript Call-centre conversation, starting at 1:50. Audio sha256 `d57a33380e57b07eae562e0c3244a2ab5930ff81266407ad582dbd02933dfb8a`. | Start | Speaker | Text | | --- | --- | --- | | 0:00 | Speaker 1 | hair shining க்கு? | | 0:01 | Speaker 2 | [filler] five | | 0:02 | Speaker 1 | hair அ dry ஆகாம வேச்சிகிரதுக்கு | | 0:04 | Speaker 2 | five தான் | | 0:06 | Speaker 1 | conditioning effect க்கு | | 0:07 | Speaker 2 | five | | 0:09 | Speaker 1 | powder ஓட துகள் hair ல தங்காம easy remove ஆகுறதுக்கு | | 0:11 | Speaker 2 | five | | 0:13 | Speaker 1 | OK [unintelligible] ஆ release பண்றதுக்கு | | 0:16 | Speaker 2 | [filler] five தாங்க | | 0:18 | Speaker 1 | நொறையோட அளவுக்கு | | 0:19 | Speaker 2 | அதுவும் five தான் | | 0:21 | Speaker 1 | hair soft ஆ smooth ஆ இருக்குறதுக்கு | | 0:23 | Speaker 2 | அதுவும் five தான் | | 0:25 | Speaker 1 | dandruff reduce ஆகுறதுக்கு | | 0:26 | Speaker 2 | five | | 0:27 | Speaker 1 | white flex லாம் remove ஆகுறதுக்கு, | | 0:30 | Speaker 2 | [filler] five | | 0:31 | Speaker 1 | hair fall control பண்ணுறதுக்கு | | 0:33 | Speaker 2 | [filler] five | | 0:34 | Speaker 1 | itching reduce ஆகுறதுக்கு | | 0:36 | Speaker 2 | [filler] five தான் | | 0:37 | Speaker 1 | fresh ஆ இருக்குற மாதிரி feel குடுக்குறதுக்கு | | 0:39 | Speaker 2 | five தான் | ## Licence Licence: custom, quoted per use (training, evaluation or both; internal or commercial; exclusive or not). Delivered in the layout your training code reads: https://kenpathlabs.com/lokah/formats. --- # Hindi general conversation > Two-speaker general conversation in Hindi, recorded on one channel with speakers labelled. 81 hours. Transcripts are time-aligned and written in Devanagari script, with English words kept as spoken. `LK-SP-HIN-002` · [Get a quote](https://kenpathlabs.com/lokah/datasets/hindi-general-conversation#contact) · [Record as JSON](https://kenpathlabs.com/api/lokah/datasets/hindi-general-conversation) · [Croissant](https://kenpathlabs.com/lokah/datasets/hindi-general-conversation/croissant.json) Two-speaker general conversation in Hindi, recorded on one channel with speakers labelled. 81 hours. Transcripts are time-aligned and written in Devanagari script, with English words kept as spoken. Measured from 6 sample conversations (25 minutes): 16 kHz, 16-bit PCM WAV, one mono file per conversation. Across them 14% of transcript words are English written in Latin script, 17% of the time is silence, and there are 359 turns. 11 distinct voices in the sample (4 male, 7 female). Every recording has a complete, segment-level transcript. Personal data: redacted. Licence: custom, quoted per use. ## Specification | Field | Value | Note | | --- | --- | --- | | id | LK-SP-HIN-002 | | | type | speech · conversational · general | | | language | हिन्दी · Hindi · hi-IN | | | hours | 81 h | | | channels | One channel, speakers labelled in the transcript | | | files | 259 files · 259 conversations | counted across the full set | | layouts | 259 mono conversation | counted across the full set | | audio | FLAC · 16 kHz · 16-bit | measured across the full set | | bandwidth | wideband (8 kHz) | no telephone-band files | | snr | 36.8 dB median | across the full set | | release | v1.0 | | | transcript | time-aligned by segment · Devanagari script · delivered as JSON | | | speakers | 177 across the full set · id and gender per speaker | | | pii | Numbers of 4 or more digits, digit strings spoken as words, and email addresses are masked as [pii] in text and replaced by a tone in audio. First names are not masked. | | | source | Recorded for the dataset | | | review | Every recording has a complete, segment-level transcript | | | personal data | Redacted | | | licence | Custom | quoted per use | ## What it is good for | Task | Fit | Why | | --- | --- | --- | | Speech recognition | yes | Time-aligned transcripts in native script, English kept as spoken. 36.8 dB median SNR across the full set, wideband (53 dB in the sample conversations). | | Full-duplex speech to speech | partly | One mixed channel. Turns are labelled, but overlapping speech cannot be separated, which these models need. | | Turn-taking and endpointing | partly | Turn boundaries come from the transcript, so gaps are approximate and overlap is marked, not separated. | | Voice agents for support | partly | Open conversation, not a support flow. Useful for language and prosody, not for task structure. | | Speaker diarisation | yes | Speaker-attributed segments across full conversations. | | Text to speech | partly | 36.8 dB median SNR across the full set, wideband (53 dB in the sample conversations): clean and wideband enough for conversational prosody data, though not a studio voice. | ## Conversation profile Measured from 6 full conversations (25 minutes, 359 turns, 3631 words). | Measure | Value | | --- | --- | | Talk time, speaker 1 / speaker 2 | 48% / 52% | | Silence | 17% | | Overlapping speech | 1% | | English words, written in Latin script | 14% | | Speaking rate | 176.8 words a minute | One mixed channel: segment edges were placed by an annotator, so turn timing is approximate and overlap is an event label (1 in this conversation). ## Audio quality Frame RMS at 20 ms on the channels mixed to mono. Noise floor is the 10th percentile, speech level the 90th; SNR is their difference. Bandwidth is the highest frequency at which speech still rises 6 dB above the recording's own noise spectrum. | Measure | Value | Reading | | --- | --- | --- | | Speech above noise floor (SNR) | 53 dB (36.2 to 107.2) | some background | | Noise floor | -75.7 dBFS | | | Effective bandwidth | 8.0 kHz | wideband (8 kHz) across the full set | | Clipping | 0.015% of samples | | | DNSMOS P.835 (1 to 5) | background 3.75, speech 3.29, overall 2.88 | listener-rated quality, estimated | ## Samples 4 public excerpts, 44.161 seconds each, from different conversations in the dataset. - Sample 1, Open conversation, from a 4-minute conversation (Speaker 1: male; Speaker 2: male): [audio](https://kenpathlabs.com/lokah-samples/hindi-general-conversation--1.m4a) · [segments](https://kenpathlabs.com/lokah-samples/hindi-general-conversation--1.segments.json) - Sample 2, Open conversation, from a 4-minute conversation (Speaker 1: female; Speaker 2: female): [audio](https://kenpathlabs.com/lokah-samples/hindi-general-conversation--2.m4a) · [segments](https://kenpathlabs.com/lokah-samples/hindi-general-conversation--2.segments.json) - Sample 3, Open conversation, from a 4-minute conversation (Speaker 1: male; Speaker 2: male): [audio](https://kenpathlabs.com/lokah-samples/hindi-general-conversation--3.m4a) · [segments](https://kenpathlabs.com/lokah-samples/hindi-general-conversation--3.segments.json) - Sample 4, Open conversation, from a 4-minute conversation (Speaker 1: female; Speaker 2: female): [audio](https://kenpathlabs.com/lokah-samples/hindi-general-conversation--4.m4a) · [segments](https://kenpathlabs.com/lokah-samples/hindi-general-conversation--4.segments.json) ### Sample 1 transcript Open conversation, starting at 0:20. Audio sha256 `0909ab3c1e2cf776a2610409d611e09665f7ab34304210efbf3aa6b05b5b8542`. | Start | Speaker | Text | | --- | --- | --- | | 0:00 | Speaker 2 | हम्म शुरू हो गया कल से शुरू हुआ है। | | 0:02 | Speaker 1 | ओह I C C ना? | | 0:04 | Speaker 2 | हम्म | | 0:05 | Speaker 1 | अच्छा ठीक है। | | 0:09 | Speaker 2 | क्या कल तो opening था ना। | | 0:11 | Speaker 1 | हम्म | | 0:12 | Speaker 2 | england और new zeeland का बीच में था। | | 0:14 | Speaker 1 | हाँ हाँ। | | 0:16 | Speaker 2 | नया एक stadium बना है इंडिया में नरेंद्र मोदी stadium ना। | | 0:19 | Speaker 1 | हाँ हाँ हाँ। | | 0:19 | Speaker 2 | हैदराबाद में। | | 0:20 | Speaker 1 | ह्म्म्म | | 0:21 | Speaker 2 | [filler] वो हमें वो था इनके साथ [unintelligible] [silence] क्या था ये उसको क्या बोलते हो। | | 0:27 | Speaker 1 | [filler] fans लोग जो? | | 0:29 | Speaker 2 | fans लोग [overlap] जो | | 0:29 | Speaker 1 | [overlap] attendance | | 0:30 | Speaker 1 | नहीं था ना ज्यादा। | | 0:31 | Speaker 2 | attendance नहीं था forty three thousand कुछ ही था। | | 0:33 | Speaker 1 | [filler] वाह forty three thousand मामूली है लेकिन बड़ा है नरेंद्र मोदी stadium जो? | | 0:37 | Speaker 2 | बहुत बड़ा है उसमें एक लाख से ज्यादा capacity है। | | 0:40 | Speaker 1 | हम्म | | 0:41 | Speaker 2 | तो उसमें इतना कम आया है। | | 0:43 | Speaker 1 | हाँ। | ## Licence Licence: custom, quoted per use (training, evaluation or both; internal or commercial; exclusive or not). Delivered in the layout your training code reads: https://kenpathlabs.com/lokah/formats. --- # Tamil two-channel customer-service calls > Scripted call-centre conversations in Tamil, telecom, delivery, e-commerce and banking, recorded with each speaker on a separate channel. 25 hours. Transcripts are time-aligned and written in Tamil script, with English words kept as spoken. Layouts across the set: 2 one side of a call, 132 two-channel. `LK-SP-TAM-001` · [Get a quote](https://kenpathlabs.com/lokah/datasets/tamil-call-centre-banking-and-retail#contact) · [Record as JSON](https://kenpathlabs.com/api/lokah/datasets/tamil-call-centre-banking-and-retail) · [Croissant](https://kenpathlabs.com/lokah/datasets/tamil-call-centre-banking-and-retail/croissant.json) Scripted call-centre conversations in Tamil, telecom, delivery, e-commerce and banking, recorded with each speaker on a separate channel. 25 hours. Transcripts are time-aligned and written in Tamil script, with English words kept as spoken. Layouts across the set: 2 one side of a call, 132 two-channel. Measured from 6 sample conversations (24 minutes): 16 kHz, 16-bit PCM WAV, two mono files per conversation. Across them 29% of transcript words are English written in Latin script, 20% of the time is silence, and there are 269 turns. 10 distinct voices in the sample (1 male, 7 female, 2 unknown). Every recording has a complete, segment-level transcript. Personal data: redacted. Licence: custom, quoted per use. ## Specification | Field | Value | Note | | --- | --- | --- | | id | LK-SP-TAM-001 | | | type | speech · conversational · call centre | | | language | தமிழ் · Tamil · ta-IN | | | hours | 25 h | | | channels | Two channels, one per speaker | | | files | 134 files · 134 conversations | counted across the full set | | layouts | 2 one side of a call, 132 two-channel | counted across the full set | | audio | FLAC · 16 kHz · 16-bit | measured across the full set | | bandwidth | wideband (8 kHz) | no telephone-band files | | snr | 31.2 dB median | across the full set | | release | v1.0 | | | transcript | time-aligned by segment · Tamil script · delivered as JSON | | | speakers | 93 across the full set · id and gender per speaker | | | pii | Numbers of 4 or more digits, digit strings spoken as words, and email addresses are masked as [pii] in text and replaced by a tone in audio. First names are not masked. | | | source | Recorded for the dataset | | | review | Every recording has a complete, segment-level transcript | | | personal data | Redacted | | | licence | Custom | quoted per use | ## What it is good for | Task | Fit | Why | | --- | --- | --- | | Speech recognition | yes | Time-aligned transcripts in native script, English kept as spoken. 31.2 dB median SNR across the full set, wideband (42.7 dB in the sample conversations). | | Full-duplex speech to speech | yes | Each speaker on a separate channel, so overlap, backchannels and turn timing survive. Moshi and PersonaPlex train on exactly this layout. | | Turn-taking and endpointing | yes | Gaps and overlaps at every change of speaker are measured from the two channels; see the profile. | | Voice agents for support | yes | Agent and customer turns in a real support flow. | | Speaker diarisation | yes | Speaker-attributed segments across full conversations. | | Text to speech | partly | 31.2 dB median SNR across the full set, wideband (42.7 dB in the sample conversations): clean and wideband enough for conversational prosody data, though not a studio voice. | ## Conversation profile Measured from 6 full conversations (24 minutes, 269 turns, 3067 words). | Measure | Value | | --- | --- | | Talk time, speaker 1 / speaker 2 | 47% / 53% | | Silence | 25% | | Overlapping speech | 4% | | English words, written in Latin script | 29% | | Speaking rate | 156.1 words a minute | ### Turn-taking, measured from the two channels Voice activity detected on each channel at 10 ms; IPUs bounded by more than 200 ms of silence; talkspurts under 90 ms dropped. | Per minute | This dataset | Fisher corpus | | --- | --- | --- | | Inter-pausal units | 21.2 | 21.6 | | Pauses | 6.5 | 7 | | Gaps | 8.9 | 7.5 | | Overlaps | 5.5 | 6.5 | | Backchannels | 4.4 | not reported | Floor-transfer offset: median +1.05 s, 10th to 90th percentile +0.07 s to +1.92 s. 8% of 240 changes of speaker were overlapped. Channel isolation: -36.3 dB. ## Audio quality Frame RMS at 20 ms on the channels mixed to mono. Noise floor is the 10th percentile, speech level the 90th; SNR is their difference. Bandwidth is the highest frequency at which speech still rises 6 dB above the recording's own noise spectrum. | Measure | Value | Reading | | --- | --- | --- | | Speech above noise floor (SNR) | 42.7 dB (28.2 to 60.7) | clean | | Noise floor | -70.5 dBFS | | | Effective bandwidth | 7.8 kHz | wideband (8 kHz) across the full set | | Clipping | 0.000% of samples | | | DNSMOS P.835 (1 to 5) | background 3.78, speech 3.3, overall 2.84 | listener-rated quality, estimated | ## Samples 4 public excerpts, 43.177 seconds each, from different conversations in the dataset. - Sample 1, Call-centre conversation, from a 4-minute conversation (Speaker 1: female; Speaker 2: female): [audio](https://kenpathlabs.com/lokah-samples/tamil-call-centre-banking-and-retail--1.m4a) · [segments](https://kenpathlabs.com/lokah-samples/tamil-call-centre-banking-and-retail--1.segments.json) - Sample 2, Call-centre conversation, from a 4-minute conversation (Speaker 1: male; Speaker 2: female): [audio](https://kenpathlabs.com/lokah-samples/tamil-call-centre-banking-and-retail--2.m4a) · [segments](https://kenpathlabs.com/lokah-samples/tamil-call-centre-banking-and-retail--2.segments.json) - Sample 3, Call-centre conversation, from a 4-minute conversation (Speaker 1: female; Speaker 2: female): [audio](https://kenpathlabs.com/lokah-samples/tamil-call-centre-banking-and-retail--3.m4a) · [segments](https://kenpathlabs.com/lokah-samples/tamil-call-centre-banking-and-retail--3.segments.json) - Sample 4, Call-centre conversation, from a 4-minute conversation: [audio](https://kenpathlabs.com/lokah-samples/tamil-call-centre-banking-and-retail--4.m4a) · [segments](https://kenpathlabs.com/lokah-samples/tamil-call-centre-banking-and-retail--4.segments.json) ### Sample 1 transcript Call-centre conversation, starting at 0:00. Stereo: left channel is speaker 1, right channel is speaker 2. Audio sha256 `2ed5586e4e3d98739f12ee23beaced07cdf0a1b6d1fadb88534bc9de532a8cb5`. | Start | Speaker | Text | | --- | --- | --- | | 0:00 | Speaker 2 | hello | | 0:00 | Speaker 1 | shree five | | 0:03 | Speaker 1 | shree five care center வாடிக்கையாளர் மையத்துக்கு உங்களை வரவேற்கிறம் நான் திவ்யா என்ன சேவை தேவை உங்களுக்கு | | 0:13 | Speaker 2 | madam வணக்கம் madam | | 0:16 | Speaker 1 | வணக்கம் வணக்கம் சொல்லுங்க உங்க பேரு | | 0:20 | Speaker 2 | madam என் பேர் காயத்திரி | | 0:22 | Speaker 1 | ஆ சொல்லுங்க காயத்திரி | | 0:25 | Speaker 2 | நாங்க சேலத்துல இருந்து call பண்ணிருக்கொங்க | | 0:28 | Speaker 1 | ஆ சொல்லுங்க சொல்லுங்க காயத்திரி | | 0:31 | Speaker 2 | madam எங்க அம்மாக்கு ஒடம்பு சரி இல்ல ஒரு ஒரு வாரத்துக்கு முன்னாடிதான் operation பண்ணிருகோம் இப்ப அவங்க விட்ல இருக்காங்க அவங்கள விட்ல வெச்சு பாத்துகோங்க சொல்லி discharge பண்ணிடாங்க | | 0:38 | Speaker 1 | ம்ம் | ## Licence Licence: custom, quoted per use (training, evaluation or both; internal or commercial; exclusive or not). Delivered in the layout your training code reads: https://kenpathlabs.com/lokah/formats. --- # Marathi general conversation > Two-speaker general conversation in Marathi, recorded on one channel with speakers labelled. 9 hours. Transcripts are time-aligned and written in Devanagari script, with English words kept as spoken. `LK-SP-MAR-001` · [Get a quote](https://kenpathlabs.com/lokah/datasets/marathi-general-conversation#contact) · [Record as JSON](https://kenpathlabs.com/api/lokah/datasets/marathi-general-conversation) · [Croissant](https://kenpathlabs.com/lokah/datasets/marathi-general-conversation/croissant.json) Two-speaker general conversation in Marathi, recorded on one channel with speakers labelled. 9 hours. Transcripts are time-aligned and written in Devanagari script, with English words kept as spoken. Measured from 6 sample conversations (25 minutes): 16 kHz, 16-bit PCM WAV, one mono file per conversation. Across them 0% of transcript words are English written in Latin script, 17% of the time is silence, and there are 428 turns. 12 distinct voices in the sample (6 male, 5 female, 1 unknown). Every recording has a complete, segment-level transcript. Personal data: redacted. Licence: custom, quoted per use. ## Specification | Field | Value | Note | | --- | --- | --- | | id | LK-SP-MAR-001 | | | type | speech · conversational · general | | | language | मराठी · Marathi · mr-IN | | | hours | 9 h | | | channels | One channel, speakers labelled in the transcript | | | files | 36 files · 36 conversations | counted across the full set | | layouts | 36 mono conversation | counted across the full set | | audio | FLAC · 16 kHz · 16-bit | measured across the full set | | bandwidth | telephone (4 kHz) | 9 telephone-band files | | snr | 37.8 dB median | across the full set | | release | v1.0 | | | transcript | time-aligned by segment · Devanagari script · delivered as JSON | | | speakers | 40 across the full set · id and gender per speaker | | | pii | Numbers of 4 or more digits, digit strings spoken as words, and email addresses are masked as [pii] in text and replaced by a tone in audio. First names are not masked. | | | source | Recorded for the dataset | | | review | Every recording has a complete, segment-level transcript | | | personal data | Redacted | | | licence | Custom | quoted per use | ## What it is good for | Task | Fit | Why | | --- | --- | --- | | Speech recognition | yes | Time-aligned transcripts in native script, English kept as spoken. 37.8 dB median SNR across the full set (60 dB in the sample conversations). | | Full-duplex speech to speech | partly | One mixed channel. Turns are labelled, but overlapping speech cannot be separated, which these models need. | | Turn-taking and endpointing | partly | Turn boundaries come from the transcript, so gaps are approximate and overlap is marked, not separated. | | Voice agents for support | partly | Open conversation, not a support flow. Useful for language and prosody, not for task structure. | | Speaker diarisation | yes | Speaker-attributed segments across full conversations. | | Text to speech | partly | 37.8 dB median SNR across the full set (60 dB in the sample conversations): clean and wideband enough for conversational prosody data, though not a studio voice. | ## Conversation profile Measured from 6 full conversations (25 minutes, 428 turns, 2877 words). | Measure | Value | | --- | --- | | Talk time, speaker 1 / speaker 2 | 49% / 51% | | Silence | 17% | | Overlapping speech | 2% | | English words, written in Latin script | 0% | | Speaking rate | 138.3 words a minute | One mixed channel: segment edges were placed by an annotator, so turn timing is approximate and overlap is an event label. ## Audio quality Frame RMS at 20 ms on the channels mixed to mono. Noise floor is the 10th percentile, speech level the 90th; SNR is their difference. Bandwidth is the highest frequency at which speech still rises 6 dB above the recording's own noise spectrum. | Measure | Value | Reading | | --- | --- | --- | | Speech above noise floor (SNR) | 60 dB (48.5 to 102.1) | some background | | Noise floor | -74.9 dBFS | | | Effective bandwidth | 8.0 kHz | telephone (4 kHz) across the full set | | Clipping | 3.004% of samples | | | DNSMOS P.835 (1 to 5) | background 3.48, speech 3, overall 2.55 | listener-rated quality, estimated | ## Samples 4 public excerpts, 44.767 seconds each, from different conversations in the dataset. - Sample 1, Open conversation, from a 4-minute conversation (Speaker 1: male; Speaker 2: male): [audio](https://kenpathlabs.com/lokah-samples/marathi-general-conversation--1.m4a) · [segments](https://kenpathlabs.com/lokah-samples/marathi-general-conversation--1.segments.json) - Sample 2, Open conversation, from a 4-minute conversation (Speaker 1: female; Speaker 2: male): [audio](https://kenpathlabs.com/lokah-samples/marathi-general-conversation--2.m4a) · [segments](https://kenpathlabs.com/lokah-samples/marathi-general-conversation--2.segments.json) - Sample 3, Open conversation, from a 4-minute conversation (Speaker 1: male; Speaker 2: female): [audio](https://kenpathlabs.com/lokah-samples/marathi-general-conversation--3.m4a) · [segments](https://kenpathlabs.com/lokah-samples/marathi-general-conversation--3.segments.json) - Sample 4, Open conversation, from a 4-minute conversation (Speaker 1: female; Speaker 2: gender unknown): [audio](https://kenpathlabs.com/lokah-samples/marathi-general-conversation--4.m4a) · [segments](https://kenpathlabs.com/lokah-samples/marathi-general-conversation--4.segments.json) ### Sample 1 transcript Open conversation, starting at 3:11. Audio sha256 `afdd448433a6ee118b59e364eea170bac31298b30a9bad1d0c3bd4234233361d`. | Start | Speaker | Text | | --- | --- | --- | | 0:00 | Speaker 1 | हा | | 0:00 | Speaker 2 | मोठं असत थोडं | | 0:01 | Speaker 1 | हा | | 0:02 | Speaker 2 | त्याच्यामध्ये पण बऱ्याच [unclear] येतात. हि दोन्ही एम प्रकार चाललेत. दोन्ही जाती उत्तमेत. | | 0:08 | Speaker 1 | हा | | 0:08 | Speaker 2 | काही प्रॉब्लेम नाही. | | 0:10 | Speaker 1 | ह्ह | | 0:10 | Speaker 2 | पण हि कशी थोडी मम जी चार नंबर जो असतो ना विषय | | 0:15 | Speaker 1 | हा | | 0:16 | Speaker 2 | ते थोडी आकाराला थोडी कमी असते. | | 0:19 | Speaker 1 | हा म्हणजे त्याची साईझ कमी असते. | | 0:21 | Speaker 2 | साईझ कमी असते आणि पण ती भरगोस लागते | | 0:25 | Speaker 1 | भरगोस लागते हा | | 0:27 | Speaker 2 | [unclear] एका घडला जवळ जवळ दहा दहा पंधरा पंधरा बिया असतात काजू मध्ये [unclear] असतात. | | 0:31 | Speaker 1 | अच्छा अच्छा हा हा | | 0:33 | Speaker 2 | आणि अअ वेंगुर्ला सातचा तर काही प्रश्नच नाही ती एकदम जाडी ब | | 0:38 | Speaker 2 | म मजबूत बी असते ती | | 0:40 | Speaker 1 | हा मोठी असते | | 0:41 | Speaker 1 | हा फल ते एकदम मोठं असतं. | | 0:43 | Speaker 2 | हा बरोबर | ## Licence Licence: custom, quoted per use (training, evaluation or both; internal or commercial; exclusive or not). Delivered in the layout your training code reads: https://kenpathlabs.com/lokah/formats. --- # Marathi call-centre conversations, banking and insurance > Call-centre conversations in Marathi, banking and insurance, recorded on one channel with speakers labelled. 4 hours. Transcripts are time-aligned and written in Devanagari script, with English words kept as spoken. `LK-SP-MAR-002` · [Get a quote](https://kenpathlabs.com/lokah/datasets/marathi-call-centre-banking-and-insurance#contact) · [Record as JSON](https://kenpathlabs.com/api/lokah/datasets/marathi-call-centre-banking-and-insurance) · [Croissant](https://kenpathlabs.com/lokah/datasets/marathi-call-centre-banking-and-insurance/croissant.json) Call-centre conversations in Marathi, banking and insurance, recorded on one channel with speakers labelled. 4 hours. Transcripts are time-aligned and written in Devanagari script, with English words kept as spoken. Measured from 6 sample conversations (25 minutes): 16 kHz, 16-bit PCM WAV, one mono file per conversation. Across them 0% of transcript words are English written in Latin script, 10% of the time is silence, and there are 309 turns. 8 distinct voices in the sample (4 female, 3 male, 1 unknown). Every recording has a complete, segment-level transcript. Personal data: redacted. Licence: custom, quoted per use. ## Specification | Field | Value | Note | | --- | --- | --- | | id | LK-SP-MAR-002 | | | type | speech · conversational · call centre | | | language | मराठी · Marathi · mr-IN | | | hours | 4 h | | | channels | One channel, speakers labelled in the transcript | | | files | 17 files · 17 conversations | counted across the full set | | layouts | 17 mono conversation | counted across the full set | | audio | FLAC · 16 kHz · 16-bit | measured across the full set | | bandwidth | telephone (4 kHz) | no telephone-band files | | snr | 35.5 dB median | across the full set | | release | v1.0 | | | transcript | time-aligned by segment · Devanagari script · delivered as JSON | | | speakers | 18 across the full set · id and gender per speaker | | | pii | Numbers of 4 or more digits, digit strings spoken as words, and email addresses are masked as [pii] in text and replaced by a tone in audio. First names are not masked. | | | source | Call-centre recordings | | | review | Every recording has a complete, segment-level transcript | | | personal data | Redacted | | | licence | Custom | quoted per use | ## What it is good for | Task | Fit | Why | | --- | --- | --- | | Speech recognition | yes | Time-aligned transcripts in native script, English kept as spoken. 35.5 dB median SNR across the full set (43.7 dB in the sample conversations). | | Full-duplex speech to speech | partly | One mixed channel. Turns are labelled, but overlapping speech cannot be separated, which these models need. | | Turn-taking and endpointing | partly | Turn boundaries come from the transcript, so gaps are approximate and overlap is marked, not separated. | | Voice agents for support | yes | Agent and customer turns in a real support flow. | | Speaker diarisation | yes | Speaker-attributed segments across full conversations. | | Text to speech | no | 35.5 dB median SNR across the full set (43.7 dB in the sample conversations): too much background for a voice model. | ## Conversation profile Measured from 6 full conversations (25 minutes, 309 turns, 3077 words). | Measure | Value | | --- | --- | | Talk time, speaker 1 / speaker 2 | 42% / 58% | | Silence | 10% | | Overlapping speech | 4% | | English words, written in Latin script | 0% | | Speaking rate | 139.9 words a minute | One mixed channel: segment edges were placed by an annotator, so turn timing is approximate and overlap is an event label. ## Audio quality Frame RMS at 20 ms on the channels mixed to mono. Noise floor is the 10th percentile, speech level the 90th; SNR is their difference. Bandwidth is the highest frequency at which speech still rises 6 dB above the recording's own noise spectrum. | Measure | Value | Reading | | --- | --- | --- | | Speech above noise floor (SNR) | 43.7 dB (39.2 to 50.6) | clean | | Noise floor | -72.5 dBFS | | | Effective bandwidth | 4.7 kHz | telephone (4 kHz) across the full set | | Clipping | 0.000% of samples | | | DNSMOS P.835 (1 to 5) | background 3.46, speech 3.23, overall 2.67 | listener-rated quality, estimated | ## Samples 4 public excerpts, 41.534 seconds each, from different conversations in the dataset. - Sample 1, Call-centre conversation, from a 4-minute conversation (Speaker 1: gender unknown; Speaker 2: male): [audio](https://kenpathlabs.com/lokah-samples/marathi-call-centre-banking-and-insurance--1.m4a) · [segments](https://kenpathlabs.com/lokah-samples/marathi-call-centre-banking-and-insurance--1.segments.json) - Sample 2, Call-centre conversation, from a 4-minute conversation (Speaker 1: female; Speaker 2: female): [audio](https://kenpathlabs.com/lokah-samples/marathi-call-centre-banking-and-insurance--2.m4a) · [segments](https://kenpathlabs.com/lokah-samples/marathi-call-centre-banking-and-insurance--2.segments.json) - Sample 3, Call-centre conversation, from a 4-minute conversation (Speaker 1: female; Speaker 2: male): [audio](https://kenpathlabs.com/lokah-samples/marathi-call-centre-banking-and-insurance--3.m4a) · [segments](https://kenpathlabs.com/lokah-samples/marathi-call-centre-banking-and-insurance--3.segments.json) - Sample 4, Call-centre conversation, from a 4-minute conversation (Speaker 1: gender unknown; Speaker 2: female): [audio](https://kenpathlabs.com/lokah-samples/marathi-call-centre-banking-and-insurance--4.m4a) · [segments](https://kenpathlabs.com/lokah-samples/marathi-call-centre-banking-and-insurance--4.segments.json) ### Sample 1 transcript Call-centre conversation, starting at 0:22. Audio sha256 `ed0f74cff2a232b8014d82089278ef6d4a10f90aa6c75727cc3941414ea04dae`. | Start | Speaker | Text | | --- | --- | --- | | 0:00 | Speaker 2 | हा कुठून विश्वास बँक का? | | 0:02 | Speaker 1 | हो सर | | 0:04 | Speaker 2 | बरं बरं बोला सर. | | 0:05 | Speaker 1 | सर तुम्हाला क्रेडिट कार्ड हवं आहे का? | | 0:09 | Speaker 2 | हम्म | | 0:10 | Speaker 1 | क्रेडिट कार्ड हवाय का सर तुम्हाला? | | 0:13 | Speaker 2 | आ हा मी केली होती रिक्वेस्ट बोला काय म्हणताय. | | 0:17 | Speaker 1 | हो सर मी तुम्हाला आमच्याकडे क्रेडीट कार्डचा काही ऑफरसये तुम्हाला सांगतो सर मी | | 0:23 | Speaker 1 | बट तरी तुमचे काही माहिती लागेल सर तुम्हाला विचारू शकतो का? | | 0:27 | Speaker 2 | हा सर ते मिनिट ह्ह | | 0:29 | Speaker 1 | हॅलो सर | | 0:30 | Speaker 2 | [unclear] तुम्ही विश्वास बँकेतून बोलतायत नाव पुन्हा एकदा सांगा [unclear] | | 0:34 | Speaker 1 | साहिल भालेराव सर साहिल भालेराव. | | 0:38 | Speaker 2 | साहिल भालेराव बर | | 0:40 | Speaker 1 | हो सर | ## Licence Licence: custom, quoted per use (training, evaluation or both; internal or commercial; exclusive or not). Delivered in the layout your training code reads: https://kenpathlabs.com/lokah/formats. --- ## FAQ Q: What is Lokah? A: Kenpath Labs' human data platform for AI. It collects speech, images, documents, human feedback and expert annotation. Lokah collects speech, images, documents, feedback and expert annotation from contributors, and delivers reviewed, catalogued datasets. Q: Can I sign up for Lokah? A: Not self-serve. Lokah is sales-led: use the enquiry form at https://kenpathlabs.com/lokah#contact. The console sign-up is for Svara TTS Turbo. Q: What is Svara TTS Turbo? A: Kenpath Labs' flagship text-to-speech model: hundreds of regional voices, 80+ languages in one model, code-switching mid-sentence, consent-gated voice cloning, inline expressive tags, input streaming, and about 80 ms to first audio. Q: Is Svara open source? A: Svara TTS v1 is — Apache 2.0, forever, with 1M+ downloads. Svara TTS Turbo's licensing has not been announced. Q: Are v1 and Turbo the same model? A: One family, two generations. Turbo is a completely new architecture and design, informed by v1's adoption — not v1 scaled up. Q: How fast is it? A: Svara TTS Turbo reaches first audio in about 80 ms and streams its input, so it starts speaking a few words into an LLM's reply. v1 was around 200-300 ms. Q: Can it clone voices? A: Yes — zero-shot from a few seconds of reference audio, consent-gated by design. Clones can speak languages the original speaker never recorded. Q: How do I get access? A: Sign up at https://platform.kenpathlabs.com/signup — self-serve, free tier, no sales call. For volume or procurement, https://kenpathlabs.com/contact-sales. Q: What does it cost? A: One rate on every plan, with no volume discount: ₹1 per 1,000 characters in India, $0.01 per 1,000 characters everywhere else. About 1,000 characters is a minute of audio, so a minute costs roughly ₹1 or $0.01. Every account gets 100,000 free characters a month; Growth is ₹1,000 a month ($10 outside India) for 1,000,000 characters; Enterprise is custom. Q: Do unused characters expire? A: It depends on where they came from. Free characters reset monthly and do not carry forward. Characters included with a paid plan roll over once into the next month, on top of that month's allowance. Top-up characters stay valid for a year. Generation spends the ones closest to expiring first. Q: What happens if I run out of characters? A: Requests fail. There is no silent overage — running out never appears as a surprise on an invoice. Topping up restores service once payment clears. Q: How do I pay? A: Through Razorpay, so any method it supports works — cards, UPI, net banking and wallets among them. International cards are accepted; charges are processed in Indian rupees and converted by the card network at its usual rate. Q: Do I own the audio I generate? A: Yes. As between the customer and Kenpath Labs, the customer owns audio generated from their own text and from reference audio they have consent to use. Kenpath Labs keeps a limited licence to process input and output only to run and support the service. Q: Is there an OpenAI-compatible API? A: Yes, and an ElevenLabs-compatible one. POST /v1/audio/speech follows the OpenAI speech shape; POST /v1/text-to-speech/{voice_id} follows the ElevenLabs shape, with /stream and /with-timestamps variants. Most existing clients work unchanged. Q: What is Kenpath Labs' relationship to Kenpath? A: Kenpath Labs originated from Kenpath and operates as its own frontier AI and data lab in Bengaluru, India. Q: Are the demos on the site real? A: Yes — the clips are model output, not recordings of people. Demo brands are fictional. --- ## Brand URL: https://kenpathlabs.com/brand The name is written "Kenpath Labs" — two words, capital K, capital L, never shortened. The page carries downloadable assets (a full ZIP plus per-file SVG/PNG): the wordmark lockup for light and dark surfaces, the dot-grid mark at three densities (master for 64px+, compact for ~32px, micro for 16px), board-backed icons and an avatar. Brand colours: Ink #08090a, Paper #ffffff, Grey #6b6f76, Marigold #e39a3b (the single accent). Usage: clear space of at least half the mark's width; don't recolour, redraw, rotate or rebuild the lockup in another typeface. --- ## Contact - Sales enquiries: https://kenpathlabs.com/contact-sales - Lokah enquiries: https://kenpathlabs.com/lokah#contact - Sign up: https://platform.kenpathlabs.com/signup - API overview: https://kenpathlabs.com/developers - Docs: https://docs.kenpathlabs.com - Lokah datasets JSON API: https://kenpathlabs.com/api/lokah/datasets (one record at /api/lokah/datasets/{slug}) - Lokah MCP server: https://kenpathlabs.com/api/lokah/mcp (search_datasets, get_dataset, get_sample, list_languages) - Open roles: https://kenpathlabs.com/careers/open-roles - Email: hello@kenpathlabs.com - GitHub: https://github.com/kenpath-labs - Hugging Face: https://huggingface.co/kenpath - LinkedIn: https://www.linkedin.com/company/kenpathlabs/ - YouTube: https://www.youtube.com/@KenpathLabs