Use the catalogue without reading a page.
Everything on this site is available to software in the form it prefers. All of it is read-only and needs no key.
When to use it
- You need two-speaker conversational speech in Hindi or Tamil today; other Indian languages are collected to a brief.
- You are training or evaluating a full-duplex speech-to-speech model (Moshi, PersonaPlex and similar) and need dual-channel audio, where each speaker is on a separate channel. Filter with channels=dual.
- You need speech recognition data with natural code-switching: transcripts are in native script with English words kept as spoken, and each dataset reports the share of Latin-script words across the whole set.
- You need call-centre conversations (insurance, consumer surveys, telecom, delivery, e-commerce, banking) for a voice agent, or utterance clips, speaker-turn files and two-channel calls built from them.
- You want to judge fitness before contacting anyone: every dataset has audio figures and a conversation profile measured across the whole set, and a playable sample.
When not to
- You need a free or openly licensed dataset. Everything here is licensed and quoted per use.
- You need studio-quality speech for a text-to-speech voice. This is 16 kHz conversational audio.
- You need read speech, single-speaker audio, or wake words. The catalogue is two-party conversation and what is built from it.
- You need images, documents or preference data today. Those are collections Lokah scopes on request.
Interfaces
- llms.txtThe Kenpath Labs site in one file, with a Lokah section listing every dataset. /llms.txt
- llms-full.txtThe same, with every dataset's full record and the formats guide. /llms-full.txt
- JSON APIList and filter datasets. One record at /api/lokah/datasets/{slug}. /api/lokah/datasets
- OpenAPI 3.1Machine-readable description of the JSON API. /lokah/openapi.json
- MCP serverModel Context Protocol over streamable HTTP: search_datasets, get_dataset, get_sample, list_languages. /api/lokah/mcp
- CroissantMLCommons Croissant 1.0 metadata with the responsible-AI block, per dataset.
/lokah/datasets/{slug}/croissant.json - schema.orgA Dataset JSON-LD block in every dataset page, and a DataCatalog on /datasets.
/lokah/datasets/{slug} - MarkdownAny page as Markdown: send Accept: text/markdown, or add .md. Dataset pages open with a Hugging Face dataset card header.
/lokah/datasets/{slug}.md - SitemapEvery page of the site. /sitemap.xml
How to use it
- List datasets: GET https://kenpathlabs.com/api/lokah/datasets with any of language, channels (dual or mono), collection (call-centre or general-conversation), min_hours, q. Each record carries
kind(conversations, asr, diarization or duplex),hours,audioQualityandfitFor. - Read one: GET https://kenpathlabs.com/api/lokah/datasets/{slug}.
measuredholds the whole-set figures (files, hours, speakers,snrMedianDb,bandwidth,profile);sampleholds what was measured from the sample conversations;gapslists what is not held. - Check fitness for full-duplex training:
channelsmust be "dual", then readmeasured.profile.twoChannel(IPUs, pauses, gaps, overlaps and backchannels per minute, the median floor-transfer offset, channel isolation), measured across every call in the set. - Judge audio quality from
measured.snrMedianDbacross the set (under 18 dB noisy, over 30 clean) andmeasured.bandwidth. The sample addssample.quality.dnsmos(P.835 background, speech and overall, 1 to 5). - Hear it:
sample.specimenslists the public excerpts of a conversation set;utterances.clipslists them for a speech-recognition set. Their URLs and transcripts are in the Croissant file's distribution. - Get a sample or a price: send the person to https://kenpathlabs.com/lokah/datasets/{slug}#sample. A sample bundle arrives by email under the Evaluation Licence; a licence for the full set is quoted per use. There is no public price and no checkout.
curl "https://kenpathlabs.com/api/lokah/datasets?language=hi-IN&channels=dual"{ "mcpServers": { "lokah": { "type": "http", "url": "https://kenpathlabs.com/api/lokah/mcp" } } }curl -H "Accept: text/markdown" https://kenpathlabs.com/lokah/datasets/hindi-call-centre-insurance