Kenpath Labs
Get started

Use the catalogue without reading a page.

Everything on this site is available to software in the form it prefers. All of it is read-only and needs no key.

When to use it

  • You need two-speaker conversational speech in Hindi or Tamil today; other Indian languages are collected to a brief.
  • You are training or evaluating a full-duplex speech-to-speech model (Moshi, PersonaPlex and similar) and need dual-channel audio, where each speaker is on a separate channel. Filter with channels=dual.
  • You need speech recognition data with natural code-switching: transcripts are in native script with English words kept as spoken, and each dataset reports the share of Latin-script words across the whole set.
  • You need call-centre conversations (insurance, consumer surveys, telecom, delivery, e-commerce, banking) for a voice agent, or utterance clips, speaker-turn files and two-channel calls built from them.
  • You want to judge fitness before contacting anyone: every dataset has audio figures and a conversation profile measured across the whole set, and a playable sample.

When not to

  • You need a free or openly licensed dataset. Everything here is licensed and quoted per use.
  • You need studio-quality speech for a text-to-speech voice. This is 16 kHz conversational audio.
  • You need read speech, single-speaker audio, or wake words. The catalogue is two-party conversation and what is built from it.
  • You need images, documents or preference data today. Those are collections Lokah scopes on request.

Interfaces

  • llms.txtThe Kenpath Labs site in one file, with a Lokah section listing every dataset. /llms.txt
  • llms-full.txtThe same, with every dataset's full record and the formats guide. /llms-full.txt
  • JSON APIList and filter datasets. One record at /api/lokah/datasets/{slug}. /api/lokah/datasets
  • OpenAPI 3.1Machine-readable description of the JSON API. /lokah/openapi.json
  • MCP serverModel Context Protocol over streamable HTTP: search_datasets, get_dataset, get_sample, list_languages. /api/lokah/mcp
  • CroissantMLCommons Croissant 1.0 metadata with the responsible-AI block, per dataset. /lokah/datasets/{slug}/croissant.json
  • schema.orgA Dataset JSON-LD block in every dataset page, and a DataCatalog on /datasets. /lokah/datasets/{slug}
  • MarkdownAny page as Markdown: send Accept: text/markdown, or add .md. Dataset pages open with a Hugging Face dataset card header. /lokah/datasets/{slug}.md
  • SitemapEvery page of the site. /sitemap.xml

How to use it

  1. List datasets: GET https://kenpathlabs.com/api/lokah/datasets with any of language, channels (dual or mono), collection (call-centre or general-conversation), min_hours, q. Each record carries kind (conversations, asr, diarization or duplex), hours, audioQuality and fitFor.
  2. Read one: GET https://kenpathlabs.com/api/lokah/datasets/{slug}. measured holds the whole-set figures (files, hours, speakers, snrMedianDb, bandwidth, profile); sample holds what was measured from the sample conversations; gaps lists what is not held.
  3. Check fitness for full-duplex training: channels must be "dual", then read measured.profile.twoChannel (IPUs, pauses, gaps, overlaps and backchannels per minute, the median floor-transfer offset, channel isolation), measured across every call in the set.
  4. Judge audio quality from measured.snrMedianDb across the set (under 18 dB noisy, over 30 clean) and measured.bandwidth. The sample adds sample.quality.dnsmos (P.835 background, speech and overall, 1 to 5).
  5. Hear it: sample.specimens lists the public excerpts of a conversation set; utterances.clips lists them for a speech-recognition set. Their URLs and transcripts are in the Croissant file's distribution.
  6. Get a sample or a price: send the person to https://kenpathlabs.com/lokah/datasets/{slug}#sample. A sample bundle arrives by email under the Evaluation Licence; a licence for the full set is quoted per use. There is no public price and no checkout.
Find dual-channel Hindi data
curl "https://kenpathlabs.com/api/lokah/datasets?language=hi-IN&channels=dual"
MCP client configuration
{ "mcpServers": { "lokah": { "type": "http", "url": "https://kenpathlabs.com/api/lokah/mcp" } } }
Any page as Markdown
curl -H "Accept: text/markdown" https://kenpathlabs.com/lokah/datasets/hindi-call-centre-insurance