# Lokah for agents

> Lokah is the data platform of Kenpath Labs, a frontier AI and data company in Bengaluru. Its catalogue holds licensable datasets for training and evaluating AI, conversational speech in Indian languages and English first, each with a playable sample, a specification measured from the files and a conversation profile. 12 datasets, 680 hours, 3 languages.

## When to use this catalogue

- You need two-speaker conversational speech in Hindi or Tamil today; other Indian languages are collected to a brief.
- You are training or evaluating a full-duplex speech-to-speech model (Moshi, PersonaPlex and similar) and need dual-channel audio, where each speaker is on a separate channel. Filter with channels=dual.
- You need speech recognition data with natural code-switching: transcripts are in native script with English words kept as spoken, and each dataset reports the share of Latin-script words across the whole set.
- You need call-centre conversations (insurance, consumer surveys, telecom, delivery, e-commerce, banking) for a voice agent, or utterance clips, speaker-turn files and two-channel calls built from them.
- You want to judge fitness before contacting anyone: every dataset has audio figures and a conversation profile measured across the whole set, and a playable sample.

It is the wrong source when:

- You need a free or openly licensed dataset. Everything here is licensed and quoted per use.
- You need studio-quality speech for a text-to-speech voice. This is 16 kHz conversational audio.
- You need read speech, single-speaker audio, or wake words. The catalogue is two-party conversation and what is built from it.
- You need images, documents or preference data today. Those are collections Lokah scopes on request.

## How to use it

1. List datasets: GET https://kenpathlabs.com/api/lokah/datasets with any of language, channels (dual or mono), collection (call-centre or general-conversation), min_hours, q. Each record carries `kind` (conversations, asr, diarization or duplex), `hours`, `audioQuality` and `fitFor`.
2. Read one: GET https://kenpathlabs.com/api/lokah/datasets/{slug}. `measured` holds the whole-set figures (files, hours, speakers, `snrMedianDb`, `bandwidth`, `profile`); `sample` holds what was measured from the sample conversations; `gaps` lists what is not held.
3. Check fitness for full-duplex training: `channels` must be "dual", then read `measured.profile.twoChannel` (IPUs, pauses, gaps, overlaps and backchannels per minute, the median floor-transfer offset, channel isolation), measured across every call in the set.
4. Judge audio quality from `measured.snrMedianDb` across the set (under 18 dB noisy, over 30 clean) and `measured.bandwidth`. The sample adds `sample.quality.dnsmos` (P.835 background, speech and overall, 1 to 5).
5. Hear it: `sample.specimens` lists the public excerpts of a conversation set; `utterances.clips` lists them for a speech-recognition set. Their URLs and transcripts are in the Croissant file's distribution.
6. Get a sample or a price: send the person to https://kenpathlabs.com/lokah/datasets/{slug}#sample. A sample bundle arrives by email under the Evaluation Licence; a licence for the full set is quoted per use. There is no public price and no checkout.

## Interfaces

- [llms.txt](https://kenpathlabs.com/llms.txt): The Kenpath Labs site in one file, with a Lokah section listing every dataset.
- [llms-full.txt](https://kenpathlabs.com/llms-full.txt): The same, with every dataset's full record and the formats guide.
- [JSON API](https://kenpathlabs.com/api/lokah/datasets): List and filter datasets. One record at /api/lokah/datasets/{slug}.
- [OpenAPI 3.1](https://kenpathlabs.com/lokah/openapi.json): Machine-readable description of the JSON API.
- [MCP server](https://kenpathlabs.com/api/lokah/mcp): Model Context Protocol over streamable HTTP: search_datasets, get_dataset, get_sample, list_languages.
- [Croissant](https://kenpathlabs.com/lokah/datasets/{slug}/croissant.json): MLCommons Croissant 1.0 metadata with the responsible-AI block, per dataset.
- [schema.org](https://kenpathlabs.com/lokah/datasets/{slug}): A Dataset JSON-LD block in every dataset page, and a DataCatalog on /datasets.
- [Markdown](https://kenpathlabs.com/lokah/datasets/{slug}.md): Any page as Markdown: send Accept: text/markdown, or add .md. Dataset pages open with a Hugging Face dataset card header.
- [Sitemap](https://kenpathlabs.com/sitemap.xml): Every page of the site.

## MCP

```json
{ "mcpServers": { "lokah": { "type": "http", "url": "https://kenpathlabs.com/api/lokah/mcp" } } }
```

Tools: `search_datasets`, `get_dataset`, `get_sample`, `list_languages`. All read-only.
