# Lokah: Human data for training and evaluating AI.

Lokah is the data platform of Kenpath Labs. Two offers: off-the-shelf datasets, ready to license, and data collection to a brief.

> Lokah is the data platform of Kenpath Labs, a frontier AI and data company in Bengaluru. Its catalogue holds licensable datasets for training and evaluating AI, conversational speech in Indian languages and English first, each with a playable sample, a specification measured from the files and a conversation profile.

12 datasets · 680 hours · 3 languages. Licensed, quoted per use. Contact: hello@kenpathlabs.com

## Choose by what the model needs

- **Full-duplex speech to speech**: each speaker on a separate channel. [Dual-channel datasets](https://kenpathlabs.com/lokah/datasets?channels=dual)
- **Speech recognition**: time-aligned transcripts in native script, English kept as spoken. [All transcribed datasets](https://kenpathlabs.com/lokah/datasets)
- **Voice agents for support**: [call-centre datasets](https://kenpathlabs.com/lokah/datasets?collection=call-centre)

## The catalogue

| Dataset | Language | Hours | Channels | English, Latin script | Sample |
| --- | --- | --- | --- | --- | --- |
| [Hindi call-centre conversations, insurance](https://kenpathlabs.com/lokah/datasets/hindi-call-centre-insurance) | हिन्दी Hindi | 356 | Two channels | 17% | 4 |
| [Hindi speech recognition utterances](https://kenpathlabs.com/lokah/datasets/hindi-speech-recognition-utterances) | हिन्दी Hindi | 350 | One channel | not measured | on request |
| [Hindi speaker diarization](https://kenpathlabs.com/lokah/datasets/hindi-speaker-diarization) | हिन्दी Hindi | 321 | One channel | 9% | 4 |
| [Hindi and Tamil full-duplex call-centre conversations](https://kenpathlabs.com/lokah/datasets/hindi-tamil-full-duplex-call-centre-conversations) | हिन्दी Hindi | 264 | Two channels | 16% | 4 |
| [Tamil speaker diarization](https://kenpathlabs.com/lokah/datasets/tamil-speaker-diarization) | தமிழ் Tamil | 229 | One channel | 39% | 4 |
| [Tamil speech recognition utterances](https://kenpathlabs.com/lokah/datasets/tamil-speech-recognition-utterances) | தமிழ் Tamil | 201 | One channel | not measured | on request |
| [Tamil general conversation](https://kenpathlabs.com/lokah/datasets/tamil-general-conversation) | தமிழ் Tamil | 105 | One channel | 36% | 4 |
| [Tamil call-centre conversations, consumer surveys](https://kenpathlabs.com/lokah/datasets/tamil-call-centre-customer-service) | தமிழ் Tamil | 99 | One channel | 39% | 4 |
| [Hindi general conversation](https://kenpathlabs.com/lokah/datasets/hindi-general-conversation) | हिन्दी Hindi | 81 | One channel | 14% | 4 |
| [Tamil two-channel customer-service calls](https://kenpathlabs.com/lokah/datasets/tamil-call-centre-banking-and-retail) | தமிழ் Tamil | 25 | Two channels | 29% | 4 |
| [Marathi general conversation](https://kenpathlabs.com/lokah/datasets/marathi-general-conversation) | मराठी Marathi | 9 | One channel | 0% | 4 |
| [Marathi call-centre conversations, banking and insurance](https://kenpathlabs.com/lokah/datasets/marathi-call-centre-banking-and-insurance) | मराठी Marathi | 4 | One channel | 0% | 4 |

## What comes with every dataset

- A sample you can play: 45 seconds of a real conversation, with its transcript.
- A measured specification: sample rate, bit depth and duration read from the files.
- A conversation profile: talk time, silence, overlap, speaking rate, share of English.
- Every recording has a complete, segment-level transcript.

## Three ways to get the data.

1. **Ready to license.** Catalogue datasets with a playable sample and a measured spec sheet. Ask for the licence and the full set follows.
2. **Ready to train on.** Languages and domains Lokah can collect on request. You set the specification and a pilot batch decides whether to continue.
3. **Expand or customise.** Take any catalogue dataset further: more hours, more speakers, a new domain, or annotation added to what is already there.

## Questions

**Can I hear the data before I ask for a quote?**
Yes, for every dataset marked with a sample. Each sample is 45 seconds of a real conversation from that dataset, with its transcript and speaker labels. A longer sample is sent on request.

**What is the difference between dual channel and mono, diarised?**
Dual channel means each speaker was recorded to a separate file, so overlapping speech, backchannels and the timing of every change of speaker are preserved. Mono, diarised means one mixed recording with speaker labels on the transcript: turns are attributed, but overlapping voices cannot be pulled apart. Full-duplex speech-to-speech models are trained on the first kind.

**Where do the measured figures come from?**
From the files. Sample rate, bit depth and duration are read from the audio itself. Talk time, silence, overlap, speaking rate and the share of words written in Latin script are computed from a full sample conversation and its transcript. Where a listed figure differs from what the audio measures, both are shown and the measured one is what you receive.

**How is personal data handled?**
Datasets are delivered with personal data redacted. Samples on this site are additionally screened: any segment in which digits or account details are read out is excluded from the excerpt, and a sample that cannot give a clean excerpt is not published.

**What licence do the datasets come with?**
A custom licence, quoted per use. Tell us whether the data is for training, evaluation or both, whether the model is internal or commercial, and whether you need exclusivity.

**What formats can the data be delivered in?**
It depends on the dataset. Conversation sets arrive as 16 kHz FLAC with segment-level transcripts; speech recognition sets as utterance clips in WebDataset shards with a manifest; diarization sets with RTTM turn files; two-channel sets as stereo files for full-duplex training. Each can also be laid out for Hugging Face datasets or NeMo. Each dataset page names its layout; the detail is on /lokah/formats.

**What if the catalogue does not have what I need?**
Then it becomes a collection. Send the language, domain and volume, and we reply with how we would run it.

## Built to a brief: beyond datasets.

The same contributor network and review process, applied to preference, safety, fine-tuning, evaluation, agent and annotation work. Tell us which of these you need and we will scope it with you.

- **Preference and alignment data.** Pairwise and ranked human judgements of model output, rubric grades and rewritten responses, in the languages your users write and speak.
- **Red teaming and safety data.** Adversarial prompts, multilingual jailbreak attempts and harm labels from native speakers, so safety holds outside English.
- **Supervised fine-tuning data.** Expert-written demonstrations and verified reasoning traces, reviewed by people who know the subject.
- **Speech and language evaluation.** Native-speaker panels that rate speech models and grade language models: naturalness, pronunciation, code-switching, correctness, and where it goes wrong.
- **Agent data and environments.** Voice-agent conversations and tool-use trajectories, and simulated callers built on Svara with a persona, an accent and a noisy line.
- **Annotation on your data.** Transcription, diarisation, text and document labelling and image annotation, delivered in the format your tooling reads.

## Next

- [Browse the datasets](https://kenpathlabs.com/lokah/datasets)
- [Request a sample bundle](https://kenpathlabs.com/lokah/samples): 10 to 20 minutes of a dataset, emailed as a 24-hour download link
- [Delivery formats](https://kenpathlabs.com/lokah/formats)
- [Get a quote or scope a collection](https://kenpathlabs.com/lokah#contact)

---

Machine-readable: [llms.txt](https://kenpathlabs.com/llms.txt) · [JSON API](https://kenpathlabs.com/api/lokah/datasets) · [OpenAPI](https://kenpathlabs.com/lokah/openapi.json) · [MCP](https://kenpathlabs.com/api/lokah/mcp) · [for agents](https://kenpathlabs.com/lokah/for-agents)
