# Delivery formats

> Each dataset is delivered in the layout its kind calls for: conversation sets as 16 kHz FLAC with segment-level transcripts, speech recognition sets as utterance clips in shards, diarization sets with RTTM turn files, two-channel sets as stereo files. Every layout below is a conversion of that source.

## Stereo conversations for Moshi and PersonaPlex

For: Full-duplex speech to speech. Needs a dual-channel dataset.

Kyutai's moshi-finetune, which NVIDIA PersonaPlex builds on, trains from stereo files: the left channel is the voice the model learns to produce, the right channel is the user it listens to. A dual-channel call maps straight onto it: agent left, customer right.

```
data/
  lokah-hi-insurance.jsonl        {"path": "stereo/HN_0001.wav", "duration": 907.2}
  stereo/
    HN_0001.wav                   2 channels: left = agent, right = customer
    HN_0001.json                  {"alignments": [["hello", [1.01, 1.53], "SPEAKER_MAIN"], ...]}
```

```bash
# two mono files per call  ->  one stereo file, agent on the left
ffmpeg -i HN_0001_1.wav -i HN_0001_2.wav \
  -filter_complex "[0:a][1:a]amerge=inputs=2" -ac 2 stereo/HN_0001.wav

# word-level alignments for the text stream (moshi-finetune's own script)
python annotate.py data/lokah-hi-insurance.jsonl --lang hi
```

- Audio is resampled to 24 kHz for the Mimi codec by the loader. The source is 16 kHz, so nothing above 8 kHz is present.
- The text stream needs word-level timestamps. Transcripts here are aligned by segment, so annotate.py, or a forced aligner run against the human transcript, produces the word times.
- PersonaPlex publishes no fine-tuning format of its own. It was trained on Fisher, which is also two-channel telephone speech.

Source: [kyutai-labs/moshi-finetune](https://github.com/kyutai-labs/moshi-finetune)

## Hugging Face datasets, Parquet

For: Any training stack that reads the Hub.

One row per conversation, audio bytes embedded in Parquet shards, with the transcript segments and speaker metadata as columns.

```
data/train-00000-of-00012.parquet
  conversation_id   string
  audio             Audio(sampling_rate=16000, num_channels=2)
  segments          list<struct<start, end, speaker, overlap, text>>
  speaker_a_id      string     speaker_a_gender
  speaker_b_id      string     speaker_b_gender
  language          string     BCP-47, e.g. hi-IN
  domain            string
```

```python
from datasets import load_dataset, Audio

ds = load_dataset("your-org/lokah-hi-insurance", split="train")
# keep both channels: older versions of datasets downmix stereo to mono
ds = ds.cast_column("audio", Audio(sampling_rate=24000, num_channels=2))

row = ds[0]
row["audio"]["array"].shape        # (2, num_samples): agent, customer
```

- num_channels arrived in datasets 4.4. Before that the Audio feature downmixed stereo to mono, which silently destroys a dual-channel dataset. Pin 4.4 or later.
- The card carries task_categories (audio-to-audio for dual channel, automatic-speech-recognition where transcribed), language, and the gating fields.

Source: [Hugging Face: audio datasets](https://huggingface.co/docs/datasets/audio_dataset)

## WebDataset shards

For: Streaming at scale.

Tar shards of about 1 GB. Files that share a prefix are one example, so each conversation is its audio plus one JSON with segments and speakers.

```
train/00000.tar
  HN_0001.wav
  HN_0001.json
  HN_0002.wav
  HN_0002.json
```

```python
from datasets import load_dataset

ds = load_dataset("webdataset", data_dir="train", split="train", streaming=True)
next(iter(ds)).keys()              # dict_keys(['__key__', 'wav', 'json'])
```

- One column per file suffix. A dual-channel conversation can also ship as HN_0001.agent.wav and HN_0001.customer.wav.

Source: [Hugging Face: WebDataset](https://huggingface.co/docs/hub/datasets-webdataset)

## NeMo manifests

For: Speech recognition, and NeMo's duplex speech-to-speech.

For recognition, a JSON-lines manifest of segments. For NeMo's speechlm2 duplex models, Lhotse cuts with the user's audio as the recording, the agent's as the target, and supervisions that carry a user or assistant role.

```
manifest.jsonl
  {"audio_filepath": "seg/HN_0001_0007.wav", "duration": 6.59, "text": "OK actually ma'am ...", "lang": "hi"}
```

```python
# one manifest line per transcript segment
import json
for seg in segments:
    print(json.dumps({"audio_filepath": seg["path"], "duration": seg["end"] - seg["start"],
                      "text": seg["text"], "lang": "hi"}, ensure_ascii=False))
```

- Transcripts keep English words in Latin script inside native-script text. Decide on a normalisation before training a recogniser on them.

Source: [NVIDIA NeMo: ASR datasets](https://github.com/NVIDIA-NeMo/NeMo/blob/main/docs/source/asr/datasets.rst)

## Turn-taking terms used on this site

- **Inter-pausal unit (IPU)**: Continuous speech on one channel, bounded by more than 200 ms of silence on both sides.
- **Pause**: Silence between two IPUs of the same speaker.
- **Gap**: Silence between IPUs of different speakers.
- **Overlap**: Both channels in speech at once.
- **Backchannel**: A short IPU that sits entirely inside an IPU of the other speaker, like a murmured yes.
- **Floor-transfer offset**: At a change of speaker, the next start minus the previous end. Negative is an overlapped transfer, positive a gap.

Published per-minute reference for the Fisher corpus: 21.6 IPUs, 7 pauses, 7.5 gaps, 6.5 overlaps.

- [Nguyen et al. 2022, Generative Spoken Dialogue Language Modeling (dGSLM)](https://arxiv.org/abs/2203.16502)
- [Heldner and Edlund 2010, Pauses, gaps and overlaps in conversations](https://doi.org/10.1016/j.wocn.2010.08.002)
- [Roy et al. 2026, PersonaPlex: Voice and Role Control for Full Duplex Conversational Speech Models](https://arxiv.org/abs/2602.06053)

---

Machine-readable: [llms.txt](https://kenpathlabs.com/llms.txt) · [JSON API](https://kenpathlabs.com/api/lokah/datasets) · [OpenAPI](https://kenpathlabs.com/lokah/openapi.json) · [MCP](https://kenpathlabs.com/api/lokah/mcp) · [for agents](https://kenpathlabs.com/lokah/for-agents)
