Kenpath Labs
Get started

In the layout your training code reads.

Each dataset is delivered in the layout its kind calls for: conversation sets as 16 kHz FLAC with segment-level transcripts, speech recognition sets as utterance clips in shards, diarization sets with RTTM turn files, two-channel sets as stereo. These are the layouts it converts to, and what each conversion involves. Name the one you want when you ask for a quote.

Full-duplex speech to speech

Stereo conversations for Moshi and PersonaPlex

Kyutai's moshi-finetune, which NVIDIA PersonaPlex builds on, trains from stereo files: the left channel is the voice the model learns to produce, the right channel is the user it listens to. A dual-channel call maps straight onto it: agent left, customer right.

  • Audio is resampled to 24 kHz for the Mimi codec by the loader. The source is 16 kHz, so nothing above 8 kHz is present.
  • The text stream needs word-level timestamps. Transcripts here are aligned by segment, so annotate.py, or a forced aligner run against the human transcript, produces the word times.
  • PersonaPlex publishes no fine-tuning format of its own. It was trained on Fisher, which is also two-channel telephone speech.

kyutai-labs/moshi-finetune · Dual-channel datasets

layout
data/
  lokah-hi-insurance.jsonl        {"path": "stereo/HN_0001.wav", "duration": 907.2}
  stereo/
    HN_0001.wav                   2 channels: left = agent, right = customer
    HN_0001.json                  {"alignments": [["hello", [1.01, 1.53], "SPEAKER_MAIN"], ...]}
bash
# two mono files per call  ->  one stereo file, agent on the left
ffmpeg -i HN_0001_1.wav -i HN_0001_2.wav \
  -filter_complex "[0:a][1:a]amerge=inputs=2" -ac 2 stereo/HN_0001.wav

# word-level alignments for the text stream (moshi-finetune's own script)
python annotate.py data/lokah-hi-insurance.jsonl --lang hi

Any training stack that reads the Hub

Hugging Face datasets, Parquet

One row per conversation, audio bytes embedded in Parquet shards, with the transcript segments and speaker metadata as columns.

  • num_channels arrived in datasets 4.4. Before that the Audio feature downmixed stereo to mono, which silently destroys a dual-channel dataset. Pin 4.4 or later.
  • The card carries task_categories (audio-to-audio for dual channel, automatic-speech-recognition where transcribed), language, and the gating fields.

Hugging Face: audio datasets

layout
data/train-00000-of-00012.parquet
  conversation_id   string
  audio             Audio(sampling_rate=16000, num_channels=2)
  segments          list<struct<start, end, speaker, overlap, text>>
  speaker_a_id      string     speaker_a_gender
  speaker_b_id      string     speaker_b_gender
  language          string     BCP-47, e.g. hi-IN
  domain            string
python
from datasets import load_dataset, Audio

ds = load_dataset("your-org/lokah-hi-insurance", split="train")
# keep both channels: older versions of datasets downmix stereo to mono
ds = ds.cast_column("audio", Audio(sampling_rate=24000, num_channels=2))

row = ds[0]
row["audio"]["array"].shape        # (2, num_samples): agent, customer

Streaming at scale

WebDataset shards

Tar shards of about 1 GB. Files that share a prefix are one example, so each conversation is its audio plus one JSON with segments and speakers.

  • One column per file suffix. A dual-channel conversation can also ship as HN_0001.agent.wav and HN_0001.customer.wav.

Hugging Face: WebDataset

layout
train/00000.tar
  HN_0001.wav
  HN_0001.json
  HN_0002.wav
  HN_0002.json
python
from datasets import load_dataset

ds = load_dataset("webdataset", data_dir="train", split="train", streaming=True)
next(iter(ds)).keys()              # dict_keys(['__key__', 'wav', 'json'])

Speech recognition, and NeMo's duplex speech-to-speech

NeMo manifests

For recognition, a JSON-lines manifest of segments. For NeMo's speechlm2 duplex models, Lhotse cuts with the user's audio as the recording, the agent's as the target, and supervisions that carry a user or assistant role.

  • Transcripts keep English words in Latin script inside native-script text. Decide on a normalisation before training a recogniser on them.

NVIDIA NeMo: ASR datasets

layout
manifest.jsonl
  {"audio_filepath": "seg/HN_0001_0007.wav", "duration": 6.59, "text": "OK actually ma'am ...", "lang": "hi"}
python
# one manifest line per transcript segment
import json
for seg in segments:
    print(json.dumps({"audio_filepath": seg["path"], "duration": seg["end"] - seg["start"],
                      "text": seg["text"], "lang": "hi"}, ensure_ascii=False))

Turn-taking, measured.

Dual-channel datasets report these per minute, measured from the audio, so you can compare them with published corpora. For the Fisher corpus that is 21.6 units, 7 pauses, 7.5 gaps and 6.5 overlaps.

Inter-pausal unit (IPU)
Continuous speech on one channel, bounded by more than 200 ms of silence on both sides.
Pause
Silence between two IPUs of the same speaker.
Gap
Silence between IPUs of different speakers.
Overlap
Both channels in speech at once.
Backchannel
A short IPU that sits entirely inside an IPU of the other speaker, like a murmured yes.
Floor-transfer offset
At a change of speaker, the next start minus the previous end. Negative is an overlapped transfer, positive a gap.