In the layout your training code reads.
Each dataset is delivered in the layout its kind calls for: conversation sets as 16 kHz FLAC with segment-level transcripts, speech recognition sets as utterance clips in shards, diarization sets with RTTM turn files, two-channel sets as stereo. These are the layouts it converts to, and what each conversion involves. Name the one you want when you ask for a quote.
Full-duplex speech to speech
Stereo conversations for Moshi and PersonaPlex
Kyutai's moshi-finetune, which NVIDIA PersonaPlex builds on, trains from stereo files: the left channel is the voice the model learns to produce, the right channel is the user it listens to. A dual-channel call maps straight onto it: agent left, customer right.
- Audio is resampled to 24 kHz for the Mimi codec by the loader. The source is 16 kHz, so nothing above 8 kHz is present.
- The text stream needs word-level timestamps. Transcripts here are aligned by segment, so annotate.py, or a forced aligner run against the human transcript, produces the word times.
- PersonaPlex publishes no fine-tuning format of its own. It was trained on Fisher, which is also two-channel telephone speech.
data/
lokah-hi-insurance.jsonl {"path": "stereo/HN_0001.wav", "duration": 907.2}
stereo/
HN_0001.wav 2 channels: left = agent, right = customer
HN_0001.json {"alignments": [["hello", [1.01, 1.53], "SPEAKER_MAIN"], ...]}# two mono files per call -> one stereo file, agent on the left
ffmpeg -i HN_0001_1.wav -i HN_0001_2.wav \
-filter_complex "[0:a][1:a]amerge=inputs=2" -ac 2 stereo/HN_0001.wav
# word-level alignments for the text stream (moshi-finetune's own script)
python annotate.py data/lokah-hi-insurance.jsonl --lang hiAny training stack that reads the Hub
Hugging Face datasets, Parquet
One row per conversation, audio bytes embedded in Parquet shards, with the transcript segments and speaker metadata as columns.
- num_channels arrived in datasets 4.4. Before that the Audio feature downmixed stereo to mono, which silently destroys a dual-channel dataset. Pin 4.4 or later.
- The card carries task_categories (audio-to-audio for dual channel, automatic-speech-recognition where transcribed), language, and the gating fields.
data/train-00000-of-00012.parquet
conversation_id string
audio Audio(sampling_rate=16000, num_channels=2)
segments list<struct<start, end, speaker, overlap, text>>
speaker_a_id string speaker_a_gender
speaker_b_id string speaker_b_gender
language string BCP-47, e.g. hi-IN
domain stringfrom datasets import load_dataset, Audio
ds = load_dataset("your-org/lokah-hi-insurance", split="train")
# keep both channels: older versions of datasets downmix stereo to mono
ds = ds.cast_column("audio", Audio(sampling_rate=24000, num_channels=2))
row = ds[0]
row["audio"]["array"].shape # (2, num_samples): agent, customerStreaming at scale
WebDataset shards
Tar shards of about 1 GB. Files that share a prefix are one example, so each conversation is its audio plus one JSON with segments and speakers.
- One column per file suffix. A dual-channel conversation can also ship as HN_0001.agent.wav and HN_0001.customer.wav.
train/00000.tar
HN_0001.wav
HN_0001.json
HN_0002.wav
HN_0002.jsonfrom datasets import load_dataset
ds = load_dataset("webdataset", data_dir="train", split="train", streaming=True)
next(iter(ds)).keys() # dict_keys(['__key__', 'wav', 'json'])Speech recognition, and NeMo's duplex speech-to-speech
NeMo manifests
For recognition, a JSON-lines manifest of segments. For NeMo's speechlm2 duplex models, Lhotse cuts with the user's audio as the recording, the agent's as the target, and supervisions that carry a user or assistant role.
- Transcripts keep English words in Latin script inside native-script text. Decide on a normalisation before training a recogniser on them.
manifest.jsonl
{"audio_filepath": "seg/HN_0001_0007.wav", "duration": 6.59, "text": "OK actually ma'am ...", "lang": "hi"}# one manifest line per transcript segment
import json
for seg in segments:
print(json.dumps({"audio_filepath": seg["path"], "duration": seg["end"] - seg["start"],
"text": seg["text"], "lang": "hi"}, ensure_ascii=False))Turn-taking, measured.
Dual-channel datasets report these per minute, measured from the audio, so you can compare them with published corpora. For the Fisher corpus that is 21.6 units, 7 pauses, 7.5 gaps and 6.5 overlaps.
- Inter-pausal unit (IPU)
- Continuous speech on one channel, bounded by more than 200 ms of silence on both sides.
- Pause
- Silence between two IPUs of the same speaker.
- Gap
- Silence between IPUs of different speakers.
- Overlap
- Both channels in speech at once.
- Backchannel
- A short IPU that sits entirely inside an IPU of the other speaker, like a murmured yes.
- Floor-transfer offset
- At a change of speaker, the next start minus the previous end. Negative is an overlapped transfer, positive a gap.