Sample bundles
Hear the data before you license it.
Each bundle is 10 to 30 minutes of a dataset, cut the way the full set is delivered. Choose one, and we email you a download link that works for 24 hours.
What is in a bundle
- A README with the dataset card excerpt for that dataset.
- The audio, in the same layout and format as the full dataset.
- Transcripts and metadata in the same files and fields you would receive with the full set.
- LICENSE-SAMPLE.txt: the bundle is for evaluation only.
Whole conversations
Complete recordings with speaker-labelled transcripts.
Full-duplex speech to speech
Each speaker on a separate channel, timing preserved.
Speech recognition
Audio with time-aligned transcripts in native script.
Diarisation
Single-channel recordings with speaker turns marked.