Evaluation by the people the model is for.
Listening tests and grading panels run with native speakers. For speech models: naturalness, pronunciation, code-switching and intelligibility, per voice and per language. For language models: correctness, tone and usefulness against your rubric, with the failures written up.
When you need it
- You are choosing between voices, models or vendors for an Indian language and automatic metrics do not agree with what users say.
- You need a scorecard you can put in front of a customer, with the raters, the method and the per-item scores behind it.
- You want the failure cases collected as data, so the next training round has them.
What one record looks like
An example task, as the rater or reviewer sees it. Try it.
Select a dimension to see the individual ratings behind the mean.
- item
- audio or text
- The utterance or response under test, with the system that produced it masked.
- scores
- object
- Each dimension on your scale, per rater.
- issues
- labels
- Mispronunciation, wrong language, robotic prosody, wrong fact, and the rest of the checklist.
- note
- text
- The rater's words on what went wrong, when something did.
- rater_id
- pseudonymous id
- Language, region and agreement in the rater table.
How Lokah builds it and checks it.
- 01
Design with you
Dimensions, scale, item count and the comparison design (absolute, pairwise or MUSHRA-style) agreed before recruitment.
- 02
Panels by language and region
Raters recruited for the language and, when it matters, the region; calibrated on anchor items before the test.
- 03
Blind and balanced
Systems are masked and item order is randomised per rater. Anchor items catch inattentive raters.
- 04
Scorecard and raw data
A scorecard with confidence intervals, plus every rating as data, so you can re-analyse.
How it arrives
- CSV scorecards
- Per system and dimension, with intervals.
- Per-rater JSON
- Every rating, item and note.
- Failure set
- The items that failed, with their labels and rater notes.
Other layouts on request. Catalogue layouts are described on formats.
Pairs with it
From Lokah, and why.
- Hindi speech recognition utterances
Held-out clips for a recogniser test, with transcripts.
- Full-duplex call-centre conversations
Material for judging a speech-to-speech model's turn-taking.
- Tamil general conversation
Natural speech to set the bar for naturalness ratings.
Questions, answered.
Also from Lokah
Start a collection.
Tell us what you want to collect and from whom. We reply with how we would run it.
- Speech, images, documents, feedback or annotation
- A new collection, or a dataset from the catalogue
- Licensing and data handling
Prefer to reach out directly? Write to hello@kenpathlabs.com.