Kenpath Labs
Get started

Evaluation by the people the model is for.

Listening tests and grading panels run with native speakers. For speech models: naturalness, pronunciation, code-switching and intelligibility, per voice and per language. For language models: correctness, tone and usefulness against your rubric, with the failures written up.

When you need it

  • You are choosing between voices, models or vendors for an Indian language and automatic metrics do not agree with what users say.
  • You need a scorecard you can put in front of a customer, with the raters, the method and the per-item scores behind it.
  • You want the failure cases collected as data, so the next training round has them.

What one record looks like

An example task, as the rater or reviewer sees it. Try it.

Listening testExample scorecard, Kannada, one voice
Naturalness: five raters gave3.544.54.55

Select a dimension to see the individual ratings behind the mean.

item
audio or text
The utterance or response under test, with the system that produced it masked.
scores
object
Each dimension on your scale, per rater.
issues
labels
Mispronunciation, wrong language, robotic prosody, wrong fact, and the rest of the checklist.
note
text
The rater's words on what went wrong, when something did.
rater_id
pseudonymous id
Language, region and agreement in the rater table.

How Lokah builds it and checks it.

  1. 01

    Design with you

    Dimensions, scale, item count and the comparison design (absolute, pairwise or MUSHRA-style) agreed before recruitment.

  2. 02

    Panels by language and region

    Raters recruited for the language and, when it matters, the region; calibrated on anchor items before the test.

  3. 03

    Blind and balanced

    Systems are masked and item order is randomised per rater. Anchor items catch inattentive raters.

  4. 04

    Scorecard and raw data

    A scorecard with confidence intervals, plus every rating as data, so you can re-analyse.

How it arrives

CSV scorecards
Per system and dimension, with intervals.
Per-rater JSON
Every rating, item and note.
Failure set
The items that failed, with their labels and rater notes.

Other layouts on request. Catalogue layouts are described on formats.

Pairs with it

From Lokah, and why.

Questions, answered.

Also from Lokah

Talk to us

Start a collection.

Tell us what you want to collect and from whom. We reply with how we would run it.

  • Speech, images, documents, feedback or annotation
  • A new collection, or a dataset from the catalogue
  • Licensing and data handling

Prefer to reach out directly? Write to hello@kenpathlabs.com.