Kenpath Labs
Get started

Preference data in the languages your users write and speak.

Human judgements of model output: which of two answers is better, how a response scores against a rubric, and what a good answer would have said. Raters are native speakers who use the product's language every day, so the preference reflects the user, not a translator.

When you need it

  • You are aligning a model for Hindi, Tamil or another Indian language and English-only preference data leaves it stiff, formal or wrong on dates, names and numbers.
  • You want a reward model to learn from people who notice code-switching, honorifics and register, not from people reading a translation.
  • You need rewritten responses, not only a choice, so the model sees what better looks like.

What one record looks like

An example task, as the rater or reviewer sees it. Try it.

Pairwise comparison 3 raters

मेरी पॉलिसी कब रिन्यू होगी?

Pick the better answer.
prompt
text
The user turn, in the language and script it was written in.
responses
list
Two or more model outputs, shuffled before rating.
choice
index or ranking
The rater's pick, or a full ranking.
rubric_scores
object
Helpfulness, correctness, tone and safety, each on the scale your rubric sets.
rewrite
text
The rater's improved response, when the brief asks for one.
reason
text
One sentence on why, in the rater's words.
rater_id
pseudonymous id
Stable across the collection; language and agreement rate in the rater table.

How Lokah builds it and checks it.

  1. 01

    Brief and rubric

    You send prompts or we write them to your domain. We turn your rubric into rater instructions with worked examples, and test it on a pilot of a few hundred items before anything scales.

  2. 02

    Raters who speak the language

    Native speakers recruited from the contributor network, screened on a calibration set. Every item is rated by at least three people; check items with a known answer are mixed in throughout.

  3. 03

    Agreement measured, not assumed

    Agreement is tracked per rater. Items below the agreement floor are rated again; raters who drift are retrained or replaced. The agreement figures ship with the data.

  4. 04

    Delivery

    Pairs or rankings as JSONL with the rater table, or loaded into Argilla for your own review. Personal data in prompts is masked before rating begins.

How it arrives

JSONL pairs
One line per comparison, with prompt, responses, choice, scores and reason.
Rankings
Full orderings when the brief compares more than two responses.
Argilla
A workspace you can open, filter and re-rate.
Rater table
Pseudonymous ids, language, agreement rate.

Other layouts on request. Catalogue layouts are described on formats.

Pairs with it

From Lokah, and why.

Questions, answered.

Also from Lokah

Talk to us

Start a collection.

Tell us what you want to collect and from whom. We reply with how we would run it.

  • Speech, images, documents, feedback or annotation
  • A new collection, or a dataset from the catalogue
  • Licensing and data handling

Prefer to reach out directly? Write to hello@kenpathlabs.com.