Preference data in the languages your users write and speak.
Human judgements of model output: which of two answers is better, how a response scores against a rubric, and what a good answer would have said. Raters are native speakers who use the product's language every day, so the preference reflects the user, not a translator.
When you need it
- You are aligning a model for Hindi, Tamil or another Indian language and English-only preference data leaves it stiff, formal or wrong on dates, names and numbers.
- You want a reward model to learn from people who notice code-switching, honorifics and register, not from people reading a translation.
- You need rewritten responses, not only a choice, so the model sees what better looks like.
What one record looks like
An example task, as the rater or reviewer sees it. Try it.
मेरी पॉलिसी कब रिन्यू होगी?
- prompt
- text
- The user turn, in the language and script it was written in.
- responses
- list
- Two or more model outputs, shuffled before rating.
- choice
- index or ranking
- The rater's pick, or a full ranking.
- rubric_scores
- object
- Helpfulness, correctness, tone and safety, each on the scale your rubric sets.
- rewrite
- text
- The rater's improved response, when the brief asks for one.
- reason
- text
- One sentence on why, in the rater's words.
- rater_id
- pseudonymous id
- Stable across the collection; language and agreement rate in the rater table.
How Lokah builds it and checks it.
- 01
Brief and rubric
You send prompts or we write them to your domain. We turn your rubric into rater instructions with worked examples, and test it on a pilot of a few hundred items before anything scales.
- 02
Raters who speak the language
Native speakers recruited from the contributor network, screened on a calibration set. Every item is rated by at least three people; check items with a known answer are mixed in throughout.
- 03
Agreement measured, not assumed
Agreement is tracked per rater. Items below the agreement floor are rated again; raters who drift are retrained or replaced. The agreement figures ship with the data.
- 04
Delivery
Pairs or rankings as JSONL with the rater table, or loaded into Argilla for your own review. Personal data in prompts is masked before rating begins.
How it arrives
- JSONL pairs
- One line per comparison, with prompt, responses, choice, scores and reason.
- Rankings
- Full orderings when the brief compares more than two responses.
- Argilla
- A workspace you can open, filter and re-rate.
- Rater table
- Pseudonymous ids, language, agreement rate.
Other layouts on request. Catalogue layouts are described on formats.
Pairs with it
From Lokah, and why.
- Call-centre conversations, Hindi and Tamil
Real customer questions in Hindi and Tamil, with the phrasing and code-switching that prompts should carry.
- Hindi general conversation
Everyday register for prompts that should not sound like a form.
- Speech and language evaluation
The same rater panels grade the aligned model afterwards.
Questions, answered.
Also from Lokah
Start a collection.
Tell us what you want to collect and from whom. We reply with how we would run it.
- Speech, images, documents, feedback or annotation
- A new collection, or a dataset from the catalogue
- Licensing and data handling
Prefer to reach out directly? Write to hello@kenpathlabs.com.