Annotation on the data you already have.
Your recordings, documents or images, labelled by native speakers to your guidelines: transcription with speaker turns and tags, entity and intent labels, document fields, image boxes and masks. The same review process as the catalogue, applied to your data.
When you need it
- You have call recordings, chats or documents in Indian languages and no transcripts or labels.
- Your existing labels were made by people who do not speak the language, and the model shows it.
- You need a consistent tag vocabulary across languages, with agreement you can measure.
What one record looks like
An example task, as the rater or reviewer sees it. Try it.
A bracket tag marks what is not a spoken word; the audio here was too faint to transcribe.
- item_id
- id
- Your identifier, unchanged.
- annotations
- list
- Segments, spans, boxes or fields, each with its label and time or position.
- tags
- labels
- From one vocabulary in square brackets: [pii], [unclear], [overlap] and the ones your guideline adds.
- annotator and reviewer
- pseudonymous ids
- Agreement per annotator in the table.
How Lokah builds it and checks it.
- 01
Guidelines with examples
Your guideline, or ours, turned into worked examples in each language and tested on a pilot batch.
- 02
Annotators who speak the language
Recruited from the network for the language and, where it matters, the dialect; calibrated before the batch.
- 03
Review and agreement
A share of items agreed in the brief is double-annotated; agreement is tracked per annotator and items below the floor are redone.
- 04
Delivery
In your tool's format: Label Studio, WebVTT, RTTM, COCO, or plain JSON, with the tag vocabulary and agreement report.
How it arrives
- Label Studio
- Projects you can open and continue.
- WebVTT and RTTM
- Transcripts and speaker turns for audio.
- COCO JSON
- Boxes and masks for images.
- JSON
- Any schema you specify.
Other layouts on request. Catalogue layouts are described on formats.
Pairs with it
From Lokah, and why.
- Hindi speaker diarization
The turn-marking convention our annotators follow, on a set you can inspect.
- Delivery formats
The layouts and tag vocabulary the catalogue uses, which annotation follows.
- Hindi speech recognition utterances
What a finished transcription set looks like, field by field.
Questions, answered.
Also from Lokah
Start a collection.
Tell us what you want to collect and from whom. We reply with how we would run it.
- Speech, images, documents, feedback or annotation
- A new collection, or a dataset from the catalogue
- Licensing and data handling
Prefer to reach out directly? Write to hello@kenpathlabs.com.