← All work

AI Infrastructure · Odia & Indic Data

A data annotation platform for Odia & Indic AI

An in-house labelling platform and expert workforce that turns raw Odia text, speech and images into high-quality training data.

Try a text annotationOdia · ଓଡ଼ିଆ

Choose a label, then select a word.

“Ravi works in Bhubaneswar.”

Select a word to begin

Interactive example. No data is submitted.

Timeline
10 weeks to v1 · ongoing operations
Our role
Data Platforms · Odia LLM & NLP · MLOps
Focus
Data Platform · Odia NLP · Human-in-the-loop

Faster labelling throughput vs. spreadsheets

98%

Inter-annotator agreement after QA gates

40k+

Odia data points labelled in first quarter

Built around the people using it.

Good Indic-language AI is bottlenecked by good Indic-language data. We built an end-to-end annotation platform - tooling, workflows, quality gates and a trained local workforce - so teams can produce gold-standard Odia and Indic datasets at scale, with human review baked in.

The challenge

Off-the-shelf annotation tools have poor support for Odia script, speech and code-mixed Indic text - leading to inconsistent, low-trust labels.

Labelling was happening in spreadsheets with no audit trail, no reviewer workflow and no measure of quality.

There was no trained local workforce or guideline set to produce data that models could actually learn from.

What we built

  1. Purpose-built labelling tooling

    A web platform with first-class Odia/Indic rendering and input, supporting text classification, NER, span highlighting, transcription and image bounding-box tasks in one place.

  2. Human-in-the-loop quality gates

    Every item flows through label → peer review → adjudication. Inter-annotator agreement, gold-set spot checks and reviewer scorecards make quality measurable, not assumed.

  3. Guidelines & workforce training

    We authored task-specific annotation guidelines and trained a local Odia-fluent workforce, so linguistic nuance and cultural context are captured correctly.

  4. Pipeline & model-ready exports

    Versioned datasets export straight into training pipelines with provenance and consent metadata attached - ready for fine-tuning Odia LLMs and NLP models.

Technology behind the work

  • Next.js
  • Python
  • PostgreSQL
  • Label Studio (extended)
  • Hugging Face
  • Object storage
  • Airflow
The quality gates changed everything - we finally trust our Odia training data instead of hoping it's right.
Data & AI LeadIndic language AI initiative

Have a similar challenge?

Let’s map it in a free 30-minute call.

Book a discovery call