Evaluating Indian-language retrieval when the benchmark does not exist
You cannot improve what you cannot measure, and for Odia retrieval there was nothing to measure against. So we built the ruler first.
Redrob
·
·
9 min read

We could not tell whether our Odia retrieval was any good.
Not in the sense of it being borderline. In the sense that there was no instrument. No public benchmark, no labelled query set, no relevance judgements, and no baseline to be better or worse than. Every change we made was measured against a feeling.
This note is about building the ruler, and about the ways a ruler you built yourself will lie to you.
Why the existing benchmarks do not help
There are multilingual retrieval benchmarks. Several include Indian languages. They were not usable for our purpose, for three reasons.
The queries are translated, not written. A benchmark built by translating English queries into Odia produces well-formed Odia sentences that no Odia speaker would type. Real queries are shorter, frequently code-mixed, often in Latin script, and full of entities that a translation pipeline mangles. A system tuned on translated queries is tuned for a distribution that does not occur.
The corpora are encyclopaedic. Retrieval over an encyclopaedia is a different task from retrieval over job postings, government scheme documents or product listings, which is what our users are actually searching. Document length, vocabulary, redundancy and the base rate of relevant documents all differ.
Relevance is binary and shallow. Most judgements mark a document relevant or not. For our purposes partial relevance is the interesting case, because the difference between a good system and a bad one is mostly in what it does with documents that are nearly right.
What we built
Four decisions shaped it.
Queries came from logs, not from translation. We sampled real queries, stratified by script, by whether they were code-mixed, and by length. This immediately made the set harder and less tidy, which is the point.
Judgements were made by native speakers who use the products. Not annotators working from a rubric in English. The distinction matters most for partial relevance, where the judgement is about whether a result is useful rather than whether it is topical.
Relevance was graded, not binary. Four levels, with the two middle ones defined by what a user would do next rather than by semantic similarity.
We wrote down the query distribution before looking at any results. This is the discipline that stops a benchmark from quietly becoming a description of the system that produced it.
How our own benchmark lied to us
Three ways, all discovered late.
Sampling from logs samples from what already works. Users do not repeat queries the system failed at. They rephrase, or they leave. A query set drawn from logs is therefore biased towards queries the current system handles, and it will make any successor system look adequate. We partially corrected for this by oversampling sessions that ended without a click, which is a proxy for failure and not a good one.
Annotator agreement was highest where the task was easiest. We were pleased with our agreement figures until we conditioned them on grade. Agreement on clearly relevant and clearly irrelevant documents was high. Agreement on the two middle grades, which is the entire reason we used four, was much lower. A headline agreement figure computed across all grades conceals this completely.
We optimised against it and it stopped measuring. Within a few months of the benchmark existing, changes were being evaluated against it and it was no longer independent of the system. This is not avoidable in principle. It is manageable by holding a portion out, refreshing the query sample periodically, and treating any large jump as a suspect rather than a success.
What we would tell someone starting
Build the evaluation before the system, not after. Everyone says this and almost nobody does it, including us, and the cost of doing it late is that you cannot interpret any of the work that came before.
Sample from reality, and then go and find the failures reality does not record.
Grade partial relevance, and report agreement per grade rather than in aggregate.
Assume your benchmark will be gamed by your own team, without anyone intending to, and design for that from the start.
Where this leaves us
We now have instruments for the languages we serve, of varying quality. They are better than nothing, which was the previous state, and they are worse than a well-constructed public benchmark maintained by people with no stake in the results.
The honest summary is that Indian-language retrieval evaluation is under-served, that every serious group working on it is building its own ruler, and that none of those rulers agree. That is a poor foundation for a field, and it is not something any single company should be fixing on its own.
Which is the argument for doing this work in the open. Our image model's weights are public under Apache 2.0, and the research behind the stack runs jointly with Seoul National University and Yonsei on Korean government grants. Evaluation is the piece that most needs the same treatment, and it is the piece nobody is funded to build.
ECOSYSTEM
SOLUTIONS
BACKED BY
Korea Investment Partners
KB Investment
Kiwoom Investment
KDB Capital
DS&Partners
Murex Partners
Daekyo Investment
Wanted Lab
© 2026 Redrob. All rights reserved.
Privacy
Terms
Security
English
