1,284

Measurement

How accurate is the capture, and does the error fall evenly?
Why this screen exists

In a service agent, a misheard word is recoverable. The agent asks again, the caller repeats, the transaction completes, and the error never reaches the output. In a survey the captured response is the output, so there is nothing downstream to correct it. Word error rate also is not uniform: it rises on noisy telephony, dense dialect, older speakers and low-literacy respondents, which are precisely the cells whose voting behaviour differs most. Error correlated with the variable being estimated is bias, not noise, and bias does not shrink as the sample grows. Three and a half lakh interviews make it worse, because they buy confidence in it. No Indian agency publishes a measurement error rate. Human enumerators have one too, and it is larger and unmeasured.

Human-audited subsample
5.0%
10,730 interviews re-coded blind
Closed-response agreement
98.6%
Machine capture against human coder
Re-ask rate
4.1%
Low confidence, question repeated
Items voided
0.9%
Never resolved, left as missing
How a response is captured

The answer is never taken from open transcription. Every substantive item is closed, the options are read aloud, and the reply is matched against a constrained grammar built for that question in that language. Below the confidence threshold the question is repeated once, then offered as a keypad response, then left as missing. Open transcription runs alongside for the single unprompted concern question and for the audit record, and it never sets a coded value. An unresolved item stays missing. Nothing is inferred to fill it.

Capture accuracy by languageRolling 14 days
LanguageCompletesClosed agreementRe-askKeypad fallbackOpen-text WERVoidedState
Hindi81,62098.9%3.6%1.9%11.2%0.7%In tolerance
Bengali19,33098.7%4.0%2.1%13.4%0.8%In tolerance
Marathi17,18098.8%3.8%2.0%12.1%0.7%In tolerance
Telugu17,18098.4%4.4%2.4%15.8%1.0%In tolerance
Tamil15,03098.5%4.3%2.3%14.9%0.9%In tolerance
Urdu12,88098.6%4.1%2.2%14.1%0.9%In tolerance
Gujarati12,88098.8%3.7%1.8%12.6%0.7%In tolerance
Kannada12,88098.3%4.6%2.6%16.4%1.1%In tolerance
Odia8,58098.2%4.9%2.8%18.7%1.2%In tolerance
Malayalam8,58097.1%6.2%3.9%21.3%1.9%Below floor
Punjabi6,44098.5%4.2%2.2%14.6%0.9%In tolerance
Assamese2,15098.0%5.1%3.0%19.8%1.3%In tolerance

Open-text word error rate is shown because it is the number vendors publish, and it is the wrong number to judge a survey instrument by. Closed agreement is the one that governs the estimate. The two diverge by roughly an order of magnitude, which is the whole reason the capture is built the way it is.

Agreement by respondent cell
Metro, graduate, 25-4499.4%
Reference cell
Small town, secondary, 25-4499.1%
Within 0.5 of reference
Rural, secondary, 25-4498.8%
Within 0.5 of reference
Rural, primary, 45-5998.5%
Within 1.0 of reference
Rural, below primary, 60+97.6%
1.8 below reference, correction applied
Rural women, below primary, 45+96.9%
2.5 below reference, cell flagged
Why the spread matters more than the level

A uniform 98% loses precision. A 99.4% in metro graduate cells against 96.9% in rural low-literacy cells moves the estimate, because those cells vote differently. The flagged cell is corrected in the weighting and the residual is carried into the published interval rather than absorbed silently.

What is published with the estimate
Agreement, overall98.6%Agreement, worst cell96.9%Agreement, worst language97.1%Audited subsample5.0%, blind double-codedCoder agreement, human to human97.4%Items voided0.9%, left missing
The comparison that matters

Two human coders working the same audio agree with each other 97.4% of the time. Machine capture agrees with a human coder 98.6% of the time. The instrument is not being asked to be perfect. It is being asked to be measured, uniform across cells, and better than the alternative on all three counts.

Status of this panel

Figures on this screen are illustrative and demonstrate the reporting design, not a measured DeployOne wave. The per-language and per-cell accuracy numbers become real after a pilot wave with a blind double-coded subsample. That measurement is the first deliverable of any engagement, ahead of any published estimate.