1,284

Glossary

What does any of this actually mean?
Why this screen exists

Survey methodology has a large vocabulary and most of it is doing real work rather than dressing up something simple. Every term used anywhere in this console is defined here in one plain sentence, with a second line on why it changes the answer. The same definitions appear on hover wherever a term is used, so nobody has to leave the screen they are on.

Where a poll goes wrong

The six ways a survey can be off, and which ones the industry actually reports.

Total survey error
Every way a poll can be wrong, added together.
Six sources: coverage, nonresponse, measurement, adjustment, processing and sampling. Indian polling reports only the last one, which is the smallest.
Sampling error
The wobble you get purely from asking some people instead of everyone.
It is the only error that shrinks reliably when you interview more people. It is also the only one most published polls admit to.
Margin of error
How far the true figure could sit either side of the number shown.
Usually quoted as a plus or minus. It covers sampling error only, so the real uncertainty is always wider than the number printed.
Coverage
Whether the people you can reach look like the people who vote.
If a group cannot be reached by phone at all, no amount of extra dialling fixes it. It has to be filled another way or declared.
Nonresponse
People who could have been interviewed but were not.
Dangerous when the people who refuse differ from the people who answer, which they usually do.
Measurement error
The answer recorded is not the answer the person gave or meant.
Caused by bad wording, leading questions, a mishearing, or a respondent saying what sounds acceptable.
Design effect
How much precision you lose because the sample is not a clean random draw.
A design effect of 1.55 means your sample behaves like one about a third smaller than its headline size.
Effective sample
The size a clean random sample would need to be to match your precision.
The honest number. Always smaller than the headline count once weighting and clustering are accounted for.
Nominal
The raw headline count of interviews completed.
What gets printed on screen during a broadcast. It overstates precision on its own.
Kish
The standard formula for turning a weighted sample into an effective size.
Named after Leslie Kish. It is the arithmetic behind the effective sample figure.
Running the field

How interviews are drawn, reached and completed.

Stratum
A defined slice of the population that gets its own interview target.
Splitting the country into strata stops a sample from over-filling in easy places and under-filling in hard ones.
Quota
The number of interviews owed in a particular slice.
Quota fill tells you where the field is running behind before the wave closes, not after.
Random digit dialling
Generating numbers to call from allocated blocks rather than buying a list.
A purchased list carries whoever the seller happened to collect. Generated numbers give every active line a chance of selection.
Dual frame
Running a second, different method alongside the phone sample.
A small face-to-face sample reaches the households a phone never will, and calibrates the rest.
Next-birthday selection
Asking for whichever adult in the house has the next birthday.
Stops the sample filling up with whoever answers the phone, which skews male and older.
Break-off
A respondent who starts the interview and hangs up partway.
Where they hang up tells you which question is driving people away.
Straight-lining
Giving the same answer to every question to get through quickly.
A sign the interview was not taken seriously. Those records are held out.
Continuum of resistance
The pattern that hard-to-reach people answer differently from easy-to-reach ones.
Track the answers by attempt number and you can extrapolate towards the people you never reached at all.
Callback
Trying an unanswered number again rather than replacing it.
Cheap here and expensive for a field team, which is why human surveys give up early and inherit the resulting bias.
Response rate
The share of eligible people approached who completed an interview.
RR3 is the standard AAPOR definition. Publishing it is normal abroad and rare in India.
Contact rate
The share of numbers where a real person was reached at all.
Separates a dialling problem from a persuasion problem.
Cooperation rate
Of the people actually reached, the share who agreed to take part.
This is the number a persona or an opening line moves.
Refusal rate
The share who were reached and declined.
Rising refusal on one question usually means the question is the problem.
Attrition
Panel members dropping out over time.
A panel that is not replenished slowly stops resembling the country.
Propensity
How likely a given panel member is to answer the next call.
Used to plan dialling load. Never used to decide whose opinion counts.
Turning answers into an estimate

Everything that happens between the raw responses and a seat number.

Interviewer variance
Answers clustering by who asked the question rather than who answered it.
Every human interviewer has a personal effect on responses, and it multiplies across their whole caseload.
rho
The measure of how much answers cluster within one interviewer.
At rho of 0.02 and 120 interviews each, a 3.5 lakh sample carries the precision of about 1 lakh.
Persona effect
The measurable shift in answers caused by which voice asked.
Randomising the voice within each slice is what makes this effect measurable and correctable instead of baked in.
Post-stratification
Reweighting the sample afterwards so it matches known population totals.
Fixes a sample that came in with too many of one group and too few of another.
MRP
A model that borrows strength across similar places to estimate every seat.
Multilevel regression with post-stratification. One national sample yields a read for all 543 seats without needing deep samples in each.
Raking
Adjusting weights repeatedly until the sample matches every target margin at once.
Standard practice. The margins used must be declared in advance or the weighting becomes a place to hide choices.
Trimming
Capping how heavy any one respondent weight is allowed to get.
Without a cap, a handful of rare respondents can quietly drive a national number.
Shrinkage
Pulling a thin local estimate towards the broader pattern.
Large where the local sample is small, near nil where it is deep. Exactly the behaviour you want.
Partial pooling
Letting each seat borrow information from similar seats.
The mechanism behind shrinkage. It is why a national sample can speak to a single constituency.
Correlated swing
Assuming errors move together across seats rather than cancelling out.
This is the 2024 lesson. Independent errors gave a range of 12 seats. A shared swing term gives 48, and contains the actual result.
Turnout model
The assumption about which respondents will actually vote.
One of the largest hidden levers in any poll, which is why it is sealed before field opens.
Imputed
Filling in an answer the person did not give.
Not done here. A refusal is reported as a refusal, because guessing hides the most interesting group in the sample.
Weighted
Counting some respondents more than others to correct sample imbalance.
Necessary and also the easiest place to quietly manufacture a result, which is why the recipe is pre-registered.
Proving it

The checks that make the estimate auditable rather than assertable.

Closed response
A question with fixed options read aloud, not an open conversation.
The reply is matched against a fixed list rather than transcribed freely, which is far more accurate.
Constrained grammar
A fixed set of expected replies the system listens for.
Recognition against ten known options is a much easier problem than transcribing open speech.
Word error rate
How often open transcription gets a word wrong.
The number voice vendors publish. It is the wrong number for a survey, because survey answers are never taken from open transcription.
Blind double-coded
Two coders working the same audio without seeing each other, or the machine.
The only honest way to measure whether the automatic capture is right.
Keypad fallback
Offering the respondent number keys when the spoken answer is unclear.
Turns an unreliable item into a certain one instead of a guess.
Quarantine
Holding a suspect interview out of the analysis entirely.
Held before analysis, not after. An interview removed after seeing the result is not a quality rule.
Suppression
Refusing to compute an estimate for a group that is too small.
Below the floor a group stops describing a segment and starts describing identifiable people.
Brier score
A score for how well-calibrated a forecast was, not just whether it was right.
A confident wrong call is punished harder than a hedged one. It rewards honest uncertainty.
Interval coverage
How often the true answer actually landed inside the stated range.
If your 80% intervals contain the truth 80% of the time, your uncertainty is honest.
Pre-registration
Sealing every analysis choice before the first call is made.
Removes the ability to pick the method after seeing the numbers, which is the main way polls go wrong without anyone lying.
Spec hash
A fingerprint of the sealed method file.
Proves afterwards that nothing in the plan was changed once results started arriving.
The legal edges

Two rules that govern when anything can be published.

Section 126A
The law barring publication of exit poll results during the polling window.
Exposure here is criminal, not commercial, so the restriction is built into the system rather than left to a person to remember.
Model Code of Conduct
Election Commission rules governing conduct once an election is announced.
Poll publication timing sits inside it, which is why publication authority stays with the client and never with the vendor.
A note on the vocabulary

None of this is proprietary. It is the standard working language of survey research at Pew, YouGov, Ipsos and the academic literature. It is unusual in Indian election polling only in the sense that it is rarely used out loud. Every term here corresponds to a decision that is either made deliberately and declared, or made by default and left unstated. The argument of this whole console is that the first is worth paying for.