No online sample matches its population — weighting is how honest research corrects for that. Here's what weighting actually does, how raking works underneath, and the reporting conventions that separate defensible weighted results from cooked ones.
By the Wavefield Research team · Published Aug 4, 2026 · Updated Aug 13, 2026
Whoever is easiest to reach answers most: online panels over-represent the extremely online, customer lists over-represent the engaged, and almost every self-served survey skews female and under-65. If your sample is 68% women and your market is 51%, every number correlated with gender is wrong until you correct the mix.
Weighting is that correction: each respondent gets a multiplier so the sample's composition matches known population targets — census distributions for age, gender, region, and education in general-population work; your own customer-base figures for customer studies; past-vote recall in political polling. Quotas manage the mix during fielding; weights fix what fielding still missed — serious shops use both, and treat them as different tools.
Each 65+ respondent counts ~1.8×; the report shows the weighted share with the unweighted base, and the margin of error widens to reflect the effective base.
The standard algorithm — raking, rim weighting, or formally iterative proportional fitting — is simpler than its names suggest. Take your first target dimension, say age: multiply every respondent's weight so each age band matches its population share. Now do gender the same way — which slightly breaks the age match. Then region, which disturbs both. Then repeat the whole cycle. Because each pass disturbs the others less than the last, the weights converge within a few dozen iterations to a set that matches every dimension's marginals simultaneously.
A toy example — two dimensions, one cycle at a time
Say a 500-person sample lands 60% women / 40% men against targets of 51 / 49, and 30% under-45 against a target of 45%. Raking proceeds:
Real studies add region and education the same way. The finished weights here run from roughly 0.7 to 1.8 — comfortably inside conventional trimming caps.
Raking's practical advantage over exact cell weighting is its appetite for data: it needs only the marginal distributions — what share of the population is 65+, what share is female — not the population count of every combination (65+ women in the Midwest with a college degree). Marginals are published; full crosses usually aren't. The trade-off is that raking matches the margins without guaranteeing the crosses, which is almost always acceptable and occasionally worth checking.
Choosing targets is where judgment enters. Weight on variables that are (a) reliably known for the population, (b) measured the same way in your survey, and (c) actually correlated with your topic — age, gender, region, education for gen-pop work; add past vote for polling. Weighting on a dimension that doesn't relate to your subject adds variance and buys nothing.
The classic target-selection mistakes are worth naming. Deriving targets from the sample itself — circular, since the weights then just confirm the skew. Weighting on more dimensions than a small sample can support: every added dimension spreads the weights wider and shrinks the effective base. Using targets whose categories do not match the questionnaire wording, so respondents land in bands the targets never defined. And copying targets from a years-old study — census distributions drift, and stale targets quietly reweight to a population that no longer exists. The professional bodies publish good guidance here: AAPOR in the US and ESOMAR internationally, whose transparency standards are the reference point serious research buyers check against.
Percentages use the weights; counts never do — a weighted count doesn't correspond to actual people. Any report hiding its unweighted bases deserves suspicion.
Uneven weights cost precision: 500 interviews might carry the power of 400 (the Kish effective sample size). Compute MoE and significance tests on that number, not the raw count — otherwise weighting silently manufactures certainty.
A respondent counted 8× is one person impersonating eight. Caps around 4–6× are conventional; if trimming pushes a dimension off target, the sample was too skewed to fix with math — fix the fieldwork instead.
Two terms you will meet in the methodology literature: post-stratification is the simpler cousin of raking — it weights on the full cross-classification (age × gender × region as one grid), which needs a population count for every cell; raking needs only the margins, which is why it is the practical default. And the precision cost that effective bases measure has a formal name, the design effect — the factor by which uneven weights inflate variance compared to an equal-weight sample of the same size.
These conventions are enforced defaults on Wavefield rather than analyst discipline: reports always pair weighted percentages with unweighted bases, and significance letters and margins of error are computed on Kish effective bases automatically.
Adjusting how much each response counts so the sample's mix matches the population's known mix. If 65+ adults are 22% of your population but only 12% of your respondents, each 65+ response is counted roughly 1.8 times so the totals reflect the population rather than whoever was easiest to reach.
The standard weighting algorithm, formally iterative proportional fitting. It adjusts weights to match each target dimension in turn — age, then gender, then region — and repeats the cycle until all dimensions match their targets at once. Its advantage over cell weighting is that it needs only the marginal distributions (22% are 65+), not the full cross of every combination (what share are 65+ women in the Midwest).
Not the count of interviews — but they reduce the effective sample size. Uneven weights add variance, so 500 responses might carry the statistical power of 400 (the Kish effective base). Honest reporting computes margins of error and significance tests on the effective base, not the raw count.
When a group is so under-represented that its members would carry extreme weights — weighting 15 respondents up to stand for 22% of the population manufactures precision that isn't there. Standard practice caps (trims) weights around 4–6× and treats a needed-but-missing group as a fieldwork problem to fix with quotas on the next wave, not a math problem to hide.
Weighting is step 2 of the full workflow — how to analyze survey results walks all six steps. Related: quota sampling · sample size calculator · data collection methods · question types · weighting multi-option gender data · reading weighted crosstabs
Set your population targets and Wavefield rakes, trims, and reports on effective bases automatically — on every plan.
Start a project