Collecting responses is the easy half. What separates a defensible finding from a spreadsheet full of percentages is the order of operations: clean, weight, topline, crosstab, test — then report only what survives. Here's the walkthrough, the way research shops actually do it.
The sequence matters: weighting before you read numbers, cleaning before you weight. Analysis done out of order tends to get quietly redone — or worse, doesn't.
Screen for speeders (finished implausibly fast), straightliners (same answer down every grid), and failed attention checks. Flag them and review — don't silently delete. Deleting rows destroys the audit trail; flagging keeps the analysis honest and lets you show exactly what was excluded and why. And be precise about your base: screen-outs and quota terminations aren't completes and shouldn't sit in any percentage.
Raw online samples over-represent whoever's easiest to reach. Weighting (usually raking, also called rim weighting) adjusts each respondent's contribution so the sample matches known population targets — age, gender, region, education. Report weighted percentages, but keep counts unweighted: a weighted count is a fiction, and reviewers know it.
The topline — every question's overall distribution — is your map. Read it start to finish before touching subgroups, noting anything that contradicts expectations (that's either your finding or your data problem). Attach a margin of error computed on the effective base, not the raw count: weighting always costs precision, and the effective base is what you actually have.
Comparing groups (age bands, regions, customer segments) is where findings live, and where most false conclusions are born. A 6-point gap between subgroups of 80 people each is usually noise. Run pairwise significance tests on every comparison — the banner-book convention of letters marking which columns differ at 95% — and respect minimum base sizes before reading any cell.
Free-text answers carry the why behind every number, but anecdotes aren't analysis. Code them into themes — deductively from your research questions, inductively from what respondents actually said — then quantify: what share of respondents raised each theme, and how does that differ by segment? Keep a few verbatims per theme for the report; they persuade in ways percentages can't.
State weighted percentages with unweighted bases visible. Give the margin of error and note it's a conventional approximation if your sample is a quota sample rather than a probability sample. Resist the temptation to narrate every significant cell: with dozens of comparisons, some will clear 95% by chance alone. The findings worth reporting are the ones that are significant, sizable, and coherent with the rest of the data.
A crosstab worth trusting shows its work: weighted percentages, unweighted bases, and significance letters computed on effective bases (weighting reduces the information in a sample — the effective base accounts for that). In the example, the 18–44 group's higher support carries a letter because the gap clears 95% confidence; “Not sure” carries none because it doesn't.
This is the format an analyst, a journalist, or an opposing counsel will ask for — and the one generic form tools don't produce.
| Total | 18–44 (A) | 45+ (B) | |
|---|---|---|---|
| Support | 54% | 63%B | 47% |
| Oppose | 31% | 24% | 37%A |
| Not sure | 15% | 13% | 16% |
Base: 512 completes (unweighted). Letters mark differences significant at 95% on Kish effective bases.
A subgroup of 40 people has a margin of error near ±16 points. Set a minimum base (60–100 unweighted is a common convention) and refuse to interpret cells below it.
If your completed sample is 68% female and the population is 51%, every gender-correlated number is wrong until you weight. Unweighted results from an online panel aren't conservative — they're just skewed.
Hard-deleting bad responses makes your dataset unauditable. Flag, exclude from analysis, and keep the rows — a client or reviewer who asks what was removed should get an exact answer.
Run 100 comparisons at 95% confidence and roughly five will look significant by chance. Significance is a filter, not a headline generator — demand size and coherence too.
The classic margin of error formally applies to random samples. On quota samples (most online research), report it as the convention it is — computed on the effective base — not as a guarantee.
Early respondents differ from late ones — the eager, the online-all-day, the panel professionals come first. Mid-field reads are for monitoring quotas and quality, not for conclusions.
Data cleaning — before any percentages. Identify speeders, straightliners, and failed attention checks, flag them for exclusion (don't delete), and confirm your base counts only genuine completes, not screen-outs or quota terminations. Every downstream number inherits whatever quality problems you didn't catch here.
Not necessarily, but you need the same operations a statistician would run: weighting, crosstabs, significance tests, margins of error on effective bases. Some platforms build these in — Wavefield computes all of them by default on every project — while generic form tools leave you to do it in a spreadsheet, which is where most methodology errors happen. SPSS exports still matter when a client's analyst wants the raw dataset.
That a difference between two subgroups is unlikely to be chance alone — conventionally tested at 95% confidence, marked with letters showing which columns differ. It is not proof of importance: with many comparisons, some clear the bar by luck, and a significant 2-point gap can still be too small to matter. Treat significance as necessary, not sufficient.
Weighted percentages with unweighted bases — that's the research-industry convention. Weighting corrects the sample's composition to match the population, so percentages should use it; counts stay unweighted because a weighted count doesn't correspond to actual people. Any report that hides its unweighted bases deserves suspicion.
It depends on the precision you need overall and within subgroups — a ±5% read on the total needs roughly 385 completes, but a ±5% read on a subgroup needs roughly 385 in that subgroup. Work it backward from the smallest group you intend to report on.
Work out your required sample with the sample size calculator, or read quota sampling vs. random sampling and what survey scripting involves.
Wavefield runs every step above by default — quality flags, raking weights, toplines, crosstabs with significance letters, SPSS exports — on every project, from $99.
Start a project