By the Wavefield Research team · Published Sep 16, 2026
Political message testing runs a controlled experiment: respondents are randomly split into arms, each arm reads one message, and the arms are compared on vote intent or persuasion. Campaigns test because instinct fails: in a 2024 PNAS study, practitioners predicted which messages persuade barely better than chance.
What follows is the method the ranking guides gesture at and never explain: the five steps, the sample-size math per arm, the line between message testing and push polling, and what it honestly costs.
The best evidence on this is recent and brutal. Broockman, Kalla, Caballero, and Easton (PNAS, 2024) measured the persuasive effect of 172 real campaign messages across 21 issues — 67,215 respondent-message observations — then asked political practitioners to predict which messages had worked. The professionals performed barely better than chance, and no better than ordinary members of the public. Years inside campaigns, polling firms, and advocacy organizations added approximately nothing to message-picking accuracy.
That finding is the entire business case for message testing. It does not say practitioners are bad at strategy — it says persuasion is empirically unpredictable from the inside, which is exactly the condition under which measurement beats judgment. A campaign that tests its shortlist is not buying certainty; it is buying protection from confidently spending the persuasion budget on the wrong message.
The shortlist comes from strategy, not the survey: the economic frame, the public-safety frame, the contrast message. Qualitative work and canvassing conversations are good at generating this list. The experiment's job is narrower and harder — to say which one moves votes.
Random assignment is what makes it an experiment: each respondent lands in one arm by chance, so the arms differ only in the message they saw. Include a control arm that sees no message when you need to measure absolute movement, not just which message wins.
Each arm reads its one message, then every arm answers the same outcome questions — vote intent, candidate favorability, issue support. Showing one respondent several messages contaminates the measurement: reactions to the second message are reactions to the pair.
The result is a comparison of arms: message A's arm at 46% vote intent against control's 41% is a five-point effect — if it clears a significance test on the arm sizes. Differences that do not clear the test are noise, however exciting they look in a deck.
The subgroup read — did the economic message move suburban women, did the contrast message backfire with independents — is usually the real deliverable. It is also where small bases lie: an arm of 300 splits into subgroups of 60, and honest tables mark what can and cannot be tested.
On Wavefield the machinery is declarative: a randomized field assigns each respondent an arm uniformly, messages display per arm, and results crosstab by arm with significance letters at your chosen confidence level — the same table discipline described in cross tabulation. Respondents come from your own channels — volunteer lists, email, consent-based contact — or registered-voter sample we arrange per study.
The mistake campaigns make budgeting a message test: using a sample-size table built for measuring one number. A message test measures a difference between two numbers, and the margin of error on a difference is roughly 1.4 times a single arm's margin. The working table, at 95% confidence: 200 completes per arm reliably detects only differences around 10 points — fine for finding a landslide winner, useless for close calls. 300 per arm brings the detectable difference to about 8 points; 500 per arm to about 6. Message effects in the wild are often smaller than that, which is why serious tests either fund real arm sizes or accept that they are screening for big winners only.
Two budget levers help. A control arm is only necessary when absolute movement matters — a head-to-head between two messages needs one fewer arm than a test measuring lift over silence. And subgroup reads should be planned before fielding: if the deliverable is “does this move suburban women,” the arm sizes must be set for that subgroup's base, not the topline's. The single-number math is in the sample size calculator; multiply its margin by 1.4 for any two-arm comparison.
Message testing polls have drawn fire for decades, mostly because a genuinely abusive practice wears the same costume. AAPOR's statement on push polls draws the line cleanly: a push poll is political telemarketing disguised as research — thousands of calls spreading a distorted negative message, collecting nothing, measuring nothing. A legitimate message test is the opposite on every operational dimension: a few hundred respondents, real questions answered and recorded, results private to the campaign, and respondents contacted through consented channels.
The practical implication for a campaign is reputational: a test that reads negative contrast copy to respondents should be genuinely testing it — honest arm sizes, real outcome measures — because volume is the tell. If a “survey” is reaching more voters than a persuasion program would, it is a persuasion program. Our stance on respondent contact is the same one documented on the political polling page: your channels with consent, or arranged sample, never scraped contact lists.
A split-sample message test on Wavefield is an ordinary project: from $99 to $299 by tier, with randomization, arms, quotas, weighting, and significance testing included, and unlimited completes from your own channels. Arranged registered-voter sample runs $5–8 per quality-checked complete, so a three-arm, 300-per-arm test lands near $5,000–7,500 all-in — a fraction of the focus-group route, and with a defensible number at the end instead of a highlight reel. The wider price context is in political poll cost. US campaigns pay list price; as a Canadian company we do not discount US political work.
Where we point campaigns elsewhere, honestly: moment-to-moment dial testing of video creative is a genuine specialty — if the question is “where in this 30-second ad do viewers disengage,” a dial vendor answers it and a survey does not. And discovering the message language in the first place is qualitative work; a split-sample test evaluates a shortlist, it does not write one. What we do is the quantitative verdict: which of these messages moves votes, by how much, among whom.
A controlled experiment that measures which campaign message actually moves voters. Respondents are randomly split into arms; each arm reads one message (a control arm reads none); all arms answer identical questions about vote intent or issue support; and the arms are compared with significance tests. It replaces the deck-room argument about which message is strongest with a measurement — which matters because practitioners' instincts about persuasion test barely better than chance.
Enough that the difference you care about clears the noise. At 200 completes per arm, only differences of about 10 points are reliably detectable; 300 per arm brings that to about 8 points; 500 per arm to about 6. The margin of error on a difference between two arms is roughly 1.4 times a single arm's margin — a detail that catches people who budget arms using ordinary sample-size tables built for one number, not a comparison.
No, and the distinction has teeth. A message test is research: a sample of a few hundred voters, honest questions, results kept by the campaign. A push poll is telemarketing wearing a survey's costume — thousands of calls spreading a negative message with no intention of measuring anything, a practice AAPOR condemns outright. The operational tells: real research uses small samples and asks real questions; push polls use huge volumes and collect nothing.
Two components: the project and the sample. On Wavefield a split-sample test is an ordinary project — from $99 to $299 depending on tier, with the arms, randomization, and significance testing included — plus respondents. Fielding to your own lists costs nothing extra; arranged registered-voter sample runs $5–8 per quality-checked complete, so a three-arm test at 300 completes per arm lands near $5,000–7,500 all-in. US campaigns pay list price — we are a Canadian company and do not discount US political work.
Related: political polling · political poll cost · sample size calculator · cross tabulation
Brief the agent with your messages and it builds the arms, randomizes the assignment, and reports which message won — with significance letters, not vibes. From $99 per study.
Start a study