Home/Help Center/A/B testing

Broadcasts

A/B Test a WhatsApp Broadcast

Split a sample of your audience between two template versions and let a statistical significance test, not a guess, decide which one goes to everyone else.

By Chirag Darji · Updated 27 Aug 2026 · 8 min read

On this page
  1. Before you start
  2. Steps
  3. How does the significance test actually decide a winner?
  4. What you will see
  5. What happens to the arms after the test decides?
  6. Settings and options
  7. Why is the sample size a fixed set of choices rather than a free-form number?
  8. Troubleshooting
  9. Related reading
WhatsApp broadcast campaign report in VGraple CRM with delivered, read and replied stats per contact

A/B testing on a broadcast splits a sample of the audience between two template versions and picks a winner using a two-proportion z-test at 95% confidence, then sends the remaining audience to whichever version actually won. This article covers how to set up a test, how the winner is decided, and what happens when the result is not clear-cut.

A/B testing: first 5 of 7 steps

  1. 1Select your primary template as usual
  2. 2Turn on "Test two versions."
  3. 3Pick Version B's template
  4. 4Set the sample size
  5. 5Choose the winning metric and decision window
The steps on this page, in order.

Before you start

  • You need two approved templates: one for Version A (the template you would already have picked) and one for Version B, an alternative wording, header, or offer you want to compare.
  • Build your audience first; the test sample is drawn from whatever audience you have set on the Audience step.
  • Think about your metric ahead of time. Read rate tells you which subject line or opening line grabs attention; reply or click rate tells you which one actually drives action.

Steps

  1. Select your primary template as usual. On the Message step, choose the template for Version A the same way you would for any campaign, including its personalization mapping.

New broadcast composer in VGraple CRM: audience selection with segment, tag and CSV options and a live recipient count

  1. Turn on "Test two versions." Below the message settings, check Test two versions. A panel opens with three controls: Version B, Sample, and Winner by.

  2. Pick Version B's template. Choose a different approved template from the dropdown; it cannot be the same template as Version A.

  3. Set the sample size. Choose what share of the total audience the test itself uses: 10%, 20%, 30% or 50%. This portion is split between the two versions; the rest of the audience waits.

  4. Choose the winning metric and decision window. Set Winner by to Read rate, Reply rate, Click rate, or Delivery rate, and choose how many hours to wait before deciding: 1, 2, 4, 8 or 24 hours after the test sample goes out.

  5. Finish scheduling and send as normal. The rest of the composer (schedule, pacing, quiet hours) works exactly as it does for any campaign; the A/B mechanics run automatically underneath.

  6. Watch the test resolve on the campaign detail page. Once the decision window closes, a comparison table appears showing each version's sent, delivered, read, replied and clicked counts, with the winner marked, and a plain-language decision note explaining the result.

How does the significance test actually decide a winner?

Email marketing tools have run this kind of test for two decades; running two campaigns to two separate lists and comparing raw percentages by eye is how a 3-point difference on 200 recipients gets mistaken for a real result when it is well within normal random variation. VGraple CRM uses a two-proportion z-test, comparing the two versions' success rate on your chosen metric (delivered, read, replied or clicked) against each other, and only calls a winner when the gap is statistically significant at 95% confidence, meaning there is less than a 5% chance the observed difference happened by chance alone.

The test needs both a minimum sample and a real gap to say anything useful. Below 30 recipients per version, the campaign reports plainly that there were too few recipients to test rather than guessing. Above that threshold, if the two rates are close enough that the test cannot rule out chance, the decision note says the versions were statistically indistinguishable and states the actual p-value, rather than crowning a winner because one number happened to be a fraction higher than the other.

Example

A real-estate agency tests two versions of a new-listing alert on a 4,000-contact audience, sampling 10% (400 contacts, 200 per version). The version leading with the property's location gets a reply rate of 18% against 11% for the version leading with the price, a large enough and consistent enough gap to be statistically significant. The remaining 3,600 contacts automatically receive the location-led version.

What you will see

While the test is collecting results, the campaign detail page shows "Collecting results. The winner goes to the rest of the audience when the window closes." Once decided, each version's row shows its counts and rates, the winning version gets a "winner" badge, and the decision note states either a clear win with its p-value, or that the two versions were statistically indistinguishable. The remaining, un-sampled audience then receives whichever version was sent to it, without a second audience split.

What happens to the arms after the test decides?

Each version of an A/B test is a child of the one campaign, not two separate broadcasts, which is why they share one audience split rather than each drawing its own. The sample is divided once, by the share percentage you set, and each arm's funnel (sent, delivered, read, replied, clicked) accumulates independently from that point on. When the decision window closes, the remaining, un-sampled audience is released to the winning arm without re-picking who gets it, so a contact who was never part of the test sample simply receives whichever version won, and every recipient of the campaign, sampled or not, still appears in the same campaign's single recipient table and export, filterable by which version they received.

Settings and options

SettingWhat it doesOptions
Version BThe second template being tested against Version AAny approved template other than Version A's
SampleShare of the audience used for the test itself10%, 20%, 30%, 50%
Winner byThe metric the significance test is run onDelivered, Read, Replied, Clicked
Decision windowHow long to wait after the sample sends before deciding1, 2, 4, 8, 24 hours
Minimum sample per armRecipients needed per version before a significance test is possible30 (fixed)
Significance thresholdConfidence level required to call a winner95% (fixed)

Why is the sample size a fixed set of choices rather than a free-form number?

Offering 10%, 20%, 30% or 50% rather than an arbitrary percentage or a raw recipient count keeps the trade-off visible in plain terms: a larger sample reaches a significant result with more confidence and sooner, but delays more of the audience from receiving the eventual winning version until the decision window closes, since the sampled portion is what runs the test while the rest waits. A small campaign benefits from a larger sample share, since the fixed minimum of 30 recipients per arm needs a meaningful fraction of a modest audience to reach at all; a very large campaign can use a small sample share and still comfortably clear that minimum on both arms, leaving most of the audience free to receive the winner sooner rather than waiting through the full decision window.

Troubleshooting

SymptomLikely causeFix
Decision note says "too few recipients per version to call a winner"Fewer than 30 recipients received one or both versionsIncrease the sample percentage, or run the test on a larger audience next time
Both versions look close but no winner is declaredThe gap between the two rates is inside the margin of error at 95% confidenceExpected and honest; the better-performing version still goes to the rest of the audience, the note just does not claim statistical proof
Version B option is missing from the dropdownNo second approved template exists yet, or it is the same template already chosen for Version AApprove a second template, or pick a different one for A
Comparison table is empty after the decision window closedThe campaign has not actually finished sending the sample yet, or the window has not elapsedWait for the decision window to pass; check the campaign is not paused
Test says a winner was found, but the difference looks small in raw numbersA statistically significant result can still be a small numeric gap on a large enough sampleThis is expected; significance measures confidence the difference is real, not its size

Pair A/B testing with send-time optimisation to control both the message and the timing, or read seeing revenue per broadcast to judge a winning version by money made, not just engagement.

Frequently asked questions

What statistical test decides the winner?
A two-proportion z-test at 95% confidence, the standard bar used in credible marketing experiments generally. It compares the two versions' rates on your chosen metric and only calls a winner when the difference is unlikely to be chance.
What happens if the two versions perform too close to call?
The campaign says so explicitly rather than declaring a winner by rounding. It still sends the numerically better-performing version to the remaining audience, since something has to go out, but the decision note states plainly that the difference was within the margin of error.
How small can my sample be and still get a real answer?
Each arm needs at least 30 recipients before the significance test can say anything at all. Below that, the campaign reports that there were too few recipients per version to call a winner, rather than pretending a result from 8 people means something.
Can I choose what "winning" means?
Yes, choose delivered, read, replied, or clicked as the metric the test decides on. Delivered is useful for judging template quality itself; replied and clicked are stronger signals of real interest.
What percentage of my audience gets used for the test itself?
You choose, 10%, 20%, 30% or 50% of the audience, split between the two versions. The remainder waits until the test window closes, then goes entirely to the winner (or the better-performing version, if the result was not significant).
Does A/B testing cost extra or need a certain plan?
No, it is available on every plan including Free, the same as every other delivery and safety control in Broadcasts.
Can I see the two versions' funnels separately after the campaign finishes?
Yes. The campaign detail page shows a comparison table with each version's sent, delivered, read, replied and clicked counts side by side, plus the decision note explaining why the winner was chosen.

Run your WhatsApp on VGraple CRM

Free forever plan, official Meta WhatsApp Business API, set up in 15 minutes. No card needed.