A/B Testing Cold DMs: The Sample Sizes Your Tests Actually Need
Most cold DM tests get decided on 200 sends and a hunch. Here is the real math: how many DMs per variant you need, how long that takes at safe send volumes, and when a winner is actually a winner.
You ran 100 DMs on script A and 100 on script B. A got 11 replies, B got 6. You killed B and rolled A out across every account.
You just made a decision on noise. At those numbers, an 11 versus 6 split sits well inside the range you would expect from two identical scripts. Flip a coin 200 times in two batches and you will see gaps that big regularly.
This is the most common way outreach teams waste a quarter. Not by testing the wrong things, but by declaring winners on sample sizes that cannot support a decision. Here is the math, the volume it implies, and what to do when you do not have that volume.
The number that decides everything: your MDE
Minimum detectable effect is the smallest lift your test can reliably find. It is the input almost nobody sets on purpose, and it drives sample size harder than anything else.
The relationship is quadratic. Halve the effect you want to detect and you need roughly four times the sample. That single fact explains why most cold DM tests are underpowered by an order of magnitude.
Below is the sample size per variant at an 8% baseline reply rate, 95% confidence, 80% power. 8% sits in the middle of the range operators report for cold Instagram DMs to a well matched ICP, where generic blasts land nearer 5% and tight targeting pushes into the mid teens.
Read the top bar and the bottom bar together. A script rewrite that doubles reply rate needs 260 DMs per arm and you can call it inside a week. A copy tweak worth one point needs 12,209 per arm and will eat most of a quarter. Same statistical machinery, 47x the volume.
What that costs you in send days
Cold DM volume is capped by account safety, not by ambition. Operators consistently settle in the 20 to 35 unique cold DMs per day per warmed account band, with 40 to 60 possible on high trust properties at materially higher action block risk.
Take 30 per day per account as a working number and see what each test costs.
| Test you want to run | DMs per arm | Total DMs | Days on 5 accounts | Days on 15 accounts |
|---|---|---|---|---|
| Full rewrite, +8 points | 260 | 520 | 4 | 2 |
| Angle change, +4 points | 883 | 1,766 | 12 | 4 |
| Copy tweak, +2 points | 3,215 | 6,430 | 43 | 15 |
| Micro tweak, +1 point | 12,209 | 24,418 | 163 | 55 |
The practical read: with fewer than 10 accounts running, you can only honestly test big swings. That is not a limitation to route around. It is a targeting instruction. Test things that could plausibly move reply rate by half or double it, and stop A/B testing punctuation.
Test on replies, not on booked calls
Every step down the funnel costs you sample. Reply rate sits around 8%. Booked call rate off raw sends sits closer to 1.5% for most operators. Detecting a 50% relative lift on that booked rate takes 5,135 DMs per arm, against 883 for the same relative lift on replies.
So run the experiment on the reply, and validate downstream over a longer window. If a script wins on replies but the calls never show up, that is a real signal, but you will need weeks of data to see it, not days.
The trap is the opposite pattern: a script that wins on replies because it over promises, then drags show rate down. Track booked and held rate per variant as a guardrail metric. You are not powering the test on it, you are watching for a collapse big enough to be obvious without statistics.
Stop peeking
The single largest source of fake winners is checking the dashboard every morning and stopping the day it turns green.
Every look is another chance to catch random noise at its peak. Reported inflation runs from roughly 8% false positives after two peeks up to about 19% after ten, and some analyses put continuous monitoring at 30% or higher depending on how correlated the checks are. Whatever the exact number, it is several times the 5% you think you are running at.
Three fixes, ordered by how easy they are to actually enforce:
- Set the sample size before you launch and do not look at the split until you hit it. Watch total volume sent, not the win.
- Fix a stop date instead. Same effect, much easier to enforce with a team that shares a dashboard.
- Use a sequential method if your tooling supports always valid p-values. Most DM tooling does not, so treat this one as optional.
What is actually worth testing
Ranked by expected effect size, which is the same thing as ranked by how fast you can resolve them.
| Variable | Typical effect | Testable at low volume |
|---|---|---|
| Targeting source and ICP tightness | Very large | Yes |
| Opening premise, insight versus ask | Large | Yes |
| Offer framing in message one | Large | Yes |
| Sending profile quality and social proof | Large | Yes, but confounded, block by account |
| Follow-up count and cadence | Medium | Yes, over 3+ weeks |
| Message length | Medium to small | Marginal |
| Emoji, greeting, punctuation | Small | No |
Note the pattern. The high leverage variables all sit upstream of copy. Who you message and what you claim beats how you phrase it, and it beats it by enough to be measurable in days rather than months.
There is supporting evidence for the framing item specifically. Reported conversion for cold DMs that lead with an ask sits in the 8% to 12% range, against 35% to 45% for messages that lead with a specific insight. Treat those as operator reported ranges rather than lab numbers, but the direction is consistent across sources and the gap is far too wide to be sampling noise.
Guardrails that keep a test clean
- Randomize at the prospect level, not the account level. Assigning script A to account 1 and script B to account 2 tests your accounts, not your scripts. If you must split by account, rotate variants across accounts mid test.
- Run both arms in the same hours. Reply rate moves with send time. A test where A ran Monday and B ran Saturday is uninterpretable.
- Freeze the list source. A fresh scrape mid test invalidates everything before it.
- One variable per test. Two changes and a win tells you nothing about which change earned it.
- Log every send, including the ones that landed in hidden requests. Dropping non delivered sends from the denominator inflates both arms unevenly.
The realistic cadence
For a team running 10 accounts at 30 cold DMs per day, that is 300 sends a day and roughly 2,100 a week.
That budget buys you one +4 point test resolved every 6 days, so roughly five clean experiments per quarter. Five real answers beats twenty guesses, and it compounds. Two consecutive +4 point wins on an 8% baseline puts you near 16%, which is double the pipeline off the same send volume and the same account risk.
Pick the five things worth knowing. Size them before you launch. Do not look until you are done.