The Cold DM Ops Stack: The Numbers to Log So a Broken Campaign Shows Up in 48 Hours
Most DM programs do not die loudly, they rot. Here are the eleven metrics to log, the denominators that make them diagnosable, and the four trip wires that catch a broken Instagram outreach campaign inside 48 hours.
Most DM programs do not die loudly. They rot. Reply rate slides from 9% to 6% over three weeks, nobody notices because the setter is still busy, and by the time someone opens the sheet you have burned a month of sends and half your warm list.
The fix is not a better script. It is instrumentation. If you cannot see a broken campaign inside 48 hours, you are not running outreach, you are running a slot machine.
Here is the tracking stack we use to catch problems while they are still cheap.
Why most DM dashboards are useless
Three failure patterns show up in almost every agency account we audit.
They track totals, not cohorts. "We sent 12,000 DMs and booked 41 calls" tells you nothing. Sends from week one and week four are different populations, on different accounts, hitting different list sources. Averaged together, a collapsing new-account cohort hides behind a healthy old one for weeks.
They track the end of the funnel only. Booked calls is a lagging indicator with a 7 to 14 day delay in most programs. By the time booked calls drop, the cause is two weeks in the past and you have no evidence left to diagnose it.
They have no denominator discipline. Reply rate over what? Sends? Delivered? Sends that landed in a visible inbox? Teams change the denominator silently and then argue about whether the campaign improved.
Fix the denominators first, then the cohorts, then the leading indicators.
The eleven numbers worth logging
Every row in your log should be one prospect, one account, one campaign, one day. Everything below is derived from that. Do not store percentages, store events, and compute rates at read time.
| Metric | Denominator | Healthy band (reported operator range) | What a break usually means |
|---|---|---|---|
| Sends per account per day | account-day | 20 to 40 on aged accounts | Over the line, action blocks follow within days |
| Send failure rate | attempted sends | under 2% | Session death, proxy flap, or a soft block already active |
| Visible inbox rate | delivered sends | 70% to 85% | Sending profile or account trust dropped, DMs sinking into hidden requests |
| Reply rate | visible inbox sends | 6% to 12% | List quality or opener. Split by list source before touching the script |
| Positive reply rate | replies | 30% to 50% of replies | Targeting drift. You are reaching people who answer but do not qualify |
| Median first response time | positive replies | under 5 minutes in working hours | Staffing gap, not a messaging problem |
| Exchanges before ask | booked conversations | 3 or more | Setters pitching the call too early |
| Booking rate | positive replies | 20% to 30% after 3+ exchanges | Calendar friction, or the offer is not landing |
| Show rate | booked calls | 60% to 80% | Confirmation sequence broken or booking too far out |
| Account survival at 30 days | accounts started | over 85% | Warmup or volume policy is too aggressive |
| Cost per qualified meeting | qualified meetings | under $200 is the elite 2026 benchmark | Everything above, rolled up |
The last row is the only one leadership cares about. The first ten are the only ones you can actually fix.
Visualise the funnel, not the total
A worked example on a 1,000 DM cohort, using mid-range reported figures. These are illustrative, not a benchmark. Run your own numbers into the same shape.
Read it stage by stage. If 780 drops to 600, that is a deliverability and account trust problem, and no amount of script testing will save you. If 90 replies drops to 55 while the visible inbox rate holds, that is list or opener. Same top line, completely different fix.
The leading indicator nobody staffs for: response speed
Reply speed is the cheapest lever in the entire stack, and it is almost always a staffing problem rather than a messaging one.
Reported 2026 speed-to-lead data puts a 32% close rate on leads answered inside five minutes against 12% at 24 hours or more, and roughly 74% of teams miss the five minute window entirely. Average B2B lead response time is reported around 42 hours.
You do not need an AI setter to beat 42 hours. You need a rota and an alert. Log the timestamp of the inbound reply and the timestamp of your first human reply, then track the median and the 90th percentile. The median tells you if the rota works. The 90th percentile tells you what happens on Sunday night.
Your list rots while you sleep
Instagram lists rot differently from email lists. Handles get abandoned, businesses rebrand, accounts flip to private, and the bio link that qualified them six months ago now points at a dead Linktree. If your list is older than 90 days and you have not re-checked it, assume a meaningful slice of your sends are hitting nobody.
Log the scrape date on every prospect row. It costs nothing and it is the first thing you will want when reply rate drops on a source you have used for months.
The 48-hour trip wires
Four alerts. That is the whole system. Anything more and you stop reading them.
- Send failure rate over 3% on any account, rolling 24h. Pause that account. Do not restart it the same day.
- Visible inbox rate down more than 10 points versus the trailing 7 days. Stop scaling, audit the sending profile, check whether new accounts entered the pool.
- Reply rate on a single list source below half its trailing average, minimum 200 sends. Kill that source before you kill the script.
- Median first response time over 30 minutes for two days running. Staffing, not messaging.
Note the minimum sample on trip wire three. Firing an alert on 40 sends is how teams end up rewriting a script that was never broken.
The weekly review that takes 20 minutes
Open one view: cohort by week started, split by account age and list source. Then answer four questions in order.
- Did visible inbox rate hold? If no, nothing else in the review matters this week.
- Which list source moved most on reply rate, and did volume shift toward or away from it?
- What is the gap between reply rate and positive reply rate? A widening gap is a targeting problem wearing a messaging costume.
- Cost per qualified meeting, this cohort versus last. If it moved, which stage moved with it?
Twenty minutes, four answers, one decision. That beats a 40 tab spreadsheet nobody opens.
Build it in this order
If you are starting from nothing, do not build the dashboard first. Build the log.
Week one, capture the raw events: send attempted, send failed, reply received, reply sent, positive flagged, booked, showed, with timestamps and the account, campaign and list source on every row. Week two, add the derived rates and the four trip wires. Week three, add cohorting by week and source. Only then does a dashboard make sense, because only then does it have something honest to display.
Most agencies do this backwards, buy a reporting tool in week one, and end up with a beautiful chart of a number they cannot act on.