Reply Latency: The Inbox SLA That Decides Whether Instagram Replies Turn Into Calls
Your reply rate is fine. Your reply latency is not. The research on first-response windows, why Instagram decays faster than a web form, and a tiered inbox SLA your team can actually hit.
You spent three weeks rewriting the opener. You tested four hooks, two offers and a shorter CTA. Reply rate moved from 6% to 9%. Then those replies sat in a shared inbox for eleven hours because the setter who owns that account clocks off at 6pm and the prospect messaged at 9:40pm.
That is the leak. Not the copy. The gap between a reply landing and a human answering it is the single most under-measured number in Instagram outreach, and it is usually the cheapest one to fix.
The research everyone quotes, and what it really says
The famous number comes from the Lead Response Management study run out of MIT with InsideSales: contacting a lead inside five minutes makes you roughly 21 times more likely to qualify it than waiting 30 minutes, and the odds of making contact at all are about 100 times better in that first five-minute window.
Those numbers are old and they are about phone calls on web-form leads. Treat them as direction, not as a law of physics. The 2026 benchmark reporting is more useful because it is closer to how you actually operate:
- Median B2B first response to an inbound lead sits around 42 hours. Some samples put the mean near 47 hours, dragged up by a long tail.
- Roughly 35% of leads wait more than a full day for any answer.
- One study of 1,000 companies found 63.5% never replied at all.
- Teams that answer inside five minutes convert at roughly 21%. Teams that answer after a day or more convert at roughly 2.3%.
A cold DM reply is a weaker signal than a demo request, so do not expect a 21% booking rate from it. What ports cleanly is the shape of the curve. Interest decays fast, and it decays fastest in the first hour.
Instagram makes the decay worse, not better
A web form is a deliberate act. Someone typed an email address and pressed submit. They half expect a slow reply because that is what forms do.
A DM reply is not deliberate. It is a thumb twitch at 11pm on a phone that is already showing three other conversations. Two things follow:
The expectation is tighter. Reported consumer expectations for Instagram DMs cluster at one to three hours, and that window has been tightening. Anything past three hours now reads as slow. Meanwhile the typical business average on Instagram DMs is reported north of 10 hours.
The context evaporates. Your prospect replied "what is this?" to a message they have already forgotten. Answer in four minutes and you are still inside the same scroll session. Answer at 9am the next day and you are a stranger restarting a conversation they never committed to.
Where your latency actually comes from
Almost nobody is slow because their setters are lazy. They are slow because of structure. Audit these five in order.
| Source of delay | Typical cost | Fix |
|---|---|---|
| No push notification on sender accounts | 2 to 8 h | One device or session per pod with notifications on, or an aggregator that pings a shared channel |
| Replies pooled in one inbox with no owner | 1 to 4 h | Assign every sender account to a named setter with a backup |
| Setter coverage matches your timezone, not the list's | 8 to 14 h | Route accounts by prospect timezone, not by where your team sits |
| Weekend blackout | Up to 60 h | One paid weekend shift covering two check-ins per day |
| Setter has to open a doc to answer a pricing question | 10 to 40 min per reply | A one-page answer bank pinned in the setter's workspace |
The timezone one is the expensive invisible killer. If you are selling into US Eastern from Manila or Warsaw, your setters are asleep for the whole American evening, which is exactly when people answer DMs. Around 56% of inbound in service businesses arrives outside 8am to 5pm weekday hours, and evening and weekend leads carry the highest drop-off of any cohort.
The SLA that is actually achievable
Do not promise five minutes across 24 hours. You will miss it, stop measuring it, and be back where you started. Tier it by reply type, because not every reply deserves the same urgency.
| Reply type | Target first response | Hard ceiling | Who owns it |
|---|---|---|---|
| Positive or question ("what do you do?", "how much?") | Under 10 min in covered hours | 2 h | Assigned setter |
| Neutral or ambiguous ("hey") | Under 30 min | 4 h | Assigned setter |
| Objection ("not interested right now") | Under 2 h | Same day | Assigned setter |
| Anything landing outside covered hours | Under 15 min of shift start | 12 h | Opening shift |
| Hard no or abuse | No reply, tag and close | n/a | Automated tagging |
Two rules make it stick. First, the clock starts when the message lands on Instagram, not when your tool syncs it. If your aggregator polls every 20 minutes, that is 20 minutes of your budget gone before a human sees anything. Second, measure the median and the 90th percentile, never the mean. A mean of 45 minutes can hide a third of your replies sitting for six hours.
Coverage math, done honestly
Work out how many hours you need before you hire.
Take your daily sends, multiply by reply rate, and that is your daily reply volume. At 400 sends a day and an 8% reply rate you get 32 replies. Assume 6 to 9 minutes of handling per conversation thread across the day including follow-ups, and that is roughly 3 to 5 hours of actual work. That does not mean three hours of coverage. It means someone has to be present across a 14 to 16 hour window doing three to five hours of work inside it.
That is the trap. Reply handling is low-volume and high-latency-sensitive, which is the worst possible shape for scheduling. Your options, in rising order of cost:
- Notification-driven part-time cover. One setter on call across a long window, paid for availability plus volume. Cheapest, works up to about 40 replies a day.
- Split shifts across timezones. Two setters, one covering the prospect morning and one the prospect evening. Necessary once you are past roughly 60 replies a day or two timezones.
- AI first-touch with human takeover. An assistant acknowledges and asks one qualifying question inside seconds, then hands to a human. This buys you the first-response window without buying 16 hours of headcount.
Option three is where most operators are landing in 2026, and it is also where most of them get burned. The rule that holds: the AI may acknowledge, ask one clarifying question and offer a time. It must never negotiate price, handle an objection, or claim to be a person. Anything past the second message goes to a human. A bot that oversteps costs you more than a slow reply does.
What to log
Add four fields to whatever you already track:
- Reply timestamp (platform time, not sync time)
- First human response timestamp
- Latency bucket (under 10 min, 10 to 60 min, 1 to 4 h, 4 to 12 h, 12 h+)
- Outcome (booked, dead, no answer)
Then run one query a week: booking rate by latency bucket. You will get your own version of the curve above within about three weeks of data, on your list, in your niche. That number beats every benchmark in this article, including the ones with a citation attached.
If your under-10-minute bucket does not outperform your 4-to-12-hour bucket, you have a different problem, and it probably is the offer.
The week-one rollout
Monday: turn on notifications for every sender account and assign each one a named owner. Tuesday: instrument the four fields. Wednesday: publish the SLA table above to your team with the ceilings, not just the targets. Thursday: pull your last 200 replies and bucket them by latency so you know your starting median. Friday: schedule one weekend check-in block.
None of that costs money. Most teams find their median first response sitting somewhere between four and nine hours, and cut it to under an hour with scheduling alone. That is usually worth more than the next three copy tests combined.