Home / Evidence

What actually moves
cold outbound.

Six claims we run the engine on. Each is sourced, and each is inconvenient for somebody. Disagree with one and that is a useful call.

Last reviewed 27 August 2026 · figures are published industry benchmarks, not Miles results

3–4× lift from account selection Unify · 25M+ outbound emails

Choosing the right accounts beats rewriting the email, by three to four times.

Unify ranked four variables across 25 million outbound emails. Cohort selection moved reply rate 3 to 4 times. Rewriting the call-to-action moved it 1.3 times.

Almost nobody spends their week on the top of that list. A rewritten subject line is a finished artefact by lunchtime; rebuilding an account universe takes a fortnight and produces a spreadsheet.

We re-sort every account each morning on fit, signal and recency. The list is the product, and the copy is downstream of it.

26% below no personalization at all Unify · 25M+ outbound emails

Personalizing on job title is worse than not personalizing at all.

Merging in a role performs about 26% worse than sending the same email with nothing personalized. It is the clearest negative result in the data, and almost every template on the market does it.

The mechanism is simple. "As a VP of Engineering, you probably..." announces the mail merge in the first line.

It also tells the reader you know nothing about them except their title. Personalization only pays when the detail was clearly hard to find.

One real variable beats two synthetic ones. Tier 1 gets a researched line about something that actually happened. Everyone else gets a clean segment email with no fake intimacy.

5% the reply rate where research stops paying Unify · cross-checked against our own send economics

Deep personalization is a losing trade below a 5% reply rate.

Research-heavy sending lifts replies, and it costs real time per prospect. Below roughly a 5% expected reply rate the arithmetic inverts: research per booked meeting costs more than the meeting returns.

This is why "personalize everything" is advice sold by people who do not carry the delivery cost. Spend research where the prior is already high, and stop everywhere else.

We tier every list before writing anything. Deep research goes to accounts carrying a live signal. Tier 3 gets a clean segment email and is then left alone.

Held the only meeting we count Mechanism, not a benchmark

Optimizing reply rate suppresses qualified meetings.

Reply rate is easy to move in the wrong direction. Curiosity-bait openers and "are you the right person?" both lift replies, and neither produces pipeline.

An agency reporting reply rate is reporting the metric it can most easily flatter. The honest one is qualified meetings held, which is slower, noisier and much harder to fake.

We bill in held meetings with a decision-maker present. Not replies, not leads, not booked calls that no-show. That choice costs us the easiest number to win on.

Week 6 when the window actually opens Mechanism, verifiable in your own reply data

The best time to reach a funded company is week six, not week one.

A funding announcement triggers a stampede. Every vendor in the category emails the founder within days, and the inbox is unusable for a fortnight.

The useful window opens later. By week six the noise has cleared and budget has moved from announced to allocated.

The executives hired with the round have also started. A VP in their first month is the most motivated buyer in the company, because they are looking for something to change.

We treat a funding round as the start of a watch window, not a send trigger. The round tells us who to watch. The hire that follows tells us when to write.

Sept 2025 when Google deleted the score Google Postmaster Tools v1 deprecation notice

Google deleted the domain reputation score. Dashboards still report it.

Google removed the Domain Reputation and IP Reputation panels when the v1 Postmaster dashboard was retired on 30 September 2025. The v1 API shut down entirely by the end of that year.

Their stated reason: the low, medium, high and bad labels were routinely misread and lagged real sending behaviour by weeks.

The replacement reports spam complaint rate, authentication status, delivery errors and compliance. Those are the numbers that now decide whether mail lands.

If a deliverability report still shows a reputation grade, ask where the number comes from. We report complaint rate, delivery rate and authentication status, because those are what Gmail still acts on.

If one of these is wrong,
we would rather know.

Every claim here is checkable, and every one of them changes what we do on a Tuesday. Bring the disagreement to the call.

Book a call See what you receive