← All resources

Personalization at scale: what works, what is noise

Quick answer

Personalization that works is about relevance, not about proving you looked someone up. Reference a real trigger tied to why you are reaching out, right now. Skip the compliments and the personal trivia. Tier the effort by account value, automate signal detection rather than warmth, and keep a person on the top accounts and every reply. Relevance scales. Flattery does not.

The personalization myth

I'm Hlib Storchak. I build and run outbound systems for B2B founders and sales teams, and this is one of the few topics where the common advice actively costs people meetings.

Somewhere along the way, personalization came to mean mentioning a detail about the person. Complimenting a post, noting their city, referencing their alma mater. Buyers see through it instantly, because none of it changes the reason you are writing. It is effort they can see, spent on something they did not ask for.

The second problem is that at volume, visible effort becomes a tell. When a hundred people receive "loved your post on X", the pattern is obvious to anyone who compares notes. Templated warmth is worse than no warmth, because it signals that the sender is running a machine and hoping you will not notice.

Relevance beats flattery

The message that lands says something true about their situation and connects it to a problem you solve. That is relevance. It answers the only question a cold reader actually has: why are you talking to me, and why now?

Relevance has two halves and most teams only do one. The first is that the observation is true about them specifically. The second is that it leads somewhere: it has to connect to what you do, or it is just an interesting fact you found. A trigger with no bridge to your offer reads as research theatre. A pitch with no trigger reads as spam.

Signals that work

The signals worth building on are the ones that change someone's priorities. Hiring for a role that implies your problem. A funding round that raises the pressure you relieve. A tech stack change, a leadership move, an expansion into a new market, a public complaint about the thing you fix.

These work because they are reasons to reach out, not decorations on a reason you already had. A hiring signal tells you the team is under-resourced in a specific area this month. That is a different message from the one you would send the same company a year earlier. I went deeper on sourcing and wiring these in my piece on turning buying signals into meetings.

One caveat: signals decay. A funding round is interesting for a few weeks and stale after a quarter. If your list build and your send are six weeks apart, the trigger you paid for is no longer a trigger. Freshness is part of the signal's value, and most stacks lose it in the handoff.

Personal touches that are noise

Complimenting their podcast, mentioning their marathon, praising a post. These signal effort but not relevance. At scale they read as templated effort, which is more off-putting than no personalization at all.

The deletion test. If the personalized line could be deleted without changing why you are writing, it is decoration, not relevance. Cut it and see whether the message still makes sense. If it reads better without, you did not personalize, you padded.

Tiers of personalization

Think in tiers and match the tier to what the account is worth. Tier one is deep manual research for a short list of top accounts. Tier two is signal-based lines applied across a segment that genuinely shares a trigger. Tier three is a clean, relevant, non-personalized message to a well-chosen list.

Tier three is the one people are embarrassed about and should not be. A tight segment with a message written for exactly that segment outperforms a broad list with fake-personal openers, because the relevance lives in the targeting instead of the first line. Most of the lift people attribute to personalization is really segmentation, which is why segmenting a B2B list properly is usually the higher-return fix.

How to write the relevant line

A structure that holds up across segments: state the observation in plain language, name what it usually means for someone in their position, then ask about it. Three short sentences, no adjectives, no flattery.

Two failure modes to avoid. Do not explain the research you did, because the reader does not care how you found it. And do not overstate what the signal means: "I saw you are hiring two AEs, usually that means pipeline coverage is the constraint" is honest, while asserting you know their problem is not. Hedged and specific beats confident and generic.

What to automate

Signal detection and insertion automate well. Pulling a hiring post or a funding event, mapping it to a segment, and wiring it into a sentence is exactly what tooling is for, and it holds quality at volume because the underlying fact is real.

What does not automate is the judgement about whether the signal applies. A tool will happily find a hiring post for a role that has nothing to do with your offer and drop it into the opener anyway. Build the filter before you build the automation, and check a sample of the output by hand every time the list changes.

Where AI helps and where it makes slop

AI is good at reading a page and extracting a fact, summarising what a company does, and turning a structured signal into a sentence that does not read like a mail merge. Those are genuinely useful and they save real hours.

AI is bad at manufacturing warmth. Ask a model to sound friendly and interested with nothing real to work from and you get fluent, confident, empty text, at volume, under your name. That is the slop everyone is complaining about, and it is a briefing problem more than a model problem. I looked at the evidence on this in whether AI personalization actually lifts reply rates. The short version: it lifts when it is fed real signals and hurts when it is asked to invent rapport.

What to keep human

The top accounts and every reply. A person should review the highest-value messages before they go, because those are the ones where a wrong detail costs you a relationship rather than a send.

Replies matter more. Automation gets you to the inbox; judgement closes the loop. The pattern I see most often when I take over an account is a well-built sequence with a neglected inbox: people replied, nobody answered fast or well, and the pipeline leaked at the one point where a human was unambiguously required.

Works vs noise

ApproachLifts replies?Scales?Effort per account
Trigger-based relevanceYesYesLow once wired
Segment pain pointsYesYesNear zero
Deep account researchYes, for top accountsOnly in tier oneHigh
Post complimentsRarelyPoorlyMedium
Personal triviaNoNoMedium
AI-written rapport with no signalNo, often negativeYes, unfortunatelyNear zero

Everything that works is tied to why you are reaching out. Everything that is noise is tied to proving you did homework. Note the last row: the cheapest option to scale is also the one most likely to hurt you.

What a research minute actually costs

Personalization debates go in circles because nobody prices the effort. Here is a model you can run in ten minutes. Every input is an assumption I am showing you deliberately, not a researched figure, so replace all of them with your own before you act on the output.

Assumptions. Tier one deep research takes 10 to 20 minutes per account. Tier two signal lines take 1 to 3 minutes once the pipeline exists. Tier three takes zero. The fully loaded cost of whoever does the work is 25 to 50 per hour. Baseline positive reply rate on a clean list is whatever yours is: use your own number, and run the model per 100 contacts.

Tier (assumptions, swap in your own)Minutes per accountCost per 100 contactsWhat has to happen to pay for itself
Tier 1, deep research10 to 20~420 to ~1670Needs high deal value, so reserve it for named accounts
Tier 2, signal line1 to 3~40 to ~250A small lift in positive replies usually covers it
Tier 3, segment message00Nothing, this is the baseline you compare against

Run it with your own average deal value and your own positive reply rate and the answer usually comes out the same shape: tier two pays for itself easily, tier one only pays on accounts big enough to justify an hour of someone's attention, and the mistake most teams make is running tier one effort on a tier three list. That is the expensive version of being busy.

How to test whether it pays

Run the same segment, same offer, same sequence, and vary only the personalization tier. Anything else you change at the same time makes the result unreadable. Send both versions in the same week, because deliverability and market conditions drift.

Measure positive reply rate, not reply rate, and not opens. A personalized opener can lift replies while attracting more polite declines, which looks like a win in the dashboard and is not one. Give it enough volume that a handful of replies cannot swing the result, and read it per segment rather than blended.

Scaling without slop

Scale the signals, not the flattery. A segment that genuinely shares a trigger can get a relevant message at volume, because the relevance is real for every one of them. That is how you stay personal at scale.

Practically, that means investing in the list and the signal pipeline rather than in the sentence generator. A well-built segment with one true shared trigger will beat a broad list with clever openers, every time, and it costs less to run once it exists.

The rule I use

Every message has to pass one test: does it say something true and relevant about their situation, and does that thing lead to why I am writing? If yes, it can scale. If it is just proof I looked them up, it gets cut.

After 2000+ meetings booked for B2B clients, relevance has always outperformed decoration. Not marginally, and not occasionally. The campaigns that worked were the ones where I could have shown the prospect exactly why they were on the list and they would have agreed it was reasonable.

Key takeaways

  • Personalization means relevance, not proof that you researched someone.
  • Trigger and pain signals lift replies and scale. Compliments and personal trivia are noise at volume.
  • Tier the effort by account value, and never run tier one research on a tier three list.
  • Automate signal detection, not warmth. AI helps when fed real signals and hurts when asked to invent rapport.
  • Most of the lift people credit to personalization is really segmentation.
  • Test by varying only the tier, and measure positive reply rate rather than raw replies.

FAQ

Does no personalization ever work?

Yes. A clean, relevant, well-targeted message with no personal line regularly outperforms a fake-personal one, because the relevance lives in the targeting. If the segment is tight enough that one message is true for everyone in it, you do not need a custom opener.

Can AI personalize well?

It is good at extracting a fact from a page and turning a structured signal into a natural sentence. It is bad at manufacturing warmth from nothing, which produces fluent empty text at volume. Feed it real signals and constrain the output, or do not use it for this.

How much research should I do per account?

Match it to account value. Deep research for a short named-account list, signal-based lines for the middle, tight segmentation for the rest. Price the minutes and compare against your average deal value before you commit to a tier.

Is mentioning a recent post always bad?

No. If the post is the trigger, meaning it reveals a priority or a problem you address, referencing it is relevance. If you are complimenting it to show you did homework, skip it. Apply the deletion test.

What is the fastest way to add real relevance?

Pull one buying signal per account and reference it in the first line, then make sure it connects to your offer. It is the highest-return personalization there is, and it costs a few minutes per account once the pipeline exists.

Want this run for you?

I build and run outbound that books meetings, and leave you the system to keep.

Book a call