Outbound 8 min read

Ninety Days of AI-Generated First Touches: What We Learned

Marcus Dele
Marcus Dele Head of Product

When we shipped the first-touch drafting feature in Leadbay, we made a bet: if you give the model a specific buying signal and the rep's actual sent-email history, the output will be good enough to send with minimal editing. That bet took ninety days to test properly. Here is what we found, including the parts that did not work.

We are not sharing reply rate percentages here. Sample sizes from a small team working with early-access users do not produce numbers that are meaningful to generalize. What we are sharing are the patterns we observed: which drafts got edited heavily before send, which got sent nearly as-is, which generated positive replies, and which got ignored.

The signal quality problem

The most predictive variable for draft quality was not the sophistication of the model. It was the specificity of the input signal. Drafts anchored to a specific, recent signal ("posted 5 backend engineering roles in the last 3 weeks," "news coverage of new data center expansion") were consistently better than drafts anchored to a general ICP observation ("growing mid-market SaaS company").

This makes sense in retrospect. The model is doing pattern-matching against the rep's past sent emails. A rep who writes good cold emails always opens with something specific: a detail that demonstrates they actually paid attention to the account. When the signal is specific and recent, the model has material to work with. When the signal is "they look like a good fit," the model defaults to the same openers you would get from any generic AI email tool.

The practical implication: the Leadbay draft is only as good as the highest-confidence signal attached to the account. This is not a criticism of the approach, it is a constraint that every AI drafting tool shares. The ones that pretend otherwise are generating generic output and calling it personalized.

What reps edited and why

We tracked the edit patterns across first-touch drafts. Three patterns showed up consistently:

First, reps adjusted the ask. The model would often draft a soft discovery ask ("Would you be open to a quick call?") when the rep preferred a more specific framing ("I have 15 minutes on Thursday at 2pm. Want to connect?"). This is a voice preference, not a model failure. We made it easier to specify ask format as a user preference.

Second, reps adjusted signal framing when they had additional context the model did not. A rep who had already spoken to someone at the account, or who remembered a conversation with a contact at that company a year prior, would reframe the opening line to acknowledge that history. The model does not have access to that context. It is writing cold; the rep knows it is warm. The right behavior there is to edit, and reps did.

Third, when the draft referenced a signal the rep felt was too obvious or had been referenced too often in their territory, they would swap it out for a different angle. "Everyone is sending emails about job postings right now," one rep told us during testing. Fair point. The model does not know your competitive context or what messaging your territory has already been saturated with.

Where the drafts were good enough to send as-is

Drafts performed best in a specific scenario: the account had multiple active signals, the rep had a consistent writing style in their sent history, and the ask was something the rep sent often enough that the model had a good sample of how they phrased it. Under those conditions, the draft-to-send rate with minimal or no editing was notably higher.

The voice matching was the part that surprised even us. Reps who write short, direct emails got short, direct drafts. Reps who tend toward a more conversational opener got drafts that matched that register. We did not expect the model to pick up on sentence rhythm patterns from a 30-email history, but it does. Several reps said independently, in early feedback, something to the effect of "this sounds like something I would actually write." That is the outcome we were trying for.

Where the drafts needed more work

The model underperforms in three situations. One: accounts where the only available signal is weak (a single news mention from six months ago, a single web visit). Two: reps with inconsistent or sparse sent email histories, where the model does not have enough signal to calibrate tone. Three: highly technical buyers who expect precision in the opening line. A draft that says "I noticed Coro Systems is expanding their engineering team" works fine for a VP of Sales. It does less work for a CTO who expects you to know which team is expanding and why it is relevant to your product.

For the third case, the right fix is not better AI. It is better research. We are not trying to replace the rep's judgment about which accounts warrant deep research. We are trying to eliminate the research tax on accounts where the signal is clear and the outreach decision is obvious.

The prep time question

One of our original goals was to reduce per-account research time. Before Leadbay, reps in our test group reported spending anywhere from 8 to 20 minutes per account before writing a first touch. That time went to: finding something current and relevant about the account, figuring out the right angle, and then drafting the message.

Leadbay collapses the first two steps. The signal is surfaced automatically. The angle is embedded in the draft. What reps are left with is reviewing the draft, editing where needed, and hitting send. That review-and-edit cycle averaged under 3 minutes for drafts that eventually got sent.

We are not claiming that is a universal truth. It is what we saw in our own use. The variable that matters most is signal richness in your account base. If your ICP naturally generates a lot of observable signals (hiring, news, web activity), the prep compression is real. If your accounts are quiet, the model will tell you that too. The honest answer sometimes is "this account does not have enough signal to make outreach worth it today."

What we are building next

Ninety days in, the thing we want to improve most is context persistence. The model does not currently know if a rep has already reached out to an account, what the response was, or how many touches deep they are. That means a draft for a second touch in a sequence looks the same as a draft for a first touch. Reps catch this, but it is extra work. We are building conversation history into the context window so subsequent touches feel like they are aware of the prior sequence.

The other area is signal interpretation. Right now the model reports what the signal is. What it does not do is explain why that signal is meaningful for this specific account, given what the rep's closed-won history looks like. A rep with a history of winning deals where the trigger was a technology stack change should get a draft that explicitly connects that pattern to the current account. That requires the scoring engine and the drafting model to share more context. That integration is in progress.

More from the Leadbay blog

Back to blog