Peer Book a call Plug Peer in

AI Call Summaries Don't Record Deal History. They Rewrite It.

AI meeting summaries aren't neutral transcriptions. They're interpretations that compound into false deal records, and your forecast trusts them.

Three months into a rollout of an AI notetaking tool, a VP Sales I know felt good about pipeline hygiene for the first time in years. CRM notes were thorough, reps were compliant, and forecast calls moved faster. Then she pulled a deal that had slipped twice and compared every CRM note against the actual transcripts. The notes described a champion, clear business pain, and an agreed timeline. The transcripts showed a contact who said "I'd have to check with finance" on two separate calls, a competitor mentioned once and never followed up, and a closing question the rep asked that the prospect answered with a subject change. None of that was in the CRM. The deal had been forecast at 70% confidence for six weeks. It died in legal.

That is not a data-quality problem in the ordinary sense. The notes were not blank or sloppy. They were polished, coherent, and wrong in a specific direction.

AI Summarisers Are Editorial Tools, Not Administrative Ones

Every AI summariser makes choices. It decides what to foreground, what to compress, and what to drop. Those choices are not random. The models are trained on outputs that reward coherence and narrative progression, which means they systematically produce summaries that read like a deal is moving forward, because forward motion is what a coherent sales narrative looks like.

Ambiguity does not survive this process well. When a prospect says "that's interesting, though the timing isn't ideal," a human rep might flag that as a yellow light. The AI summary is more likely to record "prospect expressed interest and noted timing considerations." That is not false, exactly. But it has been processed through a frame that prefers resolution over unresolved tension, and what comes out the other side is a note that coaches, managers, and forecasters will read as softer concern rather than a stall signal.

Do this four or five times across a deal's lifecycle and the CRM record of that deal will have drifted substantially from what was actually said. Not through anyone's dishonesty. Through systematic optimisation for readable summaries.

The Compounding Problem No One Is Talking About

Single-call drift is annoying. Compounded drift across a seven-stage enterprise deal is a forecasting hazard.

Here is how it accumulates:

  1. Discovery call. Prospect raises a concern about incumbent vendor switching costs. AI summary: "prospect acknowledged current vendor relationship and is evaluating alternatives." The switching cost concern disappears.
  2. Technical demo. Their IT lead goes quiet for the last twenty minutes and asks only one question. AI summary: "IT stakeholder engaged with the technical walkthrough and raised a question on integration." Silence reads as presence.
  3. Commercial conversation. Prospect says "we'd need to get procurement involved, which takes time." AI summary: "prospect outlined internal procurement process as next step." A warning becomes a milestone.
  4. Champion call. Your champion says "I'm pushing for this but I don't have final sign-off." AI summary: "champion confirmed internal support and discussed approval pathway." The uncertainty is edited out.

By the time this deal hits a forecast call at Stage 5, the CRM shows a champion, a clear pain point, an engaged technical team, and an agreed procurement path. The manager asks a few questions, the rep sounds confident, and the deal goes in the commit column. None of the four actual risk signals appear in the notes your manager read this morning.

The Deal Health / Risk Scorecard is useful here, but only if the signals going into it are accurate. If your inputs are AI-smoothed narratives, the scorecard will reflect the narrative, not the deal.

Why Managers Are Coaching to Fiction

This matters beyond forecasting. When a manager reviews calls, they usually start with the CRM note, not the transcript. That is rational: a transcript of a forty-minute discovery call is 6,000 words, and a manager running eight direct reports cannot read 48,000 words of transcript a week. So they read the summary, form a view, and then either skip the transcript or skim it looking for confirmation.

That sequence, summary first then selective transcript review, means the AI's framing is the lens through which the manager interprets everything else. If the summary says the prospect is engaged, moments of hesitation in the transcript get read as normal sales friction. If the summary says there is strong business pain, the manager does not go hunting for evidence that the pain was actually vague and speculative.

Coaching conversations built on this foundation will address the wrong things. The rep gets feedback on their discovery questioning technique when what actually needs addressing is that they do not know how to re-engage a silent technical evaluator, because the summary said that evaluator was engaged and neither the rep nor the manager revisited it.

The Drift Test You Can Run This Week

This is the practical test for any enablement or RevOps team that has deployed AI notetaking. It takes about three hours and it will tell you exactly how bad the problem is in your environment.

Step one. Pull five closed-lost deals from the last quarter that were forecast at 50% or higher in the thirty days before they died.

Step two. For each deal, list every risk signal or negative indicator mentioned in the CRM notes across all stages. Count them.

Step three. Pull the actual transcripts for the same calls. Have a human reviewer (not the rep who ran the calls) identify every instance of ambiguity, objection, hesitation, competitor mention, stakeholder concern, or stall signal. Count those too.

Step four. Calculate the ratio. If your CRM notes captured 30% of the negative signals in the transcripts, that is your drift rate. In the teams I have seen run this, the number is usually somewhere between 20% and 40% capture. That means 60-80% of deal risk indicators exist only in transcripts that no one reads after the call ends.

If you want a structured way to track whether your AI outputs are actually reliable, the Human vs AI Scoring Agreement Checker was built for exactly this kind of calibration exercise.

What To Do With This

The answer is not to stop using AI notetaking tools. The administrative lift reduction is real, and asking reps to write their own CRM notes after every call has its own accuracy problems, mostly in the direction of optimistic omission rather than systematic positive framing, but the result is similar.

The answer is to stop treating AI summaries as the record and start treating them as a first draft that requires a specific type of review: one focused on what is missing rather than what is there. That means training managers to ask "what isn't in this summary" rather than "does this summary look complete." It means building into your pipeline review a standing question: when did we last look at the transcript, not the note?

It also means being honest that your forecast is currently built, in part, on a set of narratives that were optimised for coherence. The Win Rate & Forecast Accuracy Tracker will show you the output of that problem. The drift test above will show you the source.

The tool is not lying to you. It is just telling you a tidier version of the truth, and tidy is exactly what a stalling deal does not look like.

New posts

Get new posts in your inbox.

A fresh post most mornings. No digest spam, no course funnel, just the post, and one click to stop. Prefer a reader? Subscribe by RSS.

Alerts start right away. Unsubscribe any time.