// Guides · Data quality
Can you trust a CRM record the AI wrote?
Trust it the way you trust a citation, not the way you trust a calculator. Every value the AI wrote should link back to the sentence that produced it, and the system should be willing to leave a field empty when the signal is thin. Judge the tool on how much of the record you can check, not on the accuracy number it quotes you.
What changed when the machine started filling the fields?
For most of the years I spent building CRM software at HubSpot, the data quality problem had one shape: the fields were empty. A rep had not typed the close date, so the close date was blank. That was maddening if you were trying to run a forecast off it, and it was also useful, because a blank field is a piece of information. It told you the rep had not committed to anything yet, and you could read it off a pipeline report from across the room.
An AI-maintained CRM takes that signal away. Every field is full, every record looks tended, and the pipeline report is the best-looking one your team has ever had. The failure mode has moved rather than disappeared: it is no longer the missing value, it is the value that arrived complete, plausible and unattributed. A blank close date announces itself; a wrong one sits there looking exactly like a right one, and it will still be sitting there when somebody builds a forecast on top of it.
This is the part I would push back on when a vendor tells you their extraction is 95 percent accurate. Suppose it is. On a pipeline of two hundred open deals carrying six AI-written fields each, five percent is sixty wrong values, and nothing about the number tells you which sixty. That is the trouble with an error rate as a purchasing criterion: it describes the system in aggregate at the exact moment you need to know about one deal in front of you.
What can a transcript actually prove?
Not everything on a record is the same kind of claim, and lumping them together is what makes this question feel unanswerable. It helps to pull them apart.
| Kind of field | Can the source prove it? | Who should decide |
|---|---|---|
| Observation: a meeting happened, who was on it, that a competitor was named | Yes. It is in the artifact, word for word | The machine, unsupervised |
| Inference: the buyer sounds worried about implementation time | Sometimes. It is a reading of the artifact, not a quote from it | The machine proposes, a person glances |
| Judgment: stage, amount, close date, forecast category | No. It is a prediction the artifact does not contain | A person, every time |
The rule that falls out of that is short. The further a field sits from the words actually spoken, the more a person needs to be standing next to it. And the fields carrying the most revenue, the ones that end up in a board deck, sit furthest away of all, which is an uncomfortable arrangement given they are also the fields everybody would most like to automate.
There is a second property that matters as much as provenance and gets talked about far less. A system that always fills the field is telling you nothing by filling it. If it sometimes declines, if "we could not tell from this call" is an outcome the software is allowed to produce, then a filled field starts carrying information again. Abstention is what makes the rest of it mean anything.
What happens when a later conversation contradicts an earlier one?
Here is the case I would put to any vendor, because it is completely ordinary and it separates the systems in about a minute.
On a call in early September the buyer says: "We're probably going to revisit this in Q4."
A manual CRM does nothing with that. The rep means to update the deal, gets pulled into the next call, and the record still says the deal closes on September 30. The failure is at least visible: by mid-October the close date is in the past, and anyone scanning the pipeline can see something is off.
A naive AI CRM writes the close date to December 15 and moves the stage to something like Delayed. Both are reasonable readings of the sentence. Neither is actually in it. The word doing the most work in that quote is "probably", and it has just been quietly resolved into a date that will be treated as a fact by everyone downstream.
A system built for checking proposes the same change, shows the sentence it came from, and leaves the hedge where the rep can see it, so the answer can be yes, no, or "she said probably, leave it where it is and set me a reminder."
Then two weeks later the champion emails to say budget freed up after all and can we pick this back up. Watch what the naive system does with that. It writes the close date back to something in October, and the record now reads as though the wobble never happened at all. The current value is right and the history is gone. But the fact worth having here was never the date. It was that this buyer moved twice in three weeks, which is exactly the sort of thing that decides whether you call the deal on a forecast. A record that keeps only the latest value has thrown away the signal and kept the noise.
How do you audit a week of AI-written records?
You do not need a program for this. You need an afternoon and ten deals.
- Pick ten recently closed deals, five won and five lost. Closed rather than open, because you already know how each one ended and can tell a good reading from a lucky one.
- On each deal, list the fields the machine wrote rather than a person. If you cannot tell which is which, stop there. That is your finding, and it is a serious one.
- Try to get from each value to its source in one click. Count the times you can. This is the whole exercise, and everything else is commentary on it.
- Where the source language was hedged, check whether the hedge survived. "Probably", "we might", "I'd need to ask" are the words that get resolved into confident values, and they are where the expensive mistakes live.
- For anything that turned out wrong, ask who would have caught it. If the answer is nobody until the forecast call, that field needs an approval step rather than better prompting.
The number that comes out of this is not an accuracy rate. It is the share of AI-written values you could check at all. I would want that above eight in ten before I cared much what the accuracy was, because underneath that the accuracy is unknowable and therefore not a number anyone can use.
What does this look like on a small team?
The numbers below are illustrative, chosen to show the proportions rather than measured from any particular team, and yours will differ.
Take five sellers with twenty-five active deals between them. In a normal week the system writes roughly a hundred and forty changes across those records. Around a hundred and ten are observation: this email belongs to this deal, this meeting happened, this person is new and works here. Twenty-five or so are inference: a topic, a competitor named, the temperature of a thread. Five are judgment, which in practice means a few stage moves and a couple of close dates.
So the risky slice is about four percent of the volume. That proportion is the useful thing to carry away, because it rules out both of the tempting responses. Blanket suspicion is unaffordable; nobody is checking a hundred and forty things a week. Blanket trust is how the five get through. What you actually need is to look at five, and to have the other hundred and thirty-five linked to something, so that on the day one of them looks strange you can find out why in a few seconds instead of going through somebody's inbox.
What to look for in a tool
Every vendor in this category will tell you there is a human in the loop. The questions below are the ones that establish what that actually means in the product, and I would ask them with a real deal on the screen rather than a demo record.
- Can you get from any AI-written value to the message or call behind it in one click, and what does that look like when the source is a forty-minute recording rather than a two-line email?
- Does the record show whether a person or the machine last touched each field, and what the value was before?
- Is there any circumstance in which the system declines to fill a field, and can you see the list of what it skipped?
- When two conversations disagree, does the later one simply overwrite the earlier, or does the disagreement itself become visible?
- Which fields can it change with nobody approving? Ask for that list in writing. It is the shortest honest description of how much you are delegating.
We built Ahoy so the machine does the capture and a person keeps the judgment: email, calendar and meetings land on the right record by themselves, while a change to stage, amount or close date is prepared with the evidence behind it and held for a one-tap approval rather than written quietly. A do-not-track list keeps the addresses and domains you would rather it never read out of the sync entirely. If you want to know how your current records would score on the ten-deal test above, the free CRM audit takes 45 minutes, walks your pipeline with you, and ends in a written summary you keep. It is a conversation, not an install.
Frequently asked questions
Can you trust AI-generated CRM data?
For observation it is reliable, for judgment it is not, and most of the disappointment comes from treating those as one thing. Activity capture, record matching and meeting attribution are checkable against an artifact and are generally right. Stage, amount, close date and forecast category are predictions no transcript contains. Trust the first group by default, require a person on the second, and insist every value links back to its source so you can settle any individual case yourself.
How do you check whether an AI CRM got a field right?
Take ten recently closed deals, half won and half lost, and list the fields the machine wrote rather than a person. Try to get from each value to the email or call behind it in one click, and count how often you can. That share, rather than the vendor's accuracy percentage, is the number that matters, because an error you cannot locate is one you cannot fix.
What is provenance in a CRM record?
Provenance is the link from a value back to the evidence that produced it: this close date came from this sentence, in this email, on this date. It is the difference between a record that asserts something and a record that can show its work. Without it, a field the AI read correctly and a field the AI invented look identical, and a rep checking a deal five minutes before a call has no way to tell them apart.
Which CRM fields should AI never change on its own?
The ones carrying a prediction rather than a fact: deal stage, amount, close date and forecast category. Nothing in a transcript establishes them, and they are the fields that flow into the forecast and the board deck, so an error in them travels furthest before anyone notices. The workable pattern is that the AI prepares the change with its evidence attached and a person commits it.
What happens when a later conversation contradicts an earlier AI update?
In many systems the newer value simply overwrites the older one and the record reads as though the change of direction never happened. That is the wrong outcome, because a buyer who moves twice in three weeks is telling you something about the deal that neither value captures on its own. Ask a vendor to show you a field that reversed, and see whether the earlier reading, the later one and the reason for each are all still there.
Related guides: How do you stop logging CRM activity by hand? · Which deals are at risk before the forecast call? · Why don't sales reps update the CRM? · CRMs that update themselves · All guides