Most of the calls I get these days open the same way: A senior marketer at a B2B SaaS company wants to know how they can show up more often in AI search.
It's the wrong question. Or at least the easy half of one.
The brands actually getting hurt right now are easy to find. The model mentions them but says one thing that quietly kills the deal. A head of growth told me last week that the way they get talked about isn't the way they'd like to be talked about, because the company's old positioning was still the version the LLMs repeated back.
We call this Factual Drift and it's costing companies deals that are hidden away in the dark funnel.
Here's the whole method we use to find it and a version you can run on your own brand this week.
The challenge: factual drift
When a buyer asks an AI engine about you it doesn't read your homepage and stop. Google's own documentation describes generative search using retrieval and query fan-out, issuing multiple related searches across subtopics to build a single answer. So as far as the model is concerned your brand doesn't live on your homepage. It lives across every surface the model can reach.
Your site, your docs, your G2 page, a directory listing, an old comparison post, a Reddit thread from eighteen months ago.
Most of the time those sources don't fully agree. Your docs are current, your marketing site is half a rebrand behind, the directory still has the old category label and a comparison article you forgot existed still ranks.
When the sources disagree, the model doesn't pause and ask you which version is correct. It guesses. It picks whichever fact appears most often, or most confidently, and repeats that to your buyer as if it were settled.
This matters more than people expect because so little of the input is yours to begin with. When we run a FACTUAL audit, the wrong answers almost always trace back to surfaces the brand doesn't own. Reviews, directories, comparison pages, community threads. Your own site is a minority shareholder in the story AI tells about you.
A whole layer of our work is mapping which third-party sources the models actually pull from in a given category and how a brand gets framed once it shows up in them. We publish some of our research on this on the website. And the lesson keeps repeating: whether your surfaces agree is what decides which version of you the model trusts. Get the same facts showing up everywhere the model reads or it settles the tie with whatever appears most.
So the real problem is consistency. The facts AI uses to describe you have drifted out of sync and the model papers over the gaps with its best guess.
Common misconceptions
Two things I hear constantly when I raise this. Both are reasonable but can cost money.
"It's a visibility problem"
Visibility and accuracy are different measurements, and people collapse them into one.
Visibility tells you if you get mentioned. Accuracy tells you whether the mention helps you win. You can be in 80% of the answers and lose every one of them on the framing.
Recently, a head of growth at a mid-sized Series A company (around a $250m valuation) told me the way they're talked about isn't the way they'd like to be talked about. They had rebranded to move upmarket six months earlier. Six months on, the LLMs were still pulling the old "easy to use" positioning, the exact framing they'd spent two quarters trying to leave behind. The model mentioned them fine. It just described last year's company.
That's not a visibility gap. Their visibility was good. It's an accuracy gap, and no amount of "show up more" fixes it.
"Our website is accurate, so we're fine"
Your website being accurate is necessary and nowhere near sufficient.
The model reads around twenty other surfaces before it answers, and it weights the consensus across all of them, not your homepage alone. Your site is one vote. If your G2 profile, your old comparison pages and a stale partner listing outvote it, the model goes with the crowd. A perfect homepage loses to four mediocre third-party pages saying the same wrong thing.
The business impact
The worrying part about this is silent disqualification. Your prospects' conversations with AI are happening in the dark funnel. There is no feedback mechanism telling you when or why a buyer disqualified you from the shortlist.
Take an integration for example: A buyer asks an AI which tool in your category works with their stack. The model, reading stale information, says you don't integrate with the thing they need but you do. That buyer has already crossed you off and moved to the next name on the list. Sales never logs the loss because the conversation never started.
Now take a security-conscious buyer: They ask whether you're SOC 2 compliant and you are. But the proof lives inside a gated PDF the model can't read, so there's no clear, crawlable answer anywhere it can reach. The model hedges, or worse, says it can't confirm. To a buyer who treats compliance as a hard gate, "can't confirm" reads as "no." They quietly drop you, and again, nobody on your side ever hears about it.
Both buyers were qualified and winnable. And both decided on a wrong or missing fact before anyone at your company knew they existed. The cause is usually a directory listing or an old review nobody has looked at in a year so the pipeline you lose this way is both invisible and almost impossible to trace back.
This is what FACTUAL is built to surface before a buyer does.
The FACTUAL framework
FACTUAL by Discovered Labs is a seven-step audit for working out whether AI systems describe your brand with accurate, current, consistent facts, and fixing the sources when they don't.
Seven steps: Form, Ask, Compare, Trace, Unify, Anchor, Loop. Here's what each one involves, the mistake people make on it and what it looks like in practice.
F: Form the factbase
Write the verified source of truth before you test anything. Who you serve, your three real differentiators, your three late-stage deal criteria, your real competitors (the ones you actually lose to) and your approved facts on pricing, integrations and compliance.
The common mistake is writing aspirational positioning instead of what your sales team says on calls. And if your factbase says you serve mid-market and enterprise, that single line is what every later score gets judged against, so when the model calls you "best for small teams" you've got a documented deviation, not an opinion.
When we first start working with a client at Discovered, we build a knowledge base on the company by indexing their entire website, consuming first-party information provided to us (e.g. call transcripts) and building a set of documents that we update regularly: company overview, writing guidelines, etc.
A: Ask the engines
Build a buyer-grounded prompt set and run it across ChatGPT, Claude, Perplexity, Gemini and Google AI. Three runs of each prompt per engine, because the variance between runs is itself the signal.
The common mistake is asking generic prompts ("what is [brand]") that no buyer types, which produces clean answers and a false sense of safety. A real prompt looks like "best [category] tool that integrates with Salesforce for a 200-person RevOps team," because that's the question that decides a shortlist.
We recently wrote about our measurement methodology for AI search. We believe companies are being fooled by randomness.
C: Compare the claims
Score each answer against your factbase. Accuracy, completeness, freshness, framing quality on a 1-5 scale and whether competitors are winning on criteria you should own.
The common mistake is checking only for falsehoods and missing the omissions, which are often more damaging. If the model lists four of your six integrations and skips the two a buyer cares about, nothing it said was false, but you still lost the comparison.
T: Trace the drift
For every gap, find the source the model is pulling it from. Look at the citations, click through and categorise what's there. Your own stale page, an old review, a comparison site, a Reddit thread that's the old version of you frozen in amber.
The common mistake is fixing the answer instead of the source, so the drift comes straight back next run. If Perplexity keeps understating your enterprise fit, the trace might land on a three-year-old G2 description that still calls you an SMB tool.
U: Unify the sources
Make the correct facts appear consistently across owned and earned surfaces, in the same language. Update the pages, the docs, the schema, the profiles, the directories, the partner listings. Get critical facts into crawlable text, not locked inside a PDF or an image.
The common mistake is fixing only your own website, the one vote you already controlled, and ignoring the third-party surfaces where most of the conflicting signal lives.
Accuracy alone isn't enough. The model wants corroboration, so add proof around the facts you want repeated. Named customers, review data, certifications, comparison pages, structured evidence.
The common mistake is making the claim without the proof. "We integrate with Salesforce" is weak. "We integrate with Salesforce, documented here, listed on the AppExchange, last updated this quarter" is something a model can stand on.
L: Loop for freshness
This isn't a one-time cleanup. Re-run the prompts on a cadence, update the factbase after any product, pricing or positioning change and keep stale third-party references from polluting the answers again.
The common mistake is auditing once, fixing it and assuming it stays fixed while your product ships and your old content keeps ranking. Drift keeps happening, so the audit has to keep happening too.
We run this for clients as a regular cadence, including the fixes.
Diagnostic questions
Four questions to ask yourself before you decide you're fine. If any answer is "not really," you've got drift.
Can you trace where an AI answer about you came from? If you can't point at the source behind a wrong claim, you can't fix it. You can only hope the next model update is kinder.
Do your owned and third-party surfaces tell the same story? When your site and your G2 profile disagree, the model resolves the tie for you, and it doesn't ask which one you'd prefer.
When did you last check what your G2 page and your old comparison articles actually say? These are the surfaces that quietly outvote your homepage, and they're the ones nobody owns.
Could a buyer disqualify you today on a fact that's simply wrong? If the honest answer is "possibly," that deal is already at risk, and you won't get a notification when it dies.
Run it this week
You can run a lite version of this yourself in four to eight hours. It won't be statistically tight, but it'll show you direction, and direction is usually enough to act on.
Start with the prompt set. Build 30 to 50 prompts across four categories: category ("best tool for X"), comparison ("your brand vs competitor"), capability ("does your brand integrate with X") and buyer-journey ("how do i solve X," "your brand pricing"). Mine the wording from real sales calls and reviews. If you can't imagine a buyer typing the exact string into ChatGPT in the next month, rewrite it or drop it.
Run those across the five engines, three times each, and score every response on the 1-5 framing rubric. 5 means the model describes you exactly how you'd want, 1 means it actively contradicts your current positioning, 3 is a flat neutral mention with none of your differentiators. Accuracy and competitor-positioning get scored alongside.
Then build a 30-day fix list, and keep it to three items. One with big pipeline impact, one cheap fix you can ship this week and one with a compound effect (a single comparison page can move five to ten prompts at once). Three is the working number. More than three and you'll start everything, finish nothing and never know what moved the score.
Where to start, by stage
Not everyone starts from the same place. Find yours.
If you've never measured AI perception at all, keep it small. Twenty prompts, one score: is the framing right or wrong. Don't try to build a full taxonomy on day one. You're looking for the gut-punch examples that tell you whether you have a problem, and the first wrong answer you find usually settles the argument internally.
If you've rebranded or repositioned in the last year, run it now. This is the highest-drift state there is. The old version of you is still the one the models repeat, because the old reviews and comparison pages still exist and still rank.
Why we launched Pulse
The audit above is the lite version and you can run it by hand. Running it properly is harder: across every major model, on a repeating schedule, with the noise smoothed out. That's why we built Pulse and its AI perception module. It runs this continuously, checks each claim against your source of truth, traces it to the page that caused it and tracks whether your perception is improving or slipping over time.
A handful of prompts run once is a snapshot with a lot of noise in it, so we run statistically-modelled prompt sets, a few hundred per brand to a known margin of error, and re-run them on a cadence.
It's the difference between a single lucky run and a number you can actually trust.
Note: we're not a SaaS company so we don't sell this as software. We couldn't find reliable tools on the market so we decided to build them in order to give clients an advantage.
TL;DR
Most brands chase AI visibility. The bigger risk is being visible and described wrong, because AI assembles you from conflicting, stale sources. That's Factual Drift. FACTUAL is a seven-step audit you can run on your own brand this week: form a factbase, test the engines, trace the drift and fix the sources.
Liam
P.S. Reply to this with "perception" and i'll run a free audit for your company. It's not very scalable so i'll limit this to the first two people :)
P.P.S We run SEO and AEO for SaaS companies like Instantly, Granola, Gladia, incident and others. Learn more on our website.
