Back to insights

How to measure AI visibility without fooling yourself

By LeadSpark Marketing·Sep 15, 2026

A useful AI visibility report tells you what was asked, which product answered, what it said, and how much of the planned test was completed. It should help you decide what to investigate next. It should not turn a few favorable answers into a claim about every customer's experience.

Begin with a small, repeatable set of customer-like questions. Save the whole answers, separate mentions from recommendations and citations, and keep business results on their own line.

The method below is LeadSpark's suggested observation process. It is not a platform-issued ranking metric or a representative survey of all AI users.

Before collecting answers

Prepare an approved record of your business name, services, actual geography, current affiliations, and important claims. You need it to check whether an answer describes the right business accurately.

Choose one person to collect observations and someone to review ambiguous results. A spreadsheet or document is enough for a modest manual exercise. The important part is preserving the evidence and using consistent rules.

Decide the questions and repeat policy before seeing the results. Do not keep asking until you get an answer you like. That produces a collection of highlights, not a baseline.

1. Separate recognition from discovery

Put the questions into groups according to their purpose.

A branded question such as "What services does Example Company provide?" checks recognition and accuracy after you supply the name. An unbranded question such as "Who repairs residential gates in this area?" checks which providers appear without that help. A named comparison is another group. An educational question may seek an explanation rather than a provider at all.

These are example prompts, not observations about a real company. Use questions that fit the service and customers you actually serve.

Keep the groups separate in the report. Supplying your name and website changes the task. A high recognition rate should not be presented as evidence that customers will discover you without knowing the name.

For local questions, use explicit location wording where it helps interpretation. "Near me" can involve context you do not fully control.

2. Freeze the collection plan

Record the exact prompts, product names, intended modes, language, location wording, repeat count, and collection window. Keep the plan small enough to finish and review.

For example, you might choose a few discovery questions for one product and repeat the same set on another date. That is an editorial learning exercise, not a statistically validated minimum sample. More questions are not automatically better if nobody checks the answers.

OpenAI's Search documentation says location information and saved memories can influence searches. Record known account context, and use fresh conversations to avoid carrying one test answer into the next. Do not call the test fully unpersonalized merely because the chat is new.

Name the actual product and mode. Google Search AI Overviews, AI Mode, and the Gemini app should not be treated as one interchangeable test surface. If a model label is not shown, record "not shown" rather than guessing.

3. Save the complete observation

For every planned answer, retain:

  • The unchanged prompt and its group.
  • Date, time, timezone, product, mode, and known context.
  • The full answer and visible source URLs.
  • Whether collection completed, failed, or captured only part of the answer.
  • Any retry and the reason for it.

If a technical failure requires a retry, retain the failed attempt and link the retry to the same planned observation. Do not count it as an extra opportunity for success.

A complete answer that recommends nobody is still an answer. It should not disappear from the report because it is disappointing. A failed page load is different: there was no complete answer to classify.

If you cannot tell whether search was used, record that as unknown. A missing citation does not reveal every hidden retrieval step.

4. Classify the answer, then inspect its sources

Count a mention only when the answer identifies the correct business. Count a positive recommendation when it offers the business as a viable option for the question. A warning about the company is a mention, not a positive recommendation.

Record supporting citations to your own site separately from supporting citations to outside pages about the business. Open each relevant source to check what it supports. A link elsewhere in the answer should not count as business support merely because it is visible.

If identity or source support is ambiguous, mark that field unresolved. The answer may be complete while a particular assessment is still pending. Report unresolved fields rather than silently treating them as verified successes.

Our citation-versus-recommendation explainer gives examples of the difference. Use the same definitions in every collection wave.

5. Keep the denominator visible

Here is a fictional teaching example, not a LeadSpark client result.

You plan five discovery observations for one product and mode. Four yield complete answers that are reviewed. One fails to load. Two of the four valid answers recommend the business.

The recommendation rate in that valid panel is two out of four, or 50%. Completion coverage is four out of five, or 80%. Report both. The failed load is not an answer in which the business was absent.

A complete answer that says it cannot identify a suitable provider stays in the denominator. If there are no valid answers, the rate is unavailable, not zero.

Now suppose two of those four valid answers cite your site, and one of the same answers also cites an outside source supporting the business. Owned-site support is 50%, and outside support is 25%. Any supporting citation is still 50%, not 75%, because the outside citation appeared in an answer already counted.

Keep counts beside percentages. They make the size of the exercise and overlapping outcomes much easier to understand.

6. Compare matching observations

Use the same prompts, groups, products, modes, location wording, and review rules in the next wave. When something changes, show the break in the record rather than drawing a continuous trend through it.

A move from two recommendations to three in a small panel is an observation worth investigating. It does not prove a durable effect or establish that your latest website edit caused it. Related prompts and repeated answers are not automatically independent samples of all customers.

Keep a separate change log for website edits, profile corrections, and new evidence. That helps you ask better questions about what happened without pretending that timing alone proves causation.

7. Keep native reports and business outcomes separate

A platform report has its own scope and definitions. Do not merge it with your manual panel and give the combined number a name such as "total AI ranking."

For example, Google's AI-features documentation says traffic from AI Overviews and AI Mode is included in Search Console's overall Performance report under the Web search type. That is not the same measurement as your manually collected provider recommendations.

For business outcomes, record the available referrer information and what the customer says about finding you. Those can differ. Someone may see an AI answer and later search for the business directly, while a tagged visit may never become an inquiry.

Deduplicate the same opportunity across a form, call, and follow-up message. A citation is not a lead; a captured inquiry is not a sale. Preserve uncertainty where you cannot attribute the result reliably.

When the report looks better but the business facts are wrong

Do not celebrate the percentage first. Review whether the answer describes services you actually provide and locations you actually serve.

Correct public facts you control and record where the inaccurate claim appeared. If the source is outside your control, request a legitimate correction where appropriate. Preserve the original answer so the next review can see what changed.

The search-readiness guide helps turn those findings into information improvements. Avoid diagnosing a hidden retrieval failure from an answer alone.

How often should a small business repeat this?

Choose a cadence you can sustain and repeat the same conditions as closely as possible. Recheck after material information changes, but do not promise yourself a particular response deadline. A consistent modest panel is more useful than an ambitious test abandoned halfway through.

What belongs in the summary?

State the purpose, product and mode, dates, planned observations, valid reviewed answers, results, unresolved fields, and the next question to investigate. Include access to the full evidence.

If that review exposes work beyond your website foundation, start a System Audit conversation. Custom measurement or integration work needs its own scope. A better reporting process should make the uncertainty visible, not hide it behind a more impressive score.

Make the business easier to run.

Bring your questions and the evidence behind them. We’ll review what needs attention before scoping custom work, without promising a search ranking or AI recommendation.

AI supports your process · you stay in control

More in SEO & Search Visibility