Monitoring
How to establish an AI visibility baseline
An AI visibility baseline is a dated, reproducible record of how selected assistants answer a fixed set of business-relevant questions before an intervention. Establish it by defining eligible prompts, markets, models, modes and scoring rules; collecting repeated observations under matched conditions; and saving complete answers with visible citations. The baseline should report mentions, citations, recommendations, sentiment, accuracy and competitor presence separately. It is not an industry benchmark, a permanent rank or a promise that identical wording will return tomorrow.
Start with decisions customers actually make
Build prompts from discovery, shortlist, comparison, verification and objection stages. Include unbranded questions such as "Which tools fit a small European support team?" alongside branded fact checks and direct comparisons. Record audience, geography, language, constraints and intended stage for each prompt. Exclude questions the business cannot honestly serve. A baseline dominated by branded prompts will exaggerate visibility, while a panel of impossibly broad questions will understate it. Use a fixed core for longitudinal measurement and a separate exploratory set for new customer language.
- Prompt ID, exact wording, intent, audience, locale and eligibility rule.
- Provider, visible product mode, date, account context and whether web search was evident.
- Complete response, citations, named competitors and decisive answer passages.
- Mention, citation and qualified-recommendation labels with written definitions.
- Sentiment and claim-level accuracy labels, including reviewer notes.
- Known business events, site releases or source changes recorded as annotations.
Collect enough observations to expose variability
Generated answers vary, so repeat priority prompts rather than treating one run as a stable fact. Test providers within a short collection window and keep conditions as comparable as their products allow. Do not silently combine searched and knowledge-only modes, logged-in personalization, different languages or materially changed prompt wording. Show the number of eligible answers behind every rate. If a sample is small, present the raw count, such as three mentions in eight eligible answers, rather than a percentage that looks more certain than the evidence.
Score the baseline without inventing a universal index
A mention records presence. A citation records visible sourcing. A recommendation records suitability language. Sentiment records framing, and accuracy tests material claims against approved facts. Share of voice divides your eligible appearances by all eligible brand appearances within a named, frozen competitor set; changing that set creates a new baseline. These measures answer different questions and should remain available even if leaders also receive a summary index. Document any weighting so a later analyst can reproduce it and challenge the assumptions.
- Approve the prompt universe and exclusions with marketing, sales and product owners.
- Run a short pilot to remove ambiguous prompts and calibrate reviewers.
- Collect the formal baseline without changing public evidence mid-window.
- Resolve label disagreements using the written rubric and saved answer.
- Publish rates with counts, provider splits, prompt categories and material inaccuracies.
- Lock the core panel and version future changes instead of rewriting history.
Turn the baseline into a monitoring contract
Attach an owner and next action to important gaps. A missing mention may trigger category-page research; an uncited recommendation may expose an evidence gap; an inaccurate price requires an urgent source audit. Use the free AI visibility scan to orient the first review and the schema validator when inconsistent structured facts may be involved. ModelSaid's Monitor and higher plans support recurring visibility scans and retained trends, while a plan-level Google connection supports either GSC or GA4 for adjacent search or behavior context. Review current plan capabilities rather than assuming a manual baseline automatically becomes continuous monitoring. Keep the baseline immutable, annotate interventions and compare matched future windows without claiming one edit caused every change.
Protect the baseline from hindsight bias
Schedule the first comparison before the team begins optimizing. The interval should reflect how often the monitored answers, public facts and business priorities can realistically change, not a desire to manufacture a quick win. Preserve null and unfavorable results. When a new provider or market is added, run an overlap period and publish it as a distinct cohort. When prompts are retired, retain their historical observations and reason for retirement. Store the scoring guide with the dataset and periodically have a second reviewer rescore a sample. This version control prevents a quiet shift toward easier questions and makes the baseline credible to finance, leadership and future analysts. A baseline earns trust by remaining comparable, not by making every later result look positive.
Is your business visible in AI search?
Run a free check and see what ChatGPT, Claude, Gemini and Perplexity actually say about you right now.
Check your business for free