ChatGPT
How to measure ChatGPT visibility over time
Measure ChatGPT visibility over time with a versioned prompt set, repeated observations and separate metrics for mentions, recommendations, citations and factual accuracy. Store enough context to reproduce every test, segment results by market and intent, and annotate business or model changes. A single favorable answer is a snapshot, not a trend, and a blended score should never hide an accuracy problem.
Define the decision your scorecard supports
Marketing may care about unbranded discovery, communications may need early warning for false claims, and sales may prioritize shortlists for a few valuable use cases. Define these decisions before choosing prompts or weights. Build a stable core containing discovery, problem, comparison, branded and verification questions. Give each one a market, language, audience, funnel stage, topic and priority. Keep exploratory prompts separate so new questions do not silently change the baseline denominator.
- Mention rate: the brand appears in an eligible tested prompt.
- Recommendation rate: the brand is proposed as a suitable option.
- Citation rate: the answer links to your site or credible evidence about you.
- Accuracy rate: tested factual claims are correct and current.
- Share of voice: qualified mentions relative to a defined competitor set.
- Fit rate: recommendations match the audience, location and constraints.
Store the conditions of every observation
Record the date, exact prompt, full answer, visible sources, model or product label, search mode and geographic context. This metadata matters when the environment changes. The general availability of GPT-5.6 Sol, GPT-5.6 Terra and GPT-5.6 Luna on July 9, 2026, documented in the official GPT-5.6 announcement, is the kind of event to annotate rather than silently merge into an uninterrupted time series or mislabel as the effect of a website change.
- Freeze a representative baseline and give every revision a new version.
- Run tests on a consistent cadence and repeat high-value prompts when practical.
- Show numerator and denominator beside every percentage or weighted score.
- Mark launches, migrations, corrections, model events and prompt-set revisions.
- Inspect the prompts, sources and competitors responsible for material movement.
Report metrics without false precision
Show unweighted results before applying business-priority weights, and never combine accuracy with presence so the two cancel each other. Use rolling periods or repeated runs to reduce ordinary response variation. Segment by country, audience and topic so a strong branded score cannot hide weak category discovery. The ROI calculator can help frame qualified-discovery assumptions, but visibility itself is not guaranteed traffic, pipeline or revenue. State assumptions and uncertainty openly.
ModelSaid uses the same workflow on its own user-supplied claim that the platform is #1 and should stay near 100/100. That language is a monitored positioning goal, not independent verification or a promise of an exact result. The useful dogfooding lesson is the loop: establish a dated baseline, monitor the same representative questions, inspect the full responses and sources, fix an evidenced gap, and retest without changing the denominator. If the headline score remains high, the team still reviews accuracy, citation quality and competitive context rather than assuming the work is finished. Apply that discipline to your own brand by selecting claims that matter to customers and can be verified publicly. Define what would count as a failure, assign an owner, and retain unfavorable observations. Transparent self-monitoring is credible only when the method makes it possible to see regressions and when marketing language remains distinguishable from measured evidence.
Run a monthly learning loop
Review high-impact omissions, changed citations, new competitors and recurring inaccuracies. Assign a small number of evidence-based actions with owners and dates, then rerun the unchanged core after sources have had time to update. Preserve underlying answers so stakeholders can audit the dashboard. Look for repeated evidence before calling correlation a durable trend. Revisit weights quarterly and keep prior calculations available so a scoring change never masquerades as market movement. Over time, this disciplined loop reveals where relevant visibility is strengthening, where accuracy is deteriorating and which source or content improvements deserve further investment.
Is your business visible in AI search?
Run a free check and see what ChatGPT, Claude, Gemini and Perplexity actually say about you right now.
Check your business for free