Monitoring
How to monitor AI visibility through model updates and deprecations
To monitor AI visibility through model updates and deprecations, preserve raw answers, capture version and surface metadata, overlap old and new configurations when possible, and mark every discontinuity on the trend line. Never splice results from a replacement model into history as though nothing changed. The business questions can remain stable while the measurement instrument evolves, but stakeholders must be able to distinguish brand movement from a provider transition.
Maintain a model and surface registry
For each monitored run, store provider, product surface, visible model or API identifier, mode, tool access, date, locale, account context, system prompt version where controlled and measurement code version. Record "unknown" for hidden consumer metadata. Add provider announcements and internal configuration changes as dated annotations. A marketing chart without this operational record cannot explain whether a jump coincided with a new model, browsing change or prompt edit.
Freeze a durable core prompt panel
Keep core customer questions, eligibility rules and scoring definitions stable across the transition. Version corrections rather than silently editing prompts. Maintain a separate exploratory panel for new capabilities so experimental tasks do not contaminate the historical baseline. Direct brand prompts, unbranded discovery and verification questions should remain distinct because updates may affect them differently.
Use an overlap window when access permits
- Run the old and new model against the same prompt set and time window.
- Match accessible settings, tools and source conditions as closely as possible.
- Collect multiple samples and preserve complete answers and errors.
- Score both with the same human-reviewed rubric.
- Report agreement, directional shifts and materially different cases.
- Retire the old series with a labeled final date and archived configuration.
If a provider removes access without overlap, begin a new baseline and state the gap. Do not backfill results or create a conversion factor from intuition. A short break in the chart is more honest than false continuity. Retain old evidence according to policy so later reviewers can inspect the pre-change state.
- Instrument change: model, routing, tool, interface or evaluator changed.
- Evidence change: the web, brand facts or cited sources changed.
- Outcome change: mentions, recommendations, accuracy or competitors changed.
These events can coincide without one proving another. Use annotations and targeted retests to narrow hypotheses, but avoid causal claims from a before-and-after screenshot. Review both aggregate rates and prompt-level transitions. A flat mention rate can conceal that valuable prompts were lost while low-intent prompts improved.
An API deprecation is a controlled integration event; a consumer app may change routing or personalization without exposing equivalent identifiers. Maintain separate series for API measurements and sampled app observations. Test your extraction and scoring pipeline too: a new response format can create apparent visibility loss when the real failure is a parser that no longer recognizes citations or lists.
ModelSaid helps maintain recurring evidence and trend context across supported assistants. Start with the AI visibility scan, compare operational value through the ROI calculator, and choose a cadence on pricing. Extensible coverage allows new models and providers to enter as labeled baselines rather than overwriting the historical record.
Assign an owner for provider notices, a migration checklist, regression thresholds and stakeholder communication. Define alert levels before the event: parser failures and identity errors need immediate attention, while moderate mention movement may require another sample. Before a cutoff, export configurations, complete overlap runs and validate dashboards. Test error handling, structured outputs and citation extraction as well as answer quality. Notify report consumers which metrics remain comparable and which restart. After migration, review high-impact prompts manually and document known limitations. Run an early stability check, followed by a normal-cadence review, without assuming the release caused every change. Keep the deprecated series readable but visually separate. If costs or rate limits change, adjust sampling transparently and preserve raw denominators; a smaller panel should not masquerade as the old program. Keep a rollback plan for internal evaluator changes even when the provider model itself cannot be restored. The trustworthy program does not promise an unbroken score forever. It preserves enough context to explain clearly exactly where comparability ends and what the new baseline means.
Is your business visible in AI search?
Run a free check and see what ChatGPT, Claude, Gemini and Perplexity actually say about you right now.
Check your business for free