Measurement
How to design before-and-after AEO experiments
A credible before-and-after AEO experiment changes one defined evidence set, preserves a fixed measurement panel and compares matched pre- and post-periods with an untreated control where possible. Pre-register the hypothesis, primary metric, sample rule, intervention, lag window and exclusions. Collect repeated answers because outputs vary, and save complete responses with citations. Treat a post-release movement as causal only to the extent that the design rules out model changes, seasonality, campaigns, competitor actions and measurement drift.
Write a falsifiable hypothesis
State the page, audience question and observable outcome. For example: clarifying verified compatibility information on selected product pages will increase accurate citations for the matched compatibility prompt cluster relative to unchanged product pages during the evaluation window. This is stronger than "improve AI visibility" because it identifies the evidence, affected prompts, primary metric and comparison. Choose one primary outcome; mentions, recommendations, sentiment, referral sessions and pipeline can remain secondary measures rather than opportunities to cherry-pick success.
- Unit: topic, page group, market or product that can receive or withhold the intervention.
- Baseline: repeated matched observations collected before the release.
- Treatment: a versioned content, source or technical change with deployment proof.
- Control: comparable untreated unit observed over the same period.
- Outcome: one pre-declared metric with numerator, denominator and scoring rule.
- Lag: a justified waiting and observation window, not a date chosen after seeing results.
Build comparable treatment and control groups
Match groups on intent, demand, existing visibility, page type, geography and business importance. Random assignment is preferable when practical, but phased rollouts or matched topic groups can be useful. Do not select controls that were already declining or unable to rank by the eligibility rubric. Keep prompt wording, provider mix, locale, mode, repetition count and collection cadence stable. If a provider changes its product during the test, annotate it and consider sensitivity analysis or a restarted window.
Measure each layer without mixing outcomes
Use mention rate for presence, citation rate for visible sourcing, qualified recommendation rate for suitability, sentiment for framing, accuracy for material truth and share of voice for competitive presence in a fixed universe. Referral sessions, assisted conversions and pipeline belong to downstream layers and often require longer windows. A citation improvement can validate the primary hypothesis even when no immediate pipeline movement is observable. Conversely, a sales increase during the period does not validate the AEO change without supporting journey or control evidence.
- Freeze the protocol, prompt panel, rubric and analysis plan before deployment.
- Collect enough baseline repetitions to understand ordinary answer variation.
- Archive the released content, structured data and technical state.
- Avoid unrelated changes to treatment and control evidence during the window.
- Run matched post-period collections and review labels without knowing group where feasible.
- Report effect estimates, raw counts, exceptions, confounders and null results.
Operationalize experiments without promising rankings
Use the AI visibility scan to pilot prompts, the schema generator to draft faithful markup where structured data is part of the treatment, and the schema validator before release. ModelSaid Monitor and higher plans provide recurring scans and trend monitoring; a plan-level Google connection can add either GSC or GA4 context, subject to current pricing and limits. ModelSaid observations show what sampled assistants returned, not their private ranking logic. Preserve unsuccessful experiments, distinguish correlation from causation and use the result to decide whether to roll out, revise or stop the intervention.
Analyze the registered outcome before exploring subgroups
Report treatment and control changes, their raw denominators and the range of plausible interpretations using the pre-registered rule. Exploratory findings, such as a movement limited to one provider, can generate the next hypothesis but should not replace the primary result. Replicate promising effects on another eligible topic before scaling widely, especially when implementation cost or factual risk is high. Review whether scorers agreed, whether citations actually supported the claims and whether any treatment page changed unexpectedly. Share the protocol and archived evidence with a reviewer who did not design the intervention. If the experiment is inconclusive, improve power, measurement consistency or intervention clarity; do not keep extending the window until a favorable answer appears. Publishing a clean null result protects the next test from repeating an unsupported tactic.
Is your business visible in AI search?
Run a free check and see what ChatGPT, Claude, Gemini and Perplexity actually say about you right now.
Check your business for free