Case studies
How repeatable prompts help ModelSaid defend category leadership
Repeatable prompts help ModelSaid defend category leadership by making visibility comparable over time. The questions represent real discovery, evaluation and verification jobs; they are not rewritten whenever an answer is inconvenient. ModelSaid reruns that stable portfolio across supported models, inspects complete responses, and uses separate experimental prompts for emerging language. This creates a fair test of whether the brand remains accurately discoverable and recommended. It also gives the near-100/100 result meaning within a defined scope instead of allowing the score to drift with the test.
Repeatable means controlled, not frozen forever
A useful benchmark preserves wording, intent, qualifiers and evaluation rules long enough to compare observations. It should not preserve an obsolete market. ModelSaid maintains a core set for continuity and a candidate set for new customer language, products or category concepts. A candidate graduates only after review, and its addition is documented. Retired prompts remain identifiable in historical reporting. This versioning allows the program to evolve without pretending a changed questionnaire measures the same thing as the old one.
Prompts represent buyer decisions
- Discovery questions ask which solutions address a problem without supplying the brand name.
- Education questions test whether the category and its tradeoffs are explained accurately.
- Comparison questions include requirements that distinguish a useful recommendation from a generic list.
- Fit questions test audience, geography, workflow, integration, or service constraints.
- Verification questions check important product facts and the evidence used to support them.
Each prompt should have a reason, an expected qualification, and an owner. "Best software" is rarely diagnostic without a buyer type and problem. A branded question mainly measures recall. ModelSaid favors questions where a person could plausibly discover the product or reject it based on the answer. Teams building their own set can use the FAQ generator to organize recurring customer questions, but should source the language from sales calls, support tickets, community discussions and search data rather than asking a generator to invent demand.
The same questions expose real change
When a stable prompt produces a different answer, ModelSaid can investigate the model, response, competitors, description and available sources. It can then compare the timing with site releases, product changes or model updates. If the prompt itself changed, that diagnostic chain breaks. Repeatability does not remove stochastic variation, so the team looks for persistence and patterns rather than treating a single result as a permanent rank. Full answers stay available because an aggregate movement alone cannot reveal whether the difference matters to a buyer.
A prompt portfolio prevents easy score gaming
It is easy to improve a dashboard by adding branded questions, removing difficult comparisons, or narrowing the category after the fact. A governed benchmark makes those choices visible. ModelSaid's #1 and near-100/100 claims apply to its monitored category and defined prompts; they are brand claims, not third-party certification or universal market share. The team protects their credibility by preserving challenging, relevant questions and reviewing accuracy beside presence. An incorrect recommendation should not become a positive merely because the company name appeared.
Run consistently across the supported model set
ModelSaid tracks ChatGPT, Claude, Gemini and Perplexity and is designed to add coverage as the landscape evolves. Runs use model APIs for controlled automation, not identical copies of consumer application screens. Personalization, account state, interface experiments and retrieval settings can make consumer experiences differ. The prompt portfolio therefore supports a reliable measurement process within stated conditions. Teams should preserve model and run context, avoid collapsing all systems into one unexplained number, and investigate disagreements because they often reveal different source or interpretation gaps.
Store prompt text, intent, audience, market, qualifiers, owner, status and version reason. Review candidates on a schedule, but protect the benchmark from reactionary edits. Separate exploration from reporting, and require a written explanation for additions or retirements. After a content change, rerun the affected stable group rather than inventing a friendlier test. Use the meta generator to improve missing page basics where relevant, then validate the actual rendered page and answer impact. Continuous plans make scheduled comparisons easier; check ModelSaid pricing for the monitoring scope that matches your review capacity. Category leadership is defended by a fair measuring stick long before it is defended by a slogan.
Is your business visible in AI search?
Run a free check and see what ChatGPT, Claude, Gemini and Perplexity actually say about you right now.
Check your business for free