Research/Method note
Published
Published
20 August 2026
Author
Kristopher York
Reading time
2 min
Version
Method note 1.0

Method note / Measurement

Why one AI visibility score is not enough

A useful AI visibility measurement needs its evidence, sample, method, and uncertainty—not just a number that went up.

Working conclusion

A single score can be useful as a summary, but not as the evidence. Without observations, denominators, coverage, and uncertainty, a movement may reflect prompt mix, collection changes, model drift, or ordinary sampling variation.

A score is attractive because it compresses a complicated system into something that fits on a dashboard. The compression is not the problem. Forgetting what was compressed is.

AI visibility is observed across prompts, surfaces, markets, times, repetitions, and collection methods. Two scores can be numerically identical while resting on very different evidence.

A score needs a denominator

"Visible in 40% of answers" is interpretable only when the reader can see how many prompts, which prompts, how many successful observations, and what counted as visible. Ten mentions from 25 observations are not equivalent to 400 mentions from 1,000, even if both produce 40%.

Missing observations matter too. If one engine failed for half the collection window, silently calculating the score from what remains can create an artificial improvement.

Unlike observations should not be blended silently

An official model API, a web-search-enabled API, a licensed consumer-surface feed, and a manually captured consumer answer are different observation methods. They may all be useful. They are not automatically interchangeable.

A defensible system keeps the method on every evidence record and makes cross-method aggregation an explicit analytical choice.

Movement needs context

A score can move because the brand changed. It can also move because the prompt set changed, the surface changed, the model changed, retrieval selected different sources, a provider changed its product, observations failed, or a probabilistic system returned another plausible answer.

That is why York Studio's measurement hierarchy starts with the underlying answer and works upward:

  1. raw observation and collection metadata;
  2. extracted entities, citations, claims, and recommendations;
  3. comparable samples and coverage;
  4. uncertainty and change analysis;
  5. summary metrics.

The dashboard is the last layer, not the source of truth.

Observed, inferred, and attributed

These labels answer different questions.

Observed: the recorded answer contained a brand, claim, citation, or recommendation under stated conditions.

Inferred: an analysis estimated a relationship, theme, sentiment, or likely explanation from one or more observations.

Attributed: an experimental or causal method supports the claim that an intervention contributed to a change.

Most monitoring is observational. That does not make it worthless; it makes the claim narrower. The honest statement is often "the measured answers changed after the intervention" rather than "the intervention caused the change."

What a trustworthy summary should reveal

Every aggregate view should make it easy to reach the evidence. At minimum it should show the time window, sample size, successful coverage, surfaces and methods, prompt-set version, metric definition, and uncertainty or stability signal.

The goal is not to make every user into a statistician. It is to prevent precision from being used as theatre.

Required reading

Limitations

  • This note sets a measurement standard; it does not validate a particular scoring formula.
  • The most appropriate interval or uncertainty method depends on the sampling design and metric.
  • Some consumer surfaces expose limited model or retrieval metadata, which constrains comparability.

Primary references

Sources

  1. Google: CausalImpact
  2. OpenAI: web search tool
  3. Google: grounding metadata

Found an error? Research should be correctable.

Submit a correction