The AI Answer Variability Study
A null-condition study designed to measure how much AI answers move when the world being measured has not deliberately changed.
Purpose
Before claiming that optimisation caused an answer to improve, we need to know how much the answer changes on its own. This protocol is designed to estimate that background variation and the effect sizes a realistic client programme could detect.
AI visibility products make comparison look precise: a score moved, a citation appeared, or a brand changed position. The unresolved question is how much of that movement would have happened without an intervention.
The AI Answer Variability Study is York Studio's first flagship Lab project. It is a null-condition study: hold the tracked prompts and deliberate site changes steady, repeat observations, and measure the variation that remains.
Research question
How variable are commercially relevant AI answers under repeated, comparable observations when no deliberate optimisation change is introduced?
The practical follow-up is more important: at realistic prompt-set sizes, what change could a brand or agency detect with enough confidence to act on?
Why this comes first
If baseline variation is small, a measurement product can detect modest changes with relatively few observations. If baseline variation is large, before-and-after dashboards can mistake ordinary movement for an optimisation effect. Both findings are commercially useful. The study is designed to find out which world we are in, not to defend a preselected product story.
Provisional design
The study will use a fixed, versioned prompt set across multiple web-connected AI systems or authorised observation methods. Comparable prompts will run on a fixed schedule with recorded surface, model or product identifier where available, market, language, collection method, timestamp, request settings, and raw response evidence.
The analysis will examine more than exact wording. Planned outcomes include entity presence, explicit recommendation, citation presence and overlap, source-domain substitution, ordinal appearance, extracted claims, semantic answer similarity, and cross-surface agreement.
Parameters that could contaminate the experiment—such as indexable prompt and brand details—may be committed through a timestamped cryptographic hash and revealed after collection. The final protocol will state the sampling frame, exclusions, missing-data handling, analysis plan, and stopping rule before results are inspected.
What the study will not prove
It will not show that one provider is universally more accurate, that API output reproduces every consumer experience, or that a particular content change causes visibility. It measures the stability of the selected observation panel under the documented conditions.
Planned outputs
York Studio intends to publish the protocol, methodology version, aggregate findings, limitations, correction history, and enough structured evidence for a qualified reader to audit the reasoning. Negative or inconclusive findings will be published alongside positive ones.
This page describes work in preparation. It does not contain study results yet.
Limitations
- The protocol is still a draft; final sampling and analysis parameters have not been preregistered.
- API observations and consumer-product observations will be labelled separately and cannot be assumed equivalent.
- A bounded prompt panel cannot represent every query, user, market, or model configuration.
- The study measures natural variation; causal testing of optimisation interventions is a later research question.
Sources
Found an error? Research should be correctable.
Submit a correction