Why tracking starts with variance
Generated answers are drawn from a distribution, not read from a database. Ask the same engine the same question five times and the brand list can change between draws, which means a single check is not a measurement, it is an anecdote.
Everything else in the methodology exists to deal with that fact: fixed questions, repeated draws, and results reported with the uncertainty they actually carry.
The tracking pipeline, step by step
Step 1
Fix the prompt basis
Choose the buyer questions worth winning and store them. The exact set used for each run is retained, so editing prompts later does not quietly rewrite history.
Step 2
Sample repeatedly per engine
Each prompt runs several times on each tracked surface. Failures, refusals, and rate limits get recorded as errors rather than papered over, so the sample count you see is the sample count that existed.
Step 3
Extract the signals
Every answer is scored for brand presence, first-mention position, sentiment, and the domains it cited. The raw material is kept, not just the verdict.
Step 4
Aggregate with uncertainty
Mention rate is reported with a 95% confidence interval for the observed sample size, alongside comparative share of voice across the tracked brand set.
Step 5
Keep the evidence attached
Each aggregate opens back down to its samples: prompt, engine, model, date, mentions, and citations. A number you cannot audit is a number you will eventually argue about.
What changes a trend line and what should not
A trend is only comparable while its basis holds still. Good tracking marks the runs where prompts were added or removed, so a jump in the line can be traced to a measurement change instead of being mistaken for a visibility change.
Cadence matters less than consistency. Weekly runs on a stable basis beat daily runs on a shifting one, because the question a tracker answers is always relative to its own history.
What to look for in the output
- A visible sample size next to every rate, never a bare percentage.
- A confidence range on mention rate, so ordinary wobble does not read as movement.
- A per-engine split, because visibility on one surface says nothing about another.
- Cited domains listed per answer, since citations are where the influence map lives.
- Disclosed errors and cached samples, so the method stays checkable.