Start from the decision, not the demo
Name the decision the data has to serve: which content to ship next, whether the brand program is working, what to tell the board about AI presence. Tools that dazzle in a demo can still fail the specific decision, and the shortlist exists to filter for it.
This page is the evaluation companion to the AEO tools landscape guide: that one maps what exists, this one is the protocol for choosing.
Shortlist criteria
| Topic | Criterion | What to verify |
|---|---|---|
| Evidence | Any number opens to the exact samples, prompts, citations, and dates behind it. | |
| Uncertainty | Mention rate carries a visible sample size and confidence interval, not a bare score. | |
| Coverage | The measured AI surfaces match where your buyers actually ask their questions. | |
| Competitor basis | Rivals are measured on the same prompts at the same time, so comparisons are real. | |
| Method disclosure | Models, sample counts, errors, and search activity are visible per run. | |
| Recommendations | Next steps trace to observed gaps and are framed as tests, not guarantees. |
Run the trial
Step 1
Fix the basis
One prompt set, one competitor list, entered identically in every candidate. A trial with a moving basis proves nothing about any of them.
Step 2
Let it run unchanged
Two to three weeks of scheduled runs. Resist editing prompts midway; you are testing the tool's ability to hold a stable measurement.
Step 3
Audit three numbers
Pick a good one, a bad one, and a surprising one, and drill each to its samples. This is where chart generators and measurement tools part ways.
Step 4
Judge a change against its interval
Find one week-over-week movement and check whether the tool helps you tell real shift from ordinary wobble.
Step 5
Decide with the user in the room
The person who will run this monthly should make the call. Tools get abandoned by owners, not by committees.
Judging what you saw
Weight evidence quality over interface polish and coverage breadth over feature count, because the failure mode of this category is confident numbers nobody can defend in a meeting. A smaller tool with auditable results beats a wide one with unexplainable scores.
AeoWatch is built to pass exactly this protocol, which is also the fastest way to test the claim: run the trial and audit the numbers.