How to Build an AEO Baseline Before You Spend Anything
Before changing the site, hiring an agency, or buying software, freeze the current AI-search answer. If the company does not record the before state, every later claim of improvement becomes guesswork.
AEO teams often begin with a list of changes. Publish more pages. Add schema. Rewrite product copy. Earn outside coverage. Those may become the right moves, but none of them establishes what AI search says today.
A baseline does. It records the questions tested, the answers returned, the companies mentioned, the pages cited, the facts that were right or wrong, and the traffic already reaching the site.
You can build the first version with a browser, a spreadsheet, Search Console, and the analytics already installed on the site. The point is not to avoid software forever. The point is to know what needs measuring before a vendor defines the measurement for you.
AEO work without a baseline can produce activity. It cannot show what changed.
Start With a Fixed Set of Commercial Questions
The unit of measurement should be a buyer question, not a keyword count. Use the first 20 questions your company should own across need, approach, fit, proof, and adoption.
Keep the wording stable. If the question changes between the before and after tests, the answer may change because the prompt changed. That is not evidence that the company's AEO work had an effect.
Add the answer engine, date, access mode, location when relevant, account state, and whether the prompt began in a fresh conversation. Generated answers can vary. The test record needs enough context for another person to repeat it.
For the highest-value questions, run a second fresh session. If the result changes materially, mark the question as unstable instead of choosing the output the team prefers.
Record Six Things for Every Question
Do not reduce the answer to a single visibility score. A company can be named without being cited, cited without being recommended, or accurately described while a competitor controls the decision.
- MentionsRecord whether the company or product appears in the answer and the context around the mention. Neutral inclusion, recommendation, and warning are different outcomes.
- CitationsSave every linked source, not only links to the company site. Record the exact company URL when an owned page is cited.
- Answer accuracyGrade the answer against current company truth. Capture the exact claim that is correct, incomplete, outdated, or wrong.
- CompetitorsRecord which alternatives appear, what position they occupy in the answer, and whether they are supported by citations.
- Source patternsClassify the cited domains as owned, competitor, independent editorial, review, community, documentation, or another useful group.
- Referral trafficSave current sessions, landing pages, engagement, and key events from identifiable AI-search referrals. Keep visibility and traffic as separate records.
Accuracy Needs a Rubric, Not a Feeling
"Looks good" is not an accuracy standard. Define the company facts that matter before reviewing the answers: category, primary use case, best-fit customer, product availability, pricing model, integrations, implementation requirements, proof, and known limits.
Then use a small rubric:
Correct means the answer represents the relevant company truth without a material error. Incomplete means it omits something that could change fit or the next decision. Wrong means it states a false or outdated fact. Unverifiable means the team cannot support or reject the claim with current evidence.
Save a short explanation for every incomplete, wrong, or unverifiable grade. That note becomes the repair brief later. It also prevents a favorable mention from hiding a factual problem.
Mentions and Citations Measure Different Things
A mention shows that the company entered the answer. A citation shows that a source was used or offered as support. Neither one proves that the answer is accurate or commercially useful.
Count them separately by question. Then look at coverage. Which need questions mention the company? Which comparison questions cite it? Where does the company appear but an outside source supplies the evidence? Where is a competitor named while the company is absent?
This produces a more useful work queue than a blended score. A missing mention may require clearer category relevance. An inaccurate mention may require correcting owned information and outside evidence. A mention without an owned citation may expose a source gap.
Source Patterns Explain Why the Answer Looks the Way It Does
A citation list is not a trophy count. It is a map of the public evidence supporting the answer.
Group recurring sources by type and question. If comparison answers repeatedly cite review sites, those sources influence the evaluation frame. If implementation answers rely on documentation, weak or inaccessible documentation becomes a larger problem than blog volume. If the company's own pages appear only for branded questions, the site may not explain the broader category well enough.
Do not assume a cited source caused the entire answer. Record the observable pattern. The pattern tells the team where to inspect and what evidence may be missing.
Capture Referral Evidence Without Pretending It Is Complete
Referral traffic is the clearest evidence that a person clicked through. It is not the complete value of AI-search visibility because many answers affect consideration without producing a visit.
OpenAI says ChatGPT referral URLs include utm_source=chatgpt.com, which makes those sessions identifiable in analytics when tracking is working. In Google Analytics, use session-scoped source and medium dimensions to examine visits, landing pages, engagement, and key events from referral sources.
Google includes appearances and clicks from AI features in the overall Web performance reporting in Search Console. If the dedicated generative AI report is available in the property, save that view too. Keep the export date and reporting window with the baseline.
Record zero when there is no identifiable traffic. Do not convert zero clicks into zero influence, and do not convert a handful of referrals into a revenue claim the evidence cannot support.
Save the Raw Record Before You Summarize It
The spreadsheet should link to the raw answer capture, citation URLs, analytics export, and Search Console export. Summaries are useful for decisions, but they should never replace the evidence.
A minimal record includes:
Question ID, exact prompt, answer engine, test context, timestamp, raw answer, mention status, citation URLs, accuracy grade and note, competitors, source types, and referral evidence.
Once the raw rows exist, summarize question coverage by buying stage. Show where the company is absent, where it is present but wrong, which competitors dominate, and which source types recur. Those findings should decide the first AEO work, not a generic checklist.
Repeat the Same Test After the Work Has Had Time to Appear
The after test should use the same questions, answer engines, session rules, and grading rubric. Log any model or platform change you can identify. If the testing method changes, separate the new series instead of pretending it is a clean comparison.
Do not rerun the full baseline after every page edit. Repeat it after a meaningful group of changes has been published, crawled, and given time to affect the public answer. Use a smaller watchlist for urgent inaccuracies between full comparisons.
If ongoing monitoring becomes too slow for the team, that is the moment to evaluate software. The baseline has already defined the questions, fields, and decisions the tool must support.
The One-Sentence Version
Freeze the questions, answers, citations, accuracy, competitors, sources, and referrals before changing anything, then compare the same test after the work.
Sources
Frequently Asked Questions
What Should an AEO Baseline Measure?
Measure whether the company is mentioned, whether its pages are cited, whether the answer is accurate, which competitors appear, which sources shape the response, and what referral traffic is already visible.
How Many Questions Should an AEO Baseline Include?
Start with a small fixed set of commercially important questions. Twenty questions covering need, approach, fit, proof, and adoption are enough to expose major gaps without creating an unmanageable test.
Can a Company Build an AEO Baseline Without Paid Software?
Yes. A browser, a spreadsheet, website analytics, and Search Console can produce a useful first baseline. Software becomes valuable when recurring monitoring and team workflow justify the cost.
How Often Should an AEO Baseline Be Repeated?
Repeat the same test after a meaningful set of changes has had time to be crawled and reflected. Keep the questions, platforms, session rules, and grading method as consistent as possible.
