An AI visibility baseline is a dated sample of how selected answer systems represent a business for a fixed set of customer questions. It is not an AI ranking, national market share, or proof that anyone saw the answer.
That definition sounds cautious because the measurement needs caution. Change the prompts, market, language, competitor set, model, search setting, or run date and you have changed the test. A polished score cannot repair an undocumented setup.
For a Canadian small or midsize business, we would put the measurement contract before the dashboard: which buyers, which region, which language, which questions, which answer surface, what counted, what failed, and when the same cohort will run again. The report becomes useful when another person can inspect those choices and trace a finding back to the answer and source that produced it.
Start with a measurement contract, not a score
AI visibility tools do not all measure the same object. Some use prompts derived from search demand. Some generate prompts from a domain. Some run one answer per question, while others collect repeated samples. Some count the brand name, some count linked citations, and some combine several fields into a private score.
Those approaches can support different decisions. They cannot be compared cleanly until the method is visible. Ahrefs, for example, labels its estimated impressions as modelled visibility rather than audience reach. Semrush reports mentions and citations separately and says its cross-platform research keeps platforms apart when aggregation would hide differences. These are useful disclosure habits, not proof that either vendor has access to an answer engine's private ranking system.
Our minimum contract records the business name and aliases, owned domains, frozen competitor set, target market, language, prompt cohort, prompt intent, platform and surface, model or interface when visible, search or retrieval setting, session state, run timestamp, valid or failed response, raw answer, cited URLs, and the person who reviewed the extraction.
Seven states must stay in separate fields
A report becomes misleading when it collapses several stages into “visibility.” We keep seven states apart:
- Access: the relevant crawler or user-directed fetch can reach the page under the platform's documented controls.
- Retrieval: the system used or exposed a source for the tested answer.
- Citation: the answer visibly linked to the business's site or another source.
- Mention: the brand appeared in the answer, with or without a link.
- Recommendation: the answer actively placed the brand in a suggested choice set, rather than naming it incidentally.
- Referral: an identifiable visit arrived from an answer surface.
- Outcome: the visit led to a completed tool action, contact, sale, or manually qualified enquiry.
One state can move while the others do not. A third-party article can be cited without the business being recommended. A brand can be mentioned from remembered or retrieved context without its own domain receiving a link. A cited page can receive no measurable visit. Report unavailable data as unavailable, not zero.
Build the Canadian prompt cohort before opening a platform
The prompt list should describe customer decisions, not flattering questions about the brand. Begin with the offer, buyer, service area, buying stage, constraints, and reasons a customer would reject an option. Then write a small cohort that stays stable long enough to support comparison.
For an initial baseline, 12 to 20 questions is often enough to expose obvious gaps without creating a false claim of market coverage. Include category discovery, problem diagnosis, evaluation criteria, use-case fit, comparison, local or national availability, and a small branded verification group. Keep branded questions out of the unprompted mention rate because naming the company in the question almost guarantees a mention.
Canada is not one language and one local context. A Canada-wide English service should be tested as Canada-wide English. A Halifax provider needs a Halifax or service-area cohort. A business that sells in Quebec needs a French-Canadian cohort written for that buyer task, not an English list translated word for word. Record the locale setting as a test condition, not a guarantee that every platform delivered a perfectly geo-fenced answer.
Do not turn every prompt into a new page. Several questions may belong to one strong owner. The baseline helps map a question family to an existing service page, guide, profile, review source, or missing proof. It does not grant permission to publish twenty near-duplicate articles.
Run each answer surface as its own sample
Use a fresh, documented session where the interface permits it. Record whether web search was enabled, which account or personalization state was used, the market and language settings, the model or product surface shown, and the exact time. Save the raw answer and cited URLs before asking a parser to label anything.
Failures do not count as non-mentions. A timeout, empty answer, blocked request, unavailable feature, or parser error belongs in a failure field and leaves the denominator. The report should show how many planned runs became valid observations.
Repeated runs matter because answer systems can produce different wording and sources for the same question. Two 2026 research papers on AI-search measurement argue against reading a single run as a fixed rank and recommend treating visibility as a sample with uncertainty. A practical SMB baseline does not need to imitate a large academic study, but it should disclose whether a result came from one run or several. Small changes inside the observed noise should not become a victory claim.
When a simple rate helps, publish the denominator:
- Mention rate: valid non-branded answers that name the business divided by valid non-branded answers.
- Citation rate: valid answers citing an owned domain divided by valid answers.
- Recommendation rate: valid non-branded answers that actively recommend the business divided by valid non-branded answers.
- Competitive share in this sample: the business's mentions divided by all mentions among the frozen competitor set.
The last metric is sample share, not Canadian market share. Changing the competitors or prompts changes the denominator.
Read the baseline as a decision ledger
A useful report lets the owner move from a finding to a bounded piece of work. Each material row should contain the question, platform, observed answer, cited sources, competitor context, accuracy check, alternative explanation, recommended change, owner, verification signal, and review date.
Suppose a commercial cleaning company serving the Greater Toronto Area is absent from “commercial cleaning for medical offices in Toronto,” while two competitors appear. The answer cites a trade directory and each competitor's healthcare service page.
The baseline has not proved why the company was absent. It has produced four testable possibilities: the company's service page does not clearly cover medical-office work, the public business facts conflict, relevant third-party sources do not mention the company, or the sampled answer simply varied. The next work might be a page correction, a profile and directory fact review, a source-development task, or a repeat run. Publishing a generic “best cleaners Toronto” article is not the automatic response.
Platform-owned reports answer only part of the question
Google's current guidance keeps normal SEO foundations in place for generative AI features. It also warns that third-party tools do not have access to Google's internal ranking or AI systems. In June 2026, Google announced a limited rollout of dedicated generative AI performance reports in Search Console with impressions, pages, countries, devices, and dates. Those fields describe visibility within the documented Google reports. They do not replace cross-platform prompt samples or prove a lead.
Bing Webmaster Tools' AI Performance public preview reports total citations, average cited pages, sampled grounding queries, page-level citation activity, and trends across supported Microsoft AI experiences. Microsoft explicitly says these fields do not indicate placement, authority, ranking, or the role of a page in an individual answer.
OpenAI separates OAI-SearchBot for ChatGPT search inclusion from GPTBot for potential model training and ChatGPT-User for user-directed actions. OpenAI also documents utm_source=chatgpt.com on referral URLs. That parameter can support referral analysis. It does not count unclicked mentions or prove how ChatGPT selected a source.
Perplexity and Anthropic publish distinct crawler and user-agent controls for search discovery, user-directed retrieval, and training-related access. Those controls help diagnose eligibility. They are not publisher performance dashboards. The baseline should state which evidence came from a platform report, a raw answer sample, crawler logs, analytics, or a human qualification step.
What MAXUOD would deliver
MAXUOD's GEO Proof Lab is the evidence layer behind this work. For a baseline, we freeze an approved question set, market, language, competitor group, and run protocol. We preserve the answer and citations, separate mentions from recommendations, check business facts, map source gaps, and attach each proposed change to a page or public record.
The deliverable is not a mystery score. It is a scope card, prompt register, answer and citation ledger, competitor comparison for that sample, fact-accuracy review, prioritized action register, and a dated re-test plan. Where Search Console, Bing Webmaster Tools, analytics, or server logs are available with permission, we add those records without merging them into the prompt sample.
Self-service works when the team can hold the cohort stable, inspect every answer, approve the business facts, and implement the changes. A focused GEO project is useful when several sources conflict, the site and profiles need coordinated edits, or the report needs to reach engineering, content, operations, and leadership with one owner per task.
What the baseline does not prove
A baseline does not prove a platform-wide rank, total audience, Canadian market share, model-training inclusion, cause of a mention, future answer, referral, lead, or sale. It cannot see every personalized prompt. It cannot make an opaque answer system deterministic.
It can show what a declared method observed on a declared date. That is enough to find inaccurate descriptions, missing source support, weak page ownership, competitor patterns, and measurement gaps. It is also enough to decide that no change is justified yet.
The next run should be boring
A good follow-up uses the same cohort and definitions, records what changed on the business side, repeats the sample, and reports movement with the same limits. If the business changes market, language, offer, or competitor set, start a new cohort instead of splicing it into the old trend.
We would review monthly for an active program and after a material release when a faster check serves a clear decision. The cadence matters less than comparability. A quiet, repeatable report is more useful than a dramatic score nobody can reproduce.
Method note: Platform documentation and the benchmark sources were checked on August 27, 2026. Answer systems, reports, crawler controls, models, and interfaces can change. This article states MAXUOD's professional method for educational purposes and does not promise indexing, rankings, AI mentions, citations, referrals, enquiries, or revenue.
Related reading and sources
Read next on MAXUOD
External references
Need a baseline that your team can inspect?
Start with the free SEO and GEO audit. MAXUOD can then scope the Canadian market, language, question cohort, source evidence, action owners, and re-test without selling an invented AI rank.



