A finance brand we spoke with had been tracking its own citation rate across ChatGPT and Perplexity for three months. The number was climbing. Nobody could say whether it was good. Climbing from what, compared to whom, against what standard? That is the problem with a blank baseline: it measures change but says nothing about position. An AI visibility benchmark built from category data fixes that, and the data to build one is already public.
Key Takeaways
- A single brand’s AI citation trend means little without a category benchmark to compare it against.
- Public ranking data, existing SEO rankings, review sites, industry lists, gives you a category set without a cold start.
- Share of Citation and Share of Voice are different measurements, and a benchmark must track both to be useful.
Why a Blank Baseline Fails You
Start from zero and every early result looks like progress. A brand gets cited twice in a week where it was cited zero times before. That reads as momentum. It might just be noise, or it might be the entire category getting more AI coverage that month for reasons that have nothing to do with your work.
The fix is not more frequent measurement. It is a wider frame. A benchmark answers the question a raw trend line cannot: relative to the five or ten brands a buyer would actually compare you against, where do you sit right now, this week, on this engine? Without that frame, a marketing lead reporting “up 40% month over month” to a CFO is reporting a number that could mean anything.
Public Ranking Data Is a Head Start, Not a Detour
Most mid-market and enterprise brands already have years of category intelligence sitting in tools they use for something else. Google’s organic rankings for category terms tell you who ranks for the questions buyers ask. Industry award lists and analyst reports name the players a category considers credible. Review platforms like G2 or Trustpilot rank competitors by volume and sentiment. None of this was built for AI visibility tracking, but all of it defines the competitive set an AI model is drawing from when it answers a category question.
Building a benchmark starts with assembling that set, not with guessing at competitors from memory. Pull the top 10 to 15 organic results for your three or four highest-intent category queries. Cross-reference against review site leaderboards. The overlap, brands that show up in both organic search and independent review rankings, is your benchmark cohort. This is the group whose AI presence you now track alongside your own.
growth in monthly AI mentions for Moburst’s own brand, with 42,435 total AI citations and the number 1 position in category.See the case study
That number came from tracking Moburst’s own position against a defined competitive set on the same engines, on the same cadence, not from watching a single trend line in isolation. The category frame is what turned a mention count into a claim worth reporting.
Share of Citation Is Not Share of Voice, and Your Benchmark Has to Track Both
A brand can dominate organic search, get quoted constantly in trade press, and still show up nowhere when an AI model answers a category question. Share of Voice measures how often you are mentioned across the web. Share of Citation measures how often a model names you as a source in its answer. These move independently, and a benchmark that conflates them will mislead the team reading it.
A brand can have high Share of Voice and near-zero Share of Citation, and a benchmark that does not separate the two will hide the exact gap it exists to find.
Build two columns for every competitor in your set, not one. The first tracks general mention volume across the web, standard PR and SEO monitoring. The second tracks how often each brand actually gets named and linked when your target queries are put to ChatGPT, Gemini, Perplexity, Claude and Grok. A competitor with strong Share of Voice and weak Share of Citation has a content and structure problem, not a visibility problem. A competitor with the reverse is winning on trust signals you have not built yet. The gap between the two columns is where the actual work is.
What to Actually Measure Once the Cohort Is Set
A benchmark is only as good as its query set. Pick the questions your buyers actually ask, not the keywords you already rank for. A buyer researching a VA loan does not type “va lending rates,” they ask something closer to “what is the best lender for a VA loan with no down payment.” That is the phrasing to test across engines, because that is the phrasing the model is answering.
For each query, log four things per engine: whether your brand appears at all, its position when it does, whether it is cited with a link or just mentioned by name, and which competitors appear alongside it. Run this weekly, not monthly. AI answers shift faster than search rankings because models refresh retrieval more often than they refresh training data, and a benchmark measured monthly will miss the movement that matters.
Different engines pull from different layers, so expect the benchmark to disagree with itself across platforms. ChatGPT pulls from across the open web, so broad coverage moves the needle. Gemini cross-references Search, YouTube and Scholar, so video and academic-adjacent content matter more there. Claude leans on high-authority publications and documentation, rewarding structured, well-sourced content over volume. Perplexity cites its sources inline, making it the easiest engine to audit directly. Grok reads the real-time social web, so a benchmark on Grok needs to include social mention data the other four do not.
Keeping the Benchmark Honest Over Time
Competitive sets shift. A brand that enters the category through an acquisition or a funding round can jump into AI answers within weeks if its coverage is wide enough, even without deliberate AEO work. Revisit the cohort quarterly using the same public-data method that built it: rerun the organic query check, recheck the review site leaderboards, and add or drop competitors based on what actually shows up now, not what showed up when you started.
The other honesty check is against your own SEO and PR teams’ existing work. If those teams have historical ranking data going back further than your AI tracking, use it to sanity-check whether a competitor’s AI citation gains track with a known content push, a PR campaign, or a structured data rollout. Citations compound: once a model cites a brand, the next citation becomes more likely. If a competitor’s Share of Citation jumped suddenly, there is usually a traceable cause in the public record, and finding it tells you what to replicate.
Where to Start This Week
Do not wait for a full audit before building the first version of this benchmark. Pull your top four category queries today, run them through the five major engines by hand, and log who gets cited and who does not. That single pass, done in an afternoon, already beats a baseline of one.
From there, formalize the cohort using the organic and review-site overlap method above, split tracking into Share of Voice and Share of Citation columns, and set a weekly cadence. If the manual version becomes too heavy to sustain across five engines and a growing query list, that is the point at which structured tracking through a tool built for this, like the Answerburst approach to AI visibility auditing, starts paying for itself instead of costing a Friday afternoon.
FAQs
What Counts As Public Ranking Data For This Purpose?
Organic search results for category queries, review platform leaderboards such as G2 or Trustpilot, industry award or analyst lists, and any existing SEO ranking history your team already tracks. None of it was built for AI visibility, but all of it defines who the category considers credible, which is the same signal AI models draw on.
How Often Should a Category Benchmark Be Updated?
Track query results weekly, since AI answers shift faster than organic rankings. Revisit the competitor cohort itself quarterly, rerunning the same public-data method used to build it, since new entrants can appear in AI answers within weeks of a funding round or acquisition.
Why Track Share of Citation Separately From Share of Voice?
They measure different things and often move in opposite directions. A brand can be mentioned constantly across the web, high Share of Voice, while an AI model rarely names it as a source, low Share of Citation. Tracking them separately shows whether a competitive gap is about content and trust signals or general visibility.
Does This Replace Working With an In-House SEO Team?
No. The benchmark method draws directly on data an SEO team already owns, organic rankings, historical performance, existing keyword research. It extends that work into the AI layer rather than starting a separate process alongside it.