You can't manage what you can't measure, and right now most brands can't measure what AI engines say about them. Traffic analytics don't see AI conversations; rank trackers watch the wrong surface entirely. Measuring AI visibility means answering four questions with numbers: how often do engines mention you, how much of your category's attention do you own, how favorably are you framed, and is what the engines say actually true?
This guide covers the metrics that answer those questions — mention rate, share of voice, sentiment, and accuracy — plus the method that makes them trustworthy: fixed prompt panels, multiple engines, scheduled sampling, and consistent scoring. It closes with benchmarks for reading your numbers once you have them.
Why AI visibility is hard to measure
Three properties make AI answers a genuinely different measurement problem than search rankings:
Answers are non-deterministic. The same prompt, on the same engine, on the same day can name different brands across runs. Generation is probabilistic; when several brands have comparable standing, which ones surface varies. Any single answer is a coin flip; only the rate across repeated runs is a measurement.
Engines disagree. ChatGPT, Google Gemini, Perplexity, and Claude have different training corpora, different retrieval pipelines, and different citation habits — the mechanics are unpacked in how AI engines recommend brands. A brand can dominate Perplexity's answers and barely exist in Gemini's. One engine is a sample of one; you need at least the four majors.
There's no query-volume data. SEO measurement starts from keyword volumes; nothing equivalent exists for prompts. You can't measure "all the questions people ask" — you can only define a representative panel of the questions that matter commercially and track it consistently. Panel design is therefore the foundation everything else stands on.
The four core metrics
Mention rate: are you in the answer?
Mention rate is the fraction of answers, across your tracked prompt set, in which your brand is named at least once. If you track 40 prompts across 4 engines and your brand appears in 48 of the 160 answers, your mention rate is 30 percent. It's the atomic metric of AI visibility — every brand mention either happens or doesn't, and the rate aggregates those binary events into a number that can trend.
Mention rate is most useful sliced: per engine (revealing which engines "know" you), per prompt theme (revealing which buyer questions you win), and over time (revealing whether anything you shipped moved the needle).
Share of voice: how much of the category do you own?
Mention rate treats your brand in isolation; share of voice puts it in context. It's the percentage of all brand mentions in your category's answers that belong to you. If answers across your panel name brands 500 times and 75 of those are you, your share of voice is 15 percent.
Share of voice matters because AI answers are a zero-sum surface: a typical answer names three to five brands, and every mention a competitor earns is a slot you didn't. The derived metric — the competitor gap, your visibility minus a rival's on the same prompts — is the single clearest signal of who is winning a category's AI answers, and its trend tells you whether you're closing or falling behind. The mechanics of running this comparison are in how to track competitors in AI answers.
Sentiment: how are you framed?
Two brands with identical mention rates can be in very different positions: one recommended first and enthusiastically, the other listed last with a caveat about pricing. Brand sentiment captures that framing — typically graded as recommended, neutral, or mentioned-with-caveats — along with prominence signals like whether you're named first.
Sentiment is where qualitative reading matters. A falling sentiment trend with a stable mention rate usually means engines are picking up critical third-party material — worth finding before it hardens into the default framing.
Accuracy: is any of it true?
The metric SEO never needed. Engines state facts about brands — prices, features, availability, history — and some of those facts are hallucinations. Accuracy measurement checks the claims in each answer against a set of verified brand facts and flags affirmative false claims, graded by severity. A wrong founding year is cosmetic; a wrong price or a "discontinued" product that's very much for sale costs revenue directly.
Note the word affirmative: an engine omitting a feature isn't an error, but an engine asserting something false is. That distinction keeps the metric honest and actionable — every flag corresponds to a claim you can trace to a source and fix. Brandflare's implementation of this check is documented in accuracy monitoring and flags.
Supporting metric: citation rate
If your content strategy targets being quoted as a source, track citation rate — how often your domain appears among an answer's cited sources — separately from mention rate. Brand named but never cited means engines learn about you secondhand; cited but never named means your content works harder for the category than for you.
The method: how to measure reliably
Metrics are only as good as the sampling behind them. The method has four parts.
1. Build a fixed prompt panel
Define a panel — typically 20 to 50 prompts — representing the questions your buyers actually ask: best-in-category ("best X for Y"), comparisons ("A vs B"), use-case fits, and direct brand questions ("is X any good?"). Phrase them the way people talk to a chatbot, not the way they type keywords.
Then freeze it. The panel must stay unchanged across measurement periods, or your trend line measures your edits instead of your visibility. Add prompts deliberately and version the change. This is prompt tracking — the AI-era descendant of keyword rank tracking — and the panel-design details are in prompts and the prompt matrix.
2. Run every prompt against every engine, on a schedule
One complete pass — every prompt against every tracked engine, with answers stored and scored as a dated snapshot — is a sweep. Scheduled sweeps are what turn anecdotes into trends: weekly cadence is the workable minimum for stable averages; daily cadence catches shifts (a model update, a new competitor entering answers, a hallucination appearing) within a day or two instead of a month. How Brandflare schedules and executes these passes is covered in how sweeps work.
3. Score answers consistently
Every answer needs the same evaluation: is the brand mentioned, in what position, with what framing, and are the claims true? At panel scale that's hundreds of answers per sweep — far past human review, and beyond what keyword matching can score, since answers refer to brands obliquely and sentiment lives in phrasing. The workable approach is LLM-as-judge: a strong model evaluating each answer against explicit rubrics and your verified fact set, so the scoring is uniform across engines, sweeps, and time.
4. Aggregate into a headline score
Executives won't read a prompt matrix. A composite AI Visibility Score — Brandflare's is a 0 to 100 number built from mention rates across engines over a trailing window, with the full methodology public — gives you one trendable headline, with the per-prompt, per-engine detail underneath for diagnosis. The score is the dashboard; the matrix is the debugger.
Can you do all this manually? For a one-time baseline, yes: pick ten prompts, run them across four engines in fresh sessions, log mentions in a spreadsheet. It's a worthwhile afternoon — and the fastest version of it is Brandflare's free audit, which runs a category panel across all four engines without a signup. What's impractical manually is the ongoing part: same prompts, every engine, every week, scored the same way, forever. Consistency is the product.
Benchmarks: what do the numbers mean?
Absolute numbers mean little outside your category; context comes from two comparisons.
Against your competitors. The only benchmark that's always valid, because it's measured on the same prompts, engines, and window. Directional patterns worth knowing: recognized category leaders tend to score high (often 70-plus on a 0 to 100 visibility score), established players land mid-range, and challengers commonly start below 30 — sometimes near zero on memory-heavy engines. A fuller discussion of ranges is in what is a good AI visibility score.
Against your own history. Trend beats level. A challenger moving from 12 to 25 over two quarters is a success story; a leader drifting from 80 to 70 is an early warning no absolute threshold would catch.
Reading the slices diagnostically: strong on retrieval-heavy engines (Perplexity) but weak on memory-heavy ones usually means your live web presence outruns your training-data footprint — expect the gap to close over model release cycles if coverage continues. Strong on brand-name prompts but absent on category prompts means engines know of you but don't associate you with the category — a positioning and coverage problem, not an awareness one.
The bottom line
Measuring AI visibility is a sampling discipline: a frozen panel of buyer questions, run against multiple engines on a schedule, scored consistently for mention, share of voice, sentiment, and accuracy, aggregated into a trendable score. Skip any part and the numbers stop being comparable — and comparability over time is the entire point. Get the measurement loop running first, even crudely; every improvement tactic you try afterward becomes an experiment with a readout instead of a hope.