The AI Visibility Score answers one question: when AI engines are asked the questions that matter in your category, how often are you part of the answer? This page documents exactly how Brandflare computes it, so the number is never a black box.
The AI Visibility Score is now the Presence pillar of the Flare score, which is the headline number on your dashboard and across Explore. Everything on this page still applies — Presence is this measurement, unchanged, contributing 30% of Flare. The other four pillars cover how prominently, how accurately, how evenly, and how favourably engines present you.
The pipeline in brief
Every score is the product of a four-stage pipeline:
- Ask — a panel of category-relevant prompts is run against four engines: OpenAI (ChatGPT), Google Gemini, Perplexity, and Anthropic's Claude.
- Measure — each answer is parsed for brand mentions: yours and every tracked competitor's.
- Judge — a frontier-model judge evaluates each mention's presence and framing, and checks factual claims against your verified brand facts.
- Aggregate — per-engine, per-prompt results roll up into daily metrics and the headline 0–100 score.
The prompt panel
Measurement starts from a fixed panel of prompts — the questions a real buyer in your category asks: "best X for Y", "X vs Y", "is X worth it", recommendation and comparison queries. Panels are generated per brand from its category and competitors when the brand is onboarded, and stay fixed thereafter, because trend lines only mean something on a stable yardstick. Plan tier sets the panel size (see plans and limits).
Each sweep runs the full panel against all four engines — details in how sweeps work.
Mention detection
Raw engine answers are unstructured prose, so mention counting is itself an AI task. Brandflare uses an LLM-as-judge — a frontier model prompted narrowly — to determine, per answer:
- whether the brand is named (a brand mention), including aliases and product names registered for the brand;
- which competitors are named alongside it;
- how the brand is framed when it appears.
From these judgments come the two building-block metrics:
- Mention rate — the fraction of answers naming the brand.
- Share of voice — the brand's percentage of all brand mentions in the panel's answers.
The score
The headline AI Visibility Score condenses recent mention rates into a single 0–100 number: the average mention rate across engines over the trailing window, expressed as a percentage. A score of 62 means that across the last window of sweeps, the brand appeared in roughly 62% of engine answers to its panel.
Design choices worth knowing:
- Recency-weighted window. The score reflects the trailing sweeps (roughly the last week of measurements), so it responds to real change within days while smoothing single-answer noise.
- All engines count equally. Engines disagree — that disagreement is signal, and per-engine breakdowns are always visible alongside the composite.
- Non-determinism is averaged out, not hidden. The same prompt can yield different answers run-to-run; scoring across a full panel × four engines × multiple sweeps is what makes the number stable enough to act on.
Public catalog pages on Explore show this same score for every brand Brandflare tracks openly, along with its 14-day delta.
Accuracy is measured separately
Visibility and truthfulness are separate measurements. Alongside mention detection, the judge compares every claim an engine makes about your brand against your verified fact sheet, and raises severity-graded flags for affirmative false claims — wrong prices, discontinued products described as current, invented features. Omissions are never flagged; only things AI states that are wrong. Full detail: accuracy monitoring and flags.
They remain separate measurements, but both now feed the headline: accuracy is its own pillar of the Flare score, carrying 22% alongside Presence's 30%. A brand that is named constantly but described wrongly no longer looks identical to one that is named constantly and described correctly.
Interpreting your score
There is no universal "good" score — a household name and a two-year-old challenger occupy different ranges (see what is a good AI visibility score?). The three readings that matter:
- Direction — is the trailing trend up after the work you shipped?
- Gap — how far are you from each tracked competitor on the identical panel (competitor gap)?
- Distribution — which engines and which prompts drive the number? The prompt matrix shows exactly where you're absent.
Reproducibility
Because the panel is fixed, the engines are versioned, and every answer is stored, any score can be decomposed into the specific answers that produced it — from the dashboard, down through per-engine results, to the verbatim response text. If a number surprises you, the receipts are one click away.