The Flare score answers one question: when AI engines are asked the questions that matter in your category, how well do they present you? Not just whether you appear — how prominently, how accurately, how evenly across engines, and how favourably. This page documents exactly how it is computed, so the number is never a black box.
Why one number instead of four
Changed 11 August 2026. Two corrections were applied to every score and to the full history. Questions that name your brand no longer count toward any metric — an engine naming you when the question asked about you measures nothing. And a mention is attributed to your brand by matching its name and aliases, not by the judge's per-answer label, which was wrong often enough to matter. Across the catalog this moved average share of mentions from 0.54 to 0.26 and mention rate from 0.91 to 0.72. Scores fell accordingly; nothing about the brands changed, only what we were willing to count.
Mention rate alone rewards being talked about. It says nothing about whether you were named first or fifth, whether what was said was true, whether you show up on all four engines or only one, or whether the framing helped you. Brands with identical mention rates can be in completely different positions, and a single-metric leaderboard hides that.
Flare combines five measurements Brandflare already takes into one 0–100 number, and always shows its own working.
The five pillars
Each pillar is scored 0–100 by a fixed formula, then weighted:
| Pillar | Weight | What it measures |
|---|---|---|
| Visibility | 30% | How often engines name you at all — this is the AI Visibility Score |
| Prominence | 22% | Where you land in the answer, how forcefully you're recommended, and your share of the category's airtime |
| Accuracy | 22% | How much of what engines say about you is true |
| Consistency | 16% | How evenly you appear across all four engines |
| Sentiment | 10% | How favourably you're framed |
Visibility
Your mention rate across engines over the trailing window, as a percentage. A Visibility of 62 means you appeared in roughly 62% of engine answers to your panel. This is the metric Brandflare has always published as the AI Visibility Score, unchanged.
Prominence
Three sub-signals, blended 45/30/25:
- Rank — where your first meaningful mention falls. Position 1 scores 100, 1.3 scores 75, 2 scores 35, 2.6 scores 19, decaying smoothly and never quite reaching zero, because being named last beats not being named. The curve is steep because real average ranks are tightly bunched between 1.0 and 2.6.
- Strength — how forcefully the answer recommends you, graded by the judge: none 0, weak 25, moderate 60, strong 100.
- Share of voice — measured against your fair share, not raw. In a category with four tracked competitors, fair share is 20%. Tracked brands routinely take between 1.2x and 3.1x their fair share, so the scale runs from 0.8x (scoring 0) to 3.2x (scoring 100). Raw share of voice would punish brands in crowded categories for nothing more than having many competitors.
Accuracy
The severity-weighted share of judged answers carrying no open flag. A high-severity false claim costs five times what a low-severity one does. Dismissed flags don't count.
Accuracy abstains when it was never measured. Lite-tier catalog brands are judged for mentions only — their answers are never checked against a fact sheet. Having no flags there means nobody looked, not that everything was true, so the pillar drops out entirely rather than awarding an unearned 100. See Abstaining pillars below.
Consistency
The mean per-engine mention rate as a fraction of your best engine's, scaled so that 0.65 or below scores 0 and perfect evenness scores 100. A brand at 90% on ChatGPT and 10% everywhere else scores near zero; a brand at 60% across all four scores 100.
Consistency deliberately measures evenness, not level — Visibility already covers how much you're mentioned. An earlier version counted how many engines named you at least once, which turned out to be a bar every tracked brand clears, making the pillar a constant that told nobody anything.
Sentiment
Average sentiment toward your brand, scaled so +0.2 scores 0 and +0.9 scores 100.
The bounds are not a typo. AI engines are essentially never negative about tracked brands — across the catalog, average sentiment runs from about +0.34 to +0.84. Mapping the theoretical −1…+1 range onto 0–100 would make this pillar incapable of scoring below 67, so a tenth of the score would say the same thing about every brand.
Why the curves are calibrated, not textbook
Each pillar's formula is anchored to the range its signal actually occupies, measured across the live catalog, rather than to its theoretical bounds.
This is the difference between a score that ranks and a score that doesn't. Sentiment never goes negative. Average rank is never worse than about 2.6. Engine evenness is rarely below 0.7. Scored against textbook bounds, all three become near-constants, and averaging five signals that each waste two thirds of their range yields a composite where the weakest brand in the catalog and a merely average one look almost identical.
Two things this is not:
- It is not a percentile. The endpoints are fixed constants. Your score moves only when your own metrics move, never because a competitor improved — which is what makes the trend line mean anything.
- It is not adjusted quietly. Recalibration changes every published number, so it bumps the algorithm version recorded on every stored score and triggers a full recomputation of history. You can always tell which formula produced a given number.
Visibility is deliberately left uncalibrated: mention rate is a published figure with an external meaning — "you appear in 62% of answers" — and rescaling it would silently redefine a metric customers already track.
Abstaining pillars
A pillar with no evidence does not score zero, and it does not score 50. It abstains, and the remaining weights re-normalise to cover its share.
So a brand whose accuracy was never assessed is scored on the other four pillars, with their weights scaled up from 30/22/16/10 to 38/28/21/13. The breakdown on your brand page shows this — the percentages next to each pillar are always the weights actually used.
This also means a brand that was measured for accuracy and came back clean scores slightly higher than one never measured at all. That is intentional: a verified result is worth more than an absent one.
Thin data and confidence
Sweep cadence varies by plan — daily, weekly, or fortnightly — so the same window can hold two hundred judged answers for one brand and twenty-five for another. Scoring both the same way would let a single lucky sweep produce a 97.
Every pillar is therefore pulled toward the midpoint in proportion to how little evidence backs it. A pillar backed by twelve judged answers carries half its distance from neutral; one backed by a full 40-prompt panel across four engines is essentially untouched. Brands with very few answers are marked provisional until enough sweeps land.
The correction is deliberately mild at realistic sample sizes. Pitched too aggressively it compresses every brand toward the middle, which changes no rankings and just makes the scale less useful.
The window
Flare uses a trailing 16 days. That is long enough for fortnightly-cadence brands to have at least one sweep inside every window, and short enough to respond to real change within days.
Interpreting your score
There is no universal "good" score — a household name and a two-year-old challenger occupy different ranges. Scores in the live catalog run from about 39 to 86, with most brands between 60 and 80. Bands: 80+ is Blazing, 70–79 Bright, 60–69 Warm, 50–59 Faint, below 50 Dark.
A perfect 100 is unreachable in practice, and deliberately so — it would require being named in every answer, first, on every engine equally, with flawless accuracy and uniformly glowing framing.
Three readings matter:
- Direction — is the trend up after the work you shipped?
- Which pillar is dragging — the breakdown shows exactly where the points are going. A brand at 100 Presence and 20 Accuracy has a completely different problem from one at 40 Presence and 100 Accuracy.
- Gap — how far you sit from each tracked competitor on the identical panel.
Reproducibility
Because the panel is fixed, the engines are versioned, and every answer is stored, any Flare score decomposes to the pillar, to the per-engine metric, and down to the verbatim answers that produced it.
Every stored score records the algorithm version that produced it, so when weights are revised it stays clear what any historical number meant.
One known limitation. Flags carry only their current status, with no history. Backfilled scores from before Flare launched therefore use today's dismissals when reconstructing past accuracy, which can make a historical Accuracy pillar read slightly higher than it would have at the time. Scores computed from launch onward are unaffected.