Why the old metric doesn't stretch
Share of voice was built for media you could sample: press mentions, social posts, search impressions. AI answers break the sampling model in two ways. First, they're generated fresh per conversation — there's no archive to count. Second, they're winner-take-most: an answer names three brands and omits thirty, with no long tail of page-two visibility to partially credit. You can't survey what customers were told; you have to ask the same questions yourself and count who shows up.
Building the metric properly
Fix the question set
Pick 10–30 questions your customers actually ask: category recommendations, "alternatives to X," "is Y worth it," the works. This set is your panel — changing it mid-stream breaks your trend line, so draft it carefully and version it when you must change it.
Ask all four engines, repeatedly
One engine is a sample size of one, and engines genuinely differ — Perplexity might cite you into every answer while Claude has never heard of you. Ask ChatGPT, Gemini, Perplexity, and Claude the same questions on a schedule. At Brandflare we call each run a sweep, and the differences between engines are routinely the most useful finding.
Count presence, then quality
The core number is presence: the share of answers that name you. That's the AI Visibility Score — averaged across engines over a trailing window (we use 7 days of sweeps, all four engines, 0–100). Layer quality on top: When named, are the facts right? Are you first-mentioned or an afterthought? Named enthusiastically or with caveats?
Reading the number like an analyst
- Level tells you standing. A 40 means customers asking AI get told about you a bit under half the time. Your competitor at 70 is on shortlists you're not.
- Trend tells you whether work is working. Published the comparison content, fixed the stale sources? The next sweeps are your verdict. Retrieval-driven movement shows in weeks.
- Per-engine spread tells you where the problem lives. High on Gemini, low on Claude usually means your ranked content is fine but your broader footprint is thin — a memory problem wearing a metrics costume.
The honest caveats
Engines are stochastic; the same question can produce slightly different answers on different days. That's why you average over a window instead of panicking over a single run. And share of answer measures visibility, not revenue — though a metric that counts "was my brand on the shortlist customers saw" sits about as close to revenue as brand measurement gets.
