AI engines state things about brands with total confidence — and some of those things are wrong. A discontinued product recommended as current, a price that's two years stale, a feature you never built. Brandflare's accuracy monitoring exists to catch these hallucinations in the answers real buyers are seeing, before customers act on them.
The fact sheet
Accuracy monitoring is only as good as its ground truth, and yours is the brand fact sheet: a list of short, checkable statements about your brand — what it sells, what it costs, what's current, what's discontinued, who founded it, where it operates.
Facts arrive in two states. When your brand is set up, Brandflare drafts facts automatically; these are used for checking from day one. As you review them, you verify the ones that are correct — and correct or remove any that aren't. Both automatically drafted and verified facts are checked against answers, so the practical difference is confidence: a flag against a verified fact is one you can act on immediately.
Keep the sheet current. When your pricing changes or a product ships or sunsets, update the fact sheet first — otherwise the judge is checking answers against yesterday's truth.
How checking works
During every sweep, after each engine answers each prompt, an LLM-as-judge reads the full answer and compares every claim it makes about your brand against your fact sheet. When a claim contradicts a fact — or asserts something false about your brand outright — the judge raises a flag recording:
- The claim — what the engine actually said, quoted from the answer.
- An explanation — why the judge considers it false, and which fact it contradicts, when it contradicts a specific one.
- A severity grade — how damaging the claim is. A slightly outdated founding year is not the same as "this product has been discontinued."
Every flag is tied to the run that produced it, so you can always read the verbatim answer in full context — which engine said it, in response to which prompt, on which date.
Only affirmative false claims — never omissions
This is the design decision worth understanding. Brandflare flags only affirmative false claims: things an engine positively states about your brand that are wrong. It never flags omissions.
If ChatGPT fails to mention your enterprise tier, or leaves you out of a listicle entirely, that is a visibility problem — it shows up in your mention rate, your prompt matrix, and your score. It is not an accuracy problem, and treating it as one would bury real hallucinations under an infinite list of things AI didn't say. Visibility and truthfulness are deliberately separate measurements: the AI Visibility Score tracks how often you appear; flags track whether what's said about you is true.
Working with open flags
New flags surface on your brand dashboard as open flags, ordered so the severe ones are impossible to miss. For each flag, the useful questions are:
- Is it real? Judges are strong but not infallible, and a stale fact in your sheet can produce a false flag. Read the answer, check the fact.
- Is it recurring? A claim that appears once may be sampling noise. The same false claim across multiple sweeps or engines points at a stale or wrong source the engines keep drawing from.
- Where did it come from? Persistent hallucinations usually trace to an incorrect page the engines retrieve or memorized. Finding and fixing that source is the actual remediation — the playbook is in finding and fixing AI hallucinations, and the realistic expectations are covered in can I correct what AI says about my brand?
Because sweeps keep running, you don't have to guess whether a fix worked: watch whether the claim stops appearing in subsequent sweeps.
Alerts
On plans that include email alerts, new high-severity flags trigger an email after the sweep that found them — so an engine asserting something damaging and false about your brand reaches your inbox the same day, not whenever you next check the dashboard. Alert availability by tier is listed in plans and limits.
What accuracy monitoring is not
It isn't sentiment policing — an engine being lukewarm about your brand is not a false claim. It isn't a guarantee of exhaustiveness — the judge checks the answers your panel elicits, not every conversation happening on every engine. And it isn't an edit button for AI: Brandflare finds the false claims and gives you the receipts; correcting the sources that feed them is work the platform helps you target, verify, and confirm.