How to read AI visibility scores by engine — without a blended score
Short answer: A single blended “AI visibility” number hides the real pattern: a brand can be strong in ChatGPT and empty in Gemini or AI Overview. Read per-engine columns — appears/not, position, cited sources — on a frozen prompt set. Then compare engine to engine. Do not average into one board slide.
Quick facts
- One blended score ≠ an audit. Averaging GPT+Gemini+AIO can look “fine” while two engines that matter to your buyers stay empty.
- Brand Score ≠ your content got cited. A brand name can appear without your owned URL being extracted; those are different signals.
- Freeze prompts first. Without a versioned, dated question set, you are comparing weather, not trend.
- Six columns beat one donut. GPT, Gemini, Perplexity, Copilot, AI Overview, AI Mode — report separately.
- Citation volume can thin upstream. If total sources in your monitor drop sharply week to week, SoV “up/down” can be a denominator artifact — flag it on the slide; do not stay silent.
- Fix = pages + corroboration, not swapping dashboards. Changing the chart without changing sources rarely moves answers.
Why blended scores mislead
Marketing teams are used to one KPI: rank, impressions, or “visibility %.” In AI answers, different surfaces read different sources. One engine may cite your owned FAQ; another cites an industry list or platform docs; a third shows nothing for the same prompt.
Folding that into one “AI SoV” lets the board feel progress — or panic — without knowing which engine is still dead. For content decisions, what helps is: on prompt X, on engine Y, is the brand named, in what position, citing which sources.
This extends general AI-visibility measurement: freeze prompts, read sources, repeat. The difference here is how you report so page decisions are not blurred by averages.
How to read a per-engine report (working steps)
1. Freeze 20–40 buyer prompts
Buyer language, not internal jargon. Mix: category without brand names, comparisons, and “service/product … in [market].” Save version + date. An audit without a version cannot be repeated.
2. Fill per-engine columns — do not average first
For each prompt × each engine, record at least:
| Column | Content |
|---|---|
| Appears? | Yes / no |
| Position / mention order | Number or “buried at the end” |
| Accuracy | Name & claims correct / wrong / partial |
| Source URL / domain | Owned, media, third-party list, other |
Only then, if needed, compute SoV per engine on the same prompt set. Do not blend first.
3. Separate three patterns before “action”
- Strong on one engine, empty on another — extraction/corroboration gap by surface, not “brand does not exist in AI.”
- Named without owned URL — the name appears; your page is not yet a source. Fix retrievability + pinned facts.
- Empty everywhere — no answer page yet, or third-party consensus is not aligned.
4. Decide on 3–5 pages, not 50 tips
Pick prompts with the largest gap and business value. Fix extractable structure (answer first, facts, table, FAQ). Re-measure the same columns +7–14 days after pages are live and crawlable — not same-day drama.
5. Put a volume caveat on the board slide
If your monitor shows total citations / sources shrinking week to week while a score “rises,” footnote it. Without that caveat, the board chart is cosmetic.
Table: engine × what to measure × typical reporting mistakes
| Engine / surface | Useful column read | Typical slide mistake |
|---|---|---|
| ChatGPT (GPT) | Appears / position / sources on chat prompts | Treating GPT position as “AI overall” |
| Gemini | Present/absent + whether the answer cites URLs | Hiding 0% Gemini inside a “good” average |
| AI Overview (AIO) | Summary present + domains in the Search layer | Equating AIO with classic blue-link rank |
| AI Mode | Position in AI Search mode (when measured) | Merging AI Mode with AIO without labels |
| Perplexity | Explicit citations / source lists | Using Perplexity cites as proof for every engine |
| Copilot | Present/absent on work/enterprise prompts | Ignoring Copilot as “not consumer” while B2B buyers use it |
One truth: there is no guarantee a brand enters any column next week. What you can control: clear facts, consistency, page structure, and aligned third-party traces.
Dual Legibility in one line
Dual Legibility Tension (White Wood): humans read “we are visible in AI”; machines read whether a claim can be retrieved again on that surface. A blended score serves slide ego; per-engine columns serve page fixes.
What this is not
- Selling or buying “one blended AI score” as a contract KPI.
- Promising “in Gemini / AIO after X articles.”
- Equating Brand Score overview with count of owned URLs cited.
- Filling slides with Prompt tips / “ChatGPT jobs” without a frozen prompt set.
- Chasing GEO procurement as a substitute for weak answer pages.
If you need the from-scratch measurement frame (freeze prompts, read sources, page decisions), that lives in the general how to measure AI visibility journal. Use this piece for how to report without hiding engine gaps.
FAQ
Can we show one SoV number to the board?
Yes as a secondary summary, after per-engine columns are visible. Do not let one number become the only KPI. A board that only sees a donut cannot decide which page to fix.
What does “visible in GPT, empty in Gemini or AIO” mean?
Usually: one surface already found enough sources; another has not extracted them or chose a different consensus. Next diagnosis: which sources GPT cites, and whether your page (or external corroboration) can be taken on Gemini/AIO with clearer structure — table, FAQ, pinned facts, crawlable text.
High Brand Score but our URLs are rarely cited — is that a win?
Not necessarily. Brand named in an answer ≠ owned page as source. For studio craft, track both: brand mention and extracted domain/URL. Without the second, you are praised without an auditable foundation.
Citation volume in our tools dropped — did the campaign fail?
Not necessarily. Upstream volume thinning (fewer source rows in the harvest) can move SoV without your content changing. Log total citations/sources week to week beside the score. If the denominator shrinks, discuss the caveat before celebrating or punishing the team.
Do podcasts help per-engine scores?
Only if there is a stable text page that can be cited — not video alone. Audio can strengthen human corroboration; machines still need text. Example of communications craft structured as a category article: https://proxemicspodcast.com/article-podcast-komunikasi.html
Do we need paid tools to read per engine?
Tools speed capture. What is mandatory: a versioned prompt set, column discipline, and page decisions. A dashboard with no action is just a new cosmetic score.
Checklist before the board slide
- Prompt set has version + date
- Separate columns: GPT · Gemini · AIO · AI Mode · (optional Perplexity/Copilot)
- Each large gap has a page URL to fix — not a tip list
- Brand Score and own-URL cites are not merged into one “win” claim
- Volume/denominator note if the harvest shrank
- No citation promises or blended score as a contract KPI
