
How to Track AI Visibility: Methods, Metrics, and Cadence
By be—recommended Team
TL;DR: AI visibility tracking means running a fixed set of realistic buying-intent prompts across the major answer engines on a schedule, scoring each response for mentions, recommendations, sentiment, and competitor presence, and following the trend. Spot checks mislead because answers vary between sessions; the distribution over repeated runs is the real signal.
Why spot-checking your brand in ChatGPT misleads
The instinctive check — open ChatGPT, ask "best [category] tools," see if you appear — produces a coin-flip observation. AI answers vary with phrasing, session context, sampling randomness, and whether browsing triggered. You can be recommended at 9:00 and absent at 9:05 on the same prompt.
Meaningful tracking treats AI answers the way pollsters treat voters: many samples, fixed methodology, trend over time.
The tracking method
1. Build a prompt set that mirrors real buyers
15–25 prompts covering:
- Category prompts: "best accounting software for freelancers"
- Problem prompts: "how do I keep track of invoices as a contractor"
- Comparison prompts: "[you] vs [competitor]"
- Brand prompts: "is [your brand] any good?"
- Local variants if you sell geographically: same questions with a city or country attached, in the local language.
Write them the way customers type, not the way marketers do. Keep the set frozen — changing prompts mid-stream destroys your trend line.
2. Run across engines, not just one
Visibility is engine-specific. In our audits it is routine for a brand to be recommended by Perplexity (which retrieves live pages aggressively) while being invisible to ChatGPT (whose answer may draw on model knowledge), or the reverse. Track at minimum ChatGPT, Gemini, and Perplexity — with web search active where the engine supports it, because that is how real users increasingly encounter answers.
3. Score each response consistently
For every prompt × engine run, record:
- Outcome tier: recommended (named as a pick) / mentioned (appears neutrally) / absent
- Position: first pick or fifth?
- Sentiment and framing: confident recommendation, hedged mention, or a caveat ("some users report…")
- Competitors present: who occupies the slots
- Sources cited: which pages put those brands there — this is your action list
4. Compress into a score
Individual data points are noisy; a weighted aggregate is stable. That is the idea behind an AI Visibility Score: recommendations weigh more than mentions, buying-intent prompts weigh more than informational ones, and the result is a 0–100 number you can report monthly and hold against interventions.
5. Set a cadence and stick to it
- Monthly full runs suit most brands — enough to catch model updates and competitor moves without drowning in noise.
- After major events, run off-cycle: a big PR hit, a product launch, a known model release.
- Weekly only if you are actively campaigning and expect fast movement (e.g. a roundup-placement push).
Supporting signals worth watching
Prompt tracking is the core, but corroborate it with:
- AI referral traffic: chatgpt.com, perplexity.ai, and gemini.google.com referrers in your analytics. ChatGPT''s link formatting now makes this measurable in GA4.
- AI crawler activity: GPTBot, OAI-SearchBot, PerplexityBot hits in your server logs — evidence your content is being read at all.
- Citation inventory: the third-party pages engines cite in your category, and whether you are on them yet.
Interpreting movement
- Score up after a roundup placement: the placement is being retrieved — double down on that publisher class.
- Score down with no action on your side: check for a model update, a competitor''s new placement, or a lost citation (page updated, you edited out).
- Mentioned but never recommended: engines know you exist but lack evidence of excellence — typically a review-depth or sentiment problem, not an awareness one.
- Absent on one engine only: engine-specific gap; diagnose that engine''s cited sources rather than changing global strategy.
The most valuable diagnostic remains the direct follow-up: when an engine omits you, ask it why. The answers are specific and usually actionable.
Build vs. buy
You can run this manually — a spreadsheet, a monthly hour, discipline. It works at small scale but erodes: manual runs drift in phrasing, skip months, and nobody scores sentiment consistently at 75 prompt-engine pairs.
Purpose-built tracking automates the runs, keeps the methodology frozen, asks the follow-up questions, and turns the output into a score and a gap list. Our free analysis is the fastest way to get a baseline: real responses from ChatGPT, Gemini, and Perplexity to the questions your customers ask, scored into one number.
FAQ
How many prompts do I need for a reliable signal? 15–25 well-chosen prompts across three engines gives 45–75 data points per run — enough for a stable aggregate while staying reviewable by a human.
Should I track with or without web search enabled? With, primarily — it reflects how the engines actually answer for most users now. Tracking both modes separately is a bonus: the gap between them tells you whether your visibility lives in the model''s knowledge or only in live retrieval.
What is a good AI Visibility Score? Category-dependent — the honest benchmarks are your own trend and your named competitors measured the same way. Beating last quarter and closing on the category leader matter; an absolute number without context does not.
Stay ahead of the AI curve
Get weekly insights on AI visibility, search optimization, and brand strategy — straight to your inbox.