How to audit brand mentions on ChatGPT, Gemini, and Perplexity
Most teams asked to audit AI brand visibility pick the wrong prompts and conclude the models do not know them. The manual version takes an afternoon, the right prompt design takes five minutes.
Most teams asked to audit their AI brand visibility open ChatGPT, type their company name, and report back that the model knows them. That is not the audit. That is the echo test. The actual audit asks the question a buyer would ask, then watches whether your firm surfaces unprompted in the response.
The gap between those two tests is the whole problem. An organisation that scores 9 out of 9 on the echo test ("tell me about [name]") can score 0 out of 12 on the prompts that drive shortlist decisions ("best [category] in [city]"). That second number is the one that decides whether you appear when nobody types your name into the bar. Call it the prompt floor. Most audits never measure it.
Here is the manual version, broken into four parts, runnable on any business in an afternoon.

Design the prompts a buyer would actually use
Pick eight prompts. Three industry-level open-ended ("best [industry] in [city]"). Three niche-level ("best [specific subcategory] for [audience]"). Two echo prompts ("tell me about [name]", "what does [name] do?"). The first three are where shortlists get made. They are also where competitor names surface in your place. The middle three filter broad presence from real category authority. The last two confirm the model knows you exist at all, which is the floor below which any fix is structural, not editorial.
Avoid leading words. "Best" is acceptable. "Most trusted", "top-rated", "leading" are not. They lean the model toward marketing language and bias the answer.
Run each prompt three times across four engines
ChatGPT, Gemini, Perplexity, and Claude. Run each prompt three times per engine because the models are stochastic. A name that surfaces in run 1 may vanish in run 2, so the majority vote across attempts becomes your data point. The arithmetic works out to 8 prompts multiplied by 4 engines multiplied by 3 runs. Ninety-six API calls in total. An afternoon if you do this manually with copy-paste. A single minute if you use the free AI Brand Mention Checker Foreground published for exactly this workflow.
Log three things per response. Whether your firm appeared at all (yes or no). The competitors named in your place. Sentiment of the mention: positive, neutral, negative, or absent.
Pure mention count is the wrong metric
A company named in 24 out of 24 industry-level prompts but never on the niche-level prompts has broad presence and no category depth. Another rated highly on niche but absent on industry has the opposite problem. Score them separately. The open AI brand visibility dataset Foreground publishes for the GCC region uses a weighted score (50% industry, 30% niche, 20% echo) that exposes which axis is broken inside the diagnostic.
Co-citation is the other signal worth tracking. When a rival is named in your place, that is data. When the same trio of competitors shows up across all four assistants, you are looking at the cluster the algorithms have learned to represent the category. Your job is to get into that cluster, not to displace any specific incumbent.
From audit output to repair playbook
The diagnostic output is a four-engine table with one of three failure modes per platform. The name is unknown (entity foundation problem, fix with the visibility stack). It is known but never recommended (citation density issue, fix with third-party placements and content built to be quoted on G2, Reddit, or trade press). It appears only in adjacent categories (positioning gap, fix with brand-disambiguation content via Wikidata claims and About-page rewrites). The recommendation gap post breaks down which signals each engine actually weighs, and the AI Visibility Foundation Fix maps the 10-deliverable repair sequence end-to-end.
The reason this exercise takes an afternoon and not a week is that the mechanics are public. The retrieval layer ranks candidates based on entity grounding, third-party citation density, sameAs networks, and schema-readable identity claims (AI engines retrieve, they do not rank). Nothing about the diagnostic requires proprietary access. Nothing about the remediation does either.
If you have run the echo test and stopped there, you have completed roughly a quarter of the work. The other three quarters are where the real recommendation gap lives, and where shortlist economics quietly shift.
Written by Foreground Digital. Start a project →