Why the obvious methods fail
Two things businesses try first, both of which mislead:
Asking "do you know BAIDLABS?" This tests whether a name appears in training data. It tells you nothing about whether you'd be recommended to a customer, which is the question that matters commercially.
Checking once and drawing conclusions. AI answers are non-deterministic and retrieval runs live. The same question can name different businesses on Tuesday and Thursday. A single check is noise, not a measurement.
The prompt panel method
Step 1 — Build the panel
Write 20–50 questions a genuine prospect would ask, in their words, before they knew your name. Cover the range:
- Category discovery: "who can help me get my brand recommended by ChatGPT?"
- Local intent: "GEO agency in Jaipur," "AI search consultant India"
- Problem-shaped: "my organic traffic is falling but rankings are the same — why?"
- Comparison: "GEO agency vs SEO agency, which do I need?"
- Commercial: "how much does generative engine optimization cost?"
Freeze this list. Its value comes from being run unchanged over time — a panel you keep editing produces trends you can't interpret.
Step 2 — Define what counts
Score each result on three separate levels, because they mean different things:
| Level | Definition | What it tells you |
| Mentioned | Your name appears in the answer text | The engine knows you exist in this context |
| Cited | Your URL appears in the sources | Your content was actually retrieved and used |
| Recommended | You are named as a suggested option | The commercially meaningful outcome |
Also log sentiment (how you're characterised) and competitor set (who else appears). The competitor set is often the most actionable column — it shows you exactly who owns the answer you want.
Step 3 — Run it per engine, on a schedule
Run the full panel against each engine you care about — ChatGPT, Perplexity, Google AI Overviews, Gemini, Copilot, Claude — on the same day each week or fortnight. Use a fresh session with no personalization or history, since prior conversation contaminates results.
Keep engines in separate columns. Studies have found the domains cited by different engines overlap surprisingly little, so a blended average hides the thing you need to see.
Step 4 — Compute share of voice
For each engine: share of voice = prompts where you are recommended ÷ total prompts. Do the same for your top three competitors. That single ratio, tracked over time, is the clearest signal of whether your programme is working.
Step 5 — Connect it to money
Visibility that doesn't reach revenue is vanity. In your analytics, segment referral traffic from AI sources — chatgpt.com, perplexity.ai, gemini.google.com, copilot.microsoft.com and similar. Expect the volume to look trivially small and the conversion rate to look implausibly high; that's the documented pattern, not an error.
What a healthy baseline looks like
For most businesses running this for the first time, the honest starting picture is: mentioned on a handful of prompts, cited on almost none, recommended on zero. That is normal and it is not a failure — it's the baseline that makes every later improvement provable.
Document the zero. The most valuable thing about a first panel is that it captures where you started. Six months later, that record is the difference between "we think it's working" and a defensible before/after. Screenshot the answers, not just the scores.
Tooling
A spreadsheet run manually is entirely adequate to start, and forces you to actually read the answers — which is where the insight lives. Commercial AI-visibility platforms automate the runs and are worth it once you're tracking many prompts across many engines. Start manual, automate when the manual version becomes the bottleneck.
Frequently asked
How many prompts do I need? +
Twenty is enough to be meaningful; fifty is better. Below about fifteen, the noise from non-deterministic answers swamps the signal. Coverage across question types matters more than raw count.
How often should I run the panel? +
Weekly during an active optimization programme, monthly for maintenance. The key is consistency — same prompts, same day, same conditions — so that changes reflect reality rather than method drift.
Should I use a paid AI visibility tool? +
Eventually, yes, if you're tracking dozens of prompts across six engines — the manual work becomes impractical. But run it manually first. Reading the actual answers teaches you why you're absent, which a dashboard number never will.
Does personalization affect results? +
Yes, significantly. Always test in a logged-out or fresh session with no prior conversation, and be aware that location affects local queries. Note the conditions alongside the results.
Keep reading
Want this done for your brand?
Book a free GEO audit: where you appear across 6+ AI engines today, the prompts competitors own, and a prioritized 90-day plan you keep either way.
Book a free audit →