How to Audit Your Brand’s AI Visibility (10-Step Guide 2026)

Last updated: August 10, 2026

TL;DR
An AI visibility audit checks how a defined prompt set describes and cites your brand across selected AI surfaces, then reviews the content, source, and technical factors behind those results. LLM Pulse offers a free directional report for a small prompt sample; a full audit adds broader prompt coverage, competitor comparisons, source analysis, and a prioritised action plan.

Your buyers are asking AI assistants the questions they used to type into Google. The catch is that you cannot see those conversations. There is no Search Console for ChatGPT, no rank report for Perplexity. The only way to know what those systems say about your brand is to ask them, capture the answers, and turn the patterns into a plan.

This guide is the structured version of that work. Ten steps, a fixed set of prompt categories, a simple tracking sheet, and a clear deliverable at the end. The whole thing fits in one day if you stay disciplined. We end with an honest comparison against the LLM Pulse free AI Visibility Report, which automates the same checks if you want a baseline in five minutes instead of eight hours.

What an AI visibility audit answers (and what it doesn’t)

An AI visibility audit is a structured analysis of how your brand surfaces across large language model outputs. It is not the same as an SEO audit, an AEO audit, or a PR sentiment report, although it touches all three. It answers four specific questions:

  • Presence: When a buyer asks a question that should surface your brand, does it appear?
  • Position: When you appear, where in the answer? First mention, middle of a list, or buried at the bottom?
  • Accuracy: Is what the model says about you correct? Founding year, pricing, feature set, headquarters, founder names.
  • Sentiment and framing: Are you described as the leader, the cheaper alternative, the enterprise option, the niche player? Who do you get compared to?

What the audit does not do: it does not predict revenue, it does not tell you exactly which content asset to publish next, and it does not give you Google-style keyword volume. AI outputs are probabilistic, so a single bad answer is not a crisis. Patterns across dozens of prompts are.

One more honest caveat. LLM responses are partially personalised based on session history and account context. Two people asking the same question can see different answers. The fix is to run audits in private or incognito sessions, logged out where possible, and to run the same prompt at least twice on different days before declaring a result. That way you measure the model’s default behaviour, not your own search history.

Step 1: Define your prompt universe

This is the step most teams skip, and it is the one that determines whether your audit is useful or theatre. You need a prompt set that mirrors how your buyers actually talk to AI assistants. Generic “what is the best CRM” prompts will give you generic results.

Pull from five categories. Aim for 30 to 50 prompts in total for a first audit. Below 30 you miss patterns. Above 50 you will run out of working day.

Branded prompts

These check accuracy more than presence. The model already knows the brand name, so it will say something. The question is whether what it says is true.

  • What is [Brand]?
  • What does [Brand] do?
  • How much does [Brand] cost?
  • Who founded [Brand]?
  • Where is [Brand] headquartered?
  • What are the main features of [Brand]?

Navigational prompts

Buyers who already know you but want a fact. These are how you spot wrong pricing tiers, outdated team pages, and stale integration lists.

  • [Brand] pricing
  • [Brand] login
  • [Brand] integrations
  • [Brand] free trial
  • [Brand] cancellation policy

Transactional prompts

The high-intent ones. A buyer with budget, ready to choose. These show whether you make it into the consideration set at all.

  • Best [category] software for [audience]
  • Top [category] tools in 2026
  • Recommended [category] platforms for small teams
  • Affordable [category] solutions under [budget]

Comparison prompts

Where positioning gets decided. If you never show up in head-to-heads, you are not in the conversation when buyers shortlist.

  • [Brand] vs [Competitor A]
  • [Competitor A] vs [Competitor B] (yes, audit prompts where you are not the subject)
  • Alternatives to [Competitor A]
  • Best [Competitor A] competitors

Problem-driven prompts

The questions a buyer asks before they know which category they need. These are the highest-leverage prompts because almost nobody optimises for them.

  • How do I [solve specific problem your product addresses]?
  • What tool can help me [specific job]?
  • Why is my [metric] declining and what can I do about it?
  • I am a [role] at a [company type], what software do I need for [problem]?

Checkpoint: by the end of this step you should have a numbered list of 30 to 50 prompts in a spreadsheet, tagged by category. Save it. You will reuse this prompt set every time you re-audit.

Step 2: Pick the AI engines and locales to test

The minimum viable set in 2026 is five engines, because that is where the volume is.

  • ChatGPT (default model with web browsing on). Highest volume, most general-purpose use.
  • Perplexity. The most citation-heavy engine. If your domain shows up here, you are doing something right.
  • Gemini. Especially relevant if your buyers use Google Workspace, Android, or Pixel devices.
  • Google AI Mode. The new conversational layer over Google Search. Different ranking from classic blue links.
  • Google AI Overviews. The snippet at the top of regular Google results. Still the biggest source of AI-driven impressions for most B2C brands.

If your audience skews toward developers, software engineers, or technical buyers, add Claude as a sixth engine. If you sell into China or Russia, add the relevant local engines (Yandex Neuro, Baidu, DeepSeek) and accept that the tooling around them is thinner.

Locales matter. AI answers in en-US, en-GB, and other markets diverge more than people expect. If your TAM is multinational, pick your two or three most valuable locales and run the audit in each. Use a VPN or change your Google account region setting. Document the locale next to every result, because the same prompt will produce different answers in Berlin, London, and New York.

Checkpoint: a 30-prompt audit across 5 engines and 2 locales generates 300 data points. That is the floor. Block out four to six hours of focused execution.

Step 3: Manually test in ChatGPT, Perplexity, Gemini, AI Mode and AI Overviews

This is the slow, manual part. There is no way around it if you want a real baseline. Set yourself up properly before you start:

  • Open each engine in a separate incognito or private browser window. Log out where the engine allows it.
  • Disable browser-level personalisation and clear cookies between sessions.
  • Take a screenshot of every response, named with the prompt number and engine (e.g. p07-perplexity.png). You will need these for your action plan and for stakeholder presentations.
  • Run each prompt twice, ideally on different days. AI outputs are non-deterministic. A single answer is a sample, two answers is a pattern.
  • Disable any browser extensions that inject prompts or content into the chat interface.

Pace yourself. Plan for roughly two to three minutes per prompt-engine combination. That includes the response time, the screenshot, and the sheet entry. For a 30-prompt audit across 5 engines, expect three to four hours of execution work, ignoring re-runs.

One trick that saves time: open all five engines side by side on a wide monitor. Paste the prompt into each, wait for all five to finish, then screenshot each answer in sequence. You move through the prompts faster, and you can spot differences between engines visually as you go.

Step 4: Capture mentions, citations, and sentiment

Use a single tracking sheet with one row per prompt-engine-locale combination. The minimum column set:

  • Prompt ID (matches your prompt list from Step 1)
  • Prompt category (branded, navigational, transactional, comparison, problem-driven)
  • Engine (ChatGPT, Perplexity, Gemini, AI Mode, AI Overviews, Claude)
  • Locale (en-US, en-GB, etc.)
  • Run date
  • Brand mentioned? (Yes / No)
  • Position (first mention, middle, last, not mentioned) on a 1 to 5 scale where 1 is “named first” and 5 is “not present”
  • Accuracy (correct, partial, incorrect, with a note on what was wrong)
  • Sentiment (positive, neutral, negative)
  • Citations (list of URLs the engine cited)
  • Competitors mentioned (comma-separated list)
  • Screenshot file name
  • Notes (anything qualitative worth remembering)

From these columns you can derive every metric you need without re-running prompts:

  • Visibility rate: percentage of prompts where you were mentioned at all.
  • Weighted visibility: position-weighted version. Position 1 counts 100%, position 2 counts 50%, position 3 counts 33%, and so on. This is closer to how LLM brand visibility scoring works in production tools.
  • Share of voice: your mentions divided by total mentions for you plus your competitors.
  • Sentiment ratio: positive minus negative mentions, divided by total.
  • Accuracy rate: share of mentions where every factual claim is correct.

Checkpoint: by the end of Step 4 your sheet should be fully populated. Resist the temptation to start analysing while you capture. Capture everything first, analyse in one focused block.

Step 5: Benchmark vs. competitors

Pick the three to five competitors you actually lose deals to. Not the entire market. Not the brands your sales team is afraid of. The names that show up in your CRM as “lost to” or “considered alongside”.

For each prompt where your brand was mentioned, log every other brand mentioned. For prompts where you were not mentioned, log the brands that took your spot. This gives you a competitor share of voice ranking.

The output of this step should be a small comparison table you can show a stakeholder in 30 seconds. Something like this:

Brand Visibility rate Avg. position Sentiment Most common framing
Your brand 42% 2.8 +0.4 “affordable alternative”
Competitor A 78% 1.4 +0.6 “market leader”
Competitor B 61% 2.2 +0.3 “enterprise-grade”
Competitor C 34% 3.5 +0.1 “niche, technical”

You can build this view directly inside LLM Pulse competitor tracking if you do not want to maintain it manually, but the manual version works fine for a first audit. The point of the table is to make the gaps undeniable when you walk it through internally.

Step 6: Identify content gaps in your owned domains

Now flip the question. For every prompt where your brand failed to appear, ask: do we have content that should have been retrieved as an answer? In most audits, the answer is “yes, but”. Yes, we have a page on this topic, but it is buried, thin, or written for a sales narrative rather than an answer.

Go through every “not mentioned” prompt and run two checks.

Check one: do you have a public page that directly answers this prompt? If yes, the issue is retrievability. Your page exists but the LLM did not surface it. Causes are usually one of: weak topical authority around the page, missing structured data, no third-party citations pointing in, or content that does not answer the question in the first 200 words.

Check two: do you have any page at all on the topic? If no, you have a content gap. Add the prompt to your editorial calendar and write an answer-first page targeting it. “Answer-first” matters here. Lead with the answer in the first paragraph, then expand. LLMs extract the top of pages disproportionately.

For each gap, score the lift as small (existing page needs a rewrite), medium (new page needed on existing topic), or large (entire content cluster missing). Save that scoring for Step 9 when you build the action plan.

Step 7: Audit third-party citation sources you appear in (and don’t)

This is the step that most marketing teams forget. LLMs do not just rank your owned content. They rank what other people say about you. If review sites, comparison roundups, and editorial publications do not mention you, no amount of on-site optimisation will fix the visibility problem.

Open your Step 4 spreadsheet and pull every citation URL into a single list. Group by domain. You will end up with something like:

  • Domains that cite you frequently (your owned domain, friendly review sites, partner blogs)
  • Domains that cite competitors but not you (the gap list)
  • Aggregator domains that AI engines love (G2, Capterra, Reddit, Wikipedia, industry-specific listicles)

The “competitors but not you” list is gold. Every domain on that list is a publication that has decided to write about your category and chose to mention other brands. Reach out, contribute, pitch a guest piece, get included in the next refresh of their roundup. Citation sources analysis in LLM Pulse does this aggregation automatically if you want to skip the manual grouping, but it is doable in a sheet for a first pass.

Two patterns worth watching for:

  • Reddit and community forums. AI engines weight community discussion heavily. If you have zero presence in your category’s relevant subreddits, that is a flag. Not a reason to spam, but a reason to participate honestly.
  • Wikipedia. If you have a category but no entry, and your competitors do, fix it. If you have an entry that is out of date, fix it. Wikipedia is one of the most consistently cited sources across all five engines.

Step 8: Audit structured data + llms.txt

Two technical checks that take an hour and matter more than people think.

First, structured data. Open three or four of your key pages (home, product, pricing, a top-of-funnel article) and inspect them for schema markup. The schema types LLMs care about most are Organization, Product, FAQPage, HowTo, Article, and Review. Missing schema is a common reason LLMs fail to extract pricing, founders, or features accurately, even when the information is on the page.

Use the LLM Pulse schema analyzer or Google’s Rich Results Test. Both will tell you what schema is present and what is missing. Prioritise fixing Organization schema first (it tells LLMs the basic facts about your company) and Product schema second (for pricing and feature accuracy).

optional manifest for AI-oriented consumers, not a replacement for robots.txt or sitemap.xml

If you do not have one, generate it with the LLM Pulse llms.txt generator, deploy it, and verify it is reachable. If you already have one, check that it points to your current pricing page, current product pages, and your most important pillar content. Outdated llms.txt files are common and they actively work against you.

Checkpoint: at this stage your spreadsheet should have a “technical” tab listing every schema gap and the llms.txt status. That tab feeds the action plan.

Step 9: Build the action plan (impact vs effort matrix)

By now you have four artefacts: a tracking sheet of mentions and citations, a competitor benchmark table, a list of content gaps, and a list of technical fixes. Step 9 is where you turn them into a plan that someone can execute on Monday morning.

Use a simple two-by-two matrix. Score every finding on impact and effort:

Quadrant Examples What to do
High impact, low effort Fix wrong founding year on Wikipedia; deploy missing Organization schema; update llms.txt; correct outdated pricing on category listicles Do this week. Most of it is editing existing pages.
High impact, high effort New pillar content for missing topic clusters; pitch and land 5 editorial citations; Reddit presence in 2 relevant subreddits Do this quarter. Assign owners and deadlines.
Low impact, low effort Minor sentiment cleanup, light pricing accuracy fixes on low-traffic third-party sites Batch into a single sprint. Do not let them block higher-impact work.
Low impact, high effort Building a presence in adjacent categories you do not actually serve; chasing engines with negligible audience overlap Park. Revisit next year.

The deliverable from this step is a one-page document with the matrix, the top five quick wins, and the top three quarterly initiatives. That is the artefact you share with the CEO, the head of content, and the SEO lead. Nothing else from the audit needs to leave your team.

Step 10: Schedule the re-audit (weekly automated baseline)

A one-time audit is a snapshot. It dates fast. AI engines push model updates every few weeks. Competitors publish continuously. Your own pricing, features, and team change. By the time you act on a quarterly manual audit, half the data is stale.

The fix is to make the baseline continuous. There are three realistic ways to do that:

  • Manual quarterly. Re-run the full audit every three months. Honest assessment, full coverage, but you will miss anything that happens in between.
  • Manual monthly subset. Pick the 10 highest-priority prompts and re-run them every month, full audit quarterly. Better, still labour-intensive.
  • Automate the process. Use LLM Pulse or a similar platform to run all 30 to 50 prompts across all five engines, with alerts when visibility moves more than a defined threshold. This is what production teams do, because the labour cost of manual monitoring exceeds the tool cost within the first month.

Whichever cadence you pick, schedule it now. Block the calendar. The reason manual audits die is not lack of value, it is lack of recurring time. Either commit to the calendar or move to automation.

Manual audit vs LLM Pulse Free AI Visibility Report

An honest comparison. The manual audit above is the rigorous version. It teaches you how the system works and gives you a baseline you can defend. It takes a working day, roughly seven to nine hours of focused execution.

The LLM Pulse free AI Visibility Report runs roughly the same set of checks automatically. You enter your domain, your three top competitors, and a handful of category keywords. The system runs prompts across ChatGPT, Perplexity, Gemini, Google AI Mode, and Google AI Overviews, aggregates the mentions, scores the visibility, lists the citation sources, and surfaces the gaps. It takes about five minutes to set up and a few hours to complete the underlying runs.

Manual audit LLM Pulse free report
Time investment 7 to 9 hours 5 minutes setup, automated run
Engines covered Whatever you have time for ChatGPT, Perplexity, Gemini, AI Mode, AI Overviews
Prompt count Capped by working hours Multiple prompts per category, automated
Updates You re-run it Scheduled refresh on paid plans
Cost Free, your time Free first report, then from €49/month
Learning value High. You see every output yourself. Lower for the first run, higher long-term because you actually keep doing it.

Pick the manual route once if you have never done this before. You should know what the underlying data looks like and how the engines behave. Then move to automation, because the value of an AI visibility audit comes from comparing this week to last week, not from a single point-in-time snapshot.

Summary

An AI visibility audit is not optional in 2026. It is the only way to see what a quarter or more of your buyer journey looks like inside ChatGPT, Perplexity, Gemini, and Google AI. The 10-step process above is repeatable in a single working day, uses tools you already have (a browser, a spreadsheet, screenshots), and produces three concrete artefacts: a tracking sheet, a competitor benchmark, and a prioritised action plan.

The biggest mistake is treating the audit as a one-off. The second biggest is doing it without comparing against competitors. The third is fixing the easy on-site issues and ignoring third-party citation sources. Avoid those three and you have a working baseline that pays for itself.

If you want to skip the spreadsheet and start with a baseline, the free AI Visibility Report covers the same ground in minutes. Use it for the baseline, then layer in the manual steps your team finds most valuable.

FAQ

How long does a manual AI visibility audit take?

For a 30-prompt audit across 5 engines and 1 locale, expect 7 to 9 hours of focused work. Add 2 to 3 hours per additional locale. The bottleneck is the manual prompt execution in Step 3, not the analysis. If you only have half a day, narrow the prompt set to 15 to 20 prompts focused on transactional and comparison categories. Those produce the most actionable insight per minute spent.

How often should I re-audit?

Quarterly is the minimum if you are running it manually. Monthly is healthier. Weekly is the right cadence for any brand actively investing in AI visibility, which is why most teams move to automation after their first or second manual run. The pace at which AI engines change their citation behaviour and content sources is fast enough that quarterly data is already stale by the time you act on it.

Do I need to test all five engines, or is ChatGPT enough?

ChatGPT alone will mislead you. Different engines cite different sources, weight different signals, and produce meaningfully different rankings. Perplexity weights citations heavily, Gemini leans on Google’s own index, AI Overviews mix snippets from organic results, and AI Mode is its own conversational layer. You can be invisible on one and dominant on another. The minimum useful set is ChatGPT plus Perplexity plus one of the Google surfaces (AI Mode or AI Overviews).

What is a good visibility score?

There is no universal benchmark, because every category has different competitive density. The honest answer is relative. If your three priority competitors have visibility rates of 70 to 80% on your transactional prompts and you have 20%, you have a problem. If you are at 55% and they are at 60%, you are competitive. The number you should care about is the trend across re-audits, not the absolute value at a single point in time.

How is this different from an SEO or AEO audit?

SEO audits measure rankings in classic search results. AEO (answer engine optimisation) audits measure whether your content gets pulled into rich snippets, featured answers, and AI overviews. An AI visibility audit is broader: it measures how your brand is named, framed, and compared inside conversational AI outputs across multiple engines, regardless of whether your own pages are the source. You can rank #1 on Google and still be invisible in ChatGPT, because ChatGPT is summarising what third parties say about you, not what your own pages claim.

Can I run this audit with ChatGPT itself by asking it to audit my brand?

You can ask ChatGPT to summarise what it knows about your brand, and the answer is informative as a starting point. What you cannot do is rely on a single conversational answer as a full audit. ChatGPT will reconstruct what it can remember in that session, which is not the same as what it would surface to a buyer running a category prompt cold. The audit has to be done with cold, untainted prompts in a logged-out session. Use ChatGPT-on-itself for accuracy spot checks (founding year, headquarters), not for the full visibility picture.

Discover your brand's visibility in AI search effortlessly

Are you tracking your AI Search visbility?

START NOW WITH A
14-DAY FREE TRIAL