How to Track Brand Mentions in ChatGPT (2026 Guide)

TL;DR
Tracking brand mentions in ChatGPT means running a fixed set of prompts on a schedule and recording how often your brand is named, where in the answer it lands and which sources the answer cited. This guide covers mentions versus citations, why one-off checks mislead, the three methods available, the prompt set they depend on, and what to report.

Someone in a marketing meeting asks whether ChatGPT recommends your product. A laptop opens, the question gets typed, and the answer is read out loud. That answer is real. It is also one sample from a distribution nobody in the room can see. Run the same prompt from a different machine an hour later and the list can come back in a different order, with different brands in it.

Closing the gap between one answer and a measurement is what this guide is about: what to count, what varies between runs, which method gives you what, and how to build the prompt set the numbers rest on. That last part is where most teams go wrong.

A mention and a citation are two different things

A mention is your brand named in the text of the answer: “for project management, teams often use Acme, Beta and Gamma.” No link needed. A citation is your URL attached as a source, inline or in the list underneath. They move independently:

  • ChatGPT recommends you by name in a paragraph whose sources are two review sites and a forum thread. Mention, no citation.
  • It cites your documentation page and answers in category terms without naming anyone. Citation, no mention.

Keep them in separate columns. Count only links and you miss the answers that sell you hardest, because a plain-text recommendation converts and leaves no traffic row anywhere. Count only mentions and you lose the lever, since the cited sources are pages you can go and influence. More on that side in our guide to tracking which sources AI answers cite.

Why checking by hand misleads people

One check is an anecdote because the answer you got was shaped by things you did not control. Concretely:

  • The run itself. Generation is sampled. Same prompt, same account, same minute, different wording and sometimes a different brand list. Three brands become five, and the order changes.
  • Your account. Memory, custom instructions and earlier turns all feed the answer. Spend a year researching your own category and your account stops being a neutral observer of it.
  • Country and language. The same question in Spanish from Madrid and in English from Chicago are different questions. Brands that dominate one market vanish in the next.
  • Whether it searched. Some answers trigger a web search, some come from the model’s own knowledge, and the two routes can name different brands. See how ChatGPT search works.
  • The model version. Routing and model updates change answers underneath you, with no release note that maps to your category.

There is also no position to read. In Google you were number four and everyone knew what that meant. In an answer you are named or you are not, which is a weaker signal per observation and the reason you need more of them. One check tells you the answer was possible. Thirty tell you it was likely.

The three methods, and what each is good for

1. Prompting ChatGPT yourself

Free and immediate, and the right first move, because it tells you whether your category questions produce brand recommendations at all. As a trend line it is useless, because you will not run it consistently and cannot reconstruct your account state from three weeks ago. Do it cleanly:

  1. Use a temporary chat, or log out: an account with memory on tests your own history as much as the model. Start a fresh conversation for every prompt, because follow-ups inherit everything above them.
  2. Set the country and language you care about, and note which. Three markets means three runs of everything.
  3. Run the same prompt at least five times and write down every brand list you get, including the disappointing ones.
  4. Record the date, whether the answer showed sources, and the raw text. In two months that wording is your only evidence of what changed.

Ten prompts done this way gives you a feel for your category and a spreadsheet you abandon by week three.

2. Analytics, for what happened after the answer

ChatGPT appends utm_source=chatgpt.com to links in its answers, which is why this works at all. GA4 reads session source from the campaign parameter when present and falls back to the referrer, so a visit with no referrer still gets attributed. GA4 also has an AI Assistant default channel, defined as the channel by which users arrive from sources like ChatGPT, Gemini, Deepseek, Copilot or Grok, matched when the medium is exactly ai-assistant or the referrer matches Google’s list of known AI assistant sources (Google’s channel definitions).

Two limits matter more than the setup. The same documentation puts clicks from Google’s AI Overviews and AI Mode inside Organic Search, so this channel is not a general “AI traffic” number. And analytics only measures clicks that happened after an answer: it says nothing about answers that named you without a link, or named a competitor instead. This is downstream confirmation, and it cannot be turned into a mention tracker.

3. A tool that runs a fixed prompt set on a schedule

The only method that produces a trend: a stable list of prompts, run on a schedule against the real assistants from clean sessions, in the countries and languages you chose, with every answer stored. What to demand:

  • The raw answer text, kept. If you cannot read the sentence you were mentioned in, you have a number with no explanation attached.
  • Position within the answer, because first brand named and fifth are different outcomes.
  • Cited sources per answer, so a drop traces back to the pages the model leans on.
  • Named competitors on the same prompts and runs. A share computed on a different sample is not a share.
  • Country and language as per-project settings, plus exports and an API, because the reporting request always lands in someone else’s dashboard.

LLM Pulse is our own product, so read this knowing that. It runs your prompt set against ChatGPT, Perplexity, Google Gemini, Google AI Mode and Google AI Overviews on every plan, stores each answer with its mentions, position, sources and sentiment, and tracks the competitors you name. Weekly is the default cadence, with daily and monthly available. The arithmetic matters when picking a prompt count: a weekly prompt produces about 4.33 scheduled answers per model per month, so 50 prompts across 5 models is over a thousand answers a month.

Tracking brand mentions in ChatGPT with LLM Pulse

For a reading without committing to anything, the free brand mention checker runs prompts against live assistants and shows whether your brand comes up. It is a snapshot, with every caveat a snapshot carries, but it tells you quickly whether you have a problem worth measuring.

Building the prompt set

Everything above is plumbing. The prompt set is the instrument, and one brainstormed in a meeting measures your imagination.

Start from demand you can evidence. Search Console queries are the best source: connect the property, sort by impressions, read what people actually type. Then add the questions that come up on sales calls before a demo and the recurring subjects in support tickets. Prompt research can extend a seed set, but the seed comes from real demand.

Cover the journey. Category questions where nobody is named (“what is the best tool for X?”) are where visibility is won and lost. Comparison questions naming two players show how the model frames you against a specific rival. Brand questions naming you are your reputation view.

Keep branded and non-branded apart in reporting. A prompt with your name in it mentions your brand nearly always. Blend those into one average and you get a flattering number that moves when you add prompts. Headline the non-branded set.

Freeze the set. Adding five prompts in March changes the denominator, and the March dip is then an artifact. If you must add, treat them as a new cohort and leave the original untouched.

Size it honestly. Twenty to thirty well-chosen non-branded prompts per market is a real measurement for a focused product, and a few hundred is for multiple product lines or several countries. More prompts than your team will read is a cost.

What to measure

Four numbers carry almost all the signal.

Mention rate is the percentage of answers that name you: (answers mentioning your brand / total answers) x 100. Compute it on the non-branded set, per model and per country. A blended figure hides the case where you are strong in one assistant and absent in another.

Position-weighted visibility credits being named first over being fifth. LLM Pulse calls this the AI Visibility Score and weights by order of mention, so a first mention counts full, a second half, a third about a third. It moves before mention rate does: climbing from fifth to second inside the same answers is progress a yes/no count cannot see.

Share of voice is your mentions as a proportion of all brand mentions across those answers: (your mentions / (your mentions + mentions of your named competitors)) x 100. The caveat is large: it depends entirely on which competitors you named. Add a dominant incumbent and your share drops without anything changing in reality, so fix the competitor list with the prompt set and state it whenever you report the number. The variants are in measuring AI share of voice.

Sentiment and cited sources are the qualitative half. Being named as the expensive option is a different result from being named as the safe default, and sentiment on branded prompts is where that shows. The cited sources are your action list: when the same review site or forum thread keeps feeding answers in your category, that page is the work, and it is usually not one you own.

Cadence and sample size

A single day is noisy for the reasons above. Weekly smooths it, and matches how fast the underlying content moves, since a page published today does not change what an assistant says tomorrow. Work out your resolution before committing. Thirty prompts across five models weekly is about 150 answers a week, so one answer flipping moves mention rate by roughly 0.7 points. That sets the change worth reacting to: a two point wobble is noise, a sustained ten point slide over a month is a finding. Four weeks is the minimum before a trend line deserves a sentence in a report.

Daily earns its cost around a launch, during a reputation problem where you want to watch a negative framing appear and recede, through a site migration, and in categories that genuinely churn. Otherwise it buys you more copies of the same week.

Common mistakes

  • Reporting a number from one run. It will read great once and terrible once, and both will be believed.
  • Letting branded prompts inflate the headline. The easiest way to fake progress is adding prompts with your name in them.
  • Treating a mention as a visit. Most answers end there. If the CFO wants traffic, show the analytics view and say plainly it is a subset.
  • Tracking only ChatGPT. Your category may skew toward another assistant, and you will not know until you run the same prompts elsewhere.
  • Never reading the answers. The number says something moved; only the text says why.

Frequently asked questions

Can I just ask ChatGPT whether it mentions my brand?

No. It has no record of what it said to other people and no view of its own output distribution. Ask and you get a plausible answer generated on the spot, which is worse than none because it sounds authoritative. Run the prompts and count what comes back.

Is there an API for tracking ChatGPT mentions?

OpenAI’s API gives you a model, not the consumer product. The ChatGPT app wraps its own behaviour, memory, personalization and search tooling around that model, so API answers and app answers to the same question are not interchangeable. You can build a sampling harness on the API, but you are measuring a different surface. Collecting consumer answers consistently, in the right country and language, at a stable cadence, is the actual work.

How many prompts do I need?

Twenty to thirty non-branded prompts per market is enough to see real movement, plus a smaller branded set for reputation. Below ten, one answer flipping swings your chart. The ceiling is how many answers your team will read.

Why do results on my laptop differ from what the tool shows?

Usually account state: your session carries memory, custom instructions and history, while a tracker such as LLM Pulse runs clean sessions. After that, country, language, which model handled the request, and whether that run triggered a web search. The tool’s number averages many runs and yours is one draw.

Does a ChatGPT mention drive traffic?

Sometimes, and less often than the mention count suggests. A mention with a link can produce a session, visible in GA4 thanks to the chatgpt.com campaign parameter. A mention with no link produces nothing, and neither does a reader who got their answer and moved on. Keep mentions and AI referral traffic as two related metrics.

How fast can I move the number?

Slower than a search ranking, with less direct control. Answers lean on sources the model already trusts, so the work is getting into and improving those: the review sites, comparison pages and community threads that keep appearing in your citations. Think in months. What you can do this week is set the baseline, because without one you cannot tell whether any of it worked.

Discover your brand's visibility in AI search effortlessly

Are you tracking your AI Search visbility?

START NOW WITH A
14-DAY FREE TRIAL