LLM visibility is how often a language model names your business when somebody asks it a question. This article shows how to measure it: a fixed list of questions, a written record of what came back, and a score you can repeat next month. There is no ranking to look up and no results page to open, so nobody can hand you the number. You have to collect it yourself.
The method is dull on purpose. Ask the same questions on a schedule, write down the answers, then count. Dull beats confident and wrong.
What LLM visibility actually means
LLM visibility covers two separate things, and mixing them is the first mistake. One is a mention, where the model names your business inside its answer. The other is a citation, where it also links to a page of yours.
A mention builds recognition. A citation can send a visit. So the same answer can be a win on one measure and nothing at all on the other, which is why the two need separate columns.
It also covers more than one system. ChatGPT, Gemini, Claude, Perplexity and Copilot each answer in their own way, and Google's overviews arrive inside a results page rather than in a chat window. Each one has its own habits, and those habits change without notice, so a single blended figure hides more than it shows.
Why LLM visibility is hard to measure
Ordinary search hands you a rank, an impression count and a click. None of those exist here.
Worse, the answers move. Ask the same question twice and the wording changes, and sometimes the sources change with it. That is how these systems work rather than a fault, yet it ruins any measurement that rests on a single run.
Personalisation can add another layer. Location, account history and earlier chats may all shape the reply. Products differ in how much weight they give each one. So a prompt typed by you and the same prompt typed by a customer are rarely the same question.
There is a third problem, and it is the quiet one. A model can describe your business without naming a source, which means no visit happens and your analytics never sees a thing.
Measuring only pays once you know why a model picks anybody at all. That mechanism sits in our plain guide to generative engine optimization, and this piece starts one step later, where you want a number instead of a theory.
Build a prompt set before you measure anything
A prompt set is your instrument, and every LLM visibility figure you ever quote rests on it. Without a fixed list of questions, each measurement compares apples with whatever you happened to type that morning.
Write the questions a customer would really ask, in a customer's words rather than yours. Twenty to forty of them is plenty for most businesses. Then keep the list fixed for at least a quarter, because changing the questions changes the score for reasons that have nothing to do with your work.
Four kinds of question worth tracking
- Buying questions, such as who does this work in a named city.
- Comparison questions, where somebody weighs one option against another.
- Problem questions, phrased as a symptom rather than as a service.
- Brand questions, which test what the model already believes about you.
Brand questions matter more than they look. If a model has your service list wrong, no amount of fresh content fixes it until the wrong version stops circulating.
What to record for every answer
Keep the LLM visibility record boring and identical each time. One row per prompt, per model, per run does the job.
| What to record | Why it earns its column |
|---|---|
| Named or not | The base measure, because everything else builds on it. |
| Linked or not | A mention is not a link, so keep the two apart. |
| Position in the answer | First on a list usually reads as a recommendation. |
| Rivals named | Shows who the model trusts when it does not trust you. |
| Accuracy of the claim | If the description is wrong, that costs more than silence. |
| Date and model version | Otherwise a later comparison means nothing. |
Two extra columns pay for themselves later. Record the exact prompt text, and note which account, if any, ran it. Both explain odd results months afterwards.
How often to sample, and how many runs
One run of one prompt tells you almost nothing. Three runs in three fresh chats tells you something. The variation between those runs is useful in itself, since a business that turns up once in three is a weak favourite rather than a fixture.
Monthly suits most businesses for an LLM visibility check. Weekly only makes sense when you have just shipped a change and want to watch it land. Anything more frequent measures noise, and noise is what makes people rewrite pages that were fine.
Run each prompt in a clean session. Sign out, clear the history, and use a plain browser profile. Otherwise you are measuring your own reading habits.
A simple LLM visibility score
Once the sheet holds a few months, one number keeps the conversation sane. Share of answers is the simplest honest one. Count the answers that named you, then divide by the answers you looked at.
Keep that score per model. A figure blended across five assistants moves for reasons nobody can trace.
A worked example, with small round numbers chosen to make the arithmetic clear rather than to suggest a typical result. Suppose your prompt set holds 30 questions and you run each one three times in a single assistant. That gives 90 answers. Your business appears in 18 of them, and 6 of those answers include a link. Share of answers is 18 out of 90, and share of citations is 6 out of 90. Next month the same 90 answers name you 27 times, so the score moved for a reason you can point at.
The half of LLM visibility you can measure directly
Sampling is not the only source. Your server logs record the crawlers that identify themselves, so you can often see when one fetched a page. Your analytics may also show visits arriving from a chat interface.
Neither is complete. A crawler visit does not prove the model used the page, and plenty of assistant referrals arrive with no useful label. Still, the two together beat either alone, and both cost nothing to switch on.
The reporting side has traps of its own, which we set out in reading AI search traffic in your reports. Read that before you tell anybody the channel is small.
What LLM visibility cannot tell you
It cannot tell you why. A model will not explain which page convinced it, and asking it to explain produces a plausible story rather than a reason.
It cannot tell you about revenue either. Somebody may read your name in an answer, sit on it for a week, then search for your brand and buy. That journey looks like branded search in every report you own.
So treat the score as a leading indicator. It moves before the money does, and it explains a rise in direct visits that would otherwise look like luck.
Tools, and when a spreadsheet is enough
Paid trackers exist and some are good. They run your prompts on a schedule, store the answers and chart the trend, which saves real time once the prompt set grows.
Below that size, a spreadsheet and an hour a month does the job. If you watch thirty prompts across one or two assistants, buying software is buying a chart you could have drawn yourself.
Whichever route you take, own the raw answers. A tool that shows a score without the text behind it gives you nothing to check, and an unverifiable number is exactly what this exercise replaces.
Mistakes that ruin an LLM visibility report
- Asking the model whether it recommends you, which invites flattery rather than data.
- Leading prompts that already carry your brand name, unless brand questions are the point.
- Measuring while signed into an account that has spent all year reading your own site.
- Comparing this month with last month after quietly editing the prompt list.
- Chasing one bad week, when the sensible move is another sample.
The last one matters most. These systems shift under you, so a month of movement is a signal and a week of movement is weather.
Turning LLM visibility into a decision
A score nobody acts on is a hobby. Three moves follow naturally from the sheet.
- Start with the prompts where a rival appears and you do not, then read whatever the model quoted instead.
- Correct wrong claims about your business at the source, since a model tends to echo whatever the wider web says about you.
- Publish the plain answer to any prompt that came back vague, because vagueness is an opening.
The work behind those moves is not exotic. Clarity, consistency and claims anybody can check are the whole game, and earning AI citations for your pages covers the first part in detail.
Assistants reward slightly different things, so read up on whichever ones your customers name. If ChatGPT is on that list, showing up in ChatGPT search is the sensible next read. Google's answer boxes behave differently again, and how AI overviews change the click explains why traffic can fall while your LLM visibility rises.
Frequently asked questions
How do I start measuring LLM visibility?
Start with a fixed list of questions your customers really ask, then run each one three times in a clean session with no chat history. Record whether the model named you, whether it linked to a page, and who it named instead. Repeat the same list every month. A spreadsheet is enough at first, and it keeps the raw answers where you can reread them later.
How often should I run the same prompts?
Monthly is a sensible default, and the number of runs per prompt matters more than the frequency. Three runs of one prompt in a month tell you more than one run a week, because a single answer hides how unstable the result is. Pick a date and keep it. If you miss a month, log the gap rather than closing it quietly. If you change the prompt list, start a fresh baseline instead of comparing two different instruments.
Why does the answer change every time I ask?
These systems generate an answer each time rather than reading one back from a fixed list. So two runs of the same prompt rarely match word for word. Your chat history, your location and the account you use may shape the reply as well. That is why a single run proves nothing. Three runs in fresh sessions give you a rough probability, which is the honest unit here.
Is a mention without a link worth anything?
Yes, although it counts for less than a citation and it belongs in its own column. A reader who sees your name in an answer may search for it later, so the visit lands as branded or direct traffic instead. That makes the effect real but hard to trace. Count mentions at the source rather than waiting for a report to reveal them.
Do I need a paid tool to track LLM visibility?
No, not at the start. A spreadsheet and an hour a month covers a small prompt set across one or two assistants. Before paying, check that a tracker shows the raw answer behind every score. Check too that you can export your prompt list and see which model version it queried. Scores from two different tools are not comparable, so switching tracker usually means starting the trend again.
Which assistants should I actually track?
Track the ones your own customers use, which usually means starting with two rather than five. Ask a few customers directly, and look at whatever referral labels already appear in your analytics. Watching every assistant produces a wide table nobody reads. Sampling two properly every month beats sampling six once and never repeating it. Add a third only when something in your own data asks for it.
What if an assistant describes my business wrongly?
Treat a wrong description as more urgent than a missing mention, because it travels further than silence. Find where the wrong version lives, such as an old directory listing or a stale profile of yours. Correct it at the source, then keep that wording consistent everywhere the same fact appears. Models lean on repetition, so one corrected page rarely outweighs five stale ones. Re-run the prompt next month and record whether the claim changed.
What counts as a good score here?
There is no published benchmark worth trusting, and anybody quoting one for your industry is guessing. Judge LLM visibility against two things instead: your own previous month, and the rivals turning up in the same answers. A score that climbs while one competitor keeps appearing ahead of you still tells you what to work on next. Consistency across models matters more than one high month.
Does LLM visibility replace ordinary SEO?
No. LLM visibility rests on the same foundations, since a model can only quote pages it can crawl, read and trust. Check your own analytics for the split between ordinary search and assistant referrals. That balance differs by industry, so move budget on your own numbers rather than on a headline. A page that answers a question plainly tends to do well in both places. Treat this as a second measure of the same work.