It is easy to make “AI visibility” numbers show whatever you like: ask a model once, catch a favorable answer and put it in a report. We do not work that way. This page explains how Brandometer measures what AI assistants say about brands, why we report mention rates instead of “rankings,” and what we refuse to promise.

One answer from an AI proves nothing

A language model generates text probabilistically. At every step it picks the next word from a distribution; it does not retrieve a ready-made answer. Ask the same question twice under the same conditions and you get two different answers. Today the model names your brand, tomorrow it lists three competitors and leaves you out — and to the model both answers are equally “correct.”

A single check is one random draw from the distribution of possible answers. Drawing conclusions from it is like judging whether a coin is fair from one toss. Any tool that reports a brand’s “position in ChatGPT” from a single query is reporting noise.

Series of independent runs in clean sessions

Instead of asking once, we run a series: every prompt is put to every engine several times. Each run must be an independent observation, so it takes place in a clean session:

  • A new conversation with no history, so earlier answers cannot influence later ones.
  • No personalization, memory or user profile, so the model does not adapt to whoever is asking.
  • The same prompt wording and the same settings from one measurement to the next. Otherwise the weeks cannot be compared with each other.

This is a deliberate simplification. Real users get personalized answers, and we cannot reproduce those. A clean session gives a reproducible baseline that is the same for everyone: what the model says by default, when it knows nothing about the person asking.

Where the answers come from

We send each prompt to the engine’s official developer interface (API) with web search switched on. The request uses the same model family as the consumer app, with no chat history and no personalization.

One exception is DeepSeek. Its developer interface has no web search, so for DeepSeek we measure what the model itself knows about your brand.

Google AI Overviews have no developer interface. We read them from the public Google results page for the country and language of your project. Google does not show an AI Overview for every query. When there is none, the observation is recorded as “no AI answer”: it is left out of the rates, and the credits for it are refunded.

Every project has a language and, optionally, a country. Prompts are asked in the project language. Where an engine’s interface accepts the user’s approximate location for web search, we pass the project country. Some engines offer no such setting. For those, localization is prompt-based: it relies on the language of the prompt and a short neutral instruction rather than on a real location.

The engines we currently track, and the price of a check on each, are listed on the pricing page.

Mention rates instead of “rankings”

An AI answer has no positions in the SEO sense. A brand either makes it into the generated text or it does not, and that changes from run to run. So our main metric is a rate: the share of answers in a series that mention the brand, and the share that explicitly recommend it. “The brand is mentioned in 60% of ChatGPT’s answers to this question” is a meaningful statement you can verify. “The brand ranks second in ChatGPT” is not.

95% Wilson confidence intervals

A rate calculated from a short series is itself subject to chance: 6 mentions in 10 runs does not mean exactly 60%. So every rate comes with a 95% confidence interval, calculated with the Wilson method. Put simply, it is the range in which the true rate very probably lies. For 6 out of 10, that range runs from about 31% to 83%.

We chose the Wilson method for a reason. Unlike the naive normal approximation, it behaves correctly on short series and on rates close to 0% and 100% — both common in brand monitoring, where a brand appears in almost no answers or in almost all of them.

In practice: a rise from 40% to 50% with overlapping intervals is not growth yet. It is noise. We show the intervals so that you do not base decisions on random fluctuations.

“The model named the brand” and “the model advised choosing the brand” are different events with different value. A mention that only completes a list, a neutral enumeration and an explicit recommendation (“I would choose X because…”) influence a buyer’s decision in different ways. A plain text search for the brand name cannot tell these cases apart. It also stumbles over inflected forms, spelling variants, transliteration and unrelated companies with the same name.

That is why every answer is labeled in a separate step by an LLM classifier: whether the brand is mentioned at all, whether it is explicitly recommended, and which competitors are named next to it.

The same step records sentiment: whether the answer speaks about your brand in a positive, neutral or negative way. Sentiment is judged for your brand specifically, not for the answer as a whole.

The classifier is a model too, and it can be wrong. So every label can be checked against the raw text of the answer (see below), and we regularly review samples of the labeling by hand.

The model behind every answer is on record

“ChatGPT” or “Gemini” is not a single model. It is a product name, and the models behind it change regularly. A model change can move your metrics overnight with no action on your part. So for every answer we store the engine and the exact model version that generated it, and show it next to the answer. If your mention rate jumps on the day a model was updated, the credit belongs to the update, not to your latest publication — and the record keeps the two from being confused.

Full raw answers as evidence

Every number in Brandometer can be traced back to primary data. We store the full text of each answer together with its date, engine, model version and cited sources. Any rate, sentiment label or competitor mention can be verified by opening the specific answers it was calculated from. Do not trust aggregates that come without primary data — ours included.

What we do not promise

Nobody can control AI answers directly. So we do not promise, and never will:

  • “Higher rankings in ChatGPT.” There are no rankings there, and nobody can guarantee that a brand will appear in text that is generated probabilistically. Anyone who gives such a guarantee is not being straight with you.
  • “Guaranteed placement in AI answers.” Nobody has access to the inside of someone else’s model.
  • Instant results. Changes in the sources reach the answers with a delay of days to months, depending on the engine.

What we do: measure the current state honestly, show the sources the models rely on, and record what changes after you act — in a way that separates a real effect from noise and from a model change. Visibility can only be influenced indirectly, through the sources, and the result is always probabilistic.

Limitations of the methodology

For a complete picture, here is where the method stops:

  • We measure answers to your prompt set. It is a sample, not every wording real users come up with.
  • Clean sessions leave out personalization: a particular user may get a different answer.
  • Answers come through developer interfaces, not through the consumer apps. The model family is the same, but an app adds its own settings, your history and your location, so the answer on your own screen can differ slightly.
  • For engines without a location setting, localization rests on the prompt alone, which is a weaker signal than a real location.
  • Rates from a series of runs are statistical estimates with an interval, not a census of everything an engine says.

We believe honest numbers with caveats are more useful than polished numbers without them. If you have questions about the methodology, write to support@brandometer.ai and we will go through the details.