Measuring Whether Answer Engines Cite You

Analytics cannot see an AI answer. A reproducible method for asking Perplexity, ChatGPT, Claude and Google AI Mode the same questions and recording who they cite.

You cannot see an answer engine in your analytics. When Perplexity or ChatGPT summarises a page and names it as a source, the reader’s question is usually answered inside the answer surface, and no referral is generated. The page was cited; the log file records nothing. Any attempt to manage visibility in these surfaces using referral data is therefore measuring the residue of the minority of sessions that clicked.

We run a small internal protocol instead: ask the surfaces the questions our audience asks, record what they say and who they cite, and treat that corpus as the measurement. This note describes the protocol, what it can and cannot establish, and the first set of numbers it produced for us. It is written up because the method is more useful than our particular results, and because we would rather publish a measurement process that can be criticised than a claim that cannot.

The visibility problem, stated precisely

Three distinct things are commonly conflated.

Retrieval. Did the surface fetch your page while composing this answer? Observable in server logs by user agent, if you log them: GPTBot, OAI-SearchBot, PerplexityBot, ClaudeBot, CCBot, Google-Extended and the rest each announce themselves.

Citation. Did the answer render a link to you? Observable only by reading the answer.

Mention. Did the answer name you in prose without a link? Also observable only by reading the answer — and materially different from citation, because an unlinked mention suggests the model holds the entity rather than having just retrieved it.

A page can be crawled constantly and never cited. It can be cited without being crawled that session, from an index built weeks earlier. It can be mentioned by name with no retrieval at all. Because these three are independent, a measurement that collapses them produces advice that does not work.

Prompt design

The unit of measurement is a prompt, and the prompts have to look like the ones people type, not like keyword strings. We use three classes per topic:

  • Class A — definitional. “What is a declarative prompt specification language?”, “Bradley-Terry vs Elo”. These pull encyclopaedic sources and official documentation. Being absent here is usually a vocabulary problem.
  • Class B — tooling recommendation. “Best open-source local LLM server”, “production-grade vector search without a hosted service”. This is the class that names competitors by name, and the class where absence costs most.
  • Class C — research and architecture. “How do people build metadata-private messaging?”, “state of the art for NUMA-aware scheduling”. These pull preprints and repositories.

Each research area gets at least one prompt per class. The classes are not interchangeable: a lab can be cited constantly in Class C, where arXiv and GitHub dominate, and be invisible in Class B, where product pages and comparison articles dominate. Those are different problems with different remedies, and running only one class hides that.

Collection discipline

Answers are collected from a real, signed-in browser session attached over the Chrome DevTools Protocol. We never bypass a login, never use a scraping proxy, and never present automated traffic as human where a service asks us not to.

The discipline that matters is rate limiting. One tab per surface, closed as soon as the answer is captured. A randomised eight-to-eighteen second pause between queries against the same surface. Rotation, so that two consecutive queries never hit the same service. If a challenge page appears, the run aborts and resumes later rather than retrying. This is slow — deliberately. A protocol that irritates the surfaces it measures stops being able to measure them.

Each capture stores the prompt, the surface, the timestamp, the full answer text, and the citation list. Perplexity renders citations as ordinary anchor elements, so they extract cleanly; other surfaces require reading the rendered message container. Everything lands in a line-delimited JSON file so that a later analysis can be re-run over exactly the same raw material.

Turning answers into a work list

Two derived artefacts come out of the corpus.

A cited-domain table. Count the distinct domains across all answers, and rank them. This tells you which sources the surfaces treat as ground truth in your space, which is the set you need to be cited alongside rather than the set you need to outrank.

A competitor-prominent phrase list. Extract noun phrases with simple frequency heuristics, dedupe case-insensitively on stems, and keep any phrase appearing in three or more answers that appears nowhere in your own copy. This is a vocabulary gap list. It is not a keyword quota: the correct response to seeing a phrase on it is to ask whether the phrase honestly describes something you do, and to use it plainly where it does.

Our first committed batch put one hundred prompts to two surfaces, Perplexity and ChatGPT. Forty-nine of the answers rendered citations, spanning 108 distinct domains when each domain is counted once per answer. GitHub was the most-cited domain, in fourteen answers, followed by YouTube in eleven and arXiv in six. That distribution is itself the finding: in developer-tooling questions, the surfaces lean on repositories and video walkthroughs far more than on vendor marketing pages, which suggests that a public repository with a legible README is worth more than another landing page.

The wider corpus showed how rarely the surfaces pull us into these conversations at all. Across 321 captures on four surfaces, including persona- and location-framed variants of the same prompts, one of our projects was named in exactly two answers, both tooling recommendations: a Perplexity answer that named sigc for a trading-signal compiler, and a ChatGPT answer that named llamafu for running Llama models from Flutter. The areas where we have published most, prompt engineering and agent memory, produced no mentions. Same lab, same site, same schema markup — and standing that tracks the narrow, specific question rather than the volume of our own writing, because the corpus of third-party writing that mentions us is thin. No amount of on-site markup fixes that.

What this method cannot tell you

It is a sample, not a census. Answer surfaces are non-deterministic, personalised to some degree, and change model versions without notice. Two identical prompts a week apart can differ for reasons unrelated to anything you changed.

It cannot attribute cause. If a page starts being cited after you added structured data, the structured data is one candidate explanation among several, including that somebody linked to you from a forum the model reads. We record what changed on our side and what changed in the answers, and we resist joining them with an arrow.

It measures citation, not consequence. Being named in an answer is not the same as being chosen by the reader, and we have no way to observe the second from here.

And the sample is small. A hundred prompts on two surfaces is enough to see a domain distribution and a vocabulary gap. It is not enough to detect a modest change over time, which is why we report counts rather than percentages.

Open Questions

How stable is a citation once earned? We do not yet have a long enough series to say whether being cited for a query persists across model updates or has to be re-won.

Does an unlinked mention convert into a citation? Our intuition is that parametric knowledge of an entity makes later retrieval more likely, but we have no evidence for the direction of that arrow.

What is the right control? Without a comparable site we do not change, we cannot separate our own edits from a shift in the surfaces. Constructing an honest control for a single organisation is genuinely hard.

Is per-surface optimisation even coherent? The surfaces disagree about which sources are authoritative. Writing for the intersection may be the only stable target.

Conclusion

The uncomfortable part of answer-engine visibility is that the standard instruments do not reach it. Analytics see clicks that mostly do not happen; Search Console does not break out the surface. What remains is the slow, unglamorous option of asking the questions yourself, recording the answers, and reading the citations — with enough rate-limiting discipline that the measurement stays possible.

We publish our prompt set, collection scripts and findings in the repository alongside this site, on the same principle as the rest of our work: a claim about visibility that cannot be reproduced is not a claim, it is a marketing statement. If you want to see how the same instinct applies to our software, the research pillars explain how each open question becomes a runnable artefact, the glossary defines the terms we use consistently across the portfolio, and Open Science in AI sets out why we publish the failures too.