Why a one-off reading is not enough to track AI visibility
Asking ten questions to ChatGPT and Perplexity on a Tuesday morning gives a result. That result is not a measurement, it is a dated observation.
Two reasons keep it from being a trend. The first belongs to the model: the same question replayed minutes apart can produce a different answer, with different sources. The second belongs to the environment: answer engines change version without notice, and the index they query moves continuously.
An AI visibility tracking protocol fixes both, and it demands nothing but discipline: the same questions, asked the same way, at the same rhythm.
The ten-question basket method, and how to build it, are described in our guide on the ten-question test. This page starts where that one stops: what to do with the second reading, then the third.
How often to replay the ChatGPT and Perplexity citation test
The cadence used is monthly, and that is a deliberate choice.
Any tighter and you mostly measure the models' variability. An answer that changes from one week to the next says nothing about your visibility: it says the assistant rephrases. You would pile up rows without being able to decide anything.
Any wider and you lose the cause. A quarter between two readings leaves room for a redesign, three publications and a model version change to happen together. When the gap appears, nothing lets you attribute it.
- Week 1The reading
The ten questions are asked verbatim, on each engine tracked, in a signed-out session. Nothing is rephrased compared with the previous month.
- Week 2The review
Every flipped question is confronted with the site pages meant to answer it. The question is not "why does the AI ignore us", but "which page should have been cited".
- Weeks 3 and 4The work
The pages identified are revised or created. Nothing is changed during the reading itself: it would make the month incomparable.
- Next monthThe confirmation
The reading is replayed verbatim. A gap that does not confirm on the second pass stays noise, however large it looked.
A note on choosing engines. ChatGPT and Perplexity do not answer the same way: the first synthesises more readily without citing, the second displays its sources by design. Neither is more reliable than the other for company research; they expose different things. Tracking both, and separating brand mentions from cited sources, costs less than arbitrating between them.
Google's AI overview deserves its own line. It appears above the classic results, and it shows up neither in Search Console nor in a rank tracking tool. It therefore belongs to the same manual reading as the other two, with a column of its own.
How to record readings: the grid to keep over time
The reading grid is only worth something if it does not change. Adding a column midway makes earlier months incomplete, and a removed column erases the one piece of information you will end up needing.
| Column | What goes in it | Why it is needed |
|---|---|---|
| Question | The exact wording, copied verbatim from one month to the next. | A rephrased question becomes another question: the series restarts from zero. |
| Engine and version | The name of the product queried and the version shown on the day. | A version change explains most abrupt gaps. |
| Passes cited out of passes asked | How many passes mention the brand, out of the total asked. | That ratio, not a yes/no, is what makes two months comparable. |
| Cited source | The exact address of the page reused, when the assistant gives one. | Separates a brand mention from a citation that brings traffic. |
| Date and mode | The day of the reading, and the note that the session was signed out. | A session personalised by a history only measures your own account. |
Two keeping rules complete the grid. An empty cell and a zero do not mean the same thing: the zero is a measured result, the empty cell says the measurement did not happen. And a reading is done in a signed-out session, otherwise answers are tuned to your history rather than to a prospect's.
What gap between two readings is significant
This is the question that decides everything, and it is settled before the first reading, never after. A rule set once the figures are visible always adjusts itself to the result you were hoping for.
The rule applied here holds in two cumulative conditions.
A question counts as flipped when its ratio of cited passes crosses the halfway mark, in either direction. A question cited on three passes out of eight, then six the following month, has flipped. A question moving from two to three has not changed camp.
A reading counts as carrying a signal when at least two questions out of ten flip, and the gap is confirmed at the next reading. A single question changing, or two questions returning to their original state a month later, stay noise.
- At least two questions cross half of the passes
- The gap points the same way on both engines tracked
- The next month confirms it, with no work in between
- The flipped questions share an identifiable theme
- A single question changes, however sharply
- One engine rises while the other falls
- The gap disappears at the next reading
- The engine version changed between the two readings
This threshold of two questions out of ten is a decision rule, not a measurement. It is chosen so that a ten-question basket produces a stable decision: lower, and every reading would trigger work; higher, and real erosion would take months to become visible. A wider basket would move the threshold accordingly.
How the workshop crosses that reading with site data
An isolated reading points at a problem without locating it. Crossing it with site data locates it.
The Atelier SEO AI visibility screen keeps every question, every engine and the answer excerpt obtained. A question left uncited is then matched against the site pages that should have answered it. Three cases appear, and they call for three different moves.
The first: no page covers the question. That is a content gap, and the work is a creation. The second: a page covers the question but fails a check that makes it unreadable for an assistant. The third: the page answers and stays readable, but the company itself is not identifiable, which stops the assistant from naming anyone.
The second case is read in the technical audit, which runs more than ninety checks, several of them specifically about how answer engines read a site. The families of checks are described on the audit page.

Iris then takes over and sets the order: which page to revise first, and why that one. It works on the same data as the reading and the audit, which lets it name an address rather than a best practice. The demonstration on a real case reproduces the same question asked without data, then with it.
Google not citing your brand is no reason to drop the tracking
The objection is real and deserves a direct answer: what is the point of tracking an environment that changes every month.
That is precisely the reason to track it. An unstable environment observed at random produces contradictory impressions, and the most recent impression always wins. The same environment observed under a fixed protocol produces a series, and a series can be read.
The second answer is cost. A monthly reading on ten questions and two engines takes under an hour once the grid exists. That is the order of magnitude of a meeting, for the only information available on a channel classic measurement tools do not reach.
The third answer is a limit, and it is better stated upfront. No protocol makes you cited. Tracking measures an exposure, it does not produce one. What it brings is an informed decision on which pages to revise, instead of an intuition.
What to do once the gap is identified: the prioritised follow-up
An identified gap becomes a plan when every page receives a rank, a reason and a measurement criterion.
The order is always the same. Questions whose answer already exists on the site come first: the page is there, only its readability is missing. Then come those requiring a new page, more expensive and slower to produce an effect. Questions where the brand itself is not identifiable belong to the site's identity, not to its editorial.
What follows depends on who executes. The plan is read and applied in your own publishing tool. It can also be applied by the team, and the scope of each plan is described on the offers page.
How often should I measure my visibility on ChatGPT and Perplexity?
Once a month, on the same questions and under the same conditions. A tighter rhythm measures the models' variability rather than yours. A wider one makes it impossible to attribute a gap to a cause.
Should the questions be asked in private browsing?
Yes, or in a signed-out session. A signed-in account personalises answers from your history, which makes the reading incomparable month to month and different from what a prospect would see.
What does share of voice in AI answer engines measure?
The proportion of basket questions where your brand is named, relative to the number of passes asked. It is only useful if the basket and the number of passes stay identical from one reading to the next.
Does a mention without a link count in the reading?
It counts, in a column separate from the cited source. A brand mention builds recognition without bringing a visit; a cited source does both. Merging them inflates the result.

Trust & E-E-A-T


