Four steps, and only the last one is unusual. Citino ranks your gaps by how movable they are, writes a brief against one named prompt set, waits while you publish, and then measures what changed against what the whole category did over exactly the same days.
Without a baseline you are comparing your page to your memory of last month, in a market where every brand's numbers move together whenever an engine changes something.
A gap is a prompt where the engines answer the question and do not name you. There are always more of them than you have weeks, so the ranking is the feature.
Movability multiplies three things the product already holds: what the prompt is worth in your own GA4; how movable the pages cited on it are, with competitor-owned pages excluded from that share; and how far your visibility on that prompt sits from the leader's.
A large deficit can rank low. If the only pages cited on a prompt are Crateline's own comparison and Northport WMS's documentation, and it sends four sessions a month, a 30-point deficit is not an opportunity — it is a fact about the category. Ranking by deficit alone produces a list sorted by how bad you feel.
Recomputed daily, from the same ledger the Sources page reads. → Sources
If a recommendation could have been written without reading your category, it is a platitude with a checkbox beside it. A brief here names five things, and it can only name them because every answer was stored whole.
A brief that cannot fill field 3 or field 4 is not issued. There is no fallback text and no generic template behind it, because the fallback is the failure mode — the moment a product starts emitting advice when it has nothing specific to say is the moment its specific advice stops being trusted.
The quote is a stored answer, not a paraphrase of one. → Answers
You rewrote the 3PL billing page. On the prompts the brief named, visibility went from 21.1% to 29.5% — a gain of 8.3 points. Over the identical days, every brand in Warehouse Management Software · US rose together by half a point, for reasons that had nothing to do with your page. That half point was going to happen to you anyway, so it comes out.
The subtraction is only possible because the category is collected whether or not anyone in it is a customer. All 17 brands are measured, so the counterfactual is a measurement rather than an estimate. A tool that samples only its own subscribers has nothing to subtract and has to hand you the 8.3.
Status is measuring. The experiment is not called a win until the effect clears its band.
There is a version of removing drift that means throwing away days. An engine has an odd week, the numbers dip, those days get flagged anomalous and dropped, and the chart afterwards shows a clean rise. That is choosing the sample after seeing the result — every individual step is defensible and the output is decided in advance.
Here it means one operation: measure what the category did over the identical window and subtract it. Every day stays in, including the ones that shrink your effect. If the category had risen 3 points, the 8.3 would read 5.3 and you would be told that in the same place. There is no exclude-this-day control in the product. It was never built, which is the only dependable way of never using it.
An experiment closes with one of three states: the effect cleared its band, it did not, or it is still measuring. The middle one prints as inconclusive, in the same place and at the same size as a success, and it is never rewritten into "early signs are positive".
The customer did the work. A product that wants to be liked will find them a number — a 1.2-point effect on a ±3.0 band can be drawn with an upward arrow and a green tint and nobody will complain that month. It is not a small win. It is an unanswered question.
An inconclusive read is usually not a failed rewrite. It is a sample problem — too few prompts in the set, too short a window, a band too wide for an effect that size. The card says which, and how much longer a read would take, so the response is to wait rather than to rewrite a page that may already be working.
Verdicts are computed, not editable, and they appear in white-label client reports unchanged. → For agencies
Citino does not generate the page, does not publish to your CMS, and does not promise placement on any prompt. The engines decide what they recommend. This product measures what they decided, before and after, against the rest of the category.
An experiment whose prompt set is too small to measure is refused rather than run, and the refusal states how many prompts it would take.
No result here is a forecast. Nothing on this page predicts what your visibility will be next month, because the honest answer is that we do not know and neither does anyone selling you a number for it.
The baseline exists before you start, because the category has been collecting all along. That is what makes the first experiment measurable rather than the fourth.