Product · Act

Publish the fix, then subtract the category.

Four steps, and only the last one is unusual. Citino ranks your gaps by how movable they are, writes a brief against one named prompt set, waits while you publish, and then measures what changed against what the whole category did over exactly the same days.

Without a baseline you are comparing your page to your memory of last month, in a market where every brand's numbers move together whenever an engine changes something.

Read the methodology →
01
Gapbest warehouse management software for a 3PLTallyfold named 4th of 4 · 186 sessions
02
BriefMulti-client billingthe claim your page does not make
03
Publishedtallyfold.com/3pl-billingfirst read 2 Sep
04
Read · 3PL billing page rewriteIllustrative
+7.8 ptsnet effect · measuring · next read 5 Oct
Which gap first

Rank the gaps by what
will actually move.

By deficitIllustrative
#PromptDeficitScore
1
2
3best warehouse management software for a 3PL−12.00.48
4
Sorted by how far behind you are. The movable gap sits third.
By movabilityIllustrative
#PromptDeficitScore
1best warehouse management software for a 3PL−9.70.31
2
3
4
Sorted by value × movable sources × distance. It moves to first.

A gap is a prompt where the engines answer the question and do not name you. There are always more of them than you have weeks, so the ranking is the feature.

Movability multiplies three things the product already holds: what the prompt is worth in your own GA4; how movable the pages cited on it are, with competitor-owned pages excluded from that share; and how far your visibility on that prompt sits from the leader's.

A large deficit can rank low. If the only pages cited on a prompt are Crateline's own comparison and Northport WMS's documentation, and it sends four sessions a month, a 30-point deficit is not an opportunity — it is a fact about the category. Ranking by deficit alone produces a list sorted by how bad you feel.

Recomputed daily, from the same ledger the Sources page reads. → Sources

What a brief has to contain

“Get more reviews” is a failure of this feature.

If a recommendation could have been written without reading your category, it is a platitude with a checkbox beside it. A brief here names five things, and it can only name them because every answer was stored whole.

  1. 1 · The prompts it targets, named individually from the 54.
  2. 2 · The URL. One page, or one proposed path.
  3. 3 · The competitor sentence that beat you, verbatim, in the serif.
  4. 4 · The claim your page does not make. Here: multi-client billing.
  5. 5 · What will settle it: prompt set, baseline, window, read date.
Brief · 3PL billing pageIllustrative
01Promptsbest warehouse management software for a 3PL · + named prompts in the 3PL segment
02URLtallyfold.com/3pl-billing
03CompetitorQuoted verbatim below, from a stored answer
04Claim not madeMulti-client billing, handled natively
05Settles itNamed prompt set · pre-change baseline · category baseline · next read 5 Oct
Field 03 · the sentence that beat you

“…the strongest options are Crateline and Northport WMS — both handle multi-client billing natively, which most general warehouse systems do not.”

ChatGPT · 21 Sep · 02:10 ET · stored answer

A brief that cannot fill field 3 or field 4 is not issued. There is no fallback text and no generic template behind it, because the fallback is the failure mode — the moment a product starts emitting advice when it has nothing specific to say is the moment its specific advice stops being trusted.

The quote is a stored answer, not a paraphrase of one. → Answers

The experiment reads out

Net of the baseline, your rewrite is worth 7.8 points.

You rewrote the 3PL billing page. On the prompts the brief named, visibility went from 21.1% to 29.5% — a gain of 8.3 points. Over the identical days, every brand in Warehouse Management Software · US rose together by half a point, for reasons that had nothing to do with your page. That half point was going to happen to you anyway, so it comes out.

The subtraction is only possible because the category is collected whether or not anyone in it is a customer. All 17 brands are measured, so the counterfactual is a measurement rather than an estimate. A tool that samples only its own subscribers has nothing to subtract and has to hand you the 8.3.

Status is measuring. The experiment is not called a win until the effect clears its band.

+7.8
You · 21.1% → 29.5%+8.3
Category, same days · subtracted+0.5
Net effect+7.8 pts
measuring · next read 5 Oct
What drift removal means here

Drift removal is subtraction,
not exclusion.

There is a version of removing drift that means throwing away days. An engine has an odd week, the numbers dip, those days get flagged anomalous and dropped, and the chart afterwards shows a clean rise. That is choosing the sample after seeing the result — every individual step is defensible and the output is decided in advance.

Here it means one operation: measure what the category did over the identical window and subtract it. Every day stays in, including the ones that shrink your effect. If the category had risen 3 points, the 8.3 would read 5.3 and you would be told that in the same place. There is no exclude-this-day control in the product. It was never built, which is the only dependable way of never using it.

3PL billing experiment · named prompts · every day plottedTallyfoldcategory baselineIllustrative
baseline windowpage published21 Sep · last run
All days present. The baseline is measured, not modelled.
The third verdict

An effect smaller than its band
is inconclusive.

Closed experimentIllustrative
effect clears band+7.8 pts
effect +7.8 · band ±3.0 · clears itClosed. The effect is larger than its interval.
Closed experimentIllustrative
inconclusive+1.2 pts
effect +1.2 · band ±3.0 · inside itSample too small for an effect this size. A longer window would settle it: about 30 more days.

An experiment closes with one of three states: the effect cleared its band, it did not, or it is still measuring. The middle one prints as inconclusive, in the same place and at the same size as a success, and it is never rewritten into "early signs are positive".

The customer did the work. A product that wants to be liked will find them a number — a 1.2-point effect on a ±3.0 band can be drawn with an upward arrow and a green tint and nobody will complain that month. It is not a small win. It is an unanswered question.

An inconclusive read is usually not a failed rewrite. It is a sample problem — too few prompts in the set, too short a window, a band too wide for an effect that size. The card says which, and how much longer a read would take, so the response is to wait rather than to rewrite a page that may already be working.

Verdicts are computed, not editable, and they appear in white-label client reports unchanged. → For agencies

The limits

What we will not write for you.

→ Read the methodology

Not written

Citino does not generate the page, does not publish to your CMS, and does not promise placement on any prompt. The engines decide what they recommend. This product measures what they decided, before and after, against the rest of the category.

Not run

An experiment whose prompt set is too small to measure is refused rather than run, and the refusal states how many prompts it would take.

Not forecast

No result here is a forecast. Nothing on this page predicts what your visibility will be next month, because the honest answer is that we do not know and neither does anyone selling you a number for it.

Ship a change and find
out whether it worked.

The baseline exists before you start, because the category has been collecting all along. That is what makes the first experiment measurable rather than the fourth.

Read the methodology →
Next — For agencies →