How the numbers are made

No certifications.
A published method instead.

Citino has no SOC 2 report, no ISO certificate and no compliance badges, and this page is not going to imply otherwise with a row of grey logos. We are new, and those take time and money we have not spent yet. What a badge buys is a stranger's assurance that we do what we say. In place of it, here is the thing itself: every decision that turns an engine's answer into a number on your chart, written out with the mechanism, so you can check the arithmetic rather than trust the auditor.

This is the second most-read page on the site, and it should be. Every other page shows you a number. This one tells you how wide the number's error is, what it excludes, when it is allowed to alert you, and what happens to your history when we improve the parser. If any of it turns out to be wrong, the corrections policy is further down and it is binding on the public benchmark pages as well as on your dashboard.

Read it adversarially. The whole point of publishing it is that you can.

The sample

What is sampled,
and how much of it.

Warehouse Management Software · US. Figures illustrative; brands fictional.

MeasuredHow
Engines sampledChatGPT · Gemini · Perplexity. Claude as an add-on.
Not sampledCopilot · Google AI Overviews. Neither offers an official answer API.
CadenceDaily, on a fixed schedule. Times shown in ET.
Session stateClean session. No account history, no personalisation, no prior turns.
Prompts per category54 in Warehouse Management Software · US, written as the questions buyers ask, then frozen.
Answers, 30 days1,785 in that category. Every one archived whole.
ScopeSegments you sell into, not the whole category, when you have set segments.
Sample behind a positionn = 184 — position is computed only over the answers that named you, so it always has a smaller sample and a wider band than visibility.

Two of those rows do most of the work. Prompts are the unit the bands are built over, which is the next section. And n for position is smaller than n for visibility in every category, for every brand, which is why the product prints the two sample sizes separately instead of one figure at the top of the page.

Uncertainty

The band is bootstrapped
over prompts, not answers.

If you resample answers — the common way, and the wrong one

Treat all 1,785 answers as 1,785 independent observations, resample them, take the middle 95%. It produces a tight, confident-looking band: on Tallyfold's 10.3%, roughly ±1.5. It is wrong because those answers are not independent. Thirty answers to "best warehouse management software for a 3PL" are near-copies of each other — if an engine names Crateline first today it very probably names Crateline first tomorrow. Repeating a question does not tell you more about the category; it tells you more about that one question. Counting each repeat as fresh evidence inflates the denominator without adding information, and because the interval narrows with the square root of the effective sample, the band comes out about half as wide as it really is.

0%10%20%10.3% ±1.5 — answers resampled, wrong
If you resample prompts — what Citino does

Draw 54 prompts with replacement from the 54, and each time a prompt is drawn, recompute the metric using all of that prompt's answers as one block. Repeat many thousands of times, take the 2.5th and 97.5th percentiles of the resulting distribution. The prompt is the independent unit, because a new prompt is a new question and a new answer to an old prompt mostly is not. On Tallyfold's 10.3% this gives ±3.0, and that is the number the product prints.

0%10%20%10.3% ±3.0 — prompts resampled
Tallyfold · visibility band as prompts are addedIllustrative
0%10%20%1 prompt54 prompts±3.0±12 and wider
Band on 10.3% as the bootstrap draws from more prompts. It narrows with the square root of the number of independent questions, not the number of answers.

How the interval is computed is the single most consequential decision in the product.

Every share Citino prints is a proportion drawn from a finite sample, so it carries a 95% interval. Most tools get this wrong in the direction that flatters them.

The difference is not academic. A band that is half as wide as it should be turns every ordinary week's wobble into a movement that clears its interval — which is exactly how a tool ends up emailing you about noise, and exactly why Tallyfold's +1.7 is grey here and would be green almost anywhere else.

Illustrative. The ±1.5 is what the wrong arithmetic prints on the same data, not a Citino number.

+1.7within noise
−3.00+3.0+1.7The delta sits inside its own ±3.0 band. Grey, no arrow, no alert.

Inside the band is not a move.

Tallyfold gained 1.7 points over thirty days on a band of ±3.0. Citino labels that within noise and does three specific things with it.

It prints it grey, with no arrow and no colour, because colour is a claim.

It does not alert on it. Alerts fire only on a delta that clears its own interval, so an empty alert inbox means the week was quiet rather than that the job is done.

And it never sorts as a movement. The biggest-movers list is computed only over deltas that cleared their bands, so a within-noise row cannot reach the top of it by being the largest of the small numbers. If nothing cleared this week, the list is empty and says so.

Drift

Your change,
net of the whole category.

Engines change. A model is updated, a retrieval policy shifts, a system prompt is rewritten, and the whole category moves on a Tuesday for reasons that have nothing to do with anything you published. Two rules handle it.

Drift is labelled, and never dropped. When a category-wide shift is detected, the date is marked on every chart in that category as a labelled drift event with the engine named. We do not smooth it, we do not restate history to make the line continuous, and we do not quietly re-baseline so that last month looks different than it did last month.

Your change is reported net of the category baseline. Rewriting the 3PL billing page took Tallyfold from 21.1% to 29.5% — a raw +8.3. Over the same window the whole category moved +0.5, so the reported effect is +7.8. Without the baseline you cannot tell a good page from a good week.

drift · engine named on the chart
Every line in the category steps on the same day. The day stays in the data, marked.
Raw changeIllustrative
+8.3
Tallyfold · 21.1% → 29.5%
Category baseline
+0.5
the whole category, same window
Net effect
+7.8
measuring · next read 5 Oct

A prompt's wording
is frozen. Two different
questions are not one trend line.

A prompt cannot be edited after it starts collecting — only retired and replaced. New wording is a new prompt with a new id, and it enters the chart as a hard break rather than a continuation, because the two wordings are different questions with different answer sets. The prompt set carries a version, and every archived answer records the version it was collected under.

This costs us something. When a prompt turns out to be badly worded, we cannot fix it and keep the history; we retire it, write the better one, and wait for the new series to be worth reading. Editing a prompt in place is how a chart silently becomes fiction.

“best warehouse management software”retired“… for a 3PL”new prompt · new idhard break
Two histories, never joined. The gap is part of the record.
Collection

Official APIs only, and what that costs us.

What we use

The engines' own official APIs, on paid accounts, at a fixed daily schedule. Every answer we report was returned to a request we made, under terms that permit it, and we can show you the response.

What we refuse

No scraping of consumer interfaces. No headless browsers driving a logged-in chat window. No third-party search proxies, no resold "AI rank" feeds, no data bought from someone who does not say how they got it.

What that costs

Real coverage. Copilot and Google AI Overviews are not sampled at all, because neither offers an official answer API — and both matter to buyers. Where they appear anywhere in Citino it is as a labelled referral source in your GA4 attribution, never as a measured answer. Adding a new engine takes as long as that engine takes to ship an API, so we are sometimes late. And the bill per answer is higher than scraping, which is part of why custom prompts are the thing the plans meter.

Why it is still the right trade

A scraped number and a sampled number look identical on a chart. If we mixed them you would have no way to tell which of your lines was built on a source that could break, change silently, or be turned off by the engine at any moment — and neither would we, six months later, when the archive is the only record of what happened. A measurement is only as durable as its least defensible input.

The record

Every number is recomputed
from the archive.

Perplexity21 Sep · 02:10 ETIllustrative
Crateline1 leads on multi-client billing. Palletworks2 and Tallyfold3 both suit smaller operations, with Tallyfold the lighter of the two. Bindley Ops4 is worth considering for cold-chain work.
Prompt · best warehouse management software for a 3PLPrompt set · frozen · daily since 24 Jun · Named 3rd of 4

Every answer is stored whole, exactly as the engine returned it, with the engine, the timestamp in ET, the prompt id and the prompt-set version attached. Not a summary, not the extracted brand list, not a sentiment score with the text thrown away. The archive is the record; every chart in the product is a derived view over it.

This is what makes the history correctable. When the parser improves — a brand name it was missing, a mention it was double-counting, a comparison table it was reading wrong — it re-reads the archive and the past improves with it. A tool that stores only extracted numbers has thrown away the evidence: it can stop making new mistakes, but it can never fix an old one, and it cannot show you the sentence a number came from. Restatements from a re-parse are dated and visible, and the changelog says what changed and which windows moved.

Corrections stay visible.

If you represent a brand on a Citino page — benchmark or dashboard — and something is wrong, tell us. Every claim we make is checkable against stored answers, which means a correction is a matter of fact rather than of goodwill.

What is wrongWhat we doWhat happens to the history
A mention misattributed to your brandRe-check against the stored answers, re-parse, restateCorrected series, with the restatement dated on the chart
A name, spelling or merged company wrongFix in the brand register, re-read the archiveHistory re-derived; nothing is re-typed by hand
A row that should not be in the tableRemove it and say the table changed, with the dateOther rows re-derived, because removing a brand changes every share
You want your brand excluded from a published tableWe do itThe page states that the table excludes a brand at its own request
We were not wrongWe show you the stored answers the number came fromNothing changes, and we explain the method rather than quietly adjust it

A correction is never silent. The page says what changed and when, and it keeps saying it — a fixed number with no note is indistinguishable from a number that was always that way, and the whole value of a public benchmark is that you can audit its past.

Silently dropping a brand row would be the worst of these, which is why it is the one we refuse to do quietly: every share in the table is a proportion of the same denominator, so one row leaving moves all of them.

Limits

What we cannot see.

Personalised answers. We sample clean sessions on a fixed schedule. Your buyer has chat history, an account, a location, a time of day and a way of phrasing things, and every one of those moves an answer. What Citino measures is what the engines say to a neutral asker of a fixed question — a stable, comparable population, and not a transcript of what any specific person saw.

Logged-in and enterprise surfaces. Not sampled. Neither are Copilot and Google AI Overviews, for the API reason above.

Twenty-nine per cent of your AI-referred sessions. Of 1,240 AI-referred sessions in thirty days, 71% can be matched to a prompt we sampled. The remaining 29% cannot, and the product reports that 29% as its own line rather than distributing it across the matched prompts. Distributing it would make every prompt's value look 40% larger and the report look complete.

Illustrative. Figures from the sample category.

29%
The honest list

What Citino does not claim.

We do not claim to know what your buyer sawonly what the engines answered our fixed questions, from a clean session, on a schedule.
We do not claim a causean experiment reports a measured change net of the category baseline, with a band and a next read date, and a page rewrite that coincides with a move is still a coincidence until the band says otherwise.
We do not claim coverage we do not haveCopilot and Google AI Overviews are unsampled, and no number on this site includes them.
We do not claim to improve your placementCitino measures and explains; it does not promise that anything you publish will make an engine name you, and anyone who does promise that cannot show you their band.
We do not claim visibility is revenueattribution reports sessions and the outcomes your GA4 records, with 29% of sessions explicitly unmatched.
We do not claim a census of your market17 brands in Warehouse Management Software · US means 17 brands the engines named in our sample, which is not the same as 17 companies existing.
We do not claim sentiment is reputation63 is an average over the sentences that named you in our sample, and it moves with wording as much as with regard.

Check it against the
benchmark before you buy.

Method changes are logged → Changelog

Everything on this page is doing its job on the public benchmark pages right now: bands over prompts, within-noise rows in grey, drift marked, Freight Audit Software published as paused because six brands and 402 answers will not carry a leader claim. No signup, no email field. If the method does not survive contact with a category you know well, that is the cheapest possible time to find out.

Claim your category →

Every figure on this page is illustrative and every brand shown is fictional. The method is not.