United Kingdom · English · GBP
Prices, terms and tax follow the country you choose.
GBP
Prices only. You are always billed in GBP.
We publish how our score works — including where it is weak.
In July 2026 an independent consultant put 31 methodology questions to every company selling an AI visibility score. Two answered. Here are ours, all 31, including the ones where the honest answer is that we do not know.
Anyone selling you a number should be able to tell you what it measures, what it does not, how much it moves on its own, and what result would make them withdraw it. Most cannot. That is the whole reason this page exists.
- Whether a named business appears in the answer to a locked set of questions, per AI system, on a given date.
- Whether that appearance is a recommendation, a mention, or an inaccurate description.
- Which sources the engine surfaced when it answered.
- How often a probe failed — shown, not hidden in the denominator.
- It does not estimate what every buyer in your market sees. It describes a panel we configured.
- It does not account for personalisation, logged-in history or multi-turn conversation. Every probe is logged-out and single-turn.
- It does not carry a confidence interval, because we have not yet measured our own run-to-run variance.
- It does not predict revenue. No study we have run shows that it does.
- Run-to-run variance large enough that ordinary movements cannot be told from noise.
- A test-retest agreement rate we would not accept from a supplier.
- A cross-vendor comparison we cannot explain in our favour.
- An outcome study finding no relationship with anything a client cares about.
The questions are David McSweeney's, published on 28 July 2026, and they are reproduced here in shortened form. The answers are ours. As at 29 July 2026 two vendors had answered him publicly — Ternith and Spyglasses — and we could not find a published set from any of the platforms we compete with.
LAST REVIEWED · 23 AUGUST 2026
01What exactly is being measured?ANSWERED
Whether a named business appears in the answer to a specific question, put to a specific AI system, at a specific time — and if it appears, whether it is recommended, merely mentioned, or described inaccurately. We do not combine those into a single weighted score with a hidden formula. Where we show a composite, the components are shown beside it.
02What population does the score claim to represent?ANSWERED
The prompt panel we configured for that business, and nothing wider. It is a description of a test panel, not an estimate of what every buyer in a market encounters. We do not present it as general AI visibility, because it is not.
03How were the tracked prompts shown to represent that population?PARTIAL
They are not, and we say so. The phrases come from the business's own trade and specialism, written as the questions a buyer would plausibly ask, and they are reviewed by a person. What we do NOT have is observed first-party data on what real users actually type into ChatGPT about that trade — nobody outside the model providers does. Profound comes closest — they license real conversations from consumer panels, which is better raw material than we have. But this question asks how a prompt set was SHOWN to represent the population, and we could not find that published either: the panels are not named, the scaling to population is described only as statistical modelling, and no figure we found carries a margin of error. On this question they have the better data and we have the same missing argument.
04How many independent user intents does the portfolio contain?PARTIAL
Fewer than the raw prompt count, and we do not currently report an effective sample size adjusted for the fact that paraphrases of one intent are correlated. This is a real gap. Until it is closed, treat the prompt count as a count of tests, not of independent evidence.
05How are prompt weights determined?ANSWERED
They are not weighted. Every phrase counts once. That is a crude choice and it has a cost — a rare phrasing counts as much as a commercially central one — but the alternative is a weighting model we could not substantiate, which would move the arbitrariness somewhere less visible.
06How do you prevent prompt selection from determining the result?ANSWERED
The phrases are recorded against the site on the first run and LOCKED on every repeat, so the second score is measured against the same questions as the first. A client cannot quietly add favourable phrases or drop hard ones between runs. If the phrase set is deliberately changed, the run starts a new series rather than continuing the old line, and the report says so.
07Does greater scale improve validity or only precision?ANSWERED
Only precision. Running more variations of an unrepresentative panel produces a more stable estimate of an unrepresentative panel. We do not sell prompt volume as a proxy for rigour.
08Under what conditions is one run per prompt sufficient?WE DO NOT KNOW
We do not know. We currently run each phrase once per audit and we have not measured the run-to-run variance that would tell us whether once is enough. This is the single largest hole in our own method and it is first on the list to close.
09What uncertainty accompanies each customer's score?WE DO NOT KNOW
None is calculated, because question 8 is unanswered — without a measured run-to-run variance there is no honest interval to publish. Until there is, no movement in our score should be read as meaningful on its own.
10What uncertainty is excluded from the statistical model?ANSWERED
All of it. There is no statistical model. We would rather show a raw count with its limits stated than a confidence interval that quietly excludes prompt selection, panel composition and model drift — a narrow interval around a biased panel is a more confident-looking wrong answer.
11Are prompt-platform observations statistically independent?PARTIAL
No, and we do not treat them as if they were. Phrases sharing an origin, and every phrase run against the same model on the same day, move together when the model changes. We do not multiply them up into an independent evidence count.
12Why is daily cadence methodologically appropriate?ANSWERED
It is not, so we do not run daily. The audit is a point-in-time reading and the retainer reports monthly. A daily line would mostly plot answer variance, and a chart that moves every day is commercially useful and methodologically empty.
13How was the methodology validated?WE DO NOT KNOW
It has not been validated against an external benchmark. We have no study comparing our score to referral traffic, conversions or survey recall. We are a young company with no such dataset, and we would rather say that than cite someone else's validation of a different system.
14Does the validation use genuinely external data?WE DO NOT KNOW
There is no validation, so no. When there is, it will use data we did not generate, select, weight or score — and we will publish the negative result if that is what it is.
15What assumptions does each estimator require?ANSWERED
The only assumption we make is that a logged-out, one-shot query is a defensible standardised probe — comparable across runs and across businesses, and NOT the same as what a logged-in user with history sees. See question 16.
16How do you account for personalisation and conversation history?ANSWERED
We do not, and we could not. Every probe is logged-out and single-turn. That means our reading is a controlled measurement, not a simulation of a real session, and personalisation is an excluded source of error rather than a modelled one. This limitation applies to every vendor in this market, including the ones that do not mention it.
17What does synthetic persona text actually validate?ANSWERED
Nothing we would rely on, which is why we do not use synthetic personas. Typing 'I am a CFO in London' into a prompt does not reproduce account history, location or memory; it produces a differently-worded prompt. Selling that as a persona segment would be inventing a variable.
18How reproducible are the results?PARTIAL
The inputs are: the phrase set is locked and recorded, so a client can rerun exactly what we ran, against the same models, outside our platform. What we have not published is a measured test-retest agreement rate — see question 8.
19What agreement exists between vendors?WE DO NOT KNOW
We have not run the same panel through a competitor's platform and compared. It would be a genuinely useful experiment and we would publish it either way.
20How are differences between models handled?ANSWERED
Per model, side by side. We do not average ChatGPT, Gemini, Perplexity and Google AI Overviews into one number, because we have no defensible basis for the weights that would require. A business invisible in one and strong in another has a specific problem, and averaging hides exactly the thing worth acting on.
21What counts as a meaningful mention or recommendation?PARTIAL
A first-place recommendation and a passing mention are recorded differently and shown separately. A negative mention does not count as visibility. Aliases, trading names and parent companies are resolved by hand at setup and a client can correct them. What we do not yet do is recompute history when a resolution rule changes — the report is annotated instead.
22What is the unit of counting?ANSWERED
One response, one observation. A business named five times in a single answer counts once. Share-of-voice denominators use every brand named in the answer, not just the competitors a client chose — otherwise the client sets their own denominator.
23How are citations interpreted?ANSWERED
A citation is a source the engine itself surfaced — a linked source or an entry in a sources panel — attributed to the domain. We do not infer citations, we do not count a plain-text URL in prose as one, and we do not reproduce verbatim AI Overview text as evidence, on a supplier restriction we chose to build around rather than argue with.
24How are failures and missing observations treated?ANSWERED
A refusal, timeout or malformed response is recorded as a FAILURE, not as an absence of the brand, and it is excluded from the denominator. The failure count appears on the report. Counting a timeout as 'you were not recommended' would manufacture a decline out of an outage.
25Can customers inspect and reproduce the raw measurement?PARTIAL
Clients get the exact phrase set, the model and the run date, which is enough to rerun it themselves. Full unedited outputs and timestamps are available on request but are not yet exported by default. That should be a button, not a request, and it will be.
26How are methodology and model changes handled historically?ANSWERED
A change to the phrase set or the scoring starts a new series rather than continuing the old line. We do not silently restate history, and we do not present readings taken with two different instruments as one continuous trend.
27What causal claims are permitted?ANSWERED
None. We do not claim a content change caused a movement. We have no control group and no counterfactual, so temporal sequence is all we have, and temporal sequence is not causation. What we report is what we did and what the reading was afterwards, labelled as exactly that.
28What safeguards exist before automated agents act?ANSWERED
We do not run autonomous agents that edit content or create pages off the back of a score. Every change is written by a person or reviewed by one before it ships. If we ever automate an action from a measurement, the repetition threshold, the minimum effect size and the human-review requirement will ship at the same time and be visible to the client — not added afterwards.
29What external outcome does the score predict?WE DO NOT KNOW
We do not know, and we will not claim otherwise. We have not demonstrated that a higher reading predicts referral traffic, pipeline or revenue while controlling for brand size. Improving the score currently predicts one thing reliably: a later improvement in the score.
30Why should the metric be treated as a benchmark?ANSWERED
It should not, and we do not describe it as one. It is a directional monitoring tool for a configured panel. That limitation is printed at the same size as the number, on the report and on this page.
31What would falsify the methodology?ANSWERED
Any of these, and we would say so publicly: a measured run-to-run variance large enough that ordinary movements are indistinguishable from noise; a test-retest agreement rate below what we would accept from a supplier; a cross-vendor comparison showing our readings disagree materially with two others and we cannot show ours is closer to what a user sees; or a first outcome study finding no relationship with anything a client cares about. If none of those would change our mind, the method has not been tested.
If any of this stops being true, it gets changed the same day — a stale methodology page is worse than none, because it is a published claim about our own conduct. If you think an answer is wrong, or evasive, tell us and we will either fix the answer or fix the method.
Get found. Fast.
The free audit usually lands in a couple of minutes, and shows you exactly what invisible is costing you.
Get your free AI visibility audit →Abbreviations used on this site
Statute, regulator and standard names are given in their own language, because that is what they are called — a translated citation would be one you could not look up.
- ABN
- Australian Business Number · Australia
- ACL
- Australian Consumer Law · Australia
- AEO
- answer engine optimisation
- AI
- artificial intelligence
- AIO
- AI optimisation
- ANPD
- Autoridade Nacional de Proteção de Dados · Brazil
- API
- application programming interface
- ASA
- Advertising Standards Authority · United Kingdom
- CASL
- Canada's Anti-Spam Legislation · Canada
- CCPA
- California Consumer Privacy Act · United States
- CMS
- content management system
- CPRA
- California Privacy Rights Act · United States
- CRO
- conversion rate optimisation
- CSS
- Cascading Style Sheets
- DMCC
- Digital Markets, Competition and Consumers Act 2024 · United Kingdom
- DPA
- data processing agreement
- DTC
- direct-to-consumer
- EEA
- European Economic Area
- EFTA
- European Free Trade Association
- ESIGN
- Electronic Signatures in Global and National Commerce Act · United States
- FTC
- Federal Trade Commission · United States
- GDPR
- General Data Protection Regulation · European Union
- GEO
- generative engine optimisation
- GST
- goods and services tax
- ICO
- Information Commissioner's Office · United Kingdom
- IDTA
- International Data Transfer Agreement · United Kingdom
- IP
- intellectual property
- KPI
- key performance indicator
- LGPD
- Lei Geral de Proteção de Dados · Brazil
- LLC
- limited liability company
- LLM
- large language model
- MCP
- Model Context Protocol
- MFA
- multi-factor authentication
- OAIC
- Office of the Australian Information Commissioner · Australia
- OPC
- Office of the Privacy Commissioner of Canada · Canada
- PECR
- Privacy and Electronic Communications Regulations · United Kingdom
- PII
- personally identifiable information
- PIPEDA
- Personal Information Protection and Electronic Documents Act · Canada
- QA
- quality assurance
- SCC
- standard contractual clauses · European Union
- SEO
- search engine optimisation
- SMS
- short message service
- TCPA
- Telephone Consumer Protection Act · United States
- UEMA
- Unsolicited Electronic Messages Act · New Zealand
- UETA
- Uniform Electronic Transactions Act · United States
- UGC
- user-generated content
- UWG
- Bundesgesetz gegen den unlauteren Wettbewerb · Switzerland
- VAT
- value added tax
- WCAG
- Web Content Accessibility Guidelines