Country
Malta · English · EUR
EUR
Currency

Prices only. You are always billed in EUR.

OUR METHOD, IN PUBLIC

We publish how our score works — including where it is weak.

In July 2026 an independent consultant put 31 methodology questions to every company selling an AI visibility score. Two answered. Here are ours, all 31, including the ones where the honest answer is that we do not know.

Anyone selling you a number should be able to tell you what it measures, what it does not, how much it moves on its own, and what result would make them withdraw it. Most cannot. That is the whole reason this page exists.

19 ANSWERED6 PARTIAL6 WE DO NOT KNOW
What our score does measure
  • Whether a named business appears in the answer to a locked set of questions, per AI system, on a given date.
  • Whether that appearance is a recommendation, a mention, or an inaccurate description.
  • Which sources the engine surfaced when it answered.
  • How often a probe failed — shown, not hidden in the denominator.
What it does not measure
  • It does not estimate what every buyer in your market sees. It describes a panel we configured.
  • It does not account for personalisation, logged-in history or multi-turn conversation. Every probe is logged-out and single-turn.
  • It does not carry a confidence interval, because we have not yet measured our own run-to-run variance.
  • It does not predict revenue. No study we have run shows that it does.
What would make us withdraw it
  • Run-to-run variance large enough that ordinary movements cannot be told from noise.
  • A test-retest agreement rate we would not accept from a supplier.
  • A cross-vendor comparison we cannot explain in our favour.
  • An outcome study finding no relationship with anything a client cares about.
THE 31 QUESTIONS

The questions are David McSweeney's, published on 28 July 2026, and they are reproduced here in shortened form. The answers are ours. As at 29 July 2026 two vendors had answered him publicly — Ternith and Spyglasses — and we could not find a published set from any of the platforms we compete with.

LAST REVIEWED · 23 AUGUST 2026

01What exactly is being measured?ANSWERED

Whether a named business appears in the answer to a specific question, put to a specific AI system, at a specific time — and if it appears, whether it is recommended, merely mentioned, or described inaccurately. We do not combine those into a single weighted score with a hidden formula. Where we show a composite, the components are shown beside it.

02What population does the score claim to represent?ANSWERED

The prompt panel we configured for that business, and nothing wider. It is a description of a test panel, not an estimate of what every buyer in a market encounters. We do not present it as general AI visibility, because it is not.

03How were the tracked prompts shown to represent that population?PARTIAL

They are not, and we say so. The phrases come from the business's own trade and specialism, written as the questions a buyer would plausibly ask, and they are reviewed by a person. What we do NOT have is observed first-party data on what real users actually type into ChatGPT about that trade — nobody outside the model providers does. Profound comes closest — they license real conversations from consumer panels, which is better raw material than we have. But this question asks how a prompt set was SHOWN to represent the population, and we could not find that published either: the panels are not named, the scaling to population is described only as statistical modelling, and no figure we found carries a margin of error. On this question they have the better data and we have the same missing argument.

04How many independent user intents does the portfolio contain?PARTIAL

Fewer than the raw prompt count, and we do not currently report an effective sample size adjusted for the fact that paraphrases of one intent are correlated. This is a real gap. Until it is closed, treat the prompt count as a count of tests, not of independent evidence.

05How are prompt weights determined?ANSWERED

They are not weighted. Every phrase counts once. That is a crude choice and it has a cost — a rare phrasing counts as much as a commercially central one — but the alternative is a weighting model we could not substantiate, which would move the arbitrariness somewhere less visible.

06How do you prevent prompt selection from determining the result?ANSWERED

The phrases are recorded against the site on the first run and LOCKED on every repeat, so the second score is measured against the same questions as the first. A client cannot quietly add favourable phrases or drop hard ones between runs. If the phrase set is deliberately changed, the run starts a new series rather than continuing the old line, and the report says so.

07Does greater scale improve validity or only precision?ANSWERED

Only precision. Running more variations of an unrepresentative panel produces a more stable estimate of an unrepresentative panel. We do not sell prompt volume as a proxy for rigour.

08Under what conditions is one run per prompt sufficient?WE DO NOT KNOW

We do not know. We currently run each phrase once per audit and we have not measured the run-to-run variance that would tell us whether once is enough. This is the single largest hole in our own method and it is first on the list to close.

09What uncertainty accompanies each customer's score?WE DO NOT KNOW

None is calculated, because question 8 is unanswered — without a measured run-to-run variance there is no honest interval to publish. Until there is, no movement in our score should be read as meaningful on its own.

10What uncertainty is excluded from the statistical model?ANSWERED

All of it. There is no statistical model. We would rather show a raw count with its limits stated than a confidence interval that quietly excludes prompt selection, panel composition and model drift — a narrow interval around a biased panel is a more confident-looking wrong answer.

11Are prompt-platform observations statistically independent?PARTIAL

No, and we do not treat them as if they were. Phrases sharing an origin, and every phrase run against the same model on the same day, move together when the model changes. We do not multiply them up into an independent evidence count.

12Why is daily cadence methodologically appropriate?ANSWERED

It is not, so we do not run daily. The audit is a point-in-time reading and the retainer reports monthly. A daily line would mostly plot answer variance, and a chart that moves every day is commercially useful and methodologically empty.

13How was the methodology validated?WE DO NOT KNOW

It has not been validated against an external benchmark. We have no study comparing our score to referral traffic, conversions or survey recall. We are a young company with no such dataset, and we would rather say that than cite someone else's validation of a different system.

14Does the validation use genuinely external data?WE DO NOT KNOW

There is no validation, so no. When there is, it will use data we did not generate, select, weight or score — and we will publish the negative result if that is what it is.

15What assumptions does each estimator require?ANSWERED

The only assumption we make is that a logged-out, one-shot query is a defensible standardised probe — comparable across runs and across businesses, and NOT the same as what a logged-in user with history sees. See question 16.

16How do you account for personalisation and conversation history?ANSWERED

We do not, and we could not. Every probe is logged-out and single-turn. That means our reading is a controlled measurement, not a simulation of a real session, and personalisation is an excluded source of error rather than a modelled one. This limitation applies to every vendor in this market, including the ones that do not mention it.

17What does synthetic persona text actually validate?ANSWERED

Nothing we would rely on, which is why we do not use synthetic personas. Typing 'I am a CFO in London' into a prompt does not reproduce account history, location or memory; it produces a differently-worded prompt. Selling that as a persona segment would be inventing a variable.

18How reproducible are the results?PARTIAL

The inputs are: the phrase set is locked and recorded, so a client can rerun exactly what we ran, against the same models, outside our platform. What we have not published is a measured test-retest agreement rate — see question 8.

19What agreement exists between vendors?WE DO NOT KNOW

We have not run the same panel through a competitor's platform and compared. It would be a genuinely useful experiment and we would publish it either way.

20How are differences between models handled?ANSWERED

Per model, side by side. We do not average ChatGPT, Gemini, Perplexity and Google AI Overviews into one number, because we have no defensible basis for the weights that would require. A business invisible in one and strong in another has a specific problem, and averaging hides exactly the thing worth acting on.

21What counts as a meaningful mention or recommendation?PARTIAL

A first-place recommendation and a passing mention are recorded differently and shown separately. A negative mention does not count as visibility. Aliases, trading names and parent companies are resolved by hand at setup and a client can correct them. What we do not yet do is recompute history when a resolution rule changes — the report is annotated instead.

22What is the unit of counting?ANSWERED

One response, one observation. A business named five times in a single answer counts once. Share-of-voice denominators use every brand named in the answer, not just the competitors a client chose — otherwise the client sets their own denominator.

23How are citations interpreted?ANSWERED

A citation is a source the engine itself surfaced — a linked source or an entry in a sources panel — attributed to the domain. We do not infer citations, we do not count a plain-text URL in prose as one, and we do not reproduce verbatim AI Overview text as evidence, on a supplier restriction we chose to build around rather than argue with.

24How are failures and missing observations treated?ANSWERED

A refusal, timeout or malformed response is recorded as a FAILURE, not as an absence of the brand, and it is excluded from the denominator. The failure count appears on the report. Counting a timeout as 'you were not recommended' would manufacture a decline out of an outage.

25Can customers inspect and reproduce the raw measurement?PARTIAL

Clients get the exact phrase set, the model and the run date, which is enough to rerun it themselves. Full unedited outputs and timestamps are available on request but are not yet exported by default. That should be a button, not a request, and it will be.

26How are methodology and model changes handled historically?ANSWERED

A change to the phrase set or the scoring starts a new series rather than continuing the old line. We do not silently restate history, and we do not present readings taken with two different instruments as one continuous trend.

27What causal claims are permitted?ANSWERED

None. We do not claim a content change caused a movement. We have no control group and no counterfactual, so temporal sequence is all we have, and temporal sequence is not causation. What we report is what we did and what the reading was afterwards, labelled as exactly that.

28What safeguards exist before automated agents act?ANSWERED

We do not run autonomous agents that edit content or create pages off the back of a score. Every change is written by a person or reviewed by one before it ships. If we ever automate an action from a measurement, the repetition threshold, the minimum effect size and the human-review requirement will ship at the same time and be visible to the client — not added afterwards.

29What external outcome does the score predict?WE DO NOT KNOW

We do not know, and we will not claim otherwise. We have not demonstrated that a higher reading predicts referral traffic, pipeline or revenue while controlling for brand size. Improving the score currently predicts one thing reliably: a later improvement in the score.

30Why should the metric be treated as a benchmark?ANSWERED

It should not, and we do not describe it as one. It is a directional monitoring tool for a configured panel. That limitation is printed at the same size as the number, on the report and on this page.

31What would falsify the methodology?ANSWERED

Any of these, and we would say so publicly: a measured run-to-run variance large enough that ordinary movements are indistinguishable from noise; a test-retest agreement rate below what we would accept from a supplier; a cross-vendor comparison showing our readings disagree materially with two others and we cannot show ours is closer to what a user sees; or a first outcome study finding no relationship with anything a client cares about. If none of those would change our mind, the method has not been tested.

If any of this stops being true, it gets changed the same day — a stale methodology page is worse than none, because it is a published claim about our own conduct. If you think an answer is wrong, or evasive, tell us and we will either fix the answer or fix the method.

READY?

Get found. Fast.

The free audit usually lands in a couple of minutes, and shows you exactly what invisible is costing you.

Get your free AI visibility audit →
Abbreviations used on this site

Statute, regulator and standard names are given in their own language, because that is what they are called — a translated citation would be one you could not look up.

ABN
Australian Business Number · Australia
ACL
Australian Consumer Law · Australia
AEO
answer engine optimisation
AI
artificial intelligence
AIO
AI optimisation
ANPD
Autoridade Nacional de Proteção de Dados · Brazil
API
application programming interface
ASA
Advertising Standards Authority · United Kingdom
CASL
Canada's Anti-Spam Legislation · Canada
CCPA
California Consumer Privacy Act · United States
CMS
content management system
CPRA
California Privacy Rights Act · United States
CRO
conversion rate optimisation
CSS
Cascading Style Sheets
DMCC
Digital Markets, Competition and Consumers Act 2024 · United Kingdom
DPA
data processing agreement
DTC
direct-to-consumer
EEA
European Economic Area
EFTA
European Free Trade Association
ESIGN
Electronic Signatures in Global and National Commerce Act · United States
FTC
Federal Trade Commission · United States
GDPR
General Data Protection Regulation · European Union
GEO
generative engine optimisation
GST
goods and services tax
ICO
Information Commissioner's Office · United Kingdom
IDTA
International Data Transfer Agreement · United Kingdom
IP
intellectual property
KPI
key performance indicator
LGPD
Lei Geral de Proteção de Dados · Brazil
LLC
limited liability company
LLM
large language model
MCP
Model Context Protocol
MFA
multi-factor authentication
OAIC
Office of the Australian Information Commissioner · Australia
OPC
Office of the Privacy Commissioner of Canada · Canada
PECR
Privacy and Electronic Communications Regulations · United Kingdom
PII
personally identifiable information
PIPEDA
Personal Information Protection and Electronic Documents Act · Canada
QA
quality assurance
SCC
standard contractual clauses · European Union
SEO
search engine optimisation
SMS
short message service
TCPA
Telephone Consumer Protection Act · United States
UEMA
Unsolicited Electronic Messages Act · New Zealand
UETA
Uniform Electronic Transactions Act · United States
UGC
user-generated content
UWG
Bundesgesetz gegen den unlauteren Wettbewerb · Switzerland
VAT
value added tax
WCAG
Web Content Accessibility Guidelines