Mentions, recommendations and citations answer different questions
A mention means the answer contains a recognised name for your business. It does not tell you whether AI offered it as a suitable supplier. “Company X does not provide this service” contains the name, but its accuracy and negative meaning matter to the business.
I assess a recommendation from the whole answer. The AI needs to offer the business as an option for the need in the question. I also check the conditions: a recommendation for the wrong branch or a service you do not provide deserves its own note. A name in a market overview would not, without more context, count as a recommendation.
A citation is a link given as a source for the answer. I distinguish between a link found during search and a link actually cited in the final text. If the provider does not return that information, I mark it as unknown. Then I open the link and check whether it supports the specific claim.
| Illustrative example | What to record | What to check |
|---|---|---|
| “You could contact company X.” | Mention and recommendation | Does the service and location match? |
| “Company X does not offer this service.” | Mention with a negative claim | Is the claim current and supported? |
| A link to a company article without naming the business | Citation of the business’s URL | Which sentence does the source support? |
| A technical error instead of an answer | Failed test | Can the test be safely repeated? |
Look for the denominator behind every percentage
Take an illustrative example. You plan twelve tests, ten return usable answers and four of those mention your business. The mention share among valid answers is 4 out of 10, or 40%. You also need to report that you obtained 10 of the 12 planned answers. Those are two different numbers.
A technical failure must not automatically look like an answer that left your business out. I look for counts of valid, failed and unassessable results. I also ask how an interrupted answer is treated. A few returned words do not necessarily make a complete answer.
I need to know the validity rules before assessing the result. If those rules change, I have the earlier answers in the comparison recalculated and describe the change. Otherwise, an apparent improvement could come from a different counting method.
What the L9 measurement covers
The L9 core set contains 30 customer questions. I ask ChatGPT, Gemini and Claude through their APIs with web search enabled: once per question in ChatGPT and Claude, and twice in Gemini. That means 120 planned answers in the main measurement. I report the actual number of usable results from the completed run.
I track questions that name the business and subsequent questions separately. I do not mix them into the main share. Google or Seznam can form separately agreed extensions where relevant to the target market; they are not part of the core set. Perplexity is not included in the current core set either. The report needs to state which part measures what.
An API test has recorded conditions and can be archived. I do not claim it reproduces every app user’s personal history, settings and environment. When comparing measurement services, ask how and where they collect the answers.
Repeated tests show variation where they have actually been run
AI can return a different output for the same input. OpenAI’s evaluation documentation also points this out. I therefore treat one answer as a particular observation at a particular time. [1]
Our core set has two runs per question in Gemini. I can show where those two observations differ. With one run per question in ChatGPT and Claude, I do not infer stability from that measurement. Two answers also do not establish how every future test will turn out.
Before comparing results, I check the exact wording, language, location, model, search setting, context and number of runs. I label a changed question as new. If a model or service availability changes, I state the limitation and compare only the part for which the conditions still make sense.
Review the answers that changed the conclusion
I inspect gained and lost mentions, recommendations and citations. I check for confusion between similar names, an old business name or a product brand mistaken for its retailer. I verify automated assessments by reading the answers. OpenAI also recommends calibrating automated metrics against human judgement. [1]
A useful conclusion is specific: the answer changed on this question, this source appeared and the service information is still wrong. I can use that finding to prepare a fix and another check. A higher overall percentage alone does not tell me whether AI offers the right service to the right customer.
Sources and checks
External claims use the sources listed below. The proposed steps and examples explain how L9 Studios approaches the work.
- OpenAI: Evaluation best practices
Variation in answers to the same input and checking automated metrics against human evaluation.
Sources checked on 2 October 2026. Features and availability may change when a service changes.