Skip to content
All writingMeasurement

How to choose an AI visibility tool: a 12-point evaluation checklist

Twelve criteria for evaluating AI visibility and citation tracking tools, plus ten questions to send a vendor before you buy. Vendor-neutral, built from how AI engines actually retrieve.

· 8 min read

Evaluate AI visibility tools on five things, in this order: coverage (which engines at the price you'd actually pay), attribution depth (domain, URL, or creator), control (can you choose what gets watched), evidence (can you see the raw answer that cited you), and commercials (what the meter counts). Homepage feature lists compare badly because vendors count engines differently and gate them by tier.

This is the vendor-neutral version. No products are named anywhere in it, which is deliberate: prices and tier contents in this category change monthly, and a checklist that names vendors goes stale faster than the criteria do. If you want the named comparison, it is in best AI citation tracking tools.

We build a tool in this category. The criteria below are the ones we would want a buyer to hold us to, including the two where we would lose.

Why this category is unusually hard to evaluate

Three structural reasons, worth understanding before you open a single pricing page.

Vendors count engines differently. One counts ChatGPT as one engine. Another counts ChatGPT, ChatGPT Search and ChatGPT Ads as three. A third counts Google AI Overviews and AI Mode separately, which is defensible, and a fourth bundles them, which is also defensible. "8+ engines" and "5 engines" may describe near-identical coverage.

Almost everyone gates engines by tier. The number on the homepage is the maximum. The number you get is whatever your tier includes. These are routinely different by a factor of five or more, and it is the single most common surprise in this category.

The thing being measured is unstable. AI answers change between runs. A tool that checks weekly and a tool that checks daily are not measuring the same phenomenon at different resolutions, they are measuring different phenomena. And engines behave very differently from each other: ChatGPT cites a mean of 6.88 sources per answer, Google AI Overviews 12.06, and Perplexity 16.35 (arXiv 2026, via AuthorityTech). A citation count blended across engines is close to meaningless.

The 12-point checklist

Grouped into five bands. Score each point 0, 1 or 2. Anything below 16 out of 24 is a tool you will outgrow inside two quarters.

Band 1: Coverage

1. Which engines are included at the tier I would actually buy?

Not the homepage maximum. Write down your tier, then the engine list for that tier. Do this first because it reorders every shortlist.

2. How many buyer queries can I track, and what happens at the limit?

Write your query list before you look at pricing. Ten questions becomes fifty once you split by persona, use case and competitor. If the plan caps you below your real list, you will quietly delete questions to fit the plan, and the deleted ones are the ones you never learn about.

3. How often is each query re-run?

Daily, weekly or on demand. Answers move. A weekly check will miss a citation that appeared and vanished, and drop detection is only as good as the sampling rate.

Band 2: Attribution depth

4. Does it resolve a citation to a domain, a URL, or a person?

Three genuinely different levels, and most tools stop at the first two:

  • Domain: "reddit.com cited 22 times." Tells you a platform matters.
  • URL: "this specific thread cited 6 times." Tells you which page to go work on.
  • Creator: "this person's three posts earned 14 citations." Tells you who to commission again.

If you pay or brief creators, level three is the one that closes your reporting loop, and it is rare.

A three-step ladder from domain to URL to creator attribution, showing that only creator-level data tells you which person to commission again
Domain tells you a platform matters. URL tells you which page. Only creator attribution tells you who to pay again.

5. Does it distinguish a citation from a mention?

Being named in an answer and being linked as a source are different outcomes with different value. A tool that blends them into one "visibility score" has destroyed the distinction you are paying to see. The definitions are in citation vs mention vs brand mention.

6. Can I see share of voice per engine, not blended?

Given 6.88 versus 16.35 mean sources per answer, a blended figure across engines is arithmetic noise. Insist on per-engine reporting. The conventions are in AI Share of Voice, off-site.

Band 3: Control

7. Can I tell it which specific URLs to watch?

Two opposite jobs hide behind one category name. Discovery answers "what is citing us?" Seeding answers "did the twelve placements I commissioned work?" Most tools do discovery. If you run an off-site programme with named creators, you need seeding, and you should ask explicitly because landing pages rarely mention it.

8. Can I track competitors alongside myself?

Your citation count means nothing without the denominator. Check whether competitor tracking is included or a higher tier, and how many you get.

9. Does it cover the markets and languages I sell into?

Answers differ by locale. If you sell in three countries, ask whether queries can be run per-market and whether that multiplies your prompt count, because it usually does.

Band 4: Evidence

10. Can I see the raw answer that cited me?

This is the one most buyers forget and most regret. When a stakeholder asks "are we sure?", you need the actual answer text with the citation in it, not a number in a dashboard. Receipts, not screenshots. The reporting format is in how to report off-site AI citations to stakeholders.

11. Is there a history I can point at, including drops?

Citations disappear. A tool that only shows current state cannot tell you that you lost something, and the loss is usually the more urgent event.

12. Can I export it?

CSV or API. Check whether export is included or gated to a higher tier. This is also your exit: if you cannot get your data out, an annual contract is longer than it looks.

Band 5: Commercials

Score these, but treat them as a gate rather than points, because they are pass/fail rather than better/worse.

What does the meter count? Prompts, credits, or nothing. This determines your real cost far more than the sticker price.

What is the billing term at the tier I want? Several vendors bill the entry tier annually. A quarter-long pilot may not be available at the price you were quoted.

How many seats? Some vendors include unlimited team members, some charge per seat. If your agency or client team is ten people, this swings the total materially.

The ten questions to send a vendor

Copy this into an email. The answers take a vendor ten minutes and save you a quarter. Any vendor who will not answer them in writing has told you something.

  1. At the tier priced at [X], exactly which AI engines are included? Please list them.
  2. How many distinct queries can I track at that tier, and what happens when I exceed it?
  3. How frequently is each query re-run, and can I change that?
  4. Do you attribute a citation to a domain, a specific URL, or an individual author or creator?
  5. Can I submit a list of specific URLs I want monitored, rather than only discovering them?
  6. Do you report citations separately from unlinked brand mentions?
  7. Is share of voice reported per engine, or blended across engines?
  8. Can I retrieve the full raw answer text in which my brand was cited?
  9. How far back does history go, and do you alert on a lost citation?
  10. Is CSV or API export included at my tier, and what is the billing term?

Question 5 is the one that most often produces a revealing pause, because discovery and seeding are different products and the category name does not distinguish them.

Three traps

Trap 1: buying on engine count. More engines is not more insight if the extra ones are not where your buyers ask questions. Two engines your buyers actually use beats nine they do not. Work out which engines matter first, from where AI actually gets its answers.

Trap 2: expecting traffic. A citation frequently produces no click at all. If you build the business case on sessions, it will fail review at the first quarterly check even when the programme is working. Report presence in the answer.

Trap 3: treating it as a rank tracker. Only 38% of the pages cited in Google's AI Overviews also rank in its top 10, down from 76% eight months earlier (Ahrefs, via Search Engine Journal), though Ahrefs credits part of that fall to better citation detection rather than to a shift in Google alone. Either way, rankings stopped predicting citations, which is why this category exists at all. Keep your rank tracker and add this alongside, rather than expecting one to replace the other. The full argument is in what off-site GEO is.

How to decide in one week

Day 1. Write your query list and your engine list. Both before you look at any pricing page. These two documents decide most of the outcome.

Day 2. Send the ten questions to three vendors. Include one you think is too expensive and one you think is too cheap, because the boundaries are where you learn what the category actually costs.

Days 3 to 5. Run the free tiers and trials in parallel against the same ten queries. Same queries, same days, or you are not comparing anything. Note which tool surfaces a source the others missed.

Day 6. Score all three against the twelve points. Write down where each one loses, not just where it wins. A vendor who cannot tell you what they are bad at has not been evaluated.

Day 7. Decide, and diarise a re-check for 60 days out. Prices and tier contents in this category move fast enough that your decision has a shelf life.

What no tool in this category does

Stated plainly, because a bad expectation produces a bad quarter regardless of which tool you pick.

None of them earn you citations. They measure. The work that gets you cited happens off-site: Reddit is referenced in roughly 40% of AI answers (Semrush, 150,000 citations, via MaxAEO), and domains with active G2 or Capterra profiles show roughly 3x higher citation probability (Leapd). No dashboard writes that content or earns that placement.

None of them make answers stable. You are sampling a system that changes. Good tools tell you the sampling rate and show you the variance. Tools that present a single confident number are hiding it.

Plugged in, or plugged out?

Find out what AI is already quoting about you.

free to start · no credit card