How do you evaluate an AI visibility platform?

Ask nine questions: what is counted, ratio or index, which engines, whether gaps arrive drafted, who approves client work, whether the archive is dated, whether white label costs extra, what a seat covers, and what a bad month looks like.

Rayhaan Rasheed9 min read

Counted ratio
A number you could rebuild by hand: how many checks named the business, out of how many were run, over a stated window. Every part of it is visible, so you are free to disagree with any part of it.
Composite index
Several different things blended into one figure, usually out of a hundred. It looks decisive, and whether you can audit it depends on whether the inputs and their weights are published. If they are not, you cannot tell which input moved when the figure moves.
Seat
The unit a platform charges for. What one seat covers is a large part of what a tool costs an agency over a year, and it is often the thing a pricing page is least specific about.

How to use this list

Tools in this category describe themselves in similar language, so the labels will not separate them for you. The nine questions below will. None of them requires you to understand how a language model works, because all nine are about what the tool produces rather than how it produces it.

Ask all nine of any vendor you are considering, including this one. Answering them plainly does not make a tool right for your agency. Being unable to answer them is the finding.

What exactly is being counted?

Start here, because everything else is downstream of it. Ask what one unit of measurement is. A defensible answer sounds like: one question, asked of one engine, at one point in time. Then ask what it takes to count as being named. The business name in the answer text is a strict definition and a good one. A link to a directory page that happens to list the business is not the same thing, and neither is the engine describing the category without naming anyone.

Then ask whether the questions stay fixed between runs. If the question set drifts, a change in the number is a fact about the questions rather than a fact about the business, and nobody will be able to tell the two apart six months later.

Is the score a counted ratio or a composite index?

Ask for the numerator, the denominator and the window. If all three come back, you are looking at a ratio and you can audit it. If what comes back is a figure out of a hundred, ask what goes into it, how each input is weighted, and which input moved the last time the figure moved. You need those answers to explain the number to a client, so get them before you buy rather than the week you are asked.

This matters most on the day a client asks where the number came from. With a ratio you show the counts and the window, and the client can redo the arithmetic themselves. That difference costs nothing until the day it is the only thing that matters.

Which engines are named, and are they reported separately?

Ask the vendor to name the engines, in writing, rather than accepting a phrase like "the major AI assistants". A named list is a commitment you can hold them to and a vague one is a moving target.

Then ask whether each engine is reported on its own line or averaged into one figure. The engines disagree with each other constantly. A business can be the obvious recommendation on one and completely absent from another, and an average across them hides precisely the fact you needed. A zero on one engine is a bigger finding than a decent overall number.

Do gaps arrive with the work drafted?

A dashboard that tells you a number fell has told you the easy half. Ask what arrives beside the gap: the specific thing to change, and some ordering that says which one to do first. Ask to see a real example, not a screenshot of an empty state.

This is the question that decides whether the tool saves your team time or adds a step to their week. A ranked queue of drafted work means somebody is editing. A chart means somebody is starting from a blank page, and that person is billing you for the hour.

Who approves anything that reaches a client?

Ask whether the tool can send, publish or push anything to your client or your client's website without a human on your side saying yes first. Then ask what happens by default, because the default is what will actually happen once your team is busy.

Automation that reaches a client directly takes over the part of the relationship you are paid for. Your brand on a message you did not approve is worse than the vendor's brand on one you did.

Is the archive dated and complete?

Ask whether every past answer is kept, with its date, and whether you can go back and read the one that produced any given point on the trend line. A tool that stores only the current number can draw you a chart, but it cannot show you the evidence under it.

The archive is what a renewal conversation runs on. "The score went up" is a claim. "Here is the answer from March, here is the answer from August, and here is what changed in between" is a record, and the second one is the reason the line item survives a budget review.

Does white label cost extra?

Ask whether client-facing branding is included in the plan you would actually buy, or whether it sits on a tier above it. Ask whether it is billed per report, per client, or as a percentage. Then ask what it covers: the dashboard is the easy part, and the artifacts a client keeps are reports, emails and alerts.

A branding settings screen in a demo is not proof. Ask for one delivered example, filled in, of every artifact a client would receive.

What does a seat actually cover?

Ask what one seat buys. One client business, or one login? Both? Then ask the two questions that decide the real annual cost: what does adding a second person from your team cost, and what happens when you get to ten clients.

Per-user pricing is the quiet one. It looks small on the proposal and it is the line that grows every time you hire, which means the tool gets more expensive precisely as the account team gets bigger.

What happens when a month goes badly?

Ask what the client report looks like in a month where the number fell. Ask whether the bad month stays in the archive. Ask whether the tool has anywhere to put a decline other than a smaller bar. A demo is likely to show a good month, so this is the question that shows you the rest of the product.

Accounts have bad months. A tool built only to display improvement will eventually make you explain a hole in the record, and you will be explaining it in front of the client. A tool that keeps the bad month and tells you why is the one you can still be using in year three.

Where we land on our own nine

It would be a strange list to publish without answering it, so here is Apex against the same nine, in the same terms used everywhere else on this site.

  • What is counted: one tracked question, asked of one engine, at one scheduled time. Named means the business name in the answer text.
  • The score: appearance rate, published as your AI Visibility Score, always beside its counts. It is a counted ratio, not a composite index, and never a figure out of a hundred.
  • The engines, named: ChatGPT, Gemini, Perplexity and Claude, reported separately rather than averaged together.
  • Gaps: each one comes back with the fix already drafted and ranked by what moves the number.
  • Approval: Apex drafts and you decide. Nothing publishes to a client site without your OK.
  • The archive: every check is kept and dated, including the months that went sideways.
  • White label: included with every seat rather than sold as an upgrade.
  • The seat: one client business, not one login. Invite whoever needs access, with no per-user pricing.
  • A bad month: it stays in the record, and the gap queue is where the reason lives.

Hold the same nine against anyone else you are looking at. That is the point of writing them down.

Questions people ask

Do I need to understand how AI models work to evaluate one of these tools?

No, and be wary of a sales process that suggests you do. All nine questions above are about what the tool produces: what it counted, what it named, what it hands your team, and who it can talk to. Nobody outside the model vendors can see how selection is weighted anyway, so a vendor explaining the mechanism in detail is describing a hypothesis.

What should I make of a vendor's case studies?

Read them for the shape of the reporting rather than the size of the result. A case study that shows counts, a window and a per-engine split is evidence about how that vendor reports. One that shows a percentage lift with no denominator tells you what their client reports look like, which is worth knowing too.

How long should a trial run before I decide?

Long enough that one unusual morning cannot decide it. Assistants are non-deterministic, and a single run can look unusually kind or unusually brutal for reasons unconnected to the business. What a short trial can settle is whether the outputs answer the nine questions, which is most of what you are buying.

Is one of these tools worth it for an agency with three clients?

That depends entirely on what a seat covers and how the vendor prices the fourth client, which is why those questions are on the list. The other half of the answer is whether the tool produces something you can put in front of a client. If it only produces something for you, it is a cost. If it produces a report the client keeps, it is a line item you can charge for.

Read next

Article

What is AI visibility, and how do you measure it?

AI visibility is how often AI assistants name your business when someone asks a question you should win. You measure it by asking the same questions on a fixed schedule and counting how many answers name you.

Article

White label AI visibility software

Check what carries your brand and what does not. Dashboards are the easy part. The tell is whether reports, alerts and client-facing pages do too, whether it costs extra, and whether anything reaches a client without your approval.

Discover where you actually stand.

A free scan puts real customer questions to every engine we measure and comes back with your AI Visibility Score and its counts, the questions we asked, and the businesses being discovered instead of you.

Run a free scan
Run a free AI visibility scanFree trial