Submit an agency
The method

How we measure

One ruler, applied to every agency, including our own. Most of it is measured — axes any outside auditor can re-run from the prompts and dates we publish. The rest is editorial — our judgment, labelled as judgment, never dressed up as a measurement. Every listed agency has a right of reply.

The gate — who enters the segment

An agency is listed only if all four hold. The gate cuts by tier, not by strength: a strong player that genuinely serves small business ranks where it earns, above us when it is better.

  • SMB-primary — serves private clients / small business as its primary segment, not enterprise-only with SMB as a token line.
  • Self-serve offer — has a transparent self-serve offer a buyer can purchase themselves — a published price or fixed package, no mandatory 'contact sales'.
  • GEO substance — names GEO / AI-visibility as an actual service (names answer engines, citation work, structured data, measures AI answers) — relabeled generic SEO is listed but flagged.
  • Live business — is a live, reachable business — real site, real contact.

Why these axes

Every axis here clears three filters — with one disclosed exception, named below.

  1. A buyer would agree it matters. Not knowing who scores well on it, would a small-business buyer weigh this when choosing an agency? An axis only we pass is off.
  2. It does not just reward age. An axis that mostly tracks how long a domain has existed rebuilds the ordinary directory, where the oldest name wins — so age-driven axes are off the ruler, with one honest exception: own AI-visibility (M1). It is partly age-driven, and we keep it only because an agency that sells AI-visibility yet has none itself is a real signal — and we weight it the lightest (9%) precisely because of that confound.
  3. Anyone can check it. A measured axis can be reproduced by an outsider; an editorial one is labelled as our opinion.

The seven axes

Every axis below clears those three filters. Here is what each rewards, how much it weighs, and why.

AxisWhat it rewardsWeight
M1 · Own AI-visibility the agency is itself findable, and described correctly, in AI answers 9%
M2 · Method transparency publishes how it works and what it measures, not a black box 16%
M3 · Evidence verifiability named, specific, checkable proof over anonymous hype 23%
M4 · Pricing openness published price / range, not 'request a quote' 12%
M5 · AI legibility the site is crawlable and machine-legible — not blocked to AI, server-rendered, with Schema.org markup, a sitemap and an llms.txt 10%
E1 · Segment fit genuinely built for the private / small-business buyer 12%
E2 · Promise cleanliness no snake-oil — no guaranteed rankings, no fakery 18%

M1 · Own AI-visibility · 9% — Does the agency itself surface, and get described correctly, in AI answers? A shop that sells AI-visibility and is invisible in it is a real signal. We weight it lightest because raw visibility is driven largely by domain age, not by how good the agency is — an old average shop outranks a new excellent one on age alone. As a "who should I hire" signal it is weak, so it stays light.

M2 · Method transparency · 16% — Does the agency publish how it works and, above all, what it measures? Measurement is the load-bearing part: a real specialist can name how it tracks visibility in AI answers — share of citations, prompt-set testing, mention tracking. "We grow your traffic and rankings" is an SEO answer to a GEO question, and scores low here.

M3 · Evidence verifiability · 23% — Is the proof named, specific and linkable — a real client, a bounded claim, a case you can open — rather than anonymous "5× traffic" hype? We score whether the evidence is checkable, not whether a private result is true; that result sits behind the client's wall, and pretending to confirm it would be the exact dishonesty we refuse. It is the heaviest axis because it is the most buyer-predictive and the hardest to fake, and it protects the honest newcomer with two real cases over the incumbent with fifty anonymous boasts.

M4 · Pricing openness · 12% — Is there a published price or range, or only "request a quote"? Small-business buyers need cost up front, which is why this is on the ruler for this segment. It sits in the middle: a real price is necessary information, but a price can still front a weak service, so it is a fit-and-hygiene axis, not a quality one.

M5 · AI legibility · 10% — Is the agency's own site actually readable by machines — open to AI crawlers, server-rendered, with Schema markup, a sitemap and an llms.txt? This is the age-independent half of M1: being cited (M1) depends partly on how old you are, but making your own site machine-legible is fully in your control from day one. It is entirely reproducible — an auditor can fetch the site and check — and a shop opaque to AI on its own property is failing its own craft.

E1 · Segment fit · 12% — Is the agency genuinely built for a private or small-business buyer — self-serve, proportionate scope — or an enterprise operation with a token SMB line? This is a judgment call on the close cases, so we label it as opinion and send borderline entries to human review.

E2 · Promise cleanliness · 18% — Are the claims careful and honest, or snake-oil — guaranteed rankings, fake reviews, "we manipulate AI"? This is the single biggest way a buyer gets burned in this category, which is why it is the heaviest editorial axis.

How each score is anchored (1–5)

Each axis is scored 1–5 against the fixed anchors below, so an outside auditor working from the same evidence lands on the same number.

M1 · Own AI-visibility — 5 surfaces in the category answer on most engines + accurate brand recall · 4 surfaces on ≥1 engine, brand recall mostly right · 3 absent from category answers but brand recall roughly right · 2 brand recall thin/partial · 1 invisible.

M2 · Method transparency — 5 step-by-step method + a named AI-visibility metric + a sample deliverable · 4 real "how we work", measurement named but light · 3 process in marketing terms, only SEO metrics · 2 vague "proven strategies" · 1 black box.

M3 · Evidence verifiability — 5 named clients + bounded claims + linkable artifacts, a sample checks out · 4 named cases, limited linkable proof · 3 some specifics, mostly anonymous · 2 vague claims, no names · 1 unverifiable hype / a sampled claim fails.

M4 · Pricing openness — 5 full published price list · 4 "from $X" or clear tiers · 3 published ranges · 2 qualitative only · 1 "request a quote".

M5 · AI legibility — 5 SSR + open to AI crawlers + Schema + sitemap + llms.txt · 4 same minus llms.txt · 3 crawlable but thin markup · 2 largely client-rendered · 1 blocks AI crawlers / JS-only.

E1 · Segment fit — (editorial) 5 clearly SMB self-serve · 4 mostly SMB · 3 mixed · 2 mostly up-market · 1 enterprise, SMB tokenistic.

E2 · Promise cleanliness — (editorial; red-flag cap) 5 careful, explicit about no guarantees · 4 confident but reasonable · 3 mild puffery · 2 overclaiming near snake-oil · 1 RED FLAG (guaranteed rankings, fake reviews, "we manipulate AI") → caps the verdict.

How the weighting holds up

Measured axes (M1–M5) are 70% of the score; editorial judgment (E1–E2) is the other 30%. The reproducible core leads by design, so most of the ranking is something an outsider can re-run. Judgment is real — you cannot reduce "is this snake-oil" to a keyword scan — but it stays a minority, so an opinion can never swing a verdict on its own.

There is a simple test for whether a ruler is rigged: is an axis weighted heavy just because its author scores well on it? Run it on us. Our own AI-visibility (M1) is low today — young domain, delisted from Google in mid-2026 — and M1 is our lightest axis. The axes we are strong on are heavy, but each is heavy for a reason a buyer would give before ever seeing our score. The one axis we currently lose on, we weighted down on principle. That is checkable, and it is the point.

Weights are starting priors, published as such, and will move as we calibrate against real runs. An axis where everyone scores the same carries no information and loses weight automatically.

How precise the composite is. The composite is a weighted average of the seven axis scores. When an outside auditor re-scores the axes independently, the arithmetic reproduces exactly, every axis agrees within one point, and the composite lands within about ±0.3. So the last decimal is not a real distinction: we treat two entries within ~0.3 of each other as effectively tied, and the per-axis profile — not the composite number — is the signal. Read the axes.

The red-flag cap

A disqualifying claim — guaranteed rankings or placements, fake reviews on offer, 'we manipulate AI' — caps the overall verdict no matter how good the rest of the site looks. Guaranteeing an outcome you don't control cannot be averaged away by polish, and the buyer most needs to be steered away from the vendor confident enough to promise it.

The focus penalty — a specialist who does everything isn't one

This is a register for a specialty. Real GEO specialists are still few, and an agency that sells everything — or an SEO shop that added 'AI' to its menu — usually carries AI-visibility as one line among many, with no method and no cases of its own. Where GEO is not run as a distinct service (its own offer, its own page) and the site names no way it measures AI-answer visibility, we apply a focus penalty and flag it: −0.5 for a general-services combine, −0.2 for an SEO-led hybrid, none for a focused specialist. A thin, unverifiable site is already handled by the evidence axis (M3); the focus penalty carries only the separate fact that the specialty isn't really practised. Because being a generalist is a strong likelihood — not a certainty — of weak GEO, it is a penalty with a right of reply, not an automatic disqualification.

  • Generalist combine — SEO + website production + everything bundled: −0.5
  • SEO-led hybrid — SEO-first, but with real, named GEO work: −0.2
  • Focused GEO / AI-visibility specialist: no penalty

This is a focus modifier, not a change to the axis weights — the ruler above is unchanged. The classification is observable from the agency's own offer and applied uniformly; UserSignals is a focused specialist and carries no penalty, by the same rule.

Reproducible by design

Measured axes carry the evidence snippet and source they came from; a missing finding is recorded as an explicit 'not found', never a guessed number. Every score is stamped 'data as of' a date, and the probe set (ChatGPT + Claude, search-grounded, clean session) is fixed so any external auditor can re-run it. Editorial axes and the top of the list go through human review before publish.

Each entry is marked up as an editorial Review with a Rating (not AggregateRating) — view page source to verify.

Where this methodology comes from — and how it's checked

Provenance. This register's methodology comes from AI100, our research project on how AI answer-engines cite and rank sources. We reuse AI100's method and standards — not its scoring engine: how an axis has to earn its place (a neutral-buyer, an age-independence and a checkability test), why a confounded signal like a firm's own AI-visibility is weighted down rather than flattered, and the rule that every score ships with its source and date so anyone can re-run it. AI100's own scoring product does not touch these ratings — each entry's visibility is measured by the same cheap, identical probe, our own entry included.

Disclosure. AI100, this register and UserSignals are run by the same team. So AI100 is not an outside seal of approval — it is our own research standard, stated in the open. We would rather you check the method than trust a badge: the axes, weights, anchors and prompts are all published, and we list ourselves on the identical ruler, shown exactly where we stand, unflattering parts and all.

Verification. The register is audited against this methodology by the AI100 team, in a neutral session, and the audit is published — including what it found. An outside auditor re-derives the arithmetic exactly, agrees on every axis within one point, and reproduces the composite to within about ±0.3 — so entries within ~0.3 are treated as tied and the per-axis profile, not the last decimal, is the signal. Where the audit flagged gaps, we fixed them; two choices we keep on purpose, with reasons published: one age-sensitive axis (own AI-visibility), weighted lightest, and our batch rank is by our own composite, labelled as such.

Audit log.

  • 2026-07-09 — first external clean-session audit (AI100 team, ChatGPT). The composite recomputed exactly from the published scores; the audit flagged measured cells that did not show their source, missing published 1–5 anchors, and an unlabelled batch-crown. All fixed.
  • 2026-07-09 — re-audit after the fixes. With the anchors public, an independent re-scoring reproduced every axis within one point and the composite within ~±0.3 (so entries within ~0.3 are treated as tied). Two design choices are kept on purpose, with reasons published: one age-sensitive axis (own AI-visibility, weighted lightest), and a batch rank stated by our own composite.
What this ruler does NOT measure It does not score depth of track record or sheer volume of named clients — both favour whoever has been around longest, the age bias we set out to avoid. It cannot verify private results behind a client's wall, so it scores whether evidence is checkable, not whether an outcome is true. And it reads what agencies publish; where we corroborate with a direct enquiry, that is a spot-check on the top of the list and on our own entry, not a hidden score applied to everyone.