Methodology
What we measure, and what we refuse to claim.
AgentSignal simulates prospective clients asking AI assistants which law firm to contact, and measures what those assistants answered. This page says exactly how: what is asked, which models answer, how an answer is read, and what each number is divided by. Everything below describes the code as it runs today.
What AgentSignal tests
One scan asks one question, many times over: when a prospective client describes their situation to an AI assistant and asks who to hire, does the assistant name your firm, and does it recommend you first? To answer it, AgentSignal puts a fixed set of realistic client situations for your practice area to AI models from OpenAI, Anthropic (Claude), Google (Gemini) and Perplexity, through each company’s API with web search available, records every answer with the searches and sources behind it, and counts the outcomes.
The calls are real and happen when the scan runs; the clients are simulated. No person typed these questions, and the answers are not what any one real client saw — they are what the models answered to the same questions a client would ask.
Coverage depends on the scan. A free scan asks one AI model up to 12 client situations and reports the headline: the three measures below, the firms recommended first instead, and every recorded answer. The comprehensive analysis asks every AI model this deployment has configured the 30-situation benchmark, adds the diagnosis and an ordered plan, and can be re-tested. Every report names the models that actually answered.
The benchmark: client situations, not search terms
A benchmark is a versioned population of client situations — who was hurt, how serious it is, what the insurer has already said, how urgently help is needed, who is asking — combined with a decision style (urgency-driven, comparison-heavy, wants direct attorney access) and circumstances (market, timing, language). Each situation renders into a natural request from fixed templates, the way a person actually describes a crash, never a keyword string and never text written by a language model.
The population is content-addressed: the same situations and sampling policy always produce the same benchmark version id. Two scans on the same benchmark version and the same scenario set asked the same questions, which is what makes a before/after comparison a measurement instead of a coincidence. The situations are weighted equally; they are a test set, not a model of search demand.
Where situations come from
Every situation traces to a registered corpus source with explicit licensing and provenance. Today production benchmarks use expert-curated situations authored in-house. Sources marked research_only — public forums, community platforms — inform our understanding of how people describe legal problems but can never enter production benchmark generation; the exclusion is enforced at build time and tested. We do not scrape user content, retain usernames, or reproduce anyone’s personal story.
A license that forbids what we do — no commercial use, no derivatives, no retention, or a share-alike clause that would propagate onto the corpus — blocks a source outright. A license that obliges something — credit the author, strip personal data — must arrive with the means to satisfy it, or the pack fails to start. Neither is a policy anyone has to remember.
How a question reaches the assistant
In the decisions that produce every headline number, the request contains nothing about the firm being tested — not its name, aliases, domain, phone, or address — and no list of competitors. A leakage guard checks every request before it is sent: if the firm’s identity would reach a request, the whole scan stops, because the machinery that built it cannot be trusted; if a phrase from the situation merely happens to appear on the firm’s own site, that one request is skipped with a recorded reason. The assistant builds the competitive set itself.
Each request carries one short instruction of ours alongside the client’s words: research the actual market, name only real firms, weigh the alternatives, end by recommending exactly one firm (or say there is none), explain why, and mention the sources relied on. A consumer assistant may hedge or list several firms where this instruction asks it to commit; that is the one way these answers are deliberately not what a person would see.
Named firms are then resolved to canonical entities so “Morgan & Morgan” and “Morgan and Morgan, PA” count as one firm, with the resolution method and confidence recorded. A name is counted as a competing firm only when it is confidently classified as a firm and backed by evidence.
Which AI models answer — stated honestly
We call models from OpenAI, Anthropic (Claude), Google (Gemini) and Perplexity through their APIs, with web search available — on this deployment, OpenAI, Anthropic (Claude), Google (Gemini) and Perplexity models. Whether a model actually searches for a given answer is usually its own decision, and the searches it ran are recorded with the answer. These are not the consumer products: we never claim to test “ChatGPT” when we are calling OpenAI’s API, and a consumer app can add its own instructions, memory and personalization that an API call does not have. Every report names the exact models and the setting each ran with.
The three measures
Every report leads with three measures. Two of them count the same thing — the answers that recommended your firm first — over different bases, which is why each one names its own.
- AI Discovery Index
- How often AI names your firm in the completed test answers. Answers naming your firm ÷ completed test answers × 100.
- AI Selection Index
- Of the test answers that name your firm, how often AI recommends it first. First recommendations ÷ test answers naming your firm × 100.
- Recommended first overall
- How often AI recommends your firm first across all completed test answers, not only the answers that name it. First recommendations ÷ completed test answers × 100.
Discovery answers does AI find you. Selection answers once it finds you, does it recommend you. A firm can be named often and recommended rarely, or the reverse, and the two call for different work — which is why neither is folded into a single score.
When an answer names several firms
Most answers name several firms. Each firm an answer names reaches one of four stages, strictly nested — a firm that reached one stage reached every stage before it:
- Surfaced.
- The answer named the business at all.
- Shortlisted.
- The answer weighed the business as one of the options.
- Recommended.
- The answer put the business in front of the client as something to do.
- Recommended first.
- The answer recommended the business first — its one named pick.
An answer has at most one firm recommended first. Fixed rules find the sentence in which the assistant commits to its pick and the firm it names there; where the reading pass is enabled, a separate reading by our own model — never the model under test — records the pick too, and the two readings are merged. If an answer commits to two different firms, neither is counted as its pick.
Every completed answer ends in exactly one outcome: your firm was recommended first, another firm was recommended first, or no single firm was recommended (the assistant declined, found nothing, or stopped at a list). The last is counted on its own and never as a loss to a rival.
What counts as a completed answer
An answer is completed when the model returned one. A call that fails after its retries, times out, or is refused by the provider is recorded as failed and excluded from every rate rather than counted as a loss; a written refusal or an answer with no pick is a completed answer with no recommendation. Every report prints three counts — what was planned, what completed, and what failed — and a scan in which a quarter or more of the answers failed is labelled partial.
Recommended first, with its denominator
The headline counts the answers in which the assistant recommended the tested firm first, divided by the completed answers — “recommended first in N of M” — and the owner report rounds it to Recommended first overall. Appearing in a shortlist is measured separately and never counted as being recommended first. Competitor shares use a narrower denominator — the answers that produced a single winner — and the report names it rather than implying it.
A rate over an empty denominator renders as not measured, never as zero.
How sources and citations are recorded
With each answer we store the source links the model returned: the pages it cited, and for some models the search results it was shown as well. The report lists them answer by answer and totals them by website, so you can see which sites the answers drew on and whether your own site was among them. A citation shows where an answer looked; it does not show why a firm was recommended, and we never present it as the reason.
Repetition and uncertainty
A single language-model run is not a precise measurement. A report’s measurement notes and its PDF carry a Wilson 95% interval on the headline rate wherever the sample supports one, and when the sample is small enough that one different answer moves the rate by several points, the report says so in plain language instead of printing fake decimals.
Every scan today runs each situation once per model. The engine can repeat a situation and report how often the answers agree — the measurement exists and is version-stamped — but no scan offers it, so nothing advertises it. Precision comes from more situations, not more repeats of the same one: repeats of one situation are correlated and buy far less.
Evidence traceability
Each decision stores the exact request, the raw response, the search queries and citations the provider exposed, token usage, cost and latency — immutably. The headline measures open into the answers that produced them. Reinterpretation (a better parser, a new metric version) writes new rows beside the old ones; history is never overwritten.
Loss reasons and recommendations are produced by fixed, versioned rules over this evidence, and the plan is ordered by a printed rule — sample size × stage of loss × fixability — not by a model’s opinion of what matters or an arbitrary score. No recommendation ever asks a firm to assert experience, outcomes, credentials or availability that its own evidence does not support; where a claim could not be corroborated, the brief says exactly that.
Business facts vs. business claims
We crawl the firm’s own site — and only the firm’s own site, never a competitor’s — to build an evidence-backed profile, and we record the difference between a checkable fact and the firm’s own claim. “$500 million recovered” on your homepage is stored as your claim, shown as your claim, and never laundered into verified truth. The same crawl runs the technical checks: whether search and AI crawlers may read the site, whether its pages carry a canonical, a title and a description, whether its content needs JavaScript to appear. Whether a page is in a search index is never checked, and the report says so.
Why results change between scans
Two scans of the same firm can differ for reasons that have nothing to do with the firm:
- Providers update their models and their search behaviour without announcement, and a model can answer the same question differently from one run to the next.
- The web the assistants search changes: new pages, new reviews, new directory listings.
- A scan run with more or fewer AI models, or with more failed answers, is a different measurement, and the report says which facets differed.
- When a practice area’s question set is improved, its benchmark version changes, and scans on different versions are never compared as though they asked the same questions.
That is why the benchmark is versioned, why every count carries its denominator, and why the product re-tests rather than extrapolating from one scan.
What a re-test actually re-runs
A re-test runs the whole benchmark again on the same benchmark version and scenario set — the same situations, asked again. There is no targeted re-run of a subset: a smaller population is a different measurement, so a “quick re-check of the six situations you fixed” would move with the sample rather than with the firm. Each recommendation names the situations it was diagnosed from so you know which answers to read afterwards, not because those answers run alone.
Re-tests are part of the comprehensive analysis; a free scan is a one-off look and is not re-tested. If the practice area’s question set has been updated since the first scan, the report says the two scans are not comparable instead of comparing them. What a re-test reports is an observed change in the same questions’ answers — never proof that a website edit caused it. Scans run when you start them; nothing is scheduled and no alerts are sent.
What we never claim
- No guarantee. AgentSignal measures and explains; it does not promise that any change will get a firm recommended, and no report predicts a score or a gain.
- A simulated decision is not a signed client. We report what the models answered, never revenue, lead counts or conversion.
- A before/after change is an observed benchmark change. We do not claim your website edit caused it.
- There is no “rank #3 on ChatGPT”. No such ranking exists to read, and we do not invent proprietary 0–100 scores.
- AgentSignal gives no legal advice and never states that one firm is objectively better than another.
What we do not recommend, and why
Each of these is sold as a way to be recommended by AI. We grade a technique by the evidence for it, never by how often it is repeated, and none of them appears in any plan we produce. The reasoning is the registry’s own sentence and every source is named with its publisher and the date a person read it, so a firm can check us — and answer whoever is selling it. Worst grade first.
- Adding many question-and-answer blocks to every page for AI
We will not recommend this.
- Adding schema to influence ranking or AI selection
We will not recommend this.
- Buying directory placements or sponsored articles for their ranking signals
We will not recommend this.
- Fabricated awards, press mentions, expert quotes or credentials
We will not recommend this.
- FAQPage markup on a law firm site
We will not recommend this.
- Generating a '[service] lawyer in [city]' page per city or neighbourhood
We will not recommend this.
- Generating headings that mirror exact query phrasings
We will not recommend this.
- Repeating the target phrase or the brand name for LLMs
We will not recommend this.
- Review or AggregateRating markup about the firm on the firm's own site
We will not recommend this.
- Seeding Reddit, forums or review sites with synthetic mentions
We will not recommend this.
- 'AI schema' or GEO-specific structured data
We will not recommend this.
- A separate sitemap for AI crawlers
We will not recommend this.
- Allowlisting AI crawlers to gain visibility
We will not recommend this.
- Publishing a markdown version of every page for AI
We will not recommend this.
- Publishing content designed to be quoted rather than to be useful
We will not recommend this.
- Adding statistics, quotations and citations to raise generative-engine visibility
Only ever as a labelled experiment, never as a fix.
- Publishing /llms.txt
Only ever as a labelled experiment, never as a fix.
- Writing every paragraph as a standalone answer for chunked retrieval
Only ever as a labelled experiment, never as a fix.
Known limitations
- Situation weights are uniform today: the benchmark measures a test distribution, not validated market demand, and no weighted figure is printed.
- The crawler reads server-rendered HTML; a fully client-rendered site looks thin to it — itself a finding about how crawlers that do not run JavaScript see the site, but a partial one.
- API answers are not consumer-app answers, and our instruction asks for one pick where an app might list several.
- Provider behavior changes without announcement. That is why the benchmark is versioned and why the product re-tests rather than extrapolating.
Run the measurement on your firm
The free scan asks one AI model your practice area’s client situations and shows whether it names your firm, whether it recommends you first, who it recommends instead, and the recorded answers. No credit card; sign in with email or Google to run it.