Skip to content

Methodology

What we measure, and what we refuse to claim.

AgentSignal simulates prospective clients asking AI assistants which law firm to contact, and measures what those assistants answered. This page says exactly how: what is asked, which models answer, how an answer is read, and what each number is divided by. Everything below describes the code as it runs today.

What AgentSignal tests

One scan asks one question, many times over: when a prospective client describes their situation to an AI assistant and asks who to hire, does the assistant name your firm, and does it recommend you first? To answer it, AgentSignal puts a fixed set of realistic client situations for your practice area to AI models from OpenAI, Anthropic (Claude), Google (Gemini) and Perplexity, through each company’s API with web search available, records every answer with the searches and sources behind it, and counts the outcomes.

The calls are real and happen when the scan runs; the clients are simulated. No person typed these questions, and the answers are not what any one real client saw — they are what the models answered to the same questions a client would ask.

Coverage depends on the scan. A free scan asks one AI model up to 12 client situations and reports the headline: the three measures below, the firms recommended first instead, and every recorded answer. The comprehensive analysis asks every AI model this deployment has configured the 30-situation benchmark, adds the diagnosis and an ordered plan, and can be re-tested. Every report names the models that actually answered.

The benchmark: client situations, not search terms

A benchmark is a versioned population of client situations — who was hurt, how serious it is, what the insurer has already said, how urgently help is needed, who is asking — combined with a decision style (urgency-driven, comparison-heavy, wants direct attorney access) and circumstances (market, timing, language). Each situation renders into a natural request from fixed templates, the way a person actually describes a crash, never a keyword string and never text written by a language model.

The population is content-addressed: the same situations and sampling policy always produce the same benchmark version id. Two scans on the same benchmark version and the same scenario set asked the same questions, which is what makes a before/after comparison a measurement instead of a coincidence. The situations are weighted equally; they are a test set, not a model of search demand.

Where situations come from

Every situation traces to a registered corpus source with explicit licensing and provenance. Today production benchmarks use expert-curated situations authored in-house. Sources marked research_only — public forums, community platforms — inform our understanding of how people describe legal problems but can never enter production benchmark generation; the exclusion is enforced at build time and tested. We do not scrape user content, retain usernames, or reproduce anyone’s personal story.

A license that forbids what we do — no commercial use, no derivatives, no retention, or a share-alike clause that would propagate onto the corpus — blocks a source outright. A license that obliges something — credit the author, strip personal data — must arrive with the means to satisfy it, or the pack fails to start. Neither is a policy anyone has to remember.

How a question reaches the assistant

In the decisions that produce every headline number, the request contains nothing about the firm being tested — not its name, aliases, domain, phone, or address — and no list of competitors. A leakage guard checks every request before it is sent: if the firm’s identity would reach a request, the whole scan stops, because the machinery that built it cannot be trusted; if a phrase from the situation merely happens to appear on the firm’s own site, that one request is skipped with a recorded reason. The assistant builds the competitive set itself.

Each request carries one short instruction of ours alongside the client’s words: research the actual market, name only real firms, weigh the alternatives, end by recommending exactly one firm (or say there is none), explain why, and mention the sources relied on. A consumer assistant may hedge or list several firms where this instruction asks it to commit; that is the one way these answers are deliberately not what a person would see.

Named firms are then resolved to canonical entities so “Morgan & Morgan” and “Morgan and Morgan, PA” count as one firm, with the resolution method and confidence recorded. A name is counted as a competing firm only when it is confidently classified as a firm and backed by evidence.

Which AI models answer — stated honestly

We call models from OpenAI, Anthropic (Claude), Google (Gemini) and Perplexity through their APIs, with web search available — on this deployment, OpenAI, Anthropic (Claude), Google (Gemini) and Perplexity models. Whether a model actually searches for a given answer is usually its own decision, and the searches it ran are recorded with the answer. These are not the consumer products: we never claim to test “ChatGPT” when we are calling OpenAI’s API, and a consumer app can add its own instructions, memory and personalization that an API call does not have. Every report names the exact models and the setting each ran with.

The three measures

Every report leads with three measures. Two of them count the same thing — the answers that recommended your firm first — over different bases, which is why each one names its own.

AI Discovery Index
How often AI names your firm in the completed test answers. Answers naming your firm ÷ completed test answers × 100.
AI Selection Index
Of the test answers that name your firm, how often AI recommends it first. First recommendations ÷ test answers naming your firm × 100.
Recommended first overall
How often AI recommends your firm first across all completed test answers, not only the answers that name it. First recommendations ÷ completed test answers × 100.

Discovery answers does AI find you. Selection answers once it finds you, does it recommend you. A firm can be named often and recommended rarely, or the reverse, and the two call for different work — which is why neither is folded into a single score.

When an answer names several firms

Most answers name several firms. Each firm an answer names reaches one of four stages, strictly nested — a firm that reached one stage reached every stage before it:

Surfaced.
The answer named the business at all.
Shortlisted.
The answer weighed the business as one of the options.
Recommended.
The answer put the business in front of the client as something to do.
Recommended first.
The answer recommended the business first — its one named pick.

An answer has at most one firm recommended first. Fixed rules find the sentence in which the assistant commits to its pick and the firm it names there; where the reading pass is enabled, a separate reading by our own model — never the model under test — records the pick too, and the two readings are merged. If an answer commits to two different firms, neither is counted as its pick.

Every completed answer ends in exactly one outcome: your firm was recommended first, another firm was recommended first, or no single firm was recommended (the assistant declined, found nothing, or stopped at a list). The last is counted on its own and never as a loss to a rival.

What counts as a completed answer

An answer is completed when the model returned one. A call that fails after its retries, times out, or is refused by the provider is recorded as failed and excluded from every rate rather than counted as a loss; a written refusal or an answer with no pick is a completed answer with no recommendation. Every report prints three counts — what was planned, what completed, and what failed — and a scan in which a quarter or more of the answers failed is labelled partial.

Recommended first, with its denominator

The headline counts the answers in which the assistant recommended the tested firm first, divided by the completed answers — “recommended first in N of M” — and the owner report rounds it to Recommended first overall. Appearing in a shortlist is measured separately and never counted as being recommended first. Competitor shares use a narrower denominator — the answers that produced a single winner — and the report names it rather than implying it.

A rate over an empty denominator renders as not measured, never as zero.

How sources and citations are recorded

With each answer we store the source links the model returned: the pages it cited, and for some models the search results it was shown as well. The report lists them answer by answer and totals them by website, so you can see which sites the answers drew on and whether your own site was among them. A citation shows where an answer looked; it does not show why a firm was recommended, and we never present it as the reason.

Repetition and uncertainty

A single language-model run is not a precise measurement. A report’s measurement notes and its PDF carry a Wilson 95% interval on the headline rate wherever the sample supports one, and when the sample is small enough that one different answer moves the rate by several points, the report says so in plain language instead of printing fake decimals.

Every scan today runs each situation once per model. The engine can repeat a situation and report how often the answers agree — the measurement exists and is version-stamped — but no scan offers it, so nothing advertises it. Precision comes from more situations, not more repeats of the same one: repeats of one situation are correlated and buy far less.

Evidence traceability

Each decision stores the exact request, the raw response, the search queries and citations the provider exposed, token usage, cost and latency — immutably. The headline measures open into the answers that produced them. Reinterpretation (a better parser, a new metric version) writes new rows beside the old ones; history is never overwritten.

Loss reasons and recommendations are produced by fixed, versioned rules over this evidence, and the plan is ordered by a printed rule — sample size × stage of loss × fixability — not by a model’s opinion of what matters or an arbitrary score. No recommendation ever asks a firm to assert experience, outcomes, credentials or availability that its own evidence does not support; where a claim could not be corroborated, the brief says exactly that.

Business facts vs. business claims

We crawl the firm’s own site — and only the firm’s own site, never a competitor’s — to build an evidence-backed profile, and we record the difference between a checkable fact and the firm’s own claim. “$500 million recovered” on your homepage is stored as your claim, shown as your claim, and never laundered into verified truth. The same crawl runs the technical checks: whether search and AI crawlers may read the site, whether its pages carry a canonical, a title and a description, whether its content needs JavaScript to appear. Whether a page is in a search index is never checked, and the report says so.

Why results change between scans

Two scans of the same firm can differ for reasons that have nothing to do with the firm:

  • Providers update their models and their search behaviour without announcement, and a model can answer the same question differently from one run to the next.
  • The web the assistants search changes: new pages, new reviews, new directory listings.
  • A scan run with more or fewer AI models, or with more failed answers, is a different measurement, and the report says which facets differed.
  • When a practice area’s question set is improved, its benchmark version changes, and scans on different versions are never compared as though they asked the same questions.

That is why the benchmark is versioned, why every count carries its denominator, and why the product re-tests rather than extrapolating from one scan.

What a re-test actually re-runs

A re-test runs the whole benchmark again on the same benchmark version and scenario set — the same situations, asked again. There is no targeted re-run of a subset: a smaller population is a different measurement, so a “quick re-check of the six situations you fixed” would move with the sample rather than with the firm. Each recommendation names the situations it was diagnosed from so you know which answers to read afterwards, not because those answers run alone.

Re-tests are part of the comprehensive analysis; a free scan is a one-off look and is not re-tested. If the practice area’s question set has been updated since the first scan, the report says the two scans are not comparable instead of comparing them. What a re-test reports is an observed change in the same questions’ answers — never proof that a website edit caused it. Scans run when you start them; nothing is scheduled and no alerts are sent.

What we never claim

  • No guarantee. AgentSignal measures and explains; it does not promise that any change will get a firm recommended, and no report predicts a score or a gain.
  • A simulated decision is not a signed client. We report what the models answered, never revenue, lead counts or conversion.
  • A before/after change is an observed benchmark change. We do not claim your website edit caused it.
  • There is no “rank #3 on ChatGPT”. No such ranking exists to read, and we do not invent proprietary 0–100 scores.
  • AgentSignal gives no legal advice and never states that one firm is objectively better than another.

What we do not recommend, and why

Each of these is sold as a way to be recommended by AI. We grade a technique by the evidence for it, never by how often it is repeated, and none of them appears in any plan we produce. The reasoning is the registry’s own sentence and every source is named with its publisher and the date a person read it, so a firm can check us — and answer whoever is selling it. Worst grade first.

  • Adding many question-and-answer blocks to every page for AIContradicted or risky

    We will not recommend this.

    Scaled, low-value content is a named spam policy; FAQ rich results are unavailable to law firms. Answering the questions clients actually asked, once, in the page's own voice is covered by the original-information technique.

  • Adding schema to influence ranking or AI selectionContradicted or risky

    We will not recommend this.

    Google's structured data documentation does not claim a ranking effect and prohibits marking up content that is not visible; 'AI schema' is not a documented concept anywhere.

  • Buying directory placements or sponsored articles for their ranking signalsContradicted or risky

    We will not recommend this.

    Site reputation abuse and link spam are named policies; Rule 7.2 limits what a lawyer may give for a recommendation.

  • Fabricated awards, press mentions, expert quotes or credentialsContradicted or risky

    We will not recommend this.

    False or misleading communication about a lawyer's services violates Rule 7.1 in every state that adopts it; fake testimonials violate the FTC rule.

  • FAQPage markup on a law firm siteContradicted or risky

    We will not recommend this.

    Google restricts FAQ rich results to well-known government and health sites; on a law firm site the markup produces nothing and pads the page.

  • Generating a '[service] lawyer in [city]' page per city or neighbourhoodContradicted or risky

    We will not recommend this.

    Doorway pages, including region and city variations funnelling to the same destination, are a named Google spam policy, as is scaled content; Bing demotes scraped and keyword-stuffed content.

  • Generating headings that mirror exact query phrasingsContradicted or risky

    We will not recommend this.

    Relevant words in a heading are a documented fundamental, and a heading that states the question a client asked needs no technique. What this names beyond that — generating headings to mirror query phrasings — is keyword stuffing and scaled content abuse, both named in Google's spam policies. No vendor documents a benefit to the generated version.

  • Repeating the target phrase or the brand name for LLMsContradicted or risky

    We will not recommend this.

    Keyword stuffing is a named spam policy at Google and a demotion cause at Bing; no vendor documents any density signal for assistants.

  • Review or AggregateRating markup about the firm on the firm's own siteContradicted or risky

    We will not recommend this.

    Google states LocalBusiness and Organization pages are ineligible for the star feature when the entity controls the reviews about itself, including embedded widgets; misleading markup can draw a manual action.

  • Seeding Reddit, forums or review sites with synthetic mentionsContradicted or risky

    We will not recommend this.

    The FTC rule bans fake reviews and testimonials and undisclosed insider posts; Google's user-content policy bans fake engagement. Correlational studies of community citations say nothing about planted ones.

  • 'AI schema' or GEO-specific structured dataUnsupported

    We will not recommend this.

    No platform documents any AI-specific structured data; schema.org defines no such type.

  • A separate sitemap for AI crawlersUnsupported

    We will not recommend this.

    No vendor documents an AI-specific sitemap; ordinary sitemaps aid discovery without guaranteeing indexing.

  • Allowlisting AI crawlers to gain visibilityUnsupported

    We will not recommend this.

    Beyond removing an exclusion, no vendor documents an effect; Google-Extended in particular does not affect Search at all.

  • Publishing a markdown version of every page for AIUnsupported

    We will not recommend this.

    No platform documents consuming such files; Google's guidance asks for text in the ordinary HTML.

  • Publishing content designed to be quoted rather than to be usefulUnsupported

    We will not recommend this.

    Nothing documents it, and it is the definition of search-engine-first content.

  • Adding statistics, quotations and citations to raise generative-engine visibilityExperimental

    Only ever as a labelled experiment, never as a fix.

    The KDD 2024 GEO paper reports gains on a constructed engine over a benchmark, not on production assistants; where a statistic is real, first-party and useful it belongs to original information, and where it is added for the model it is padding.

  • Publishing /llms.txtExperimental

    Only ever as a labelled experiment, never as a fix.

    Google states no AI system currently uses it and that consumer assistants do not request it; Google's AI features page requires no machine-readable file; llmstxt.org cites adoption by documentation tooling and the AI labs' own docs, not by any search or assistant retrieval system. Harmless, cheap, and not a visibility lever on any documented surface.

  • Writing every paragraph as a standalone answer for chunked retrievalExperimental

    Only ever as a labelled experiment, never as a fix.

    Retrieval systems do chunk text, but no vendor documents a chunking format, and writing for the chunker rather than the reader is the search-engine-first pattern Google names.

Known limitations

  • Situation weights are uniform today: the benchmark measures a test distribution, not validated market demand, and no weighted figure is printed.
  • The crawler reads server-rendered HTML; a fully client-rendered site looks thin to it — itself a finding about how crawlers that do not run JavaScript see the site, but a partial one.
  • API answers are not consumer-app answers, and our instruction asks for one pick where an app might list several.
  • Provider behavior changes without announcement. That is why the benchmark is versioned and why the product re-tests rather than extrapolating.

Run the measurement on your firm

The free scan asks one AI model your practice area’s client situations and shows whether it names your firm, whether it recommends you first, who it recommends instead, and the recorded answers. No credit card; sign in with email or Google to run it.