Guides / Methodology

How BaaDigi Collects and Checks Its Data

BaaDigi's research methodology rests on data we collect ourselves: Google Business Profile reviews and replies for the profiles we manage, repeated runs asking ChatGPT, Perplexity and Google AI Overviews which businesses they recommend, and scans of live business websites. Every published figure states its sample size and the date it was pulled, third-party benchmarks are labeled and linked, and nothing is estimated to fill a gap.

Last updated: September 2026 · Ryan Goering, CEO

What rules does every BaaDigi number follow?

  1. No invented numbers. If we did not measure it and cannot link a source for it, it is not published.
  2. Every published figure carries its sample size and the date it was pulled or checked.
  3. Vendor facts such as software pricing come from the vendor’s own website, with the date we checked it. When a vendor does not publish pricing, we say so.
  4. Small samples are stated as small. We do not split data into segments with too few businesses to mean anything.
  5. Corrections are made on the page and the page date updated. We have retracted a published claim when a re-query did not support it.

How is the google review response study collected?

Every review on the Google Business Profiles BaaDigi manages, pulled through the Google Business Profile API under our two agency logins, along with the owner reply and its timestamp. The current sample is 7,167 reviews across 78 profiles, queried 2026-09-03. Next scheduled re-pull: 2026-12-03.

What it cannot tell you

  • It is one agency’s client base, weighted toward US home service businesses. It is not a random sample of all businesses.
  • We count profiles, not businesses. Some profiles are not yet matched to a client record, so a business count would be a guess.
  • Google reports some metrics as “fewer than 15” instead of a count. Those rows are excluded from any total we publish.

Used on: Which Google reviews do businesses ignore?

How is the ai recommendation boards collected?

For a trade and city, we ask the same customer-style questions of ChatGPT (with web search forced on), Perplexity and Google AI Overviews, then extract the local businesses each engine names and which engine named them. A board is published only if the city is a recognisable "City, ST", at least 10 businesses were named, and all three engines ran. Boards refresh on a weekly schedule, oldest first, once a market is more than 21 days old; a run that returns fewer than 5 businesses is rejected as an engine failure instead of being published. Every run is kept in an append-only history, so changes over time can be compared.

What it cannot tell you

  • AI answers vary between runs. We have seen different names returned minutes apart for the same question. A board is a dated snapshot, not a permanent ranking.
  • Business names are extracted from the engines’ prose by a language model, and website matches come from the domains each engine cited. Matches can be wrong; we correct them when found.
  • Google often shows no AI Overview for local questions, so it contributes far fewer names than ChatGPT or Perplexity.

Used on: Who AI recommends, by trade and city

How is the agency visibility runs collected?

For each trade roundup, 10 buyer-style prompts ("who are the best roofing marketing agencies?") are run against Perplexity, ChatGPT (web search) and Gemini with Google Search, and we count how many of the 30 answers name each agency. Agency facts in the tables (pricing, contract terms, location) are taken from each agency’s own website, not from the AI answers.

What it cannot tell you

  • Thirty answers is a small sample per trade, and phrasing changes which agencies appear.
  • BaaDigi appears in these tables and sells a competing service. We show our own count even when it is zero, and never list ourselves first.

Used on: Marketing agency roundups by trade

How is the website agent-readiness scans collected?

Google Lighthouse’s agentic browsing checks, run against live business websites on a single day and cross-referenced with whether AI engines named each business for its trade and city.

What it cannot tell you

  • Scans are a point in time. A site can change the next day.
  • Some checks cannot be tested remotely, so the scan covers the checks that can.

Used on: We scanned 97 business websites

How is the website platform study collected?

Starting from the businesses named on our published AI recommendation boards, we fetched each business’s homepage once and identified its website platform from signatures in the raw HTML, such as WordPress content paths or a website builder’s asset host. Franchises, directories and non-contractor businesses are set aside, sites whose title no longer matches the business are dropped, and sites BaaDigi built are excluded from headline figures. A random sample of classifications is checked by hand against the saved pages, and every share is published with a 95% confidence interval.

What it cannot tell you

  • Pages are read without running scripts, so a builder that only shows itself after JavaScript loads is counted as no known builder.
  • It describes businesses AI engines named. With no comparison group of businesses that were not named, it cannot show that any platform affects recommendations.

Used on: What website platform do AI-recommended contractors use?

How is the industry benchmarks from other publishers collected?

Cost-per-lead and conversion benchmarks by trade are a synthesis of figures published by other firms, with each source named and dated on the page. They are labeled as third-party data. Where we add our own client results, they are marked as individual cases, not averages.

What it cannot tell you

  • We did not collect these figures, and publishers measure differently. Treat ranges as directional.

Used on: Contractor growth benchmarks by trade

Common questions about our data

Where does BaaDigi’s data come from?
From three first-party sources: Google Business Profile data for the profiles BaaDigi manages, repeated question runs against ChatGPT, Perplexity and Google AI Overviews, and scans of live business websites. Industry benchmarks from other publishers are used too, but always labeled as third-party and linked to the original source.
Is BaaDigi’s research a random sample?
No. The review and profile data comes from one agency’s client base, weighted toward US home service businesses, and the AI runs cover the trades and cities we chose to track. Each study states its sample so readers can judge how far it generalises.
Why do AI recommendation results change between checks?
AI engines generate each answer fresh and pull from live search results, so the same question can return different businesses minutes apart. That is why BaaDigi publishes boards as dated snapshots, re-runs them on a schedule and keeps the full history rather than presenting one run as a permanent ranking.

Browse every guide and study on the guides hub.