Douglas County, Oregon · published 17 August 2026

Our AI visibility score did not predict AI visibility

We built a 0 to 100 score for how readable a business website is to an AI answer engine, and we sold it. Then we tested whether a higher score meant an assistant was more likely to name that business. Across 34 assistant runs in five local trades, the rank correlation was 0.069 on a scale where 1 is a perfect relationship and 0 is none. Our six highest scoring sites were named zero times. A site scoring 45 was named in four runs out of six.

By Shane Hayes, founder, The Roseburg Plug LLC. The pilot design was fixed before the first query ran. The raw runs, the method and the anonymised business rows are published below under CC BY 4.0, so anyone can check this rather than take our word for it.

This pilot found no signal, and it was not powerful enough to prove there is none. Those are different claims and both matter.

The result

The Spearman rank correlation between a business score and how often an assistant named it was 0.069, with a 95 percent confidence interval from -0.256 to 0.360. That interval sits near zero and is wide enough to hold a modest effect in either direction. We treat it as a failure to validate, not as proof of no effect.

The top of our score range was the invisible end of it

Six businesses scoring 85, 84, 78, 77, 75 and 74 were named in zero of six runs each, 36 opportunities with nothing to show. A business scoring 45 was named in four of six and had its own domain cited. If the score discriminated even weakly, our best work should not be the quiet end of the distribution.

The gap between high and low scorers is smaller than our own noise floor

Businesses at or above the median score appeared in 11.4 percent of their trials. Businesses below it appeared in 8.9 percent. The gap is 2.5 percentage points. Our measurement engine ships a function whose only job is refusing to call a difference under 6 points a win. Applied to our own score, it refuses.

Appearance rate with 95 percent confidence intervals

Geometry is computed from the measurements, not drawn by eye. The horizontal axis runs 0 to 30 percentage points across 600 pixels, so every pixel is 0.05 of a point and each marker sits at 60 plus 20 times its value.

0 5 10 15 20 25 30 Score 59 and above 11.4% [6.1, 17.5] Score below 59 8.9% [3.3, 15.6] appearance rate, percent of business-run trials

What this study cannot support

A null result from an underpowered study is not a finding, it is an absence of one. So here are the bounds rather than the headline. We would rather a reader arrive at the limits from us than find them alone, and the numbers below are the ones a skeptical reader should quote against us.

The pilot could not have detected an effect at its own noise floor

Simulating the trial counts we actually ran, 114 against 90, the minimum difference detectable at 80 percent power is roughly 15 percentage points. Our engine calls 6 points the smallest difference worth reporting. So this design was blind to effects between those two sizes, and that is a real weakness.

The interval on the difference reaches to about 11 points

Bootstrapping the high minus low difference directly gives 2.5 points with a 95 percent interval from -5.7 to 10.6. It contains zero, so nothing was detected. It also reaches high enough that a real effect of 10 points would probably have been missed. This rules out a large effect and not a moderate one.

We ran 3 runs per prompt where the method asks for 385

Our own sample size function asks for 97 runs per prompt for a 10 point margin, 267 for a 6 point margin and 385 for a 5 point margin. We could drive 3 per prompt by hand in a browser. The budget went on breadth instead, 10 prompts across 5 trades, because the correlation was the question.

What we will and will not claim from this
ClaimSupported?
A higher score makes an assistant more likely to name youNo. Never was. We stopped saying it.
The score is worthlessNot shown. The pilot was too small to prove that.
Fixing what the scan finds moves your placementNever measured, for anyone, including ourselves.
We looked for the relationship we were selling and did not find itYes. That is the whole finding.
Single-run appearance claims are noiseYes, and it reproduces in about a minute.

Method

Every rule below was fixed before the first query ran, which is the part that makes the rest worth reading. Two rule bugs were caught and corrected pre-measurement, and both are disclosed here rather than buried: a cafe had been sorted into tree services, and a tattoo studio into restaurants.

Population

Douglas County businesses already in our scan database. Four of the 59 rows, about 7 percent, were dropped by script rather than judgement: three whose stored score was flagged invalid, and one outside the county. That leaves 55 analysable businesses, mean score 59.3, standard deviation 12.6, range 28 to 85.

Trades, chosen by rule

A trade qualified only if it held at least 3 analysable businesses and a score spread of at least 15 points. Six qualified and five were tested: construction, electrical, HVAC, excavation and restaurants. The sixth was nonprofits, which has no realistic consumer query, so it was left out.

Prompts and engines

Ten prompts, two per trade, plain local discovery questions such as who is a good excavation contractor in Roseburg Oregon. Thirty runs on Perplexity, each a fresh search thread, plus four on ChatGPT as a cross-engine check. Google AI Mode returned zero usable runs and is therefore not claimed at all.

The outcome, defined in advance

NAMED means the business appears in the visible answer prose as a listed or recommended provider, matched after stripping suffixes and punctuation. Appearing only in a places sidebar did not count, because that is a map widget rather than the answer. Domain citations were recorded separately as a secondary outcome.

Stopping rule and score source

Run all 30 Perplexity runs, report whatever comes out. No early stop, and no runs added after seeing results. Scores come from the latest scan per domain, because 5 of 55 stored scores, 9 percent of them, were stale. Both versions were computed and the conclusion is identical either way.

What did appear to drive placement

Both engines answered these questions out of a places index carrying star ratings and review counts. Neither gave any sign of having read the websites we grade. One of them said its criterion out loud, which is rarer than it should be and worth recording exactly.

“the largest review sample among these four”
ChatGPT, explaining which excavation contractor it picked, 16 August 2026. Recorded in the published runs.

Review counts, for the businesses that got named

Only 5 of the 34 businesses in scope were ever named, about 15 percent. Those five held 973, 29, 16, 15 and 13 Google reviews, and the competitors that beat them held 889, 85 and 35. We cannot turn that into a correlation, because a review count is only observable for a business the engine already surfaced.

The identical question gave three different answers

Run three times in a row, one excavation query returned a different top recommendation each time. That is the most reproducible thing on this page and it takes about a minute to check. Any single-run claim that a business appears in AI answers, ours or a competitor’s, is one coin flip reported as a measurement.

Citations went mostly to third parties

Across 30 Perplexity runs there were 35 citation tokens. Six belonged to a business we had scored, covering 3 distinct domains, so 20 percent of runs cited at least one, with a 95 percent interval from 7 to 37. The rest went to directories, tourism sites and competitors.

The county dataset

Separately from the pilot, we scanned local business websites to see what an AI engine can read off them. These are first scans, one per domain, and the figures below are read live from the same file that produces our own graphics, so a stale number cannot survive on this page.

58
websites scanned
57.3
mean score of 100
33
graded D or F
0
graded A

The sample is small and we are not going to hide it

58 websites is a convenience sample of the businesses we could reach, not a random draw. Treated as a sample, the D or F share of 56.9 percent carries a 95 percent interval from 44.8 to 69.0. So this supports most local sites fail this and never a precise share.

The near-unanimous results are the ones that survive the sample size

Where almost everyone fails or almost everyone passes, a small sample is still informative. 56 of 58 sites carry no review or rating markup, an interval of 91.4 to 100 percent. Every single site was indexable and allowed AI answer crawlers in robots.txt, with zero failures out of 58.

Machines get the worst of these websites

Averaged across the county, the parts a person experiences are healthy and the parts a machine reads are not. Performance and technical health sit in the 80s. Entity data and structured data sit in the low 60s, conversion in the 40s, and content in the low 30s.

Mean pillar score across the county, out of 100
Performance86.8
Technical85.7
Structured data62.0
Entity & Authority61.4
Conversion46.9
Content32.2

Each bar width equals the figure printed beside it. Read the widths out of the page source if you want to verify that nothing here was drawn by hand.

The checks we are proudest of cannot tell anyone apart

This is the uncomfortable structural finding. Our two best evidenced AI retrieval checks pass for 100 percent of the county, so they carry 16 weight points and separate nobody. The checks that actually move a local score are schema and conversion advice, which is a weaker evidence base.

One check earns its weight

Content that only appears after JavaScript runs is the exception, and 11 of 58 local sites fail it. Vercel and MERJ instrumented more than 500 million AI crawler fetches and found the crawlers downloaded JavaScript and never executed it (The rise of the AI crawler).

Threats to this conclusion

In descending order of how much they worry us. No published work we can point to has shown a stable causal effect for any of these optimisation techniques across platforms. The nearest academic anchor is the Princeton generative engine optimization study, and it measured a purpose-built benchmark rather than a live assistant.

Our trade assignment was too coarse, and we found out afterwards

Four post-hoc probes showed two of our silent high scorers were metal fabricators being asked a general contractor question. Dropping that bucket leaves 16 businesses, a rank correlation of 0.020 with an interval from -0.482 to 0.497, and a difference of -9.8 points, meaning lower scorers appearing more.

We are not claiming low scores help

That -9.8 figure comes from 16 businesses and its correlation interval spans nearly the entire possible range, so it is unstable and we treat it as noise. It is reported only because hiding a sensitivity analysis that moved is exactly the failure this document exists to avoid.

Two engines, one county, one query class

Google AI Mode is entirely unmeasured and it carries the most local query volume. Every prompt was plain local discovery, and vendor work puts AI answers on only about 15 percent of those queries, so this may be a small surface. The population is one rural county, which may say more about our market than the score.

The confound may be the entire effect

Score correlates with business size, age, review count and marketing budget. In this pilot the size-linked variable is doing visible work and the score is not. No observational design separates them. Only a randomised intervention would, and we have not run one.

The data

Published under CC BY 4.0. The JSON carries the full method, the county audit, the anonymised business rows and all 34 assistant runs verbatim. The CSV is the business rows alone, which is everything needed to recompute the rank correlation in a spreadsheet.

Cite as: GridSignal (The Roseburg Plug LLC), Douglas County AI Visibility Dataset, August 2026, https://gridsignal.app/research

What we redacted, and why we will not undo it

Business rows carry score, times named and run count. No name, domain, city, letter grade or trade. We never publish a named local business as scoring badly, and in a county this small the trade label alone re-identifies, because one tested trade holds only four businesses.

The businesses the assistants named are named

The 34 runs are verbatim, including every business each assistant recommended and every domain it cited. That information is public and favourable to those businesses, and no score of ours is attached to any name anywhere in the file. The cost of this is stated plainly below.

What our redaction costs you

The post-hoc sensitivity analysis needs the trade label, so you cannot recompute that one from the public file. Everything else reproduces: the correlation, the median split, the confidence intervals and the pooled appearance rate of 10.3 percent across 204 business-run trials.

Reproducing the analysis

Our intervals come from the same statistics code the product ships, imported rather than reimplemented, with a fixed bootstrap seed so every interval reproduces exactly. The scripts are design.mjs, analyse.mjs, power.mjs and publish.mjs.

What we changed because of this

Publishing a result you then ignore is worse than not measuring it. So this is the list, and every item is already done or already removed from what we say. It cost us the strongest sentence we had, and it was not a defensible sentence.

We stopped selling the score as a prediction

Any phrasing where the number forecasts an AI outcome is gone, including businesses scoring above 80 are more likely to be surfaced. The scan is an audit of real site defects and it is described as exactly that. It is necessary for accurate description, and it is not sufficient for placement.

We stopped implying our own fixes made us visible

Our own site score went up a lot after we fixed what the scanner told us to fix. We have never measured a visibility change for anyone, ourselves included, so that story stays a story about closing defects and never becomes evidence of lift.

We now tell local businesses the thing with no money in it

The signal that showed up in this pilot was review volume and a complete Google Business Profile. We do not sell either. Finishing that profile and asking happy customers for reviews is free, and it is what the measurement actually pointed at.

Who did this

GridSignal is a product of The Roseburg Plug LLC, a registered Oregon company with a downtown office in Roseburg. We sell website audits and rebuilds, so we had every commercial reason not to run this test and no independent party verified it. Check the raw data rather than trusting us.

Shane Hayes, founder. 537 SE Main Street, Roseburg, Oregon 97470. (541) 378-5562.

Get in touch

Corrections are welcome and will be published on this page with the date they were made. If you replicate this and get a different answer, we would rather know.

Our free scan reads your website the way an AI answer engine does and tells you what it can and cannot read off the page. It is not a prediction of placement, for the reasons above.

Run the free scan