Living case study

From GEO readiness scores to reproducible result evidence.

This case study records the completed baseline, treatment, measurement protocol, and unresolved outcome. It will become a true before/after case only after the September 14 retest.

Direct answer

AI Growth Bench is replacing a readiness-heavy GEO portfolio with a reproducible results loop. The starting evidence is mixed: Google Search Console recorded 189 property-level impressions and no clicks from 2026-08-05 through 2026-09-01; server logs confirmed seven search and AI crawler classes; and the August 31 fixed-prompt baseline completed 30 of 30 checks with three mentions and two displayed citations, all limited to the branded definition prompt. The post-baseline treatment candidate synchronizes the Results ledger, current report, entity facts, benchmark summaries, and AI-readable files; consolidates the overlapping quality-gate URL; and narrows the active index portfolio. The benchmark currently has 30 of 30 baseline checks completed, so no before/after improvement is claimed yet. The outcome will be evaluated on 2026-09-14 using the same prompts and an evidence rule that keeps mentions, citations, source URLs, referrals, and crawler visits separate.

Reviewed by Alex. Published August 17, 2026; baseline updated August 31 and treatment candidate prepared September 4, 2026. No external distribution treatment has been executed.

Before

Evidence fragmented

Readiness scores, dated reports, crawler logs, and historical prompt checks used different public structures.

Baseline

189 impressions / 0 clicks

2026-08-05 to 2026-09-01; property-level Search Console aggregation.

Treatment

Local candidate

Synchronized evidence, entity, index portfolio, and intent consolidation; deployment and external distribution remain unexecuted.

Outcome

Pending

No before/after claim until the fixed-prompt retest on 2026-09-14.

What problem did the first version have?

The site showed strong technical hygiene and clear workflow diagrams, but a reviewer had to inspect several pages to learn what had actually happened. Citation-readiness scores sat close to historical prompt checks, crawler access, Search Console impressions, and GA4 observations even though those signals measure different outcomes.

That made the portfolio useful as a process demonstration but weaker as an experiment. The repair is not another batch of pages. It is a smaller public protocol with explicit denominators, dates, and evidence gates.

What changed in the measurement design?

Ten prompts are frozen for one cohort and tested across three engines. Every engine and prompt combination has a baseline and retest field. An observation is incomplete until the result date, mention state, citation state, accuracy, and displayed source URL state are recorded.

Summary numbers are derived from those cells. The public Results page separately reports analytics referrals and crawler access, which are measured outside the answer interface.

Why is the GEO checker restricted to this site?

Arbitrary server-side URL fetching creates SSRF and abuse risk. The MVP therefore accepts only paths already present in the AI Growth Bench public route inventory, fetches only the current application, enforces timeout and size limits, and exposes a deterministic rubric.

External-site auditing remains a later capability that would need DNS and IP validation, redirect revalidation, egress controls, and rate limits.

What happens after the retest?

The same prompt wording will be rerun on September 14. Results will be entered whether they are positive or negative. The case study will compare recognized mentions, verified citations, and displayed canonical URLs without interpreting crawler hits as success.

Search Console clicks and AI referral sessions will remain separate outcome measures. If the treatment produces no change, the report will say so and identify the next variable.

Proof links

Inspect the work, not just the summary.