# SEO, AEO & AI-Crawler Optimization Plan
Requirements and rules for optimizing the BlueCap Strategies rebuild for classic search (SEO),
answer engines / AI citation (AEO/GEO), and AI-crawler acceptance. Written **2026-07-15** against
mid-2026 best practices (see Sources). Where a rule can be machine-enforced, it is defined as a
**Zod schema constraint** or a **build-time check** so the build fails when copy violates it —
this is the TypeScript/Astro equivalent of Pydantic validation (do **not** add Python/Pydantic to
a JS build; Astro Content Collections already run on Zod).
## 0. Guiding principles (2026 reality check)
1. **HTML is the source of truth, not `llms.txt`.** AI search crawlers (GPTBot, ClaudeBot,
PerplexityBot, OAI-SearchBot, Google-Extended) overwhelmingly **skip `/llms.txt` and crawl
rendered HTML directly**; Google confirmed it does not support llms.txt. Keep our `llms.txt`
(low cost, ~10% of sites have one) but invest first in semantic HTML, structured data, and
crawlable copy. Our Astro output is **fully static** — 100% of meaningful copy is in initial
HTML, which is the single biggest AEO advantage over the old JS-heavy WordPress/Elementor site.
2. **Structured data drives citation.** Schema lets answer engines verify *who we are*, *what a
page covers*, and *which facts are reliable to cite*. FAQ schema has among the highest AI
citation rates (question/answer mirrors how engines present answers).
3. **Answer-ready structure wins.** Clear headings, direct definitional first sentences, and
self-contained paragraphs get extracted; buried or paraphrased marketing prose does not.
4. **E-E-A-T via the principal.** Boutique-advisory credibility is entity-driven — a real
`Person` entity (Brad Rodgers) with `knowsAbout`, linked to the `Organization`, is a core GEO
asset. (Gated on owner-approved bio — see parity plan Phase 5.)
## 1. Programmatic enforcement (Zod + build linter)
### 1a. Content-collection SEO schema (fails the build)
Extend `src/content.config.ts`. Every collection (pages, services, focusAreas, insights) shares
an `seo` shape with enforced limits:
```ts
import { z } from "astro:content";
export const seo = z.object({
//
renders as "`title` | BlueCap Strategies" — keep the full SERP title ≤ ~60 chars
title: z.string().min(10).max(55),
// meta description + OG description; 120–158 is the safe SERP/AI-snippet window
description: z.string().min(120).max(158),
// canonical is derived from slug, but allow override for cross-posted insights
canonicalUrl: z.string().url().optional(),
// every page must ship a social share image (absolute path under /assets)
ogImage: z.string().startsWith("/assets/").default("/assets/og/default.jpg"),
// primary entity/topic — powers internal linking + schema keywords
primaryTopic: z.string(),
noindex: z.boolean().default(false),
});
```
Rules encoded here (build fails on violation): title length, description length window, OG image
present, canonical is a valid URL. This is the "rules are being followed" guarantee you asked
about — enforced on every `npm run build` in CI before ArgoCD ever sees the image.
### 1b. Build-time SEO linter (what Zod can't see)
Zod validates frontmatter, not rendered HTML. Add a small check (a Node script run in CI, and/or
Playwright assertions in `tests/`) that fails the build if any built page in `dist/` violates:
- Exactly **one `
`** per page.
- Heading hierarchy has no skipped levels (no `
` before an `
`).
- ``, ``, `` all present and non-empty.
- Every `` has non-empty `alt` (parity plan already requires this).
- At least one JSON-LD block present, and **every** JSON-LD block `JSON.parse`s.
- Internal links resolve (no 404s to our own routes).
- Canonical host matches the chosen production host (apex vs `www` — see §5).
Recommended: a `scripts/seo-lint.mjs` that walks `dist/**/*.html` with a lightweight parser, plus
keep the existing Playwright suite for link/CTA coverage.
## 2. Per-page technical SEO requirements
Already implemented in `src/layouts/BaseLayout.astro` (keep + enforce):
- ``, ``, ``
- Open Graph (`og:type/title/description/url/image`) + Twitter `summary_large_image`
- JSON-LD injection slot (Organization + WebSite globally, page schema per route)
- `lang="en"`, viewport, skip-link, RSS ``
To add/verify:
- [ ] **Per-page OG images** — currently every page falls back to `/assets/water-strategy.jpg`.
Generate topic-specific OG images (services, focus areas, insights) at 1200×630.
- [ ] **`robots` meta** wired to the `noindex` frontmatter flag (for `/thank-you/`, staging).
- [ ] **`article:published_time` / `modified_time`** OG tags on insights posts.
- [ ] **Breadcrumbs** (visible + `BreadcrumbList` schema) on services/focus/insights detail pages.
- [ ] **Sitemap** already via `@astrojs/sitemap` (excludes `/thank-you/`); add `lastmod` from
content dates for insights.
- [ ] **Trailing-slash consistency** — config uses `trailingSlash: "always"`; ensure canonical +
internal links all match to avoid duplicate-URL dilution.
## 3. Structured data (JSON-LD) matrix
| Page / type | Required schema | Notes |
|-------------|-----------------|-------|
| Global | `Organization`, `WebSite` | Already in `src/utils/schema.ts`. Add `sameAs` (real socials), `logo`, `contactPoint` (phone/email), `areaServed`. |
| Homepage | `Organization` + `WebSite` (+ optional `SearchAction`) | — |
| `/services/` + each service | `Service` (with `provider` → Organization, `serviceType`, `areaServed`) | One per detail page. |
| `/focus-areas/` + each focus | `Service` or `DefinedTerm`/`Article` | Focus areas double as topic-authority hubs. |
| About / Brad Rodgers | `Person` + `AboutPage`; `Person.worksFor` → Organization; `knowsAbout[]` | **Owner-approved facts only.** Core GEO/E-E-A-T asset. |
| Insights posts | `Article`/`BlogPosting` (author=Person, publisher=Organization, dates, `headline`, `image`) | Powers AI citation of thought leadership. |
| Insights index | `Blog` / `CollectionPage` | — |
| Contact | `ContactPage` + `Organization.contactPoint` | Use real values, never the WP placeholders. |
| FAQ sections | `FAQPage` — **only where the same Q&A is visibly rendered** | Highest AI-citation rate. Never add invisible FAQ schema. |
| Breadcrumbs | `BreadcrumbList` | Detail pages. |
Rule: **structured data must reflect content that is visibly present on the page.** No hidden-text
schema, no unsupported claims, no fabricated review/rating markup.
## 4. AEO / GEO content rules
Applied during Phase 2 copy restoration and Phase 5+ enhancements (must not violate the
no-paraphrasing parity rule on restored copy):
- **Answer-first**: each section's opening sentence should stand alone as a direct answer /
definition (e.g., "Marine conservation finance is …"). Extractable by answer engines.
- **Semantic headings** phrased as the questions/topics people ask ("What makes PFAS remediation a
capital formation problem?").
- **FAQ modules** on high-intent authority pages (PFAS economics, marine conservation finance) —
visible Q&A + `FAQPage` schema. Owner-approved answers only.
- **Entity consistency**: use canonical terminology everywhere — *catalytic capital*, *aquatic
resources*, *marine & freshwater ecosystems*, *PFAS remediation*, *marine conservation finance*,
*environmental markets*, *coastal resilience*. Feeds `primaryTopic` + internal linking.
- **Internal linking graph**: services ↔ focus areas ↔ insights ↔ principal bio. Dense, relevant
internal links are a top GEO signal.
- **Freshness**: insights carry `publishDate`/`modifiedDate`; surface "Updated" dates.
- **Statistics & specifics**: the homepage metrics (27 years, $10B+, $3.5B+, 275+) are exactly the
kind of concrete, citable facts answer engines favor — keep them prominent and, where true,
attributable.
## 5. AI-crawler policy (`robots.txt`) — **owner decision required**
Current `public/robots.txt` = `Allow: /` for everyone. 2026 best practice is **explicit per-agent
rules**, splitting *search/answer* bots (that can cite and send traffic) from *training* bots
(that ingest content into model weights). Decision matrix:
| Bot | Job | Recommendation |
|-----|-----|----------------|
| `OAI-SearchBot`, `ChatGPT-User` | ChatGPT search index + live citation fetch | **Allow** |
| `Claude-SearchBot` | Claude web-search citations | **Allow** |
| `PerplexityBot` | Perplexity answer citations | **Allow** |
| `Google-Extended` | Gemini + Google AI Overviews | **Allow** (maximizes AI visibility) |
| `Googlebot`, `Bingbot` | Classic search | **Allow** |
| `GPTBot` | OpenAI model *training* | Owner choice — allow = max reach, disallow = opt out of training |
| `ClaudeBot` | Anthropic model *training* | Owner choice (same trade-off) |
| `Bytespider`, stealth scrapers | Known to ignore robots.txt | Consider **edge/WAF block** (Cloudflare) — robots.txt won't stop them |
**Recommended default for an authority-seeking advisory firm: allow all search/answer + classic
bots (maximize citation and discoverability); training bots are a values call.** Because BlueCap
wants to be cited as an authority, allowing everything is the simplest high-visibility posture.
Non-compliant scrapers are handled at the Cloudflare edge (see gitops doc §4.2), not robots.txt.
Keep `llms.txt` (already present, low cost) but treat it as supplemental, not load-bearing.
## 6. Measurement & Monitoring — **recommendation, needs sign-off**
Three distinct layers; today none is decided.
### 6a. Uptime / infrastructure — reuse the existing `monylog` stack (do NOT deploy Uptime Kuma)
The homelab already runs a full self-hosted **LGTM + Alertmanager** stack on the `monylog` box
(`10.66.15.21`, Ansible role `pglta`): Prometheus (15d retention, remote-write receiver), Grafana,
Loki, Tempo, Alertmanager → **Pushover** (creds from OpenBao) with **ntfy** fallback → phone. It
runs **outside k8s on purpose** so it still alerts when the cluster is down. Grafana Alloy runs as
a k3s DaemonSet and remote-writes cluster metrics + ships all pod logs to monylog.
**Already flowing for BlueCap (zero work):**
- **Pod/deployment health** — Alloy scrapes kube-state-metrics + annotated pods → Prometheus.
`kube_deployment_status_replicas_unavailable`, restarts, OOMKills are queryable now.
- **nginx logs** — Alloy ships all pod logs to Loki. 5xx rate / errors available via LogQL now.
**To add (small):**
1. **blackbox_exporter** — the one gap (not in the monylog compose). Add as a compose service +
a Prometheus scrape job. Probe the in-cluster Service (app health) **and** the public URL (full
edge-path uptime + TLS expiry). monylog is off-cluster → probing the public URL is a genuine
external vantage. Metrics: `probe_success`, `probe_http_status_code`,
`probe_ssl_earliest_cert_expiry`, latency.
2. **`web-bluecap.yml` alert rules** (mirror the existing host-infra rules pattern): `SiteDown`
(`probe_success==0` for 2m), `CertExpiringSoon` (<14d), `High5xx` (Loki/blackbox),
`DeploymentDegraded`, `PodRestarting`.
3. **Alertmanager** already routes to Pushover — just add a `service: bluecap` label; no new
channel. **Grafana** already present → build a BlueCap uptime/latency/5xx/pod-health dashboard.
- **ArgoCD** health/sync status as an additional deploy-level signal.
### 6b. Visitor & behavior analytics — PostHog + Umami side-by-side trial (ADR-0005)
Decided: **run PostHog and Umami side-by-side initially** and keep whichever fits best. This is
**product/visitor analytics**, distinct from the `monylog` ops telemetry in §6a. Trial setup:
- **PostHog — Cloud free tier** (1M events/mo) for the trial, not self-host (self-hosting needs
ClickHouse/Kafka/etc.; PostHog's docs steer small users to Cloud). Reverse-proxy ingest through
our domain to reduce ad-blocker loss. Gives autocapture, funnels, session replay.
- **Umami — self-hosted** in k3s: cookieless, ~2 KB, no consent banner, always-on baseline
(pageviews, sources, devices, custom events). Aligns with the Digital-ESG positioning.
- Both wired through one config-driven `Analytics.astro` in `BaseLayout`, gated by env vars, so
either can be toggled and the winner kept without code churn. Same event set in both:
consultation-CTA clicks, per-service/focus CTA source (ties to the segmented-CTA strategy),
insights reads, contact submits.
- **Consent/privacy must cover the superset** — PostHog (esp. session replay) likely triggers
cookie-consent + privacy-policy obligations; add these **before** enabling PostHog in production.
Umami alone needs no banner.
- **Avoid GA4** unless Google Ads attribution is later required.
- Evaluate on insight value vs. weight/consent/ESG fit; record the winner in a follow-up ADR.
### 6c. Search & AI-citation visibility
- **Google Search Console** + **Bing Webmaster Tools** (free, essential) — submit sitemap, monitor
impressions/queries/coverage, Core Web Vitals.
- **AI-citation tracking** (emerging, optional): periodically test brand/topic prompts across
ChatGPT/Perplexity/Google AI Overviews and log whether BlueCap is cited. Lightweight manual
tracker to start; dedicated GEO tools are immature — don't over-invest yet.
- **Performance budget**: Lighthouse in CI; target 95+ SEO/Perf/Best-Practices/Accessibility.
Static Astro should hit this easily; guard against regressions from images (use Astro `Image`).
## 7. Pre-launch SEO checklist
- [ ] Zod SEO schema in place; build fails on title/description/OG/canonical violations.
- [ ] `scripts/seo-lint.mjs` (single H1, heading order, canonical, alt text, JSON-LD parses) green in CI.
- [ ] Every page: unique title, description in window, canonical, OG image, valid JSON-LD.
- [ ] Organization/WebSite/Service/Person/Article/Breadcrumb schema per §3, all visible-content-backed.
- [ ] `robots.txt` per §5 (owner-decided); `sitemap-index.xml` correct; `/thank-you/` + staging `noindex`.
- [ ] Canonical host decided (apex vs `www`) + redirect configured at the edge.
- [ ] Search Console + Bing verified, sitemap submitted.
- [ ] Analytics deployed (Umami) + consultation/CTA events firing.
- [ ] Uptime Kuma + external check live and alerting.
- [ ] Lighthouse ≥ 95 across the board; no CLS from unsized images.
## Sources (verified 2026-07-15)
- [The State of llms.txt in 2026 — aeo.press](https://www.aeo.press/ai/the-state-of-llms-txt-in-2026)
- [llms.txt Explained (May 2026) — Codersera](https://codersera.com/blog/llms-txt-complete-guide-2026/)
- [Structured Data for AEO and GEO: Schema Markup Guide 2026 — Kurieta](https://kurieta.com/schema-for-aeo-geo/)
- [Answer Engine Optimization: Complete AEO Guide [2026] — Frase](https://www.frase.io/blog/what-is-answer-engine-optimization-the-complete-guide-to-getting-cited-by-ai)
- [Are FAQ Schemas Important for AI Search, GEO & AEO? — Frase](https://www.frase.io/blog/faq-schema-ai-search-geo-aeo)
- [The AI User-Agent Landscape in 2026: A Complete Reference — No Hacks](https://nohacks.co/blog/ai-user-agents-landscape-2026)
- [AI Crawlers Explained: GPTBot, ClaudeBot, PerplexityBot (2026) — Anagram](https://www.anagram.ai/blog/ai-crawlers-explained-gptbot-claudebot-perplexitybot-and-how-to-let-them-in-2026)
- [Robots.txt for AI Crawlers in 2026 — Cubitrek](https://cubitrek.com/blog/robots-txt-2026-managing-ai-crawler-budgets)