Chapter 9: Resources
Primary documentation first. Vendors change crawler names and policies; always check the source before changing your
robots.txt.
Search Engine Documentation
- Google Search Central — SEO Starter Guide: developers.google.com/search/docs/fundamentals/seo-starter-guide
- Google — AI features and your website (AI Overviews, AI Mode): developers.google.com/search/docs/appearance/ai-features
- Google — JavaScript SEO basics: developers.google.com/search/docs/crawling-indexing/javascript/javascript-seo-basics
- Google — robots.txt introduction: developers.google.com/search/docs/crawling-indexing/robots/intro
- Google — Common crawlers, including the Google-Extended token: developers.google.com/crawling/docs/crawlers-fetchers/google-common-crawlers
- Google — Structured data introduction: developers.google.com/search/docs/appearance/structured-data/intro-structured-data
- Google — Spam policies: developers.google.com/search/docs/essentials/spam-policies
- Bing Webmaster Guidelines: bing.com/webmasters/help/webmaster-guidelines-30fba23a
AI Crawler Documentation
- OpenAI — Overview of OpenAI crawlers (GPTBot, OAI-SearchBot, ChatGPT-User): platform.openai.com/docs/bots
- OpenAI — Publishers and developers FAQ: help.openai.com/en/articles/12627856-publishers-and-developers-faq
- Anthropic — Crawler and web-content controls (ClaudeBot, Claude-SearchBot, Claude-User): support.anthropic.com/en/articles/8896518-does-anthropic-crawl-data-from-the-web-and-how-can-site-owners-block-the-crawler
- Perplexity — Perplexity crawlers: docs.perplexity.ai/guides/bots
- Apple — About Applebot: support.apple.com/en-us/119829
- Common Crawl — CCBot: commoncrawl.org/ccbot
- Meta — Web crawlers: developers.facebook.com/docs/sharing/webmasters/web-crawlers
If a link has moved, search the vendor's site for "crawler" or the bot name; every vendor above maintains a current page.
Standards and Proposals
- Robots Exclusion Protocol (RFC 9309): rfc-editor.org/rfc/rfc9309
- Sitemaps protocol: sitemaps.org/protocol.html
- schema.org vocabulary: schema.org
- JSON-LD 1.1 (W3C): w3.org/TR/json-ld11
- Open Graph protocol: ogp.me
- IndexNow: indexnow.org
- llms.txt proposal (not an adopted standard): llmstxt.org
Tools
| Tool | Use |
|---|---|
| Google Search Console | Indexing, URL Inspection, performance |
| Bing Webmaster Tools | The same for Bing; IndexNow integration |
| Schema Markup Validator | Validate JSON-LD against schema.org |
| Rich Results Test | Google's rich-result eligibility |
| PageSpeed Insights | Core Web Vitals and Lighthouse |
curl | See exactly what a non-rendering crawler receives |
| Browser DevTools → Network → disable JavaScript | Quick visual check of the no-JS page |
Related Workshops in This Curriculum
- Professional Output Suite — publish the site you will audit here.
- RAG System Implementation — the retrieve-then-generate pattern answer engines are built on.
- Agent Orchestration & Safety — why prompt injection in web content is an attack, not a tactic.
- Claude Code Mastery — automate the audits in Chapter 8.
The DreamLab Case Study
The DreamLab site's own AI-visibility work is recorded in ADR-043 ("AI Search Visibility — Server-HTML Depth, Schema Hygiene, FAQ Layer") in the site repository. Its live outputs are public:
- dreamlab-ai.com/robots.txt
- dreamlab-ai.com/llms.txt
- dreamlab-ai.com/sitemap.xml
- View source on dreamlab-ai.com to see the pre-hydration content and JSON-LD.
Continue Learning
This module completes the self-guided curriculum. If you want to apply it to your organisation with practitioners in the room, our residential programmes cover discoverability alongside agentic AI, product prototyping and secure distributed systems.
Explore programmes: dreamlab-ai.com/programmes
Get in touch: dreamlab-ai.com/contact
A DreamLab AI Self-Guided Workshop | Last Updated: October 2026