User-agent: * Allow: / # Keep the CMS admin and its API out of search indexes Disallow: /keystatic Disallow: /api/keystatic # The site is static with no server-side search, so `?s=` and `/search/` resolve # to whatever page the path happens to be (usually the homepage) rather than any # real results. Spam sites inject backlinks pointing at these to get their text # indexed against our domain (seen live: /?s=); GSC was crawling # them despite the self-referencing canonical already pointing at the clean URL. # Blocking the pattern stops the crawl rather than relying on canonical alone. Disallow: /*?s= Disallow: /search/ # AI answer engines and their crawlers are explicitly welcome: Putti optimises # for GEO/AEO (see /llms.txt), so we want these to index and cite the site. # Listed by name so the intent is on the record even though `*` already allows # them; tighten or remove a line here to opt a specific crawler out later. User-agent: GPTBot User-agent: OAI-SearchBot User-agent: ChatGPT-User User-agent: ClaudeBot User-agent: Claude-Web User-agent: anthropic-ai User-agent: PerplexityBot User-agent: Perplexity-User User-agent: Google-Extended User-agent: Applebot-Extended User-agent: CCBot User-agent: Amazonbot User-agent: Bytespider Allow: / # A machine-readable site overview for LLMs / answer engines (llmstxt.org). # Not a standard robots directive, kept here for discoverability. # llms.txt: https://www.puttiapps.com/llms.txt Sitemap: https://www.puttiapps.com/sitemap-index.xml