# robots.txt — sparkyintimo.com # Updated: 2026-06-02 # # POLICY: YES to AI *search & citation* crawlers — they surface us in AI # answers WITH attribution and drive discovery. NO to AI *training* crawlers # — we are not a training corpus. This mirrors our data-dignity / GDPR # posture: be findable everywhere, harvested by no one. # # Note: blocking the training crawlers below does NOT reduce our presence in # Google Search / AI Overviews (governed by Googlebot) or in AI search answers # (governed by the search bots) — those are separate product tokens. # ── Default: allow well-behaved crawlers ───────────────────────────── User-agent: * Allow: / Disallow: /node_modules/ Disallow: /*.map$ Disallow: /coverage/ Disallow: /test-results/ # ── Sitemap ────────────────────────────────────────────────────────── Sitemap: https://sparkyintimo.com/sitemap.xml # ── Search engines — explicit allow ────────────────────────────────── User-agent: Googlebot Allow: / User-agent: Bingbot Allow: / User-agent: DuckDuckBot Allow: / # ── AI SEARCH / live-citation crawlers — WELCOME ───────────────────── # These cite us live in AI search answers and do NOT train models on us. # Most sites block these; we want maximum visibility here, so say so. User-agent: OAI-SearchBot Allow: / User-agent: ChatGPT-User Allow: / User-agent: PerplexityBot Allow: / User-agent: Perplexity-User Allow: / User-agent: Applebot Allow: / # DuckAssistBot crawls live for DDG's AI answers and, per DuckDuckGo's own # bot page, is never used to train models. Answers cite their sources. User-agent: DuckAssistBot Allow: / # meta-webindexer is the Meta AI SEARCH crawler — per Meta's crawler docs it # "navigates the web to improve Meta AI search result quality". It is not the # training crawler (that is meta-externalagent, blocked below). User-agent: meta-webindexer Allow: / # meta-externalfetcher fetches one link when a user asks Meta AI about a URL — # user-directed, same class as ChatGPT-User. User-agent: meta-externalfetcher Allow: / # facebookexternalhit builds the link preview when someone shares us on # Facebook, Instagram, Messenger or WhatsApp. User-agent: facebookexternalhit Allow: / # MetaAI — allowed on Meta AI's own guidance (30 Jul 2026), which told us to # Disallow Meta-ExternalAgent and FacebookBot while allowing facebookexternalhit # and this token. CAVEAT: 'MetaAI' does NOT appear in Meta's published crawler # docs, which list only FacebookExternalHit, Meta-WebIndexer, Meta-ExternalAds, # Meta-ExternalAgent and Meta-ExternalFetcher — so this may be a token the # assistant invented about itself. An Allow for an unused agent is inert, so it # costs nothing and covers us if the token is real or becomes real. The # documented meta-webindexer and meta-externalfetcher entries above are the ones # with a paper trail; keep both. User-agent: MetaAI Allow: / # ── AI TRAINING crawlers — BLOCKED ─────────────────────────────────── # We are not a training dataset. Refusal is the brand statement. User-agent: Google-Extended Disallow: / User-agent: GPTBot Disallow: / User-agent: ClaudeBot Disallow: / User-agent: anthropic-ai Disallow: / User-agent: Claude-Web Disallow: / User-agent: Applebot-Extended Disallow: / # Meta-ExternalAgent was flipped to Allow for a few hours on 30 Jul 2026 to test # whether it was the token Meta AI's answer path actually needed — Meta's own docs # give it two jobs at once, "training foundation AI models OR improving products by # indexing content". Reverted the same day: the block-training/allow-search line # has to be true in this file for the published policy to mean anything, and a # training crawler sitting on Allow would have made it unverifiable. The four Meta # tokens that serve search and link previews stay allowed above. User-agent: Meta-ExternalAgent Disallow: / User-agent: meta-externalagent Disallow: / # FacebookBot — a legacy Meta crawler that gathered AI/speech training data. # Cloudflare classifies it 'AI Crawler' under Meta and its Block AI Bots setting # already blocks it at the edge, but robots.txt never declared the position. # Blocked on Meta AI's own guidance (30 Jul 2026), which named it alongside # Meta-ExternalAgent as the pair to Disallow. Closing the gap so the policy is # stated where crawlers actually read it. User-agent: FacebookBot Disallow: / User-agent: Amazonbot Disallow: / User-agent: Bytespider Disallow: / User-agent: CCBot Disallow: / User-agent: cohere-ai Disallow: / User-agent: Diffbot Disallow: / User-agent: ImagesiftBot Disallow: / User-agent: Omgilibot Disallow: / User-agent: Omgili Disallow: / # ── Block abusive SEO-recon scrapers ───────────────────────────────── User-agent: AhrefsBot Disallow: / User-agent: SemrushBot Disallow: / User-agent: DotBot Disallow: / User-agent: MJ12bot Disallow: /