# ============================================ # robots.txt for ankurai.com # ============================================ # --- Default: Allow all search engines --- User-agent: * Allow: / Sitemap: https://www.ankurai.com/sitemap.xml # --- AI Content & LLM Files --- # https://llmstxt.org standard # llms.txt: https://www.ankurai.com/llms.txt # llms-full.txt: https://www.ankurai.com/llms-full.txt # ============================================ # AI Search & Retrieval Bots (ALLOWED) # These serve answers to users and drive traffic. # ============================================ # OpenAI — ChatGPT browsing (real-time user queries) User-agent: ChatGPT-User Allow: / # Perplexity — AI-powered search engine User-agent: PerplexityBot Allow: / # Microsoft Bing — also powers Copilot answers User-agent: Bingbot Allow: / # Google — standard search User-agent: Googlebot Allow: / # Apple — Siri and Spotlight User-agent: Applebot Allow: / # Amazon — Alexa answers User-agent: Amazonbot Allow: / # You.com — AI search User-agent: YouBot Allow: / # Phind — developer AI search User-agent: PhindBot Allow: / # OpenAI — SearchGPT / ChatGPT search results User-agent: OAI-SearchBot Allow: / # Anthropic — Claude search & retrieval (user-facing) User-agent: Claude-SearchBot Allow: / User-agent: Claude-User Allow: / # ============================================ # AI Training & Scraping Bots (BLOCKED) # These crawl for model training data, not # to serve users. No attribution, no traffic. # ============================================ # OpenAI — GPT model training User-agent: GPTBot Disallow: / # Google — Gemini / AI training User-agent: Google-Extended Disallow: / # Apple — Apple Intelligence training User-agent: Applebot-Extended Disallow: / # Common Crawl — dataset used for LLM training User-agent: CCBot Disallow: / # Meta — AI model training User-agent: FacebookBot Disallow: / User-agent: Meta-ExternalAgent Disallow: / User-agent: Meta-ExternalFetcher Disallow: / # ByteDance — AI training User-agent: Bytespider Disallow: / # Anthropic — Claude model training User-agent: anthropic-ai Disallow: / User-agent: ClaudeBot Disallow: / # Cohere — LLM training User-agent: cohere-ai Disallow: / # AI2 — research model training User-agent: AI2Bot Disallow: / # Diffbot — web scraping / knowledge graph User-agent: Diffbot Disallow: / # Webz.io — data scraping for AI User-agent: Omgilibot Disallow: / # Semianalysis / Velenpubliciteit scraper User-agent: ImagesiftBot Disallow: / # Scrapy-based generic crawlers User-agent: Scrapy Disallow: / # iask.ai crawler User-agent: iaskspider Disallow: / # Seekr AI User-agent: Seekr Disallow: / # PetalBot — Huawei AI search/training User-agent: PetalBot Disallow: / # Timpibot — decentralized AI training User-agent: Timpibot Disallow: / # Webz.io news crawler User-agent: Webzio-Extended Disallow: / # img2dataset — image scraping for AI training User-agent: img2dataset Disallow: / # Kangaroo Bot User-agent: KangarooBot Disallow: /