# ============================================================================= # giskard.ai robots.txt — COPY/PASTE into Webflow → Site settings → SEO → robots.txt # Policy: search + AI citation YES · AI training / fine-tuning OPT OUT # Content Signals: https://contentsignals.org # Also: https://www.giskard.ai/llms.txt # ============================================================================= # As a condition of accessing this website, you agree to abide by the following # content signals: # (a) If a content-signal = yes, you may collect content for the corresponding use. # (b) If a content-signal = no, you may not collect content for the corresponding use. # (c) If the website operator does not include a content signal for a corresponding # use, the website operator neither grants nor restricts permission via content # signal with respect to the corresponding use. # Meanings: # search = search index / classic results (not AI-generated summaries alone) # ai-input = RAG / grounding / real-time generative answers (GEO citations) # ai-train = training or fine-tuning AI models # ANY RESTRICTIONS EXPRESSED VIA CONTENT SIGNALS ARE EXPRESS RESERVATIONS OF RIGHTS # UNDER ARTICLE 4 OF THE EUROPEAN UNION DIRECTIVE 2019/790 ON COPYRIGHT AND RELATED # RIGHTS IN THE DIGITAL SINGLE MARKET. # ============================================================================= # --- Default --- User-agent: * Content-Signal: search=yes, ai-input=yes, ai-train=no Allow: / # ============================================================================= # ALLOW — classic search engines # ============================================================================= User-agent: Googlebot Allow: / User-agent: Googlebot-Image Allow: / User-agent: Googlebot-News Allow: / User-agent: Googlebot-Video Allow: / User-agent: Storebot-Google Allow: / User-agent: Google-InspectionTool Allow: / # Google Gemini / AI Overviews discovery (does not affect Googlebot search indexing) User-agent: Google-Extended Content-Signal: search=yes, ai-input=yes, ai-train=no Allow: / User-agent: Bingbot Allow: / User-agent: bingbot Allow: / User-agent: msnbot Allow: / User-agent: AdIdxBot Allow: / User-agent: DuckDuckBot Allow: / User-agent: DuckAssistBot Allow: / User-agent: Slurp Allow: / User-agent: Yahoo Allow: / User-agent: Yandex Allow: / User-agent: YandexBot Allow: / User-agent: Baiduspider Allow: / User-agent: Sogou Allow: / User-agent: Applebot Allow: / # Apple Intelligence discovery (does not affect Applebot search / Spotlight) User-agent: Applebot-Extended Content-Signal: search=yes, ai-input=yes, ai-train=no Allow: / User-agent: SeznamBot Allow: / User-agent: Qwantify Allow: / User-agent: ecosia Allow: / User-agent: BraveBot Allow: / User-agent: PetalBot Allow: / User-agent: DotBot Allow: / # ============================================================================= # ALLOW — AI search / citation / user-triggered fetch (not training) # ============================================================================= # OpenAI — ChatGPT search, on-demand fetch, and discovery crawlers User-agent: OAI-SearchBot Allow: / User-agent: ChatGPT-User Allow: / User-agent: GPTBot Content-Signal: search=yes, ai-input=yes, ai-train=no Allow: / User-agent: GPTBot-1.0 Content-Signal: search=yes, ai-input=yes, ai-train=no Allow: / User-agent: GPTBot-1.1 Content-Signal: search=yes, ai-input=yes, ai-train=no Allow: / # Anthropic — Claude search + on-demand fetch User-agent: Claude-SearchBot Allow: / User-agent: Claude-User Allow: / # Perplexity User-agent: PerplexityBot Content-Signal: search=yes, ai-input=yes, ai-train=no Allow: / User-agent: Perplexity-User Allow: / # You.com User-agent: YouBot Allow: / # Amazon Alexa / answer features User-agent: Amazonbot Allow: / # Meta AI fetchers (not Facebook training crawler) User-agent: Meta-ExternalAgent Allow: / User-agent: meta-externalagent Allow: / User-agent: Meta-ExternalFetcher Allow: / User-agent: meta-externalfetcher Allow: / # Phind / Kagi / other answer engines User-agent: PhindBot Allow: / User-agent: KagiBot Allow: / # Mistral — user-triggered fetch (Le Chat browsing) User-agent: MistralAI-User Content-Signal: search=yes, ai-input=yes, ai-train=no Allow: / # LinkedIn / Twitter / social preview (not model training) User-agent: LinkedInBot Allow: / User-agent: Twitterbot Allow: / User-agent: Slackbot Allow: / User-agent: Discordbot Allow: / User-agent: WhatsApp Allow: / User-agent: TelegramBot Allow: / # SEO / monitoring (allow — useful for our own visibility) User-agent: AhrefsBot Allow: / User-agent: SemrushBot Allow: / User-agent: MJ12bot Allow: / User-agent: Screaming Frog SEO Spider Allow: / User-agent: SiteAuditBot Allow: / User-agent: Rogerbot Allow: / User-agent: ScreenerBot Allow: / # Archive (historical web, not AI training corpus scrapers) User-agent: archive.org_bot Allow: / User-agent: ia_archiver Allow: / # ============================================================================= # DISALLOW — AI training / fine-tuning / training corpora (HARD OPT-OUT) # ============================================================================= User-agent: GoogleOther Disallow: / User-agent: GoogleOther-Image Disallow: / User-agent: GoogleOther-Video Disallow: / # --- Anthropic training (legacy + ClaudeBot training fleet) --- User-agent: ClaudeBot Content-Signal: ai-train=no, search=no, ai-input=no Disallow: / User-agent: anthropic-ai Disallow: / User-agent: Claude-Web Disallow: / User-agent: claude-web Disallow: / # --- Common Crawl (feeds many LLM training datasets) --- User-agent: CCBot Content-Signal: ai-train=no, search=no, ai-input=no Disallow: / User-agent: CCBot/2.0 Disallow: / # --- ByteDance --- User-agent: Bytespider Content-Signal: ai-train=no, search=no, ai-input=no Disallow: / User-agent: ByteSpider Disallow: / # --- Meta / Llama training crawlers --- User-agent: FacebookBot Content-Signal: ai-train=no, search=no, ai-input=no Disallow: / User-agent: facebookexternalhit Disallow: / User-agent: meta-webindexer Disallow: / # --- Cohere --- User-agent: cohere-ai Disallow: / User-agent: cohere-training-data-crawler Disallow: / # --- Allen AI / AI2 --- User-agent: AI2Bot Disallow: / User-agent: AI2Bot-Domestic Disallow: / User-agent: ai2bot Disallow: / User-agent: ai2bot-dolma Disallow: / # --- Other training / scrape-for-ML bots --- User-agent: Diffbot Disallow: / User-agent: ImagesiftBot Disallow: / User-agent: img2dataset Disallow: / User-agent: omgili Disallow: / User-agent: omgilibot Disallow: / User-agent: Webzio-Extended Disallow: / User-agent: Timpibot Disallow: / User-agent: TimpiBot Disallow: / User-agent: Peer39_crawler Disallow: / User-agent: Scrapy Disallow: / User-agent: dataforseo Disallow: / User-agent: DataForSeoBot Disallow: / User-agent: Magpie-Crawler Disallow: / User-agent: TurnitinBot Disallow: / User-agent: ICC-Crawler Disallow: / User-agent: VelenPublicWebCrawler Disallow: / User-agent: Webz.io Disallow: / User-agent: Kangaroo Bot Disallow: / User-agent: Melonbot Disallow: / User-agent: NovaAct Disallow: / User-agent: Operator Disallow: / User-agent: FriendlyCrawler Disallow: / User-agent: Crawlspace Disallow: / User-agent: Querysact Disallow: / User-agent: DeepSeekBot Disallow: / User-agent: deepseek Disallow: / User-agent: xAI-Bot Disallow: / User-agent: GrokBot Disallow: / User-agent: Amazonbot-Training Disallow: / User-agent: DuckAssistBot-Training Disallow: / # ============================================================================= Sitemap: https://www.giskard.ai/sitemap.xml