# ============================================ # 常见搜索引擎 → 允许 # ============================================ User-agent: Baiduspider Allow: / User-agent: Baiduspider-render Allow: / User-agent: Googlebot Allow: / User-agent: Googlebot-Image Allow: / User-agent: bingbot Allow: / User-agent: msnbot Allow: / User-agent: Bingpreview Allow: / User-agent: Yisouspider Allow: / User-agent: ShenmaSpider Allow: / User-agent: YandexBot Allow: / User-agent: DuckDuckBot Allow: / User-agent: Sogou web spider‌ Allow: / User-agent: Sogou Pic Spider Allow: / User-agent: Sogou inst spider Allow: / User-agent: 360Spider Allow: / User-agent: HaosouSpider Allow: / User-agent: Applebot Allow: / # ------- OpenAI (ChatGPT) ------- # GPTBot:纯训练爬虫 → 屏蔽 User-agent: GPTBot Disallow: / # OAI-SearchBot:ChatGPT搜索索引,带引用链接 → 允许 User-agent: OAI-SearchBot Allow: / # ChatGPT-User:用户主动要求ChatGPT读取某页面 → 允许 User-agent: ChatGPT-User Allow: / # OAI-AdsBot:广告落地页安全检测,非训练/非引用,一般允许 User-agent: OAI-AdsBot Allow: / # ------- Anthropic (Claude) ------- # ClaudeBot:批量训练爬虫 → 屏蔽 User-agent: ClaudeBot Disallow: / # anthropic-ai:旧版训练token,历史遗留 → 屏蔽 User-agent: anthropic-ai Disallow: / # Claude-SearchBot:为Claude网页搜索建立索引,带引用 → 允许 User-agent: Claude-SearchBot Allow: / # Claude-User:用户让Claude读取某具体页面 → 允许 User-agent: Claude-User Allow: / # ------- Perplexity ------- # PerplexityBot:搜索索引,带链接引用 → 允许 User-agent: PerplexityBot Allow: / # Perplexity-User:用户提问触发的实时抓取 → 允许 # 注意:Perplexity-User通常不遵守robots.txt,这里只是声明立场 User-agent: Perplexity-User Allow: / # ------- Google ------- # Google-Extended:控制内容是否用于Gemini/AI Overviews训练 → 屏蔽训练用途 # 注:常规Googlebot索引(含AI Overviews引用来源)不受此token影响,不建议屏蔽Googlebot本身 User-agent: Google-Extended Disallow: / # ------- Microsoft (Bing / Copilot) ------- # BingBot 同时承担搜索索引与Copilot引用来源,一般建议允许 User-agent: bingbot Allow: / #专为DuckDuckGo的Assist功能服务,实时抓取网页以生成AI答案,并且会标注信息来源 User-agent: DuckAssistBot Allow: / # ------- Amazon ------- # Amazonbot:支持Rufus/Alexa+等产品的检索式引用 → 允许 User-agent: Amazonbot Disallow: / # ------- Meta ------- # meta-externalagent:抓取内容用于Meta AI回答及训练,二者未分离token # 如果只想屏蔽训练可考虑屏蔽此项,这里按"训练相关"归类为屏蔽 User-agent: meta-externalagent Disallow: / User-agent: Meta-ExternalFetcher Disallow: / # 深度求索 (DeepSeek 语料抓取) User-agent: DeepSeekBot Disallow: / # 阿里大模型抓取 (保留 Yisouspider 用于夸克搜索,单独屏蔽 AliyunSpider 训练节点) User-agent: AliyunSpider Disallow: / User-agent: FacebookBot Disallow: / # ------- 常见纯训练型爬虫(非搜索/引用用途)→ 全部屏蔽 ------- User-agent: CCBot Disallow: / #User-agent: Bytespider #Disallow: / User-agent: Applebot-Extended Disallow: / User-agent: Diffbot Disallow: / User-agent: ImagesiftBot Disallow: / User-agent: Cohere-ai Disallow: / User-agent: Omgilibot Disallow: / # ============================================ # 常见垃圾 / SEO抓取 / 采集类爬虫 → 屏蔽 # ============================================ User-agent: AhrefsBot Disallow: / User-agent: SemrushBot Disallow: / User-agent: SemrushBot-SA Disallow: / User-agent: MJ12bot Disallow: / User-agent: DotBot Disallow: / User-agent: BLEXBot Disallow: / User-agent: SEOkicks-Robot Disallow: / User-agent: dataforseo Disallow: / User-agent: MegaIndex Disallow: / User-agent: PetalBot Disallow: / User-agent: SeznamBot Disallow: / User-agent: ZoominfoBot Disallow: / User-agent: Scrapy Disallow: / User-agent: HTTrack Disallow: / User-agent: Nimbostratus-Bot Disallow: / User-agent: serpstatbot Disallow: / User-agent: BLEXBot Disallow: / User-agent: DataForSeoBot Disallow: / User-agent: SeekportBot Disallow: / Sitemap: https://www.DissertationTopic.Net/sitemap/sitemap.xml