# robots.txt for twitterapis.com # # ONE crawl group on purpose. Under RFC 9309 a crawler obeys ONLY its most # specific matching group and never inherits User-agent:*, so every named group # must repeat the whole rule set or it silently unblocks what the star group # blocks. This file used to carry a second group naming 14 crawlers; measured # 2026-08-27 across 280 crawler x path pairs, it changed exactly one thing, and # that one thing was a bug (see below). Do not add a named "Allow: /" stanza. # # Do not add any rule that can match /_next/static/. It serves the JS, CSS and # fonts a renderer needs, and Vercel appends ?dpl= to each one, so a query # rule reaches them even when a path rule does not. A sibling property carried # such a rule and lost 95% of its Googlebot crawl for 25 days (2026-07-31 to # 08-27). Enforced by the shared website-system gate next_build_assets_noindex, # which evaluates the real asset URL shapes rather than one literal rule. User-agent: * Allow: / Disallow: /api/ Disallow: /auth/ Disallow: /dashboard/ Disallow: /payment/ Disallow: /signup Disallow: /login # Query variants of the blog listing (?page=, ?tag=). Inert today: the listing # emits no ?-links and the sitemap carries 0 query URLs, so this blocks nothing # that exists. Kept as a guard in case pagination is added later. It does NOT # touch /blogs or /blogs/, both of which stay crawlable. Disallow: /blogs?* # Content Signals, per https://contentsignals.org # search: allow indexing in search results # ai-train: allow use of content in AI model training # ai-input: allow retrieval for AI-generated answers (RAG / grounding) Content-Signal: search=yes, ai-train=yes, ai-input=yes Sitemap: https://www.twitterapis.com/sitemap.xml