Crawl Rate Limiting and Robots.txt (Complete 2026 Guide)

Rate limits cap bot traffic at the origin. The cap protects bandwidth, CPU, and database connections from being exhausted by any single crawler, including...

Dilshad Akhtar
Dilshad Akhtar
Published: 10 June 2026
3 min read
TL;DRAI summary
  • Rate limits cap bot traffic at the origin.
  • Google does not honor Crawl-delay for Googlebot.
  • Deploy-day rate-limit verification, subsequent to a CDN configuration change.

Rate limits cap bot traffic at the origin. The cap protects bandwidth, CPU, and database connections from being exhausted by any single crawler, including Googlebot, GPTBot, or ClaudeBot. Without the cap, training crawlers demonstrate no upper bound on resource consumption. robots.txt is the...

What crawl rate limiting does

Rate limits cap bot traffic at the origin. The cap protects bandwidth, CPU, and database connections from being exhausted by any single crawler, including Googlebot, GPTBot, or ClaudeBot. Without the cap, training crawlers demonstrate no upper bound on resource consumption.

robots.txt is the negotiation layer. The file tells crawlers which paths to fetch and at what pace, sitting upstream of any enforcement at the CDN or WAF. The two layers must agree or the rule is purely advisory.

The majority of teams tune robots.txt and ignore the enforcement layer. The file says "Crawl-delay: 10" and the CDN serves unlimited requests anyway. The robots.txt alone does not throttle traffic in 2026. The CDN or origin does the actual enforcement.

How Crawl-delay directives behave

Google does not honor Crawl-delay for Googlebot. Per Google Search Central's crawling documentation, Googlebot utilizes dynamic throttling based on site response times, not a static directive (https://developers.google.com/search/docs/crawling-indexing/overview).

Bingbot honors Crawl-delay. Per Microsoft Bing Webmaster Tools documentation, Bingbot reads the directive and paces requests accordingly (https://www.bing.com/webmasters/help/which-robots-crawler-6302fc2e). OpenAI's GPTBot and Anthropic's ClaudeBot also honor the directive.

The Crawl-delay value is in seconds. A value of 10 means one request per 10 seconds per bot instance. The majority of training crawlers ignore values below 5 and treat anything above 30 as a soft suggestion.

The directive is rarely the rate-limiting tool of record. Effective rate-limiting in 2026 happens at the CDN layer.

Cloudflare, Fastly, and Akamai enforce per-bot request quotas at the edge, prior to requests reaching the origin. Robots.txt talks. The CDN enforces.

The 2025-2026 crawl ceiling

Cloudflare's AI Labyrinth write-up tallied more than 50 billion AI crawler requests per day across its network (https://blog.cloudflare.com/ai-labyrinth/). That figure roughly doubled through 2026 as GPTBot, ClaudeBot, and a dozen vertical bots joined the pool.

Googlebot crawled 1.70 times more unique URLs than ClaudeBot and 1.76 times more than GPTBot in early 2026, per Digital Applied's June 2026 access-control matrix (https://www.digitalapplied.com/blog/ai-crawler-access-control-2026-robots-llms-txt-decision-matrix). The crawl-volume gap narrows each quarter, but Google still leads the pack.

A median mid-sized publisher with 10,000 URLs sees 80,000 to 250,000 crawler requests per day across Googlebot, Bingbot, GPTBot, ClaudeBot, and CCBot combined.

Origin CPU load spikes when a training crawler hits a heavy path. Setting a per-bot 429 ceiling at the CDN now drops requests faster than any robots.txt edit ever will.

The crawl rate test

Deploy-day rate-limit verification, subsequent to a CDN configuration change. You pull the request log and group hits by user-agent. You check the per-bot 429 rate against the pre-deploy baseline. A spike above 3x confirms the new rule is engaging.

You execute a synthetic load test against the slowest crawler route on the property. You watch origin CPU and database pool saturation. You record the breaking point and set the ceiling 15% below it. Friday's traffic survives.

You document the ceiling in the robots.txt comment block. You link the rate-limit decision to the deployment that introduced it. The audit trail answers next quarter's "why is GPTBot at 10% of last month?" question in one click.

Note the gap. This post synthesizes 2025 and 2026 data from five sources: Google Search Central, Microsoft Bing Webmaster Tools, Cloudflare, Digital Applied, and Search Engine Land.

CDN-specific per-bot 429 ceilings remain a vendor configuration surface. Replication required on your own edge prior to declaring done. Rate-limit decisions protect the origin and preserve crawl budget. Audit subsequent to every CDN change.

Ready to Build Your Dream Website?

Let's discuss your project and create something amazing together.