Skip to content
Home Posts About us Privacy Policy
Rate Limiting in Web Scraping Explained: Server Limits, Crawl Rates and Backoff

Rate Limiting in Web Scraping Explained: Server Limits, Crawl Rates and Backoff

Share 𝕏 in f W

Rate limiting is a cap on how many requests a client sends or a server accepts within a period of time. In web scraping the term covers two related things: the limits a website enforces on incoming traffic, and the limits a responsible scraper places on itself. A common assumption is that only the first one matters, as if the site's limit were the speed to aim for. It is the other way around: a well-behaved scraper sets its own pace well below whatever the server would tolerate, so the site never needs to push back.

What rate limiting means on the server side

Websites and APIs limit traffic to protect their capacity for everyone. A server-side limit usually counts requests per client over a window, for example per minute or per hour, and identifies the client by an API key, an account or a network address. When a client goes over the limit, the server rejects further requests for a while.

The standard signal is the HTTP status code 429 Too Many Requests. Servers under general strain may answer 503 Service Unavailable instead. Either response can include a Retry-After header, which gives either a number of seconds to wait or a date and time after which to try again. Official APIs often go further and document their limits, and some return headers showing how many requests remain in the current window. Header names and limits vary from one provider to another, so the provider's developer documentation is the place to confirm them.

A concrete rule follows from this: a 429 or a 503 is an instruction, not an obstacle. The correct response is to stop, wait at least as long as Retry-After says, and come back slower.

What rate limiting means on the scraper side

On the client side, rate limiting is the set of controls a scraper uses to keep its load on a site small. Three settings do most of the work:

  • Request rate: how many requests per second or per minute go to a single host.
  • Concurrency: how many requests are in flight to that host at the same moment.
  • Delay: the pause between one request finishing and the next one starting.

These interact. Suppose, as a hypothetical, a scraper waits two seconds between requests but runs eight parallel workers against the same host. The effective rate is about four requests per second, not one every two seconds. The polite default is one request at a time per host, with a delay of a few seconds, adjusted down if the server responds slowly.

How rate limits vary

Limits come in several shapes, and the shape changes how a scraper should plan its work.

  • Fixed window: a set number of requests per calendar minute or hour, with the counter reset at the start of each window. Bursts at the edge of a window can still strain a server, so spreading requests evenly is kinder than using the whole allowance at once.
  • Sliding window: the count covers the last minute or hour from the current moment, which smooths out those edge bursts.
  • Token bucket: a client gains permission to send requests at a steady rate and can save up a small burst. Many scrapers use the same model internally to pace themselves.
  • Quota: a daily or monthly total, common in official APIs, which matters more for planning a project's scope than for second-by-second pacing.

Limits also vary by endpoint. A search endpoint or a heavy report page usually costs a site more than a static product page, so it deserves a slower pace even when no limit is published.

Related terms: crawl delay, backoff and conditional requests

Crawl delay is a line some sites put in robots.txt, such as Crawl-delay: 10, asking automated clients to wait that many seconds between requests. It is not part of the formal robots.txt standard and search engines treat it differently, but for a scraper it is a direct statement of what the owner wants. Treat it as a minimum. The walkthrough on scalable and reliable social media scraping shows how pacing fits into a larger collection plan.

Backoff is how a client slows down after an error. Exponential backoff doubles the wait after each failed attempt, for example 5, 10, 20 and 40 seconds, up to a ceiling and a maximum number of retries. Adding jitter, a small random variation, keeps many workers from retrying at the same instant. Backoff is for temporary errors such as 429, 503 and timeouts; a 403 Forbidden or a 401 Unauthorized is a refusal, and the answer to a refusal is to stop, not to retry.

Conditional requests reduce load without changing the pace. If a page sent an ETag or Last-Modified header the last time, the scraper can send If-None-Match or If-Modified-Since on the next visit. When nothing changed, the server answers 304 Not Modified with no body, which costs the site far less than a full page.

How rate limiting relates to legal and ethical questions

Pace is one of the clearest ways a scraper shows good faith. Sending requests faster than a site can comfortably handle can degrade service for its real users, and an aggressive load pattern is the kind of conduct that turns a site owner from indifferent to hostile. Website terms of service may also set rules on automated access, including speed, and those terms can matter more than any technical limit.

Rate limiting does not settle whether collection is allowed in the first place. That depends on the terms of service, on copyright in the content, and on privacy laws such as the GDPR and the CCPA when personal data is involved, and the answers differ between jurisdictions. Polite pacing is necessary, not sufficient. For a project with any doubt on those points, a lawyer familiar with data law in the relevant countries is the right person to ask. The article on emerging legal frameworks in social media scraping outlines the questions worth raising.

Common mistakes with rate limits

  • Treating the server's limit as a target. A published limit is a ceiling for the whole client, often shared with other uses. Running at the ceiling leaves no margin when the site is busy. Aim well below it.
  • Retrying immediately after a 429. Instant retries multiply the load at the exact moment the server asked for less. Read Retry-After and wait.
  • Counting delay but ignoring concurrency. As the hypothetical above showed, parallel workers multiply the real rate. Limit concurrency per host, not just per job.
  • Splitting traffic to dodge a limit. Spreading requests across many addresses or accounts to stay under a per-client limit defeats the purpose of the limit and can breach the site's terms. If the allowed pace is too slow for the project, the fix is a smaller scope, an official API, a data agreement with the owner, or a longer schedule.
  • Running at peak hours by default. A job that can run overnight in the site's time zone puts less pressure on the site during its busiest period.
  • No stop condition. A scraper should stop on its own after repeated errors from a host and alert someone, rather than keep hammering a site that is struggling. Managed services, ScrapingIsNotACrime among them, are one way to hand off this pacing and monitoring work, and the same questions apply to any provider: how it paces requests per host and what it does when a site answers 429.
Share 𝕏 in f ✈