Skip to content
Home Posts About us Privacy Policy
How to Read a robots.txt File and Apply Its Rules Before You Scrape a Website

How to Read a robots.txt File and Apply Its Rules Before You Scrape a Website

Share 𝕏 in f W

Reading robots.txt takes a few minutes and answers one practical question: which paths on this host the site owner has asked automated clients not to request. By the end of the steps below you will have a short, written list of allowed and disallowed paths for your own user agent, a crawl pace, and a rule for re-checking the file. The file is a request from the site owner, not a legal permission slip, so it is the first check before a project, not the only one.

Step 1: Decide on an honest user agent name

Before you can read the rules that apply to you, you need to know who you are in the file's terms. robots.txt groups its rules by user agent, and your scraper should send a stable, descriptive one, such as a product name plus a contact page or email address, so the site owner can identify the traffic and reach you.

Write down two names: the full user agent string your client sends, and the short token you will match against the file, for example acme-price-monitor. Matching uses the product token, and comparisons are case-insensitive. The common mistake here is borrowing a browser's user agent string. That makes your traffic harder to identify, and it also means you cannot honestly say which group of rules applies to you.

Step 2: Fetch the file from the root of the exact host

robots.txt always lives at the root of a host, at the path /robots.txt. The rules apply only to that combination of protocol, host name and port. A file on a site's main host says nothing about its shop or api subdomains; each subdomain has its own file, or none.

Request the file the same way your scraper will request pages, with the same user agent, and note the HTTP status code you get back. The status matters in Step 3, so record it alongside the date you fetched the file. A useful check: open the same address in a normal browser. If the browser shows a file and your client gets something else, your client is not requesting what you think it is, and you should fix that before going further.

Step 3: Interpret the status code before reading any rules

What the server returns for /robots.txt changes how you should behave, even before you look at the content. The current standard for the Robots Exclusion Protocol, published as RFC 9309, describes the generally accepted behavior:

  • 200 OK: parse the file and follow it.
  • A 4xx status such as 404: the file is unavailable, and crawlers may treat the site as having no robots restrictions. Your other obligations, like the terms of service and polite request rates, still apply.
  • A 5xx status or a network failure: the file is unreachable, and the standard says to assume everything is disallowed until you can fetch it.
  • Redirects: follow a small number of them; the standard suggests at least five hops before giving up.

The mistake to avoid is treating a server error as permission. A 503 on robots.txt often means the site is already under load, which is the worst moment to start sending requests.

Step 4: Find the group that applies to your user agent

A robots.txt file is a series of groups. Each group starts with one or more User-agent lines, followed by Allow and Disallow lines. Read the file top to bottom and look for a group whose User-agent line matches your product token from Step 1.

The selection rule is simple once you see it: if a group names your token, you follow that group and ignore the wildcard group marked with an asterisk. Only if no group names you do you follow the asterisk group. If several groups name your token, their rules are combined. If nothing matches at all and there is no asterisk group, no rules apply to you.

Consider a hypothetical file with three groups: one for a search engine crawler, one for acme-price-monitor, and one for the asterisk. Your scraper follows only the second group, even if the asterisk group is stricter or looser. A common error is to merge your named group with the asterisk group, which can produce rules the site owner never wrote.

Step 5: Resolve Allow and Disallow for each path you plan to request

Inside your group, each Allow or Disallow line holds a path prefix. To decide whether a URL is allowed, compare its path against every rule in the group and pick the rule with the longest matching path. If an Allow and a Disallow rule match with the same length, Allow wins. If no rule matches, the path is allowed.

Work through a hypothetical example. Your group contains Disallow: /products/ and Allow: /products/public/. A request for /products/public/shoes-123 matches both rules, but the Allow rule is longer, so the request is allowed. A request for /products/internal/report matches only the Disallow rule, so it is not allowed.

Two characters change the matching. An asterisk matches any sequence of characters, so Disallow: /*?sort= blocks any URL whose path and query contain that sort parameter. A dollar sign anchors the end of the path, so Disallow: /*.pdf$ blocks addresses that end in .pdf. Also remember that matching is case-sensitive for paths, and that a line reading Disallow: with nothing after it disallows nothing.

Write the outcome down as a short table: the URL patterns your project needs, and whether each is allowed. If a pattern you need is disallowed, that is the end of the question for that path. Look for an official API, a data feed or a direct agreement with the site owner instead.

Step 6: Check the parser your code relies on

Most languages have a robots.txt parser, and the usual pattern is the same: load the file once, then ask a can-fetch question with your user agent token and a full URL before each request. Python's standard library, for example, includes a parser in the urllib.robotparser module with a can_fetch method and helpers that read Crawl-delay and Request-rate lines.

Parsers do not all implement the standard the same way. Some ignore the asterisk and dollar-sign wildcards or resolve conflicts differently. A verifiable check: take the table from Step 5 and run each URL through your parser. Every answer should match what you worked out by hand. Where it does not, either switch to a parser that follows RFC 9309 or add your own check for those patterns, and keep the test URLs so you can rerun them after library updates.

Step 7: Turn Crawl-delay and Sitemap lines into crawl settings

Some files include a Crawl-delay line, for example Crawl-delay: 10. It is not part of RFC 9309 and search engines treat it differently, but it is a clear signal of how fast the owner wants automated clients to go. When it appears in your group, use it as the minimum gap in seconds between your requests to that host, and do not run parallel requests that add up to more.

When there is no Crawl-delay, choose a conservative pace yourself: one request at a time per host, a pause of a few seconds, and a slower schedule outside the site's busiest hours. If the server answers with HTTP 429 Too Many Requests or 503, honor any Retry-After header and back off further. The best practices for scalable and reliable social media scraping apply the same logic to platforms with heavier traffic.

Sitemap lines point to XML sitemaps that list the site's public URLs. They are often a cleaner way to find pages than following every link, and they reduce the number of requests you send. Note that the Sitemap line applies to the whole file, not just one group.

Step 8: Check the terms of service and the data you will collect

robots.txt answers which paths automated clients are asked to avoid. It does not answer whether the site's terms of service allow automated collection, whether the content is protected by copyright, or whether the pages contain personal data covered by privacy laws such as the GDPR or the CCPA. Read the terms, note anything about automated access or data reuse, and decide which fields you actually need.

A concrete test: list every field in your planned output. Any field that identifies a person, such as a name, handle, email or photo, needs a documented reason and a plan for storage and deletion. Rules differ by country and by use, so for anything beyond clearly non-personal public data, confirm the plan with a lawyer before collection starts. The overview of ethical web scraping and data privacy standards covers the questions to bring to that conversation.

Step 9: Cache the file and re-check it on a schedule

Site owners change robots.txt, and a rule that allowed a path last month may disallow it today. RFC 9309 says crawlers should not use a cached copy for more than 24 hours unless the file is unreachable. In practice, fetch the file at the start of every run, and at least once a day for jobs that run continuously.

Store each version with its fetch date and compare it to the previous one. When the file changes, rerun the table from Step 5 and pause any job whose paths became disallowed. If you rely on a managed service such as ScrapingIsNotACrime instead of running your own crawler, ask how it handles robots.txt changes and what it does when a path you depend on becomes disallowed. Keep the stored versions with your project notes: they show what the site asked for on the day you collected the data, which is useful when someone later asks how the dataset was built. For more on keeping collection runs accurate over time, see these tips for enhancing data precision in large-scale internet scraping projects.

Share 𝕏 in f ✈