Vai al contenuto principale
Torna al wiki
Web & Crawling

What Is robots.txt?

Ultimo aggiornamento il 4 agosto 2026

robots.txt is a plain-text file at the root of a website that declares which automated crawlers may access which paths. It implements the Robots Exclusion Protocol, invented in 1994 and formalized as RFC 9309 in 2022, with a deliberately small grammar: user-agent lines name a crawler, allow and disallow rules scope its access, and a sitemap line points to the site's URL inventory. Nearly every site has one, every major crawler reads it, and its single most misunderstood property is stated in its own specification: compliance is entirely voluntary. The file is a posted sign, and signs bind only those who choose to read them.

What the file actually controls

For cooperative crawlers, robots.txt is genuinely load-bearing. Search engines use it to keep out of infinite calendars, faceted-search explosions, and staging paths; site owners use it to shape crawl budget toward pages that matter. Its newest and fastest-growing role is consent signaling for AI crawlers: model operators publish user-agent tokens that sites can disallow to opt out of training collection, making a 30-year-old convention the de-facto interface for one of the web's most contested questions. Two technical corrections prevent common self-inflicted wounds: a disallowed page can still appear in search results if other sites link it (blocking crawling is different from blocking indexing, which needs noindex), and robots.txt is public, so listing secret admin paths in it hands every attacker a map of exactly where to look.

The enforcement gap

The protocol's honor-system design cleanly splits the bot world. Good bots identify themselves, honor the rules, and can be verified against published IP ranges. Hostile automation, such as scrapers, credential testers, and inventory bots, reads the same file and, at most, treats it as reconnaissance. Some crawl operators have also been caught fetching disallowed paths or crawling under generic browser identities, which is why "we have a robots.txt" appears in web scraping disputes as evidence of stated policy rather than as a control that stopped anything. The file defines what unauthorized means; it contributes nothing to preventing it.

Pairing policy with enforcement

A coherent crawl policy therefore has two layers that do different jobs. robots.txt states the rules in the standard location, such as search directives, AI-training opt-outs, and crawl-delay hints, giving cooperative operators everything needed to comply and creating the paper trail for disputes with the rest. Enforcement then falls to bot detection: verifying that self-declared crawlers originate from their operators' networks, and confronting undeclared automation with behavioral verification such as CaptchaFox at the endpoints worth defending. Sites that skip the first layer forfeit cooperation they could have had for free; sites that stop at the first layer protect themselves precisely against the bots that were never the problem.

Informazioni su CaptchaFox

CaptchaFox è una soluzione conforme al GDPR con sede in Germania che protegge siti web e applicazioni da abusi automatizzati, come bot e spam. Il suo approccio distintivo e multilivello utilizza segnali di rischio e sfide crittografiche per facilitare un processo di verifica robusto. CaptchaFox consente ai clienti di essere operativi in pochi minuti, non richiede gestione continua e offre alle aziende una protezione duratura.

Per saperne di più su CaptchaFox, contattaci o inizia a integrare la nostra soluzione con una prova gratuita.

Termini correlati

What Is Web Scraping?

Web scraping is the automated extraction of data from websites. It powers search engines and price comparison, along with content theft and competitive abuse.

Continua a leggere
What Is a Screen Reader?

A screen reader converts on-screen content into speech or braille for blind and low-vision users, and a common failure point for CAPTCHA and verification flows.

Continua a leggere
What Is EN 301 549?

EN 301 549 is the European accessibility standard for ICT products and services, the technical spec that laws like the EAA point to for what "accessible" means.

Continua a leggere
What Is the BFSG?

The BFSG is Germany's Accessibility Strengthening Act, the national law implementing the European Accessibility Act, in force since 28 June 2025.

Continua a leggere

Combatti i bot e proteggi i dati dei tuoi utenti.

Non dare ai truffatori e agli spammer alcuna possibilità e proteggi il tuo sito web con CaptchaFox oggi.

CaptchaFox protegge i siti web su desktop e dispositivi mobili