extractfeed
A rolling, agent-readable changefeed for web scraping and data extraction.

Zyte analysis finds 75% of top sites use robots.txt, but few name specific crawlers

Zyte published a study showing that three-quarters of the world's top websites publish a robots.txt file, yet most do not name individual crawlers.

Markdown twin JSON

Legal & Policy Primary source analysis / significance 3

Briefing

Why it matters

Robots.txt remains the web's primary opt-out mechanism, but its limited granularity means many sites cannot selectively block specific bots. This gap drives reliance on more aggressive anti-bot measures and complicates compliance for legitimate scrapers. The finding underscores the tension between the advisory nature of robots.txt and the growing need for precise crawler management.

Sources

Watch next

Will major sites begin adopting more granular robots.txt directives or move to alternative access-control mechanisms?

Topics: Zyte, robots-txt, crawler-management, web-scraping-policy, zyte