# Zyte analysis finds 75% of top sites use robots.txt, but few name specific crawlers

> Zyte published a study showing that three-quarters of the world's top websites publish a robots.txt file, yet most do not name individual crawlers.

- Canonical: https://extractfeed.io/story/zyte-analysis-finds-75-of-top-sites-use-robots-txt-but-few-n-e5e97be/
- JSON: https://extractfeed.io/api/v1/stories/zyte-analysis-finds-75-of-top-sites-use-robots-txt-but-few-n-e5e97be.json
- Beat: Legal & Policy · Evidence: Primary source · Type/significance: analysis/3 · First seen: 2026-09-05T14:57:46.592713+00:00 · Updated: 2026-09-05T14:57:46.592713+00:00 · Edition: 2026-09-06
- Framing: model-written (headline, standfirst, why it matters, tags); source facts deterministic

## Briefing
- Zyte: Zyte published a study showing that three-quarters of the world's top websites publish a robots.txt file, yet most do not name individual crawlers.

## Why it matters
Robots.txt remains the web's primary opt-out mechanism, but its limited granularity means many sites cannot selectively block specific bots. This gap drives reliance on more aggressive anti-bot measures and complicates compliance for legitimate scrapers. The finding underscores the tension between the advisory nature of robots.txt and the growing need for precise crawler management.

## Sources
- Primary source · Zyte · 2026-09-02 — [75% of the web uses robots.txt - here's how](https://www.zyte.com/blog/robots-txt-overview/)

## Watch next
Will major sites begin adopting more granular robots.txt directives or move to alternative access-control mechanisms?

Topics: Zyte, robots-txt, crawler-management, web-scraping-policy, zyte

---
extractfeed is an agent-readable changefeed for web scraping and data extraction. Index: https://extractfeed.io/agents.md
