MarketLatch · Crawl-Control Diagnostics

Robots.txt Checker

Inspect robots.txt syntax, user-agent groups, Allow/Disallow directives, sitemap references, duplicate entries, and common crawl-control mistakes.

CAPABILITIES

Built for focused, practical work.

D

Directive analysis

Review user-agent groups, Allow, Disallow, sitemap references, and directive placement.

H

HTTP evidence

Inspect response status, final URL, and content type.

I

Implementation warnings

Identify duplicates, invalid sitemap URLs, non-standard directives, and informational unknown directives.

J

JSON + CI friendly

Produce machine-readable output for scripts and automated workflows.

WORKFLOW

From input to useful evidence.

ROBOTS.TXT → FETCH/PARSE → GROUPS → DIRECTIVES → SITEMAPS → WARNINGS → JSON/CLI OUTPUT
SETUP GUIDE

Clear steps, no guesswork.

1

1. Download

Get the source ZIP or clone the repository.

2

2. Install

Python 3.10+ is required; no third-party runtime package is needed.

3

3. Check an authorized resource

Use a site origin, robots.txt URL, or local robots.txt file.

4

4. Interpret carefully

The tool reports observable syntax and HTTP conditions; crawler implementations can differ.

QUICK START

Run it in minutes.

python robots_txt_checker.py https://example.com
python robots_txt_checker.py https://example.com/robots.txt
python robots_txt_checker.py ./robots.txt
python robots_txt_checker.py https://example.com --json
RESOURCES

Source, documentation & downloads.

RELEASE

Stable v0.1.0 snapshot.

SCOPE

Know what the tool can—and cannot—tell you.

This project reports implementation evidence within its defined scope.

It does not guarantee rankings, indexing, traffic, search visibility, or search-engine decisions. Evaluate findings against the actual website, evidence, current authoritative documentation, and project context.

RELATED TOOLS

Explore related MarketLatch tools.