Architecture & Observability

AI Crawler Honeypot

The AI Crawler Honeypot is an automated observatory measuring how AI search engines, training scrapers, and autonomous agents interact with web infrastructure. It collects request telemetry, verifies claimed identity against published network ranges, and publishes open, transparent access logs.

Zero Raw IP Retention

Privacy is architected into the system by default. Real client IP addresses are never written to disk or long-term databases. When a request is received, the IP is hashed using SHA-256 combined with a daily rotating salt (salt:{YYYY-MM-DD}). The salt expires automatically after 24 hours and is purged daily at 00:05 UTC, making irreversible reconstruction mathematically impossible.

Vendor CIDR Verification

When a request claims a user-agent string associated with known AI crawlers (such as OpenAI, Anthropic, Google, or Perplexity), the connecting IP is checked against official published CIDR blocks and reverse DNS records. Requests are categorized as:

Robots.txt Conformance & Telemetry

The honeypot passively tracks whether agents respect robots.txt directives without dropping or resetting connections. Observational data is classified for cadence and sequential anomalies, giving researchers clear insight into automated crawler behavior.

Public Logs & Data Access

All activity is made publicly accessible across several views: