Methodology

This page documents how the AI Crawler Honeypot works. It describes the verification pipeline, the logging schema, the privacy architecture, and the public record.

— What this is —

A public AI crawler observation honeypot. It logs HTTP requests to this domain, verifies claimed identity against published vendor infrastructure, and publishes a public record.

— Zero raw IP retention —

Real client IP addresses are never written to disk or long-term databases.

When a request is received:

  1. The source IP is truncated to /24 (IPv4) or /48 (IPv6).
  2. The full IP is hashed using SHA-256 combined with a daily rotating salt.
  3. The salt is stored at salt:{YYYY-MM-DD} with a 24-hour TTL.
  4. The salt is purged automatically at 00:05 UTC each day.

After 24 hours, the salt no longer exists. The hash cannot be reversed. Historic entries remain intact with their computed hash, but no link to the original IP is retained.

— Vendor CIDR verification —

When a request claims a user-agent associated with a known AI crawler, the connecting IP is checked against the vendor's published infrastructure.

Categories:

— Robots.txt conformance —

The honeypot tracks whether agents respect robots.txt directives. It does not drop or reset connections. It serves a 200 response to every request, including those to disallowed paths.

Requests to disallowed paths are logged with the matching rule. This allows the record to show which agents respected the rules and which did not.

— DeepSeek behavioural classification —

Requests to disallowed paths and requests to /ask are classified using DeepSeek. The classification records:

Compliant requests (robots_rule == "none") do not call DeepSeek. They receive a stub classification: unclassified — skipped — routine traffic.

— Log schema —

Every entry contains the following fields:

No other fields are added. The schema is stable.

— Data retention —

Public log (log:public:{date}): 24-hour TTL. Entries expire automatically.

Archive (log:archive:{timestamp}:{uuid}): permanent, no TTL. Retained for 12 months, then reviewed under the retention policy at /privacy.

Salt (salt:{date}): 24-hour TTL. Purged daily at 00:05 UTC.

— Public views —

The record is published at the following paths:

— What the record does not do —

No judgments. No scores. No labels like “safe,” “unsafe,” “malicious,” or “threat.” Every entry is a factual field/value record.

No claims about whether any agent is safe or unsafe. No claims about intent. Verification describes the source network, not the agent's purpose.

— Contact —

Corrections, removal requests, and general enquiries: vic@fixseo.uk