A defensible method for separating verified crawler activity from user-agent claims, CDN artifacts, and bot noise

Rod Stockebrand
Co-founder, Brandleap.ai

Key Takeaways
Short on time? Here are the top things to know.
Article framework
Is a GPTBot or Googlebot user-agent enough to identify a crawler?
Does a 200 response prove that a crawler read useful content?
How should robots.txt appear in a crawler dashboard?
How much log data should a team retain?
What makes an AI crawler report defensible?
Teams often ask, “How many AI crawlers visited us last month?” The tempting answer is a user-agent filter and a line chart. That shortcut produces an impressive number, but not necessarily a trustworthy one. A client can send any User-Agent header, a CDN can satisfy a request without contacting the origin, a retry can create several records for one retrieval attempt, and a 200 response can contain a challenge or an empty shell. Measurement has to preserve those distinctions.
The useful question is narrower: which requests claimed to be which agents, which claims could be verified under the provider’s documented method, what did the edge and origin actually deliver, and how did that behavior compare with the published access policy? That produces an operational signal. It does not claim that a model trained on a page, indexed it, cited it, or used it in an answer—none of those outcomes is visible in an ordinary HTTP access log.
Separate observed request, verified identity, policy decision, delivered response, and downstream use. Conflating those five layers is the root of most crawler-reporting errors.
Before writing a query, inventory where a request can be recorded: CDN or edge, load balancer, reverse proxy, web server, application gateway, and observability pipeline. Each layer answers a different question. Edge logs can show a request that was served from cache or blocked before origin. Origin logs can show what reached the application, but they cannot count a cache hit that never traveled inward. Firewall logs may show a rejected connection with no HTTP status at all.
Do not join edge and origin data by timestamp alone. Concurrent requests, retries, clock skew, and connection reuse make that unreliable. Prefer a propagated request ID. If none exists, reconcile conservatively using a bounded time window, hostname, method, path, status, and a coarse source token, and label the result as an estimate. Never silently present a many-to-one join as a precise request count.
Preserve the raw User-Agent string for forensic work, but classify it into stable families for reporting: named search crawler, named AI crawler, browser or app, generic bot, monitoring client, and unknown. Keep provider purpose separate when documentation distinguishes retrieval, search discovery, training, or user-triggered fetches. “AI crawler” is not a protocol field; it is an interpretation that must be backed by a versioned pattern list.
raw_user_agent -> parsed_tokens -> claimed_provider -> claimed_purpose source_ip -> documented_verification -> identity_status request -> robots_decision + response_class -> measurement_event
Use exact or carefully bounded patterns rather than a substring such as “bot.” Preserve unknown strings in a review queue. Provider strings change, product names overlap, and a legitimate browser can contain automation-related tokens. Store the rule ID and version used to classify each event so a future taxonomy update does not rewrite history without explanation.
Reverse DNS can turn an IP address into a hostname, but the reverse lookup alone is not proof: the owner of an address can configure a misleading PTR record. Google’s documented verification procedure performs reverse DNS and then forward DNS, checking that the resulting address maps back to the original IP or uses its published ranges. Other providers document different approaches. Follow the provider’s current official instructions; do not invent a universal suffix check.
Verification is time-sensitive. DNS can change, proxy infrastructure can be shared, IPv4 and IPv6 can follow different paths, and the address seen at the origin may be a trusted proxy rather than the crawler. Store the observed address, lookup time, resolver context if relevant, verification method, result, and documentation URL. A failed lookup should be “unverified,” not automatically “malicious.”
Identity labels that should not be collapsed
✗ Un-optimized
Claimed: the request’s user-agent matches a documented token. Useful for policy and triage, but easy to spoof.
✓ Triple-rich rewrite
Verified: the source address passes the provider’s documented check at the recorded time. Stronger evidence, still not proof of content use or downstream citation.
RFC 9309 describes robots.txt as a convention for cooperating crawlers. A dashboard should fetch and version the policy from each public hostname, parse the effective group for the claimed user agent, and evaluate the requested path. Then compare that decision with what the request did. Count “allowed request,” “disallowed request,” “unknown policy,” and “not evaluated” separately.
A Disallow request from an unverified client is not proof that a named provider violated your policy. It may be a scraper copying the string, a stale policy cache, an alternate hostname, or a parser mismatch. Likewise, an absence of requests does not prove a provider respected the rule; it could reflect no demand, cache behavior, rate limits, or a different acquisition path.
HTTP status is a valuable first dimension, not a success metric by itself. A 2xx response says the server accepted the request; it does not say the body was complete or useful. A 3xx may be a normal canonical redirect or a loop. A 403 may be an intentional policy decision, a WAF false positive, or a missing allowlist. A 429 reveals rate limiting, while a 5xx can indicate an origin problem that affects users and crawlers alike.
SELECT claimed_provider, identity_status, response_class,
cache_status, robots_decision, COUNT(*) AS requests
FROM crawler_events
WHERE observed_at >= :start AND observed_at < :end
GROUP BY 1, 2, 3, 4, 5;Add content signals where privacy and cost permit: content type, response size bucket, compression, cache status, and a non-reversible content fingerprint. A fingerprint can reveal that repeated 200 responses are identical without retaining the body. For a controlled sample, retain a short-lived body excerpt or encrypted object with strict access controls so engineers can distinguish an article from a challenge page. Never infer readability from bytes alone.
Spoofing is not an edge case. Scrapers, vulnerability scanners, browser extensions, uptime monitors, and internal tests can claim to be a famous crawler. Some traffic also passes through shared infrastructure where IP identity is not enough to establish a particular product. Report claimed and verified identity side by side, and make “unverified claim” a useful queue rather than forcing every event into good or bad.
High-volume properties may need to sample. Sampling every tenth request is simple but can miss short bursts and underrepresent rare agents. Prefer a documented strategy: deterministic sampling by a stable request key for broad trend analysis, plus full capture for errors, policy violations, unknown identities, and a rotating sample of successful requests. Store the sampling rate and use weights when calculating estimated totals.
Do not mix sampled and unsampled populations in one KPI. A count of verified 403 events captured in full is not comparable to an estimated count of all 200 events. Show numerator, denominator, sampling rate, and confidence or estimation notes. If the question is “which URLs failed,” a targeted error sample is more useful than a statistically neat estimate of every request.
Access logs can contain IP addresses, query strings, identifiers, referrers, cookies, and paths that reveal sensitive interests. Collect the minimum fields needed for the stated analysis. Remove credentials and session tokens at ingestion, exclude query parameters that are not needed, truncate or keyed-hash addresses where exact verification is unnecessary, and restrict raw records separately from aggregate dashboards.
Write a retention schedule and deletion job. Document who may resolve a privacy-preserving token, when a sampled body is destroyed, and how incident access is audited. Aggregation is not anonymization by magic: a rare URL, timestamp, and address combination may still identify a person or organization. Consult the applicable privacy and security owners before expanding collection for an attractive chart.
A defensible dashboard starts with a definition panel: data sources, coverage window, timezone, normalization rules, taxonomy version, verification date, sampling policy, and exclusions. Give every headline a path to its underlying evidence. “AI requests” should expand into claimed provider, verification status, edge-versus-origin count, robots outcome, status class, and a representative URL sample.
Use medians and percentiles for latency, not only averages. Break out cache hits from origin fetches. Compare the same weekday windows when seasonality matters. Alert on abrupt changes in verified identity, status mix, robots fetch health, and reconciliation—not merely on a rising “AI bot” line. A spike may be useful discovery, a retry storm, a broken cache, or a spoofing campaign.
The most important conclusion may be “we cannot tell yet.” That is a successful measurement outcome when the alternative is a confident but false attribution. Add the missing request ID, document the provider’s verification rule, or capture the edge field required to resolve the uncertainty. Better instrumentation compounds; invented certainty does not.
Finally, keep the scope honest. Server logs can show requests, delivery, policy context, and operational failures. They cannot prove that an AI system retained a page, ranked it, cited it, or generated an answer from it. Pair log evidence with controlled retrieval tests and citation monitoring when you need those downstream questions. The result is a complete measurement program rather than a bot counter dressed up as attribution.
Primary references: Google Search Central crawler-verification guidance; OpenAI crawler documentation; RFC 9309 for robots behavior; RFC 9110 for HTTP semantics; and your CDN’s official request-log schema.