Agent Analytics (AI Bot Crawler Tracking)
Before an AI model can cite your site, its crawlers have to visit it. Agent Analytics is server-side AI bot traffic monitoring: it shows which AI crawlers (GPTBot, PerplexityBot, ClaudeBot, OAI-SearchBot, Google-Extended, and many more) are hitting your website and how that evolves over time.
With Agent Analytics, you can:
- See which AI bots (GPTBot, PerplexityBot, ClaudeBot, OAI-SearchBot, Google-Extended) crawl your site and how often
- Confirm that AI models are actually discovering new or updated pages before you expect them in AI answers
- Compare crawl volume by company (OpenAI, Anthropic, Google, Perplexity, and more) to see who is investing in your content
- Catch a drop in AI crawler activity that could explain falling citations
Find it under Bots & Traffic > Agent Analytics in the sidebar. Available on Scale plans and above.
In embed and full-whitelabel portals, the client connects and manages their own sources: Cloudflare, Bunny CDN, an Amazon S3 or CloudFront log bucket, and CSV or log uploads. Log drains are also available when the portal is served on the client's own domain, because a drain hands out an ingest URL and that URL has to sit on a host the client recognises. Asking us to support a provider we do not cover stays in the main workspace. Team member permissions apply as everywhere else, so a portal user who cannot create connections still sees the report and is told to contact their account manager.
What it does
- Tracks over 40 AI bots grouped by company (OpenAI, Anthropic, Google, Perplexity, Apple, Meta, Microsoft, Amazon, and others)
- Daily traffic charts per bot and per company, so you can see whether AI crawlers are discovering your content
- Multiple ways to get your data in:
- Cloudflare integration: paste a scoped API token (Zone Analytics: Read), pick a zone, and data syncs daily. Works on every Cloudflare plan, including Free.
- Bunny CDN integration: paste your account API key and pick a pull zone; data syncs daily.
- Log drains (real time): Vercel, Akamai DataStream 2, Fastly, Netlify, and GCP Load Balancer. Generate a unique ingest URL in the app and paste it into your provider's log drain settings.
- AWS CloudFront / Amazon S3 bucket: connect the S3 bucket where CloudFront or your server writes access logs and we pull new files once a day (plus a manual "Sync now"). CloudFront standard W3C logs work directly. S3-compatible stores such as Cloudflare R2 and MinIO work too.
- CSV / log upload: drop a Cloudflare CSV export, an AWS CloudFront access log, an Apache or nginx log, or any CSV with timestamps and user agents. Gzipped files work too.
- Enterprise accounts also get a Per URL tab showing exactly which pages each AI bot crawled, with hits, status codes, and first/last-seen dates, plus CSV export
How to use it
- Open Bots & Traffic > Agent Analytics
- Connect a data source (Cloudflare is the fastest if your site is behind it)
- Give the first sync a little time, then review which AI bots visit you and how often
- Watch the trend after publishing new content: rising AI crawler traffic on your site is an early signal that models are picking it up
Set up Vercel
In Vercel, go to Team Settings > Drains > Add Drain and choose Logs as the data type. Select the projects you want to track and include Production. Enable the Static, Lambda, Edge, External, Firewall, and Redirect log sources. Leave Build off because it contains deployment output, not visitor traffic. Do not add sampling rules, so Vercel forwards 100% of requests.
Use Custom Endpoint with NDJSON and paste the LLM Pulse endpoint URL. Web Analytics, Speed Insights, and Traces are not supported because they do not include the request user agent needed to identify AI bots. The Signature Verification Secret is optional in LLM Pulse. Copy Vercel's generated value into Signing secret only if you want to verify signed requests. Enter a verification code in LLM Pulse only if Vercel shows one.
Connect AWS CloudFront or another Amazon S3 log bucket
If CloudFront, your server, or another CDN already sends access logs to an S3 bucket, you can let LLM Pulse read them directly instead of uploading files by hand.
What it is: a read-only connection to a bucket of log files. Once a day (plus a manual "Sync now") we list the bucket, pick up only the log files added since the last sync, parse them, and match each request against the AI bot catalog. It also works with S3-compatible stores such as Cloudflare R2 and MinIO through an optional custom endpoint URL.
CloudFront format: legacy CloudFront standard logs sent to S3 already use the supported W3C, tab-separated, gzip-compressed format. For standard logging v2, choose Amazon S3 as the destination, W3C as the output format, and Tab as the field delimiter. Keep date, time, cs(User-Agent), cs-uri-stem, and sc-status. We also recommend sc-bytes for bandwidth reporting. JSON, Plain, Raw, and Parquet output are not supported.
Before you start: create an IAM user with a read-only policy that allows s3:GetObject and s3:ListBucket on the bucket (and prefix, if you use one), then generate an access key ID and secret access key for it. A minimal policy looks like this:
{
"Version": "2012-10-17",
"Statement": [
{
"Effect": "Allow",
"Action": ["s3:GetObject", "s3:ListBucket"],
"Resource": [
"arn:aws:s3:::your-log-bucket",
"arn:aws:s3:::your-log-bucket/*"
]
}
]
}
Set it up:
- Open Bots & Traffic > Agent Analytics
- On the AWS CloudFront card, click Connect via Amazon S3. For another S3 log source, use the Amazon S3 card. Enter your bucket name, AWS region, and (optionally) a key prefix so we only read the folder that holds your logs
- Paste the read-only access key ID and secret access key. For Cloudflare R2, MinIO, or another S3-compatible store, also add the custom endpoint URL
- Click Connect. Your credentials are stored encrypted, and the first sync starts automatically. Use Sync now any time you want to pull the latest files without waiting for the daily run
Sync cadence and limits: we sync once a day and on demand via "Sync now", reading only files that are new since the last sync. We ingest log entries up to 90 days old from files up to 256 MB each. We process the files in batches of up to 500 and start another batch automatically until the bucket catches up. We scan up to 50,000 objects in the configured bucket or prefix; use a narrower prefix if it contains more.
Supported log formats: Apache and nginx combined (CLF) logs, AWS CloudFront standard (W3C) logs, and CSV log exports. Gzipped (.gz) files are read directly.
Tips & notes
- Agent Analytics is the inverse of AI Traffic: AI Traffic measures humans arriving from AI tools, Agent Analytics measures the AI crawlers themselves visiting your site
- Verified bot detection combines the provider's bot classification with our own user-agent classifier, so spoofed user agents are not over-counted
- If your CDN or host is not listed, the CSV/log upload path accepts most standard log formats
- Per-URL crawl data requires an Enterprise plan and a connected Cloudflare source