Content Agents

How to Read Server Logs to See What's Actually Crawling You

By Roey Granot · September 28, 2026

Category: ai-transformed-workflows

How to Read Server Logs to See What's Actually Crawling You

Key takeaways

  1. The problem Most teams never look at server logs, so fake crawlers, AI bots, and crawl budget waste go undetected until rankings have already suffered.

  2. Core insight Reading server logs well means parsing who is crawling, verifying their identity with reverse DNS, and mapping what they hit - so you can act on real evidence rather than assumptions.

  3. Practical outcome After reading this, you can open your access logs, filter and verify crawler traffic by provider, spot traps and orphaned pages, and set up the baselines you need to catch indexing problems early.

Server logs tell you exactly who is crawling your site, which pages they're hitting, how often, and whether they're who they claim to be. Most teams never look at them. That's a problem - because fake crawlers are common, AI bots are multiplying fast, and crawl budget is finite.

This guide is for SEO practitioners who already understand crawl budgets, robots.txt, and the difference between a 301 and a 404. We won't explain what a crawler is. We'll show you how to read the evidence they leave behind. A quick note on scope: the parsing and verification steps below are grounded in direct practice. The migration monitoring and defense-layering sections synthesize established industry methods - we'll be clear about which is which throughout.

Step 1: Parse Your Server Logs to Identify Who Is Actually Crawling You

Industrial steel staircase steps viewed from above.
Photo by mikecook1 on Pixabay

Start with the raw log file your server generates. Apache, Nginx, and IIS each produce slightly different formats, but every line contains the same core fields: IP address, timestamp, request method, URL path, HTTP status code, bytes transferred, and User-Agent string. The User-Agent field is where crawlers identify themselves - and where spoofing begins.

A typical Nginx log line looks like this:

66.249.64.1 - - [12/Jun/2025:08:14:22 +0000] "GET /blog/content-ops HTTP/1.1" 200 4823 "-" "Mozilla/5.0 (compatible; Googlebot/2.1; +http://www.google.com/bot.html)"

The User-Agent string at the end is self-reported. Any bot can write anything there. That matters enormously, which we'll come back to in Step 2.

To isolate bot traffic, filter on known User-Agent signatures. The basic grep command:

grep -i "googlebot\|bingbot\|gptbot\|claudebot" access.log

Pipe that into a count by User-Agent to see volume per crawler. For larger log files, tools like GoAccess or AWStats give you this at a glance without writing shell scripts.

Once you have a filtered view, segment by crawler type: search bots (Googlebot, Bingbot) versus AI crawlers (GPTBot, ClaudeBot, others). This distinction matters because their purposes differ. Search bots crawl to index your content for retrieval. AI crawlers crawl to collect training data. They have different crawl depths, different request patterns, and different implications for your crawl budget and content strategy.

One number worth keeping in mind: around 51% of web traffic is now automated. Most of it isn't doing anything useful for you. Segmenting by crawler type is the first step toward understanding what portion actually matters.

What this means in practice

  • Pull your access.log for the past 30 days. Run the grep filter above and count lines per User-Agent string.

  • Sort results by volume. Anything in the top five that you don't recognize warrants investigation before you trust it.

  • Create two buckets: search bots and AI crawlers. Track them separately from this point forward.

  • If your logs are compressed (.gz), use zgrep instead of grep - same syntax, works on compressed files.

  • Set a baseline. What's normal crawl volume for each bot on your site? You need that number before anything else is meaningful.

Step 2: Verify Crawler Identity with Reverse DNS Before Trusting Any Signal

User-Agent strings are trivially easy to fake. A scraper can claim to be Googlebot in about one line of code. Verification requires checking whether the IP address that made the request actually belongs to the crawler it claims to be. The standard method is FCrDNS - forward-confirmed reverse DNS.

The process has two steps. First, reverse-DNS the IP: run host 66.249.64.1 or nslookup 66.249.64.1. If the result is a hostname like crawl-66-249-64-1.googlebot.com, that's a good sign. Second, forward-DNS that hostname: run host crawl-66-249-64-1.googlebot.com. If it resolves back to the original IP, the crawler is verified. If either step fails - no hostname, or the hostname doesn't resolve back - treat the request as unverified.

Google, Microsoft, and OpenAI all publish their official IP ranges. Anthropic does the same for ClaudeBot. These are maintained by the companies themselves. If an IP claiming to be GPTBot doesn't fall within OpenAI's published ranges, it isn't GPTBot - regardless of what the User-Agent says. That's where the 5.7% of traffic falsely claiming to be known AI crawlers gets caught.

Automating this check makes it operationally viable. Most log analysis pipelines can run a reverse DNS lookup and flag IPs that fail FCrDNS or fall outside official ranges. You don't need to do this manually for every request - you need a script that tags unverified IPs and surfaces them for review.

What this means in practice

  • Run FCrDNS on the top 20 IPs in your crawler traffic. This takes about 20 minutes manually and tells you immediately whether your logs contain impersonators.

  • Download the official IP ranges for Google, Bing, OpenAI, and Anthropic and save them as reference lists. Check these quarterly - they change occasionally.

  • Flag any IP that claims to be a known crawler but fails FCrDNS. That traffic is either malicious or misconfigured. Either way, it shouldn't be treated as legitimate crawl activity.

  • Document your verification results. If you're reporting on crawl budget or AI content usage to stakeholders, verified numbers carry more weight than raw log counts.

Step 3: Map Crawl Patterns to Find Traps and Orphaned Pages

Symmetrical concrete or metal steps photographed from a low angle.
Photo by dimitrisvetsikas1969 on Pixabay

Once you know who is crawling and that they're legitimate, look at what they're hitting and in what order. Crawl traps and orphaned pages both show up clearly in logs - but they look different, and fixing the wrong one wastes time.

Crawl traps are URL patterns that generate effectively infinite unique pages. Pagination loops are the most common: /products?sort=date&page=1, /products?sort=date&page=2, and so on, with no canonical or noindex signal to stop the crawler. Session IDs and tracking parameters create a similar problem - every visit generates a new URL, every new URL gets crawled, and your crawl budget drains on pages that are functionally identical.

The log signature is volume on a pattern. If you see 10,000 requests to /products?page=N variants in a single crawl session, Googlebot found a pagination loop. The fix is canonical tags, parameter handling in Google Search Console, or noindex on paginated pages beyond a reasonable depth - not a log-analysis task, but logs are how you find the problem in the first place.

Orphaned pages look different: low-traffic pages that appear in crawl logs but don't appear in your internal link map. Crawlers found them via external links or old sitemaps, but your site navigation doesn't connect to them. These pages consume crawl budget and rarely rank well because they have no internal equity. Cross-referencing your log data against your internal link map (export from your crawl tool of choice) surfaces them quickly.

AI crawlers add a layer worth noting here. As reported by Cloudflare, GPTBot traffic grew 305% year over year and is now the second most active crawler after Googlebot. But as one industry observer noted, "AI crawlers often operate with more conservative crawl limits than Googlebot, so take note of how deep into your site structure they'll need to go." A crawler with conservative depth limits won't find deep orphans - which means those pages may be invisible to AI training pipelines even if Googlebot finds them.

What this means in practice

  • Sort your crawler log by URL, then look for high-volume patterns with numeric or parameter-heavy suffixes. Those are your trap candidates.

  • Export your internal link map from Screaming Frog or equivalent. Cross-reference against pages appearing in logs. Anything in logs but not in the link map is an orphan candidate.

  • Check redirect chains while you're here. A crawler hitting a 301 that points to another 301 that eventually resolves is wasting hops. Three or more redirects in a chain is worth fixing.

  • Segment trap and orphan findings by crawler. Googlebot hitting a trap and GPTBot hitting the same trap have different implications and may warrant different fixes.

Step 4: Segment AI Crawler Activity by Provider to Understand Content Exposure

Different AI companies crawl at different rates, with different depth behavior, and presumably for different purposes. Treating all AI crawler traffic as a single category misses information that's operationally useful - especially if you're trying to understand which parts of your content are being ingested by which LLM providers.

Filter your logs separately for each AI crawler User-Agent. GPTBot uses GPTBot/1.0. ClaudeBot uses ClaudeBot/0.1. Create a daily count per crawler, broken down by page and content type. Run this for 30 days and you'll have a clear picture of crawl intensity per provider.

A concrete example of why this matters: if GPTBot is hitting your blog three times more than your product pages, and ClaudeBot is doing the opposite, that tells you something about what each company's systems are prioritizing. It may also inform decisions about what you allow or disallow - you might be comfortable with OpenAI training on your editorial content but less comfortable with product documentation.

Crawl intensity is worth calculating directly. Take total requests per crawler per day, average response size in bytes, and crawl depth (how many unique URL paths per session). If GPTBot is consuming 20% of your bandwidth on a given day, that's a resource decision, not just an information one. Rate limiting becomes a practical question at that point, not just a theoretical one.

What this means in practice

  • Create a simple spreadsheet: rows are crawlers (GPTBot, ClaudeBot, Googlebot, Bingbot), columns are days. Fill in request counts. Trends become obvious within two weeks.

  • Break down by content section, not just total volume. Which crawler is most active on your blog vs. documentation vs. product pages? This shapes allowlist decisions.

  • Calculate bandwidth per crawler. Multiply average response size by request count. If one crawler is consuming a disproportionate share of your server resources, that warrants a rate limit review.

  • Check crawl depth distribution. How many of GPTBot's requests are to pages three or more clicks from your homepage? That tells you whether it's getting into your content archive or staying shallow.

Step 5: Monitor Post-Migration Crawl Patterns to Catch Indexing Failures Early

Site migrations - domain changes, URL restructures, CMS switches - are where crawl problems get expensive fast. Logs are the earliest warning system you have. By the time a ranking drop shows up in Search Console, the damage has already been done for weeks. Logs show you the problem in real time.

The scenario to watch for: you migrate from /blog/post-title to /articles/post-title and set up 301 redirects. In the logs immediately post-migration, you should see crawlers hitting the old URLs, receiving 301s, and then requesting the new URLs. That chain - old URL, 301, new URL, 200 - is the healthy pattern. If you see crawlers hitting old URLs and getting 404s, your redirects aren't in place. If you see crawlers hitting old URLs and not following the 301s, something is misconfigured at the redirect level.

Segment your post-migration log view by HTTP status code. Filter for 3xx, 4xx, and 5xx responses. Count by crawler type. A spike in 404s from Googlebot in the week after a migration is a signal that needs immediate attention. A spike in 5xx errors from any crawler suggests server-side problems that could be suppressing crawl entirely.

Set a baseline crawl volume expectation before you migrate. A 30-50% drop in crawl activity in the first week post-migration is normal - crawlers are recalibrating. If crawl volume is still down 50% after four to six weeks, that's not recalibration, that's a structural problem. Logs are how you tell the difference. This is synthesized from established migration practice rather than our direct experience, but the signal logic is sound.

What this means in practice

  • Pull pre-migration

Frequently Asked Questions

How do I find out which bots are crawling my site using server logs?

Filter your access.log file using a grep command targeting known User-Agent strings - for example: grep -i "googlebot\|bingbot\|gptbot\|claudebot" access.log. Pipe the output into a count by User-Agent to see volume per crawler. For larger log files, tools like GoAccess or AWStats can give you this view without writing shell scripts. Pull at least 30 days of logs and sort results by volume so you can spot anything unfamiliar in your top five crawlers.

How can I tell if a bot claiming to be Googlebot is actually Googlebot?

Use FCrDNS - forward-confirmed reverse DNS. First, reverse-DNS the IP address using the host or nslookup command. If it returns a hostname like crawl-66-249-64-1.googlebot.com, then forward-DNS that hostname and check whether it resolves back to the original IP. If either step fails, treat the request as unverified. You can also cross-reference the IP against Google's officially published IP ranges - if the IP falls outside those ranges, it is not Googlebot regardless of what the User-Agent string says.

What does a crawl trap look like in server logs and how do I fix it?

A crawl trap shows up as very high request volume against a repeating URL pattern with numeric or parameter-heavy suffixes - for example, thousands of requests to /products?page=N variants in a single crawl session. This typically means Googlebot found a pagination loop or is generating unique URLs from session IDs or tracking parameters. The fix involves canonical tags, parameter handling in Google Search Console, or noindex on paginated pages beyond a reasonable depth. Logs are how you detect the problem; the resolution happens in your site configuration.

How do I use server logs to monitor a site migration and catch indexing problems early?

Immediately after migration, filter your logs by HTTP status code and look for the healthy redirect chain: crawlers hitting old URLs, receiving 301 responses, then requesting new URLs and getting 200 responses. If crawlers are hitting old URLs and getting 404s, your redirects are missing. If they are not following 301s at all, something is misconfigured at the redirect level. Set a baseline crawl volume before you migrate - a 30 to 50% drop in the first week is normal recalibration, but if volume is still down 50% after four to six weeks, that indicates a structural problem worth investigating.

Should I track GPTBot and ClaudeBot separately, and why does it matter?

Yes - different AI crawlers operate at different rates, reach different crawl depths, and focus on different content sections. Filtering logs separately for each User-Agent (GPTBot/1.0 and ClaudeBot/0.1, for example) and tracking daily request counts by page type over 30 days shows you which provider is most active on which parts of your site. This matters for resource decisions too: if one AI crawler is consuming a disproportionate share of your bandwidth, that is a practical trigger for reviewing rate limits. It also informs allowlist decisions if you want to treat editorial content differently from product documentation.