Tell AI Agents, Bots and Humans Apart in Your Analytics

Article Objective:
Help marketers classify sessions into humans, good bots, AI crawlers, AI agents, proxy bots, IVT and out-of-geo traffic, then measure, exclude or block each segment.
Estimated Read Time:

Your analytics reports now mix three very different kinds of visitor: people, conventional bots, and AI agent traffic that browses, reads and sometimes clicks on behalf of a person. If you treat them as one number, your conversion rates, channel comparisons and ad-platform bidding signals all drift. This guide shows marketers which traffic categories to separate, which signals identify each one, and what to do with every segment once you can see it.

Why built-in bot filtering in GA4 is not enough

Google Analytics 4 (GA4) does exclude some automation automatically. According to Google's help article on known bot traffic exclusion, known bots are identified "using a combination of Google research and the International Spiders and Bots List, maintained by the Interactive Advertising Bureau." The same page adds that you "cannot disable known bot traffic exclusion or see how much known bot traffic was excluded."

That leaves three gaps for a marketer:

  • It is list-based. A list catches bots that announce themselves. It does not catch automation that presents a normal Chrome user agent from a residential IP address.
  • It is invisible. You cannot see what was removed, so you cannot compare filtered volume against your ad spend or server logs.
  • It is binary. A visit is either excluded or counted as human. There is no bucket for "AI assistant fetching a page for a real buyer," which is neither fraud nor a normal session.

The scale of the problem is not small. Imperva's 2025 Bad Bot Report announcement found that automated traffic made up 51% of all web traffic in 2024, with malicious bots at 37%. Much of that never reaches GA4 because it does not run your tag, but the automation that does run JavaScript is exactly the kind a list will miss.

The seven traffic categories worth separating

Instead of "bot or not," sort every session into a category that implies a decision. These seven cover most marketing sites:

  1. Humans. Real people in your target markets. The only traffic that should train your bidding and your conversion-rate benchmarks.
  2. Good bots and search crawlers. Googlebot, Bingbot, uptime monitors, link checkers. You want them, but they are not audience.
  3. AI crawlers and fetchers. Declared crawlers such as OpenAI's GPTBot and OAI-SearchBot or PerplexityBot, plus user-triggered fetchers such as ChatGPT-User and Perplexity-User that load a page when someone asks an assistant a question.
  4. Autonomous AI agents. Browser-based agents that navigate, fill forms and click for a user. They often execute JavaScript, so they can land in your analytics as sessions.
  5. Residential-proxy bots. Automation routed through home internet connections so it looks like consumer traffic. The Imperva 2025 Bad Bot Report summary on Security Boulevard says 21% of bot attacks use residential proxies provided by internet service providers (ISPs).
  6. Invalid traffic (IVT). Activity with no genuine interest behind it, including click fraud, click farms and accidental clicks.
  7. Out-of-geo traffic. Real or fake visits from regions you do not sell to. Not always malicious, but never a valid signal for a campaign targeted elsewhere.

The advertising industry already uses a similar split for IVT. The Media Rating Council (MRC) invalid traffic addendum defines general invalid traffic (GIVT) as traffic caught "through routine means of filtration executed through application of lists or with other standardized parameter checks," and sophisticated invalid traffic (SIVT) as situations requiring "advanced analytics, multi-point corroboration/coordination, significant human intervention." GA4's built-in filter is roughly a GIVT tool. Residential-proxy bots and agents that mimic people sit on the SIVT side.

Signals that identify AI agent traffic, bots and humans

No single signal is reliable on its own. Strong classification comes from combining declared identity, network data and behavior, then checking what happens downstream.

Declared user agents and verified IP ranges

Well-behaved crawlers say who they are and publish where they come from. OpenAI's crawler documentation lists OAI-SearchBot (ChatGPT search), GPTBot (content that may be used for model training) and ChatGPT-User (user-initiated actions), each with a published IP range file. It also notes that robots.txt rules may not apply to ChatGPT-User because those visits are triggered by a person. Perplexity's bot documentation follows the same pattern with PerplexityBot and Perplexity-User.

A user agent string is trivial to fake, so always pair it with a network check. Google's guide to verifying Googlebot describes a reverse and forward DNS lookup, or matching the IP against Google's published crawler ranges. Note that some AI controls have no user agent at all: per Google's common crawlers page, Google-Extended is only a robots.txt token and "doesn't have a separate HTTP request user agent string."

Declared identity also only covers operators who choose to declare. In August 2025, Cloudflare reported undeclared crawling it attributed to Perplexity, using a generic Chrome on macOS user agent and IPs outside the published ranges, at 3 to 6 million requests per day. The lesson for analytics: a normal-looking browser string proves nothing.

Web Bot Auth signatures

Cryptographic signing is the newer, stronger signal for agents. Cloudflare's post on signed agents from August 2025 describes Web Bot Auth, which uses HTTP message signatures so agents can prove who operates them, and names ChatGPT agent, Block's Goose, Browserbase and Anchor Browser among the first participants. OpenAI's allowlisting help article explains that its agent requests carry a Signature-Agent header set to chatgpt.com, verifiable against a public key directory. When a signature validates, you can label the visit an AI agent with confidence. When a request claims to be an agent but carries no valid signature, treat it as unverified.

Behavior on the page

Behavior is where agents and disguised bots most often separate from humans:

  • Scroll and dwell. People scroll unevenly and pause. Scripts jump to fixed offsets or never scroll at all.
  • Interaction timing. Forms completed in under a second, identical gaps between clicks, or clicks without prior pointer movement are machine patterns.
  • Session shape. Dozens of pages in a minute, single-page sessions that fire a conversion event instantly, or many sessions sharing one device fingerprint.

Network and ASN data

The autonomous system number (ASN) tells you which network owns an IP. Cloud and hosting ASNs rarely carry consumer shoppers, which is why the MRC lists "known invalid data-center traffic" under GIVT. Residential ASNs are harder: the traffic looks consumer, so you need behavior and downstream quality to catch proxy bots on them.

Referrer and UTM anomalies

AI assistants now send measurable referral traffic. According to OpenAI's publishers and developers FAQ, "ChatGPT automatically includes the UTM parameter utm_source=chatgpt.com in referral URLs." That is useful, and also easy to copy. Watch for utm_source=chatgpt.com sessions arriving from data-center ASNs, paid-campaign UTMs on traffic that never came through the ad platform, and referrers that do not match the UTM source.

Conversion quality downstream

The final test is what the traffic produces. Leads with disposable emails, fake phone numbers, zero sales-accepted opportunities or refunds cluster by source. Tie CRM outcomes back to session source so you can see which categories ever turn into revenue.

How to segment the traffic in analytics and ad platforms

You do not need a data warehouse to start. Work with the tools you already have:

  • Create an AI channel in GA4. GA4 custom channel groups let you define channels with source and medium rules, including regex, and apply them to reports retroactively. Standard properties get two custom groups. Put AI assistant referrers such as chatgpt.com, perplexity.ai and gemini.google.com in their own channel.
  • Pass a traffic category into analytics. If a bot-classification tool labels each visitor, send that label as a custom dimension or event parameter so you can build explorations and audiences by category.
  • Review server or CDN logs for crawlers. Most declared crawlers fetch HTML without running your tag, so they never appear in GA4. Logs are where you measure their volume.
  • Check ad-platform IVT reporting. Google Ads' invalid traffic overview explains that invalid clicks found before month end are removed from billing and reports, and later detections appear as "Invalid activity" credits. Compare those credits against what your own classification flags on paid landing pages.
  • Clean what you send back. Only import conversions from sessions classified as human and in-geo, so Smart Bidding learns from real buyers.

Build a traffic-quality KPI for every channel

Once sessions carry a category, turn it into a number your team reviews next to cost per acquisition. A simple version:

  1. For each channel or UTM source, count sessions by category.
  2. Calculate the share that is clean human, in-geo traffic.
  3. Add a downstream check: the share of conversions from that channel that survive CRM qualification.
  4. Track both weekly and set a threshold that triggers a review, for example a sudden drop on one campaign or placement.

This makes quality a channel attribute rather than an occasional audit. A display campaign with a low cost per lead but a poor clean-traffic share is easy to spot, and so is a referral source whose volume comes mostly from agents rather than people.

What to do with each segment

Classification only pays off when each category has a policy. A practical starting checklist:

  • Humans: allow, and make them the only source of bidding signals and conversion benchmarks.
  • Good bots and search crawlers: allow after verification, and exclude from marketing reports.
  • AI crawlers and fetchers: decide per operator with robots.txt and your firewall. Allowing search-oriented fetchers keeps you visible in AI answers, while training crawlers are a content-licensing choice.
  • Verified AI agents: allow and measure separately. An agent comparing prices for a shopper may lead to a sale, but its clicks should not train your ad algorithms. Our article on agentic AI traffic, the good and the bad goes deeper.
  • Residential-proxy bots: block where confidence is high and exclude from bidding signals. See how residential proxy SDKs drain your marketing budget.
  • IVT: block, exclude from conversions, and keep evidence for platform credit requests.
  • Out-of-geo: exclude from campaign reporting and bidding, and block if you cannot serve those regions.

One way to do this with Ðeny

If you would rather not assemble the rules yourself, Ðeny Intelligence Hub classifies each visitor in real time as Clean, Good Bot, Residential Proxy, IVT, Out-of-Geo or AI Agent. It rolls those categories into a Traffic Health Score graded A to F, with AI summaries that explain what changed. The UTM Breakdown shows the same categories per channel, which gives you the traffic-quality KPI described above without spreadsheets. Ðeny Bot Shield can then block the segments you decide to block, and it deploys with a one-line script.

Turn one traffic number into several honest ones

AI assistants, crawlers, agents and proxy bots are now a permanent part of marketing traffic, and GA4's known-bot list only removes the easy cases. Separate the categories, verify identity with IP ranges and signatures, confirm with behavior and downstream quality, and give each segment a clear policy. For a wider view of what is coming, read our overview of 2026 bot threats to the marketing stack. When you are ready to see your own traffic split by category, request a Ðeny demo.

Receive better insights, in your inbox
Subscribe to Deny's insights & news.
Subscribe
Thank you! Your submission has been received!
Oops! Something went wrong while submitting the form.

Start protecting your funnel today

Put Ðeny to work from day one, and your boss will thank you.

$79/month
Cancel anytime
Credit card required
Get Started