How to Advertise on a Web That’s Mostly Bots

14–20 minutes
AI agent advertising

18 min read

TL;DR

  • On 3 June, Cloudflare CEO Matthew Prince published Radar data showing automated requests at 57.5% of HTML web traffic against 42.5% from humans, the first machine majority in internet history.
  • Measuring a different basket that includes app and API traffic, Imperva’s 2026 Bad Bot Report put automated traffic at 53% of all web traffic in 2025, up from 51% in 2024.
  • Roughly half of new web articles are now primarily machine written. Graphite’s Common Crawl sample put it at 50.9% in Q4 2025 and 49.9% in Q1 2026.
  • AI agent advertising is not a future line item. Adobe measured AI referred traffic to US retail sites up 393% year over year in Q1 2026, converting 42% better than non AI channels in March 2026.
  • The fix is not a new dashboard as they teach in business schools. It is five straightforward but crucial moves: separate the traffic ledger, re-anchor on money, shrink the domain list, structure pages for machine reading, and subtract spend to see what breaks.
  • Three brands already ran the subtraction test. Chase cut 395,000 domains, Uber cut $100 million, P&G cut almost nine figures and none of them lost measurable business.

The Dead Internet Stopped Being a Theory in June 2026

In 2021, a user called IlluminatiPirate posted a thread on a small forum called Agora Road’s Macintosh Cafe titled “Dead Internet Theory: Most Of The Internet Is Fake.” The claim was that the web had quietly emptied of people and filled with bots. The Atlantic’s Kaitlyn Tiffany covered it that September and called the post the theory’s ur-text. It was treated, correctly at the time, as a conspiracy theory with a good mood and bad evidence.

Interestingly the evidence arrived five years late. On 3 June 2026, Cloudflare co-founder and CEO Matthew Prince posted (see below) that automated requests had passed human requests on the web for the first time. Cloudflare Radar, which sees traffic across roughly a fifth of all websites, showed the split at 57.5% machine to 42.5% human. Prince had forecast the crossover at SXSW in March 2026 and expected it by the end of 2027. Unsurprisingly, it arrived about eighteen months early.

Every media plan signed off in the last decade rested on an assumption that no longer holds in the aggregate: that the thing on the other end of the impression is a person. AI agent advertising is what you get when you take that seriously and rebuild the buying and the measurement around it.

The Truth in the Traffic

There are two separate measurement bodies that have each independently arrived at a machine majority, though they get there using different methodologies and different underlying baskets of traffic. Let’s compare them side by side to see where they agree and where they diverge.

SourceWhat it measuresMachine shareDate
Cloudflare RadarHTTP requests to HTML content across its network57.5%3 June 2026
Imperva 2026 Bad Bot ReportAll web traffic including app and API53%Full year 2025
Imperva 2025 Bad Bot ReportAll web traffic including app and API51%Full year 2024
Imperva 2024 Bad Bot ReportAll web traffic including app and API49.6%Full year 2023

The Imperva numbers are useful because they show a gradual trend over years rather than one dramatic snapshot. Within that trend, bad bots (the malicious kind) made up 37% of all internet traffic in 2024 and grew further in 2025, and more than a quarter of bot attacks that year hit APIs directly rather than going through a normal web interface. It’s not just the amount of bot traffic that changed, but what kind of bot traffic it is.

The second part shifts to a different data point from Cloudflare: over half of AI crawler requests in May 2026 were crawlers gathering data to train models (not to power search results), and only about 9% were for search. This means the training crawlers take from your site (crawling it, costing you server resources) but send nothing back (no visitors, no referral traffic), so their only visible effect on you is added infrastructure cost, with no upside.

Reading Between the Data Points

The SEO firm Graphite sampled tens of thousands of English language URLs from Common Crawl and used Surfer’s detector to classify each article. Axios found:

  • Within a year of ChatGPT’s launch, primarily AI generated articles made up 35.9% of new online articles.
  • Within two years they reached 48%.
  • They crossed the halfway line in Q4 2025 at 50.9%, then settled back to 49.9% in Q1 2026.
  • The share has hovered near half for five consecutive quarters. The feared exponential takeover flattened.

Another study by Graphite found 86% of articles appearing in Google Search and 82% of articles cited by ChatGPT and Perplexity were human written. When machine written pages do surface, they tend to rank lower. Graphite’s own hypothesis for the plateau is that primarily AI generated articles do not perform, so publishers stopped making more of them.

Two things to note here. Graphite tested Surfer’s detector and found it labelled human written articles as machine written 4.2% of the time. And because many paywalled publishers now block Common Crawl, a body of almost certainly human writing sits outside the sample entirely. The true human share is probably higher than 50%.

The Fiction of Modern Engagement

If you work in performance marketing, you’ll notice the bot problem in your data before anyone officially acknowledges it as a strategic issue. For example, impressions and sessions keep growing, engagement rate looks stable, but revenue doesn’t move. Most teams misread this pattern, thinking their creative is underperforming (so they rewrite ads) when the real issue is that a growing share of the “traffic” was never a real customer to begin with.

As more traffic becomes automated, the usual metrics (engagement, conversion rate, demand signal) stop meaning what they used to. A spike in traffic might just be bots repeatedly hitting your pricing or stock-check endpoints, not real buyer interest, but the danger is that this fake signal still feeds into automated ad systems, which then allocate real budget based on it.

The invalid traffic tax

The cost has been measured from several directions, and the estimates broadly agree.

MetricFigureSource and period
Global digital ad fraud lossesPast $100 billionJuniper Research forecast for 2026
Same metric, three years earlier$84 billionJuniper Research, 2023
Wasted open web programmatic spend$26.8 billionANA Programmatic Transparency Benchmark, Q2 2025
Same metric at first measurement$20.0 billionANA, June 2023
Global invalid traffic rate20.64%Fraudlogix, 105.7bn impressions, full year 2025
Made-for-advertising share of spend0.8% median, down from 15%ANA, Q2 2025 vs June 2023

It’s important to note that made-for-advertising sites, the clearest and most discussed version of the problem, were largely dealt with: median MFA exposure fell from 15% of spend to 0.8%, the lowest since the ANA began tracking it. Over the same two years, total waste rose 34%, from $20.0 billion to $26.8 billion. The industry solved the part of the problem it could see and the money leaked somewhere else.

The crawl-to-refer gap

Cloudflare publishes a metric that belongs on every publisher and brand dashboard: the crawl-to-refer ratio. It divides how many pages a platform’s crawler requests by how many visitors that platform sends back, which makes it a reasonable proxy for whether the old bargain of the open web still holds.

PlatformPages crawled per referral sent
DuckDuckGoAbout 1.5 to 1
GoogleAbout 5 to 1
PerplexityAbout 111 to 1
OpenAI (GPTBot)About 904 to 1
Anthropic (ClaudeBot)About 10,300 to 1

Figures are Cloudflare Radar as of 31 May 2026, and they are more directional rather than fixed. ClaudeBot’s ratio has been reported anywhere between roughly 10,000 to 1 and 24,000 to 1 across different weeks in 2026, which tells you the metric moves enough to be read as a trend rather than a scoreboard. Cloudflare’s own post on the methodology sets out how it is calculated.

Under search, being crawled was how a site earned traffic. Under training it mostly is not, and any content strategy still built on the older exchange rate is paying for attention it will not receive.

The AI to AI Economy Is Already Trading

Not all machine traffic is fraud. A large and growing slice of it is an agent acting on behalf of a real person with a real credit card, arriving through the same pipe and looking nearly identical in your server logs. That resemblance is what makes lumping the two together expensive.

For instance, Adobe Analytics, which tracks over a trillion visits across US retail sites, published its Q1 2026 numbers in April:

  • AI referred traffic to US retail sites grew 393% year over year in Q1 2026, following a 693% year over year surge across the November to December 2025 holiday season.
  • In March 2025, AI referred traffic converted 38% worse than non AI channels. In March 2026 it converted 42% better. That is an 80 point swing in twelve months.
  • Revenue per visit from AI referrals ran 37% above non AI traffic.
  • Those visitors spent 48% longer on page and viewed 13% more pages per visit.
  • By May 2026, Adobe reported AI traffic converting 54% better than non AI sources.

These are Adobe Analytics figures, published alongside Adobe’s own LLM Optimizer product. The sample is very large and the numbers are vendor stated, and individual retailers report widely different outcomes, so the direction is more reliable than the magnitude.

The ad industry has meanwhile started building plumbing for machine buyers. OpenAI began testing ads in ChatGPT in early 2026 and opened a self-serve Ads Manager on 5 May 2026, with the pilot expanding to the UK in June and more markets flagged for later in the year. IAB Tech Lab published its Agentic Roadmap on 6 January 2026, consolidated the work under the name AAMP in February, and on 27 May 2026 released draft guidance on bot and crawler management for public comment. The standards bodies have already accepted that machines are counterparties, and AI agent advertising is being specified in public this year while plenty of brand teams are still deciding whether to block crawlers at all.

AI Agent Advertising: The Signal Ledger Framework

There are five steps, each with an action, a reason and a method. The name comes from the core move, which is bookkeeping: you stop treating traffic as one number and start keeping separate accounts for signals that behave differently.

Step 1. Separate the ledger

Split traffic into four books instead of one: human, agent acting for a human, declared crawler, and undeclared or invalid. Those four categories have opposite economics. An agent buying on someone’s behalf is your best converting visitor. A training crawler is pure infrastructure cost. Invalid traffic is a refund you are not claiming. Averaging them produces a number that describes none of them.

Google Analytics 4 added a native AI Assistant channel on 13 May 2026 covering ChatGPT, Gemini and Claude, though it misses some sources including Perplexity. Supplement with server log analysis by user agent, and follow the IAB Tech Lab crawler management guidance as it finalises.

Step 2. Re-anchor on money, not motion

Demote impressions, sessions, engagement rate and time on site to diagnostics. Promote revenue per visit, contribution margin and incremental orders to the only KPIs that appear on the board slide. Every metric that can be produced by a script is now produced by scripts at scale. Money is harder to fake because it requires a settled payment, which makes it the one signal a bot cannot cheaply manufacture.

Rebuild your reporting so that the top of every dashboard is a currency figure. If a metric cannot be traced to a settled transaction within a defined window, it goes below the fold.

Step 3. Shrink the surface

Cut your programmatic domain list hard, and cut your supply path partner count with it. The ANA found the average campaign runs across roughly 44,000 websites when a few hundred would reach most of the audience, and estimated that a more selective approach reaches about 95% of valued audiences. Breadth is where invalid traffic hides.

Pull log level data, rank domains by settled revenue rather than impressions, and cut everything with no revenue attached. Be careful about swinging too far the other way: aggressive keyword blocklists demonetise legitimate publishers and shrink your own reach, which we cover in this audit of blocklist overblocking.

Step 4. Structure the page for machines

Make your product and category pages readable by a retrieval system without JavaScript execution. Adobe’s April 2026 retail benchmark found sites were broadly not ready for machine reading, with product detail pages the weakest layer. If an agent cannot parse your price, stock status and specification, it recommends a competitor it can parse. That is a conversion loss you will never see in your funnel because the visit never happens.

Server-render critical commercial facts, apply Product and FAQ schema, keep specifications in text rather than images, and test what a plain HTTP fetch of your page returns.

Step 5. Subtract and see what breaks

Run geo-holdout tests. Turn a channel off in matched markets and measure the gap in settled revenue, not in platform-reported conversions. Attribution fraud works by claiming credit for outcomes that were going to happen anyway. No amount of verification tooling detects that. Turning spend off does, immediately and unambiguously.

Pick matched market pairs, hold one dark for a full purchase cycle, and compare total settled revenue. If the number does not move, you have your answer. All three cases below were found this way.

What Happens When Brands Do This

All three are documented and dated, and each one is a subtraction test run at scale.

BrandWhat was cutResultSource
JPMorgan ChaseProgrammatic display domains from about 400,000 a month to 5,000, a 98.75% reductionNo deterioration in performance metrics, little change in cost per impression or ad visibilityNew York Times, March 2017, CMO Kristin Lemkau
Uber$100 million of a $150 million annual ad budget, two thirds of spendNo change in rider app installs. Installs previously credited to paid channels reappeared as organicKevin Frisch, former Head of Performance Marketing, Marketing Today podcast
Procter & GambleRoughly $140 million of digital spend in a single quarter over placement concernsOrganic sales grew 2%, beating analyst forecasts and key rivals despite the cutAd Age, July 2017

Of the 400,000 domains its ads appeared on in a 30 day window, only 12,000, or 3%, produced any activity beyond an impression. An intern then manually clicked through those 12,000. About 7,000 were places the bank did not want to appear, leaving the final list of 5,000. A 99% reduction in reach, built by hand, with no measurable loss.

Uber’s case is the more alarming one. Frisch’s team found instances where a user apparently clicked an ad and was signed into Uber two seconds later, which is not physically plausible. Those fake interactions convinced optimisation systems to shift more budget toward the fraudulent sources, so the fraud was not only taking budget, it was directing where the rest of it went. Frisch has also said that at the time, saving $100 million was not treated as a win internally.

The Case Against This Case

The strongest objection to everything above comes from the person who published the 57.5% figure. Matthew Prince has explicitly rejected the dead internet reading of his own data. His argument is that generative tools lowered the barrier to publishing, so more people can now create rather than fewer. He also noted that the web actually shrank between 2015 and 2025, and that the reversal began in the six months before his June post, with what he described as exponential growth in new and creative material.

Pew Research Center has separately found that 38% of webpages that existed in 2013 were no longer accessible a decade later. There is a mechanical point underneath that. A person might visit five sites before buying; an agent might visit five thousand on their behalf. On that reading, a bot majority reflects human intent moving through a higher-volume interface rather than humans leaving.

Graphite’s plateau supports the same interpretation. If machine written content were winning, its share would not have sat near 50% for five straight quarters while 86% of Google’s results and 82% of ChatGPT’s citations stayed human written.

Where to Start This Quarter

None of this needs a transformation programme though. It needs one number you probably do not have yet. Pull your last 30 days of programmatic log level data and answer the Chase question. What percentage of the domains you paid for produced any activity beyond an impression?

Chase’s answer was 3%. If yours lands anywhere near that, you have found next quarter’s budget without asking anyone for it. Then run one geo-holdout on your weakest performing channel for a full purchase cycle. That single test will tell you more about your media than any verification vendor’s dashboard, because it measures the only thing a bot cannot fake: whether the money still arrives when you stop paying to be seen.

If you are also weighing how machine written material fits into your own publishing, the ownership and disclosure questions need to be settled first. We covered those in four questions to answer before you publish AI-generated content, and the wider defensive picture in the brand protection stack for scaling brands.


FAQ

What percentage of internet traffic is driven by bots in 2026?

Data varies by measurement method: Cloudflare Radar reports that automated requests account for 57.5% of HTML web traffic, while Imperva’s broader analysis, which includes app and API traffic, places automated activity at 53%.

Is the ‘Dead Internet Theory’ considered accurate based on current data?

The theory is only partially supported; while machines now generate the majority of web requests and a significant portion of articles, human-written content remains dominant in search results and AI citations. Experts like Matthew Prince argue that generative tools have actually expanded, rather than diminished, the human ability to publish online.

How much financial impact does digital ad fraud have in 2026?

Global digital ad fraud losses are projected to exceed $100 billion in 2026, a significant increase from $84 billion in 2023. Additionally, reports indicate that programmatic open web spend continues to face high levels of waste, with invalid traffic rates reaching over 20%.

Should brands block AI crawlers from accessing their websites?

The decision depends on the crawler’s value, as some bots provide traffic referrals while others only scrape data for training purposes. Blocking training-only crawlers can reduce infrastructure costs, but it may also limit your brand’s presence in future AI-generated model outputs.

Does traffic referred by AI tools lead to actual sales conversions?

Yes, AI-referred traffic has shown a rapid improvement in performance; by May 2026, Adobe Analytics reported that AI-referred traffic converted 54% better than non-AI channels. However, because individual retailer results vary, businesses should validate these trends against their own internal data.

What is AI agent advertising?

AI agent advertising describes two connected activities: placing paid media inside AI assistants where a person still sees the ad, and structuring your site so that autonomous agents acting on a person’s behalf can find, parse and recommend your products. OpenAI opened a self-serve Ads Manager for ChatGPT on 5 May 2026, while IAB Tech Lab is standardising the agentic layer under its AAMP framework.

Written by

Contribute

Have expertise worth sharing?

We're looking for domain experts, lawyers, ecommerce operators, finance professionals, to write for Industry Contents.

Write for IndustryContents
industrycontents logo
industrycontents

Join our private reader network to receive next deep-dive analysis directly in your inbox.

Upon subscribing, instantly receive our blueprint on the highest-performing AI stacks for marketing.

Discover more from Industry Contents

Subscribe now to keep reading and get access to the full archive.

Continue reading