AI referral went from converting half as well as other traffic to 60% better in eighteen months
| Jan 2025 | -49% |
|---|---|
| Apr 2025 | -38% |
| Jul 2025 | -23% |
| Oct 2025 | 16% |
| Nov-Dec 2025 | 31% |
| Mar 2026 | 42% |
| May 2026 | 54% |
| Jul 2026 | 60% |
The one percent is a filing error before it is a market size
The AI row in your GA4 property is a classification, and the classification was built for a web that did not have assistants in it. GA4 ships 29 default channel definitions. Until 13 May 2026 not one of them covered AI, so every assistant click fell through to Referral, or to Direct when the referrer was stripped.
The native AI Assistant channel that landed in May fixes less than the launch post implied. Google named ChatGPT, Gemini and Claude on launch day; the published definition now lists ChatGPT, Gemini, Deepseek, Copilot and Grok, and Claude has quietly gone. Perplexity, the one assistant that passes a referrer reliably, is absent and sits in Referral. AI Overviews and AI Mode clicks arrive on the ordinary google hostname, so they are Organic Search by definition and no rule you can write will ever separate them. It is forward-only too, so everything that arrived before 13 May 2026 stays wherever it fell.
That is the small leak. The big one is structural. Native iOS and Android assistant apps do not pass a referrer, in-app browsers drop it on hand-off, people copy a URL out of a chat into a fresh tab, and OpenAI's Atlas browser has been reported to strip it on outbound opens. Vendor estimates for the share of AI sessions landing in Direct run from 35% to 70%, a range wide enough to tell you nobody has measured it properly. One analytics shop layered dark-social tracking onto a SaaS property and found 47% of Direct was identifiable as LinkedIn, ChatGPT or Slack.
I have never audited a retail property where the AI row was the AI number. It has always been a floor, and on the ten-store Shopify portfolio I come back to below, Direct ran 26% bot against 4% for Google, so the floor has a hole in it.
So when the board slide puts AI referral at 1% and too small to fund, the honest reading is that 1% is what survived the filing system. Whether the real figure is 1.5% or 3% is the whole question...
Which assistants tell you who they are
The leak is not one leak. It is seven behaviours across seven assistants, and five of them can be recovered with a single regex.
| Assistant | Referrer hosts | Passes referrer | Default GA4 landing |
|---|---|---|---|
| ChatGPT web | chatgpt.com, legacy chat.openai.com | Often on desktop; tags about two thirds of links with utm_source=chatgpt.com | AI Assistant (native), or Unassigned when the tag arrives without a medium |
| ChatGPT iOS and Android, Atlas browser | None | Almost never | Direct |
| Perplexity | perplexity.ai and its www subdomain | Reliably from web and Comet; mobile app strips | Referral, because it is not in the native list |
| Claude | claude.ai | Rarely | Referral when it passes, otherwise Direct; not in the native definition |
| Gemini | gemini.google.com | Passes from web; app traffic can show as (not set) | AI Assistant, but a source-contains-google rule can misfile it as Organic Search |
| Copilot | copilot.microsoft.com, bing.com entry points | Passes | AI Assistant |
| Grok | grok.com | App-heavy, often strips | AI Assistant when it passes, else Direct |
Perplexity passes its referrer cleanly from the browser and lands in Referral only because Google left it off the native list, so its fix is one regex rule; call it invisible in a brief and someone will check. Claude is the reverse: most of its clicks are Direct and no configuration recovers them, which matters more in B2B, where its share of measurable AI referrals reached 18.5% by spring 2026.
The UTM tag carries its own trap. When a UTM parameter is present GA4 does not fall back to the referrer, and with utm_medium missing the session usually drops straight into Unassigned. Across one Shopify agency's ten-store portfolio, ChatGPT sessions carrying the UTM tag outnumber sessions carrying a chatgpt.com referrer three to one. A property with no rule for that tag is throwing away the biggest slice of AI traffic it can actually see, and filing it in the bucket everyone ignores.
Build the custom group in thirty minutes and put it above Referral
None of this needs a developer. The default group cannot be edited, so you build a parallel custom group and it labels AI Search across every report that reads it.
- Admin, Channel groups, create new. Name it AI Search.
- One condition on Session source, match type matches regex. Choose 'contains' and the pattern is read as a literal string that never matches.
- Value: chatgpt\.com | chat\.openai\.com | perplexity\.ai | claude\.ai | gemini\.google\.com | copilot\.microsoft\.com | bing\.com/chat | you\.com | meta\.ai | grok\.com. Every dot escaped, because GA4 conditions run RE2 and an unescaped dot matches any character. RE2 has no lookaheads and no backreferences, so keep it a flat alternation.
- Add ^chatgpt$ to the alternation so UTM-tagged sessions arriving with source chatgpt and no medium leave Unassigned.
- Never a bare \.ai$. It swallows every SaaS tool and country-code domain ending in .ai.
- Drag AI Search to the top, above Referral. GA4 evaluates top to bottom and assigns the first match.
- Standard properties get two custom groups in total. This is one of them.
The sources disagree on whether this is retroactive, and both are right. Custom groups are evaluated at query time, so explorations populate across your retention window, two or fourteen months. Standard reports fill from creation forward. Build the exploration first; it is where your history lives.
Maintenance is real. Quarterly, pull the top 50 referral sources and look for hostnames the regex misses. Whenever an engine launches a new surface, check for a new referral domain within 30 days. Claude-Web is retired, so any robots.txt copied from a 2024 article is already wrong.
Merchants with GTM add two layers. A variable reading document.referrer, mapped to an ai_referrer custom dimension, catches what the group cannot label. In a server container, inspect the Referer header before the event forwards. Server-side tagging also recovers 30 to 35% of the conversions ITP and blockers drop. BigQuery export at $5 to $50 a month is the cheap insurance under all of it.
What the reclassified number says, and where the evidence stops
If we look at one client who runs ten established ecom stores against a single revenue line... Through June 2026 the five assistants combined drove 0.10% of sessions and 0.10% of revenue: $75K of $76.5M. That is up 6x year on year, and ChatGPT referrals went from 59 a month to over 5,400 in under two years. Then they granted every hidden UTM-only click ChatGPT-average conversion and re-ran it. AI revenue rose to at most 0.16%.
AI-referred sessions are 0.9% of all web traffic on Similarweb's 2026 read, up 5x. Across 973 ecommerce sites ChatGPT was roughly 0.2% of sessions, about 200 times smaller than Google organic. Generative engines take 3.3% of US online discovery time and around 0.77% of US desktop activity. The ceiling shows up where the purchase is considered: one property with a signup goal had AI at 0.5% of visits driving 12.1% of signups.
So the 1% on the slide is generous for most merchants, and that is exactly why reclassification comes first. It will not reveal a hidden channel. It will tell you whether you are the median store at 0.1% of revenue or the property where half a percent of visits carried an eighth of the signups, and it is the only instrument that can.
What holds everywhere the direction was measured is quality. On the ten stores AI visitors converted at 2.66% against a 1.91% portfolio average, and 85% of AI-referred revenue came from first-time buyers, double the 42% norm. Small, better, and new to the brand. At 85% first-time the channel is acquiring people the brand had never sold to, which prices it against blended new-customer CAC rather than against the organic row, and that is the comparison a CFO can act on.
Growth has fallen from 4,700% to 62% in twelve months
| Month | |
|---|---|
| Jul 2025 | 4700% |
| Oct 2025 | 1200% |
| Dec 2025 | 1151% |
| Mar 2026 | 269% |
| May 2026 | 138% |
| Jul 2026 | 62% |
The assistant does the browsing so the shopper does not
Two thirds of ChatGPT visitors to the ten stores never see a homepage; the assistant has already picked the product and deep-linked to it. They browse fewer pages than average, 2.3 against 2.8, and buy more anyway: 4.1% of those product-page landings check out, against 1.6% for product-page landings site-wide. The comparison happened inside the chat, and the page it deep-linked to is the only impression you get.
The comparison already happened, so the product page is the landing page
Once you can see the traffic, brief for it, because it does not arrive the way anything else does.
On Shopify more than half of AI-referred sessions begin on a product detail page, against roughly 20% for organic search, and Adobe's July 2026 panel has AI shoppers generating 53% more revenue per visit, adding to cart 28% more and bouncing 33% less.
The homepage funnel is irrelevant to this person. The PDP is the landing page, and it is being read by someone who has just been told, in one sentence, why this product. So the hero's job is to confirm that sentence, in the shopper's own terms, in the first line.
The method I borrow comes from landing-page work built for ChatGPT Ads clicks, which OpenAI began testing in the US in January 2026, and I flag it as borrowed because that traffic is paid. An answer-continuation headline of 8 to 14 words that restates the claim and adds one specific benefit. A subhead of 20 to 30 words that kills the top objection. Specs, price and reviews above the fold, because the shopper arrives to verify a recommendation they already hold. Body copy in the Flesch 60 to 80 band, which is also what the assistant parses most cleanly. Reviews count twice, because the same page is what the assistant read before it sent the click: the first ten reviews on a product carry a 53% conversion uplift and customer photos a 137% purchase-likelihood lift.
What I'm still investigating is the loop from this traffic back into paid social. Nothing here tells you how a Meta buyer uses AI-referred sessions as a creative signal. My call is that the prompts people type at assistants are the best free hook research a paid social team has ever had, and on the accounts we run I would have the creative team pull the top 20 assistant prompts in their category every month and write hooks from them.
The citation is earned in editorial and review coverage first
The AEO brief is a content and PR brief, and it waits for the classified number.
Start with readability, because you cannot be cited from a page the model cannot parse. Adobe found roughly 25% of retailer homepage content not optimised for large language models and about 34% of individual product pages not accessible to AI at all. By page type in cosmetics, blog and news pages were 78% readable, category pages 71%, help pages 60%; grocery PDPs hit 70%. Readability is the floor a page has to clear.
Citation share is the number above it, and it is earned off-site. A PR-software vendor classified more than 25 million cited links from ChatGPT, Claude and Gemini across 17 industries: about 84% earned media, journalism alone 27%, paid and advertorial together 0.3%. Only about 11% of domains overlap between the top-ranking organic results and what the major assistants cite. For a retail query the footprint is editorial and review coverage first and your own site a distant second.
The mechanics that move it, as far as they have been measured:
- Content with quotations earns 40.6% higher AI visibility than unmodified content.
- Pages over 20,000 characters draw 10.18 citations each against 2.39 for pages under 500, a 4.3x multiplier.
- 44.2% of citations come from the first 30% of an article, so the answer goes at the top.
- 68.7% of ChatGPT citations follow a clean H1 to H2 to H3 hierarchy.
- ChatGPT cites 15% of what it retrieves, and triggers web search on 53.5% of commercial prompts against 18.7% of informational ones.
- Roughly 70% of pages cited in AI Overviews change over any two to three months, so the work never finishes.
This adds up to a brief for the people who write and pitch, run through editorial partners and review ecosystems, with the PDP work as hygiene. That is an eighteen-month programme carrying at least one writer and one PR lead, and I would staff it only after the classified AI row clears 1,000 sessions a month.
Grocery is the least machine-readable category and the least lifted
| Retail category | |
|---|---|
| Apparel | 76% |
| Electronics | 70% |
| Cosmetics | 68% |
| Sport. Goods | 67% |
| Furniture | 64% |
| Gen. Merch | 63% |
| Grocery | 59% |
The vendors with a mechanism and the vendors with a deck
Judge the AEO market the way a media buyer judges a publisher: does the thing have a mechanism I can see in my own numbers, or a deck?
| Claim or product | What it actually measures | Mechanism or deck |
|---|---|---|
| Daily prompt tracking on the $399 Growth tier, annual billing only | Your brand's presence across 100 prompts and 9,000 responses a month on ChatGPT, Perplexity and AI Overviews | Mechanism |
| The $99 Starter tier | ChatGPT only, 50 prompts, one seat | Mechanism, narrow |
| Agent-drafted content changes | Drafts; a person still publishes every change | Mechanism with hidden labour |
| GEO services market from $848M to $33.7B by 2034 | Vendor revenue projected at a 50.5% CAGR | Deck |
| 92% of marketers plan to optimise, 14% measure | Survey intent | Deck |
| llms.txt | Nothing; no major vendor honours it and no citation lift has been found | Deck |
| Princeton GEO study | Up to 40% visibility lift from statistics, quotations and citations in the paper's own benchmark | Mechanism, the one academic one |
The top three rows belong to the category's best-funded platform, which has the most credible mechanism in the market and, since February 2026, a $96M Series C at a $1B valuation and around 10% of the Fortune 500 as customers. It also has the episode the group chat sent round. On 21 May a screenshot circulated on LinkedIn of an email from its founding marketing engineer offering a $250 Amazon gift card for a G2 review. Its G2 count went from just over 300 in mid-May to above 1,100 by 30 August. On 24 June G2 published a $100 cap on review incentives that its guidelines had not carried the day before. Semrush had already taken the G2 AEO top spot by March, and the category now lists nearly 600 tools.
The sceptics have evidence too. Google's John Mueller called the acronym proliferation spam and scamming in August 2025; Danny Sullivan called AEO a subset of SEO. Nearly 40% of AI citations are wrong or invented, and OpenAI itself admits standard ChatGPT Search gets product information wrong roughly two thirds of the time.
My test is simple. A vendor who can show me the session count in my GA4 AI row before and after their work has a mechanism. One selling a visibility score with no sessions behind it has a deck, and the spend follows the classified number, never the other way round.
What the fix costs and the only honest incrementality design
Price the whole thing.
The channel group is under 30 minutes with no code. Wait 24 to 48 hours before comparing history, because some reports cache on that cycle, and take about a quarter of clean data before any strategic conclusion. The GTM dimension and the server-container header inspection are an analyst day each in our experience. BigQuery export is $5 to $50 a month. Two analyst days and a rounding error.
The cost of skipping it belongs on the board slide next to the AI row. One operator had 18% of paid traffic in Unassigned because the agency tagged utm_medium=paid where GA4 expects cpc. Rebuilding nine months of channel groups in BigQuery took two weeks, with roughly $400K of misdirected budget behind the gap.
If the number says go, the AEO run-rate is $1,188 to $4,788 a year for prompt monitoring plus the 20 to 30% content budget uplift one consultancy argues for. I would not sign the uplift until the group has run a quarter.
The measurement design the referrer cannot break: compare sessions with an observed AI source against Direct sessions on the same landing pages in the same weeks. The page match is the control; it removes the intent difference that otherwise dominates. Run it on a rolling window of at least eight weeks, because assistants change what they pass without notice and one month where an assistant switched will mislead the next quarter. Report three rows monthly: observed, from referrers and UTMs; probable, from a deep-Direct segment on those pages; self-reported, from a how-did-you-hear question whose under-reporting rate you calibrate from the observed group. Sample size is fixed by arithmetic: 1,000 AI sessions detects a 50% conversion difference, 5,000 detects 20%.
The evidence carries the landing-page control and no holdout. My design, if you want one: withhold the citation rewrite from half the PDPs in a category and read the page-matched AI row over eight weeks.
The identity problem has no fix. An AI session that converts on a later branded-search visit is handed to Organic by last-click and GA4's 30-day acquisition window. The three-row table is the only place it ever appears.
When not to bother, and the call I will be wrong about
Do not bother if you sell low-AOV commodity goods. The AI lift there is 1.1 to 1.4x and sometimes inside normal channel noise. Do not bother if the classified row is under 1,000 sessions a month, because you cannot detect anything from it. Do not bother much if you are grocery, the least machine-readable category at 59% and the one with least to gain.
Do not buy llms.txt. Do not buy a visibility score before you have a channel group. Do not commission a PDP rewrite for a channel at 0.10% of revenue when the Shop app on the same ten stores did $4.8M, 35x everything AI combined, and doubled. Do not let a vendor landscape shaped by $250 gift cards set your roadmap.
Do fix the channel group anyway. It is thirty minutes and it is the only way to learn you are the exception: durable goods at 1.5 to 2.0x, subscriptions 1.6 to 2.0x, B2B 1.8 to 2.3x.
Now the call, stated so it can be checked in September 2027. Correctly classified AI referral, native channel plus custom group plus UTM rule, will sit under 2% of sessions for the median US retailer and above 5% only in considered, high-AOV categories. The growth rate has already fallen from 4,700% in July 2025 to 62% in July 2026, and Semrush's projection that AI search delivers more visitors than traditional search by early 2028 does not survive that curve.
GA4 shipped the native channel on 13 May 2026, and it omits Perplexity and Claude and never reclassifies the dark share. My call is that Google adds both within twelve months and the Direct bucket never gives its AI sessions back. Which means the custom group and the exploration you build today are the only historical view of this channel you will ever have. Build them this week and then, for most of you, wait.





