11.0% of AI Overview claims are not supported by the pages cited, and better sourcing does not fix it.
Google AI Overviews cite credible sources and still put words in their mouths. A Washington University team broke 7,583 AI Overviews into 98,020 atomic claims and checked each one against the pages the summary itself pointed to. 11.0% of those claims are not supported by the cited pages. AI citation accuracy, measured this way, turns out to be independent of source quality: the correlation between how credible the cited domains are and how faithful the answer is sits at r ≈ 0.045.
For B2B teams, that breaks the assumption underneath most AI visibility dashboards. A citation counts as a win. This research says a citation tells you your URL was surfaced, not that the sentence attached to it reflects what you published.
Three arXiv papers landed between May and September 2026 measuring this. Here is what they found and what to change in your measurement stack.
Haofei Xu, Umar Iqbal, and Jacob M. Montgomery at Washington University in St. Louis ran a 40-day crawl from March 13 to April 21, 2026, issuing 55,393 trending queries pulled daily from the US Google Trends dashboard across 19 topical categories. Of those, 7,583 returned an AI Overview, an activation rate of 13.7%. The paper, Measuring Google AI Overviews: Activation, Source Quality, Claim Fidelity, and Publisher Impact (arXiv 2605.14021), published May 13, 2026, and the September aggregators picked it up as the first large-scale claim-level audit of the format.
The pipeline split each summary into self-contained factual assertions, then handed every assertion to a verifier alongside the full body text of each page the overview cited. The verifier assigned one of five labels: Clear, Vague, Ambiguous, Incorrect, or Omitted. Two annotators checked a stratified sample of 100 verdicts and agreed with the machine on 98 of them, Cohen's κ = 0.94. Weighted by how often each label occurs, verifier accuracy lands at 95.6%.
Track this metric for your brand → nobori.ai
Across 98,020 claim-level judgments, 88.97% are consistent with the cited pages and 11.03% are not. The full split:
The per-overview picture is less alarming than the aggregate suggests, and more uneven. The median AI Overview has 93.33% of its claims grounded. 3,141 of 7,491 verifiable overviews, or 41.9%, are perfectly grounded. But 205 overviews (2.74%) have fewer than half their claims supported, and 64 (0.85%) have zero supported claims at all. A reader gets no signal about which bucket the answer in front of them belongs to.
Each overview carries a mean of 12.9 claims, median 12, ranging from zero to 64. At 11% unsupported, roughly one and a half assertions in a typical AI Overview trace back to nothing in the sources shown.
Omitted means the summary attributes a fact to your page and your page does not contain that fact. The cited source does not contradict the claim, it simply never addresses it. When an AI Overview claim fails, it is 2.6 times more likely to be a fact no cited source mentions than a fact a cited source contradicts.
This matters more for B2B brands than the contradiction case, because omission is invisible to everyone involved. Your analytics show a citation. Your AI visibility tool logs a mention. The buyer reading the answer sees your domain attached to a claim about pricing, integration support, or compliance coverage that you never made. Nobody in that chain has a reason to check.
The authors call the Incorrect category, 2,609 claims, "a failure mode a reader who sees only the AIO summary has no way to detect." The same logic applies to your team reading a visibility dashboard.
Because the two move independently. The paper reports a correlation of r ≈ 0.045 between source quality and claim fidelity, and states that "improving the former is unlikely to reduce the unsupported claim rate on its own."
That finding kills a common assumption. Marketers building E-E-A-T signals, chasing authoritative backlinks, and cleaning up schema often treat those investments as a path to accurate representation. The data says those investments change whether you get cited. They do not change whether the sentence next to your citation is true.
Google is sourcing well by its own standard. Across 37,020 AI Overview references that resolve to a credibility score, mean credibility runs 0.732 against 0.645 for the matched first-page results, a gap of 0.087 on a 0 to 1 scale (95% CI 0.085 to 0.089, Welch's t = 80.9, p ≪ 0.001). The paper notes this "directly contradicts prior work suggesting that AIOs draw on lower-quality sources than traditional results." Google solved source selection. It has not solved the step where the model turns those sources into sentences.
Health tops the list at 94.77% consistent, followed by Politics at 93.65%, Science at 91.82%, Business and Finance at 91.42%, and Hobbies and Leisure at 91.40%. Jobs and Education sits at the bottom of the measurable range at 76.85%.
Health and Politics leading is not an accident. The authors suggest Google applies stricter grounding requirements to categories that fall under its own YMYL (Your Money or Your Life) content policy. Politics also draws the lowest activation rate in the study at 7.5%, alongside Law and Government at 9.6%, which suggests Google suppresses AI Overviews on sensitive topics rather than trying to get them right. Health breaks that pattern at 26.6% activation, so suppression is not applied uniformly.
After the researchers strip out real-time content like weather feeds and stock tickers, which the crawler cannot verify against a page that has since changed, 18 of 19 categories land inside an 85.9% to 94.8% band. Business and Finance, the category closest to B2B buying research, sits at 91.42%. Read that as roughly one unsupported claim in every eleven.
The 11.0% figure is conservative. The same research group calls its combined 9.64% Omitted-plus-Incorrect rate "a conservative ceiling on substantive AIO unfaithfulness rather than a precise estimate," and their own sensitivity analysis puts a floor at roughly 5.3% if you assume every claim sourced to an uncrawled social platform is in fact supported.
Other measurements run higher. Oumi, working at the request of The New York Times, ran OpenAI's SimpleQA dataset through Google Search in early 2026 and found that only 39% of Gemini 3 powered AI Overviews were both correct and fully supported by their citations. At the claim level, Oumi put support at 67%, meaning a third of claims failed. Oumi also found Gemini 3 more accurate than Gemini 2 while hallucinating more, which fits the broader pattern of reasoning models trading grounding for fluency.
Earlier benchmarks were worse still. A 2025 evaluation of multiple LLMs with web access found between 50% and 90% of response statements were not fully supported by cited sources. The Tow Center at Columbia documented confident misattribution in the majority of ChatGPT Search cases it tested in 2024.
The spread comes from query mix. SimpleQA is built from hard factual questions with verified answers. The Washington University crawl uses whatever is trending, which skews toward Sports at 51.4% of the corpus. Both designs point the same direction: citation and substantiation are different events.
Across the 7,583 overviews, 29.8% of cited domains do not appear anywhere on the corresponding first page of results. At the URL level, 28.5% (17,451 of 61,206) come from hosts the first page never surfaces for that query. Google is running a source selection mechanism distinct from its ranking algorithm.
The off-page citations are the better ones. Among those 17,451 references, mean credibility is 0.758 and user-generated content accounts for 3.4%. Among the 43,755 references that also appear on page one, credibility is 0.724 and UGC runs 18.5%. Google reaches past its own rankings to find cleaner sources.
Citation concentration also differs from search. The top five hostnames take 20.0% of AI Overview citations against 39.1% of first-page results, and 56.3% of the unique hosts cited over the 40 days were cited exactly once. A long tail exists here that does not exist in the top 10. We covered the cross-engine version of this pattern in AI Citation Overlap: Engines Share 9% of Sources, 90% of Brands.
Three costs, in ascending order of damage.
First, wasted optimization. If you are grading content changes by citation count, you are measuring retrieval and reporting it as representation. A page can gain citations while the claims attached to it drift further from what it says.
Second, misinformed buyers. Pricing is the highest-inaccuracy claim theme in AI answers, per Profound's August 2026 FactCheck analysis of more than 158,000 claims. A prospect who arrives with a stale price or a feature you deprecated costs your sales team a call.
Third, competitor capture. When an AI Overview omits a claim from its sources, it is pulling from somewhere the pipeline could not verify, often a forum thread or a comparison page you do not control. The claim lands under your citation. Your domain lends it credibility. You never wrote it.
See if AI is describing your product accurately → nobori.ai
A preregistered field experiment gives the causal number that observational studies could only estimate. Stephanie T. Wang, Jeffrey Gleason, Yakov Bart, Christo Wilson, and Danaé Metaxa recruited 1,100 US Chrome users in waves between March 17 and 19, 2026, then randomized their exposure to Google's AI surfaces. The paper, AI in Search Reduces Publisher Referrals Without Improving User Experience (arXiv 2608.18352), went up on August 18, 2026.
Assignment to an AI Mode experience reduced click-through rate by 18.8 percentage points, with a 95% confidence interval from 22.2 to 15.3. Running it the other direction, removing AI Overviews and AI Mode raised click-throughs to publishers. The experiment also found AI Mode eroded user experience and trust in information found on Google, which undercuts the argument that lost clicks buy a better product.
Two calibration numbers from the same panel. Participants saw AI Overviews on 36% of their real searches, far above the 13.7% trending-query crawl rate, because everyday browsing skews more conversational. And AI Mode accounted for 0.6% of searches, so the surface with the steepest measured click penalty is also the one users pick least often.
Yes, and no single study measures all of them the same way, which is the reporting problem. Independent tests in 2026 put unsupported claim rates for You.com and Perplexity between 23% and 47%, with citation accuracy across engines landing in a 40% to 68% range. Profound's FactCheck work found Claude 1.3 times more inaccurate than ChatGPT, with the two overlapping on only 11% of their inaccurate claims.
Engines also disagree about which pages to read. Analysis of 22.7 million citations from January to June 2026 found Perplexity never touches 89.1% of the websites ChatGPT cites. Different corpora produce different errors about the same brand.
Practical consequence: a blended accuracy score across engines hides everything useful. If ChatGPT has your pricing right and Gemini has it wrong, an averaged number tells your team nothing about which one to fix. Audit per engine. We break down the per-engine reporting framework in AI Brand Sentiment: How to Change What AI Engines Say About You.
The publishers whose pages supply the citations. More than half of AI Overview cited pages, 50.63% (30,994 of 61,212), display visible ads. Another 14.2% of references point to social and video platforms the crawler treats as ad-free even though they run large ad businesses, so the real ad-supported share is higher.
Google's own inventory stays intact. Of 7,583 AI Overview bearing result pages, 164 (2.16%) also carried a sponsored search ad, and 39 (0.51%) placed one above the AI Overview block. None appeared inside the overview container. The researchers read this as AI Overviews being additive to Google's ad inventory rather than substitutive.
Two quasi-experimental papers put numbers on the ecosystem effect. A Wikipedia difference-in-differences study (arXiv 2602.18455, version 6 dated September 2, 2026) estimates English Wikipedia lost 5.45% of search referrals against a German benchmark and 4.82% against a French one. A Reddit study (arXiv 2605.16428) found Safe-for-Work communities, which AI Overviews are allowed to surface, gained 12.0% in daily comments and 12.4% in commenting users relative to Not-Safe-for-Work communities that Google prohibits from AI Overview references. The gains concentrated in advice and personal experience threads, and the later arrival of AI Mode largely erased them.
Five steps, running about half a day for a mid-size B2B catalog.
1. Build a claim inventory. List the 20 factual assertions about your company that a buyer needs correct: pricing tiers, seat minimums, contract length, SOC 2 status, native integrations, data residency, SLA, deprecated features. Write the correct answer next to each.
2. Prompt every engine. Run each claim as a direct question across ChatGPT, Gemini, Perplexity, Google AI Overviews, and Claude. Log the answer and the citations verbatim.
3. Verify claim by claim. For each assertion in the answer, open the cited page and find the supporting sentence. Mark it supported, contradicted, or absent. The absent bucket will be your largest, matching the research.
4. Trace the absent ones. When a claim is not on the cited page, search the claim text. You will usually find the real origin: an outdated comparison site, a review left three pricing changes ago, a Reddit thread, a partner page. That source is your remediation target, not your own site.
5. Score and re-baseline. Report claim fidelity as a percentage per engine, separate from citation count. Two brands with identical citation volume can have a 30-point fidelity gap.
Days 1 to 3. Run the claim inventory and the five-engine prompt sweep. Nothing else on this list works without that baseline.
Days 4 to 6. Publish an explicit, machine-readable facts page: current pricing with an effective date, integration list, certifications with issue dates, and the deprecations you want to stop seeing quoted. Give every number a visible last-updated stamp. Omission happens when the pipeline cannot find the fact anywhere it trusts.
Days 7 to 9. Fix the off-site sources feeding the wrong claims. Update your G2, Capterra, and TrustRadius listings, correct partner directory entries, and file corrections with any comparison site carrying stale figures. Off-page sources drive more of your representation than your own domain does.
Days 10 to 12. Rewrite the pages that already earn citations so the load-bearing facts sit in short, self-contained sentences near the top. A model that has to paraphrase a buried fact introduces drift at that step.
Days 13 to 14. Split your reporting. Citation count and claim fidelity become two metrics with two owners. Re-run the sweep monthly and track fidelity as a trend line.
Scale is running ahead of grounding. Google put AI Overviews at more than 2.5 billion monthly active users at I/O on May 19, 2026, and AI Mode above 1 billion. The Washington University authors decline to extrapolate an 11% unsupported claim rate to that user base. Several news outlets ran the multiplication anyway and landed on millions of bad attributions per hour.
Three things to watch through the rest of 2026. Regulators are shifting from traffic complaints to accuracy complaints, and claim-level audits like this one hand them a measurable standard. Google's source selection operates independently of ranking, so the gap between what ranks and what gets cited should widen rather than close. Measurement vendors will ship fidelity scores once enough buyers ask, and this research gives buyers the vocabulary to ask.
Your team spent 2025 working out how to get cited. The work that pays in 2026 is checking what the citation says, and most B2B marketing teams have no process for it yet.
What is AI citation accuracy?
AI citation accuracy measures whether the claims in an AI-generated answer are actually supported by the pages that answer cites. It differs from citation count, which only records that a URL was surfaced. The Washington University study found 11.0% of 98,020 AI Overview claims unsupported by their cited pages.
What is the difference between an omitted and an incorrect claim?
An omitted claim is a fact no cited source mentions at all. An incorrect claim is a fact a cited source explicitly contradicts. Omission runs 6.98% and contradiction 2.66%, making omission 2.6 times more common.
Does getting cited by an AI Overview help my brand?
Citation raises visibility and correlates with higher click-through on the queries where it happens. It does not guarantee the claim attached to your citation reflects your content. Treat citation and representation as separate measurements.
Why do AI Overviews cite pages that don't rank on page one?
Google runs source selection separately from ranking. 29.8% of cited domains never appear on the first page for the same query, and those off-page citations score higher on credibility than the on-page ones.
How often do AI Overviews appear?
13.7% of trending queries in the Washington University crawl, rising to 64.7% for question-form queries against 9.5% for non-questions. A real-user browsing panel of 1,100 people saw them on 36% of searches, because everyday queries skew conversational.
Which AI engine is most accurate?
No study measures all engines on the same methodology, so cross-engine accuracy rankings are unreliable. Independent 2026 tests place unsupported claim rates between 23% and 47% for several engines. Audit each engine separately against your own facts.
Ready to see where you stand in AI search?
Nobori tracks your brand's visibility across ChatGPT, Gemini, Perplexity, Google AI Overviews, and Claude, updated daily. See who's getting cited, where you're missing, and what to fix.
Get daily AI visibility alerts for your brand → nobori.ai