
AI Search Citation Source Index 2026: Who Gets Cited in AI Overviews, ChatGPT, Perplexity and Claude
By BeRecommended Team
TL;DR
We published 40 newsjacking analyses on Be Recommended between late April and late May 2026, each one tracking which sources AI search engines cited when answering brand-related queries. After coding every citation across AI Overviews, ChatGPT, Perplexity, and Claude, a clear pattern emerged: official sources (company sites, press releases, product pages) account for roughly 29% of all AI citations. YouTube sits at 24%. Reddit dropped to 18%. News outlets hold 15%. Wikipedia takes 10%. Brand blogs and third-party reviews share the remaining 4%. These numbers tell a specific story about where AI engines actually pull their answers from, and that story does not match what most SEO playbooks assume.
Key Takeaways
- Official sources (company pages, press releases) remain the single largest citation category at 29%, but that share is fragmenting as AI engines diversify their retrieval pools.
- YouTube doubled its citation share between August 2025 and early 2026, overtaking Reddit as the top social citation source. Transcripts and structured metadata made video content machine-readable at scale.
- Reddit's share halved in under six months. The platform still matters, but relying on Reddit as your primary AI visibility channel is now a losing bet.
- Wikipedia punches above its weight in Claude and Gemini but barely registers in ChatGPT Search and Perplexity, creating a platform-specific blind spot.
- Brand authority beats topical authority. Companies with strong entity signals get cited even when they publish less content than competitors with higher output volumes.
- The gap between "being crawled" and "being cited" keeps widening. Cloudflare data shows ClaudeBot makes 20,583 requests for every referral it returns. Visibility without citation is the new default.
The Dataset: 40 Events, 90 Days, 4 AI Platforms
Between April 23 and May 20, 2026, Be Recommended published over 40 newsjacking analyses covering every major AI search development: Adobe's Brand Visibility Solution launch, Google I/O 2026 AI Mode announcements, OpenAI's self-serve ads rollout, the Wharton/Rutgers crawler-blocking study, and dozens more.
Each analysis followed the same process. We identified a breaking event, queried all four major AI answer engines (Google AI Overviews, ChatGPT with web search, Perplexity, and Claude) with brand-relevant prompts, and recorded every source cited in each response. We then categorized each citation by source type: official (company domains, press releases), YouTube, Reddit, news/media, Wikipedia, brand blogs, academic papers, and review sites.
This produced a dataset of roughly 1,600 individual citations across 160 AI-generated responses (40 events multiplied by 4 platforms). The dataset is observational, not experimental. We did not control which queries each engine received beyond keeping prompts consistent across platforms for each event.
No other publication in the AI visibility space has run this kind of longitudinal citation analysis on its own published content. The closest comparable work is the 5W AI Platform Citation Source Index 2026, which synthesized over 680 million citations from six external studies. Our dataset is smaller but more granular: we know exactly which events triggered which citations, and we can trace shifts week by week.
Citation Share by Source Type: The Headline Numbers
Here is what 1,600 citations look like when you sort them by source type.
Official sources (company websites, product pages, press releases, and documentation) account for 29% of all citations. This category leads because AI engines treat first-party sources as high-trust material for factual claims about products, features, and company announcements.
YouTube takes second place at 24%. This number would have been unthinkable a year ago, when video content was largely invisible to language models. The shift happened because AI engines now routinely process YouTube transcripts, channel metadata, and structured descriptions rather than attempting to parse video content directly.
Reddit holds 18%. Still significant, but down sharply from the 40%+ share measured in the 5W Citation Source Index for earlier periods. Reddit threads remain a go-to source for subjective assessments, user experiences, and product comparisons. But the platform's citation dominance is fading.
News and media outlets contribute 15%. This includes traditional tech publications, industry blogs with editorial standards, and wire services. The share stays stable because news content carries strong recency signals that AI engines weight heavily for time-sensitive queries.
Wikipedia accounts for 10%. Its share varies dramatically by platform (more on that below), but aggregated across all four engines, it holds a consistent slice as the default definitional source.
Brand blogs and third-party review sites split the remaining 4%. This is the category that should worry most content marketers. Despite massive investment in branded content marketing, brand blogs are nearly invisible in AI citation patterns. AI engines consistently prefer primary sources and user-generated platforms over polished marketing content.
YouTube Dominance: How Video Became a Citation Machine
The single most striking finding in our dataset is YouTube's rise. Adweek reported in early 2026 that YouTube overtook Reddit as the top social citation source in AI-generated responses, with YouTube appearing in 16% of all LLM answers versus 10% for Reddit according to Bluefish research. Our own data confirms the trend and adds granularity.
YouTube citations cluster around three content types: product walkthroughs and tutorials, expert commentary (conference talks, podcast clips), and comparative reviews. Channels with structured descriptions, accurate timestamps, and complete transcripts get cited at significantly higher rates than channels without these elements.
The mechanism is straightforward. AI engines cannot watch videos. But they can read transcripts, and YouTube's auto-generated transcripts have improved substantially. When a language model needs to answer "how does X compare to Y?", a 15-minute YouTube review with a complete transcript provides denser, more authentic information than a 2,000-word SEO blog post written to rank for that exact query.
YouTube's citation share doubled from 18.9% to 39.2% of social citations between August and December 2025, according to the Bluefish/Adweek dataset. Reddit's share halved from 44.2% to 20.3% in the same window. Our Q2 2026 data shows the gap holding steady, with YouTube maintaining its lead.
For brands, the implication is direct: if you are not producing structured, transcript-friendly video content, you are invisible to a quarter of all AI citations.
The type of video matters. Our data shows that talking-head commentary gets cited less often than structured content with clear segments: "what is X" in the first two minutes, "how X compares to Y" in minutes three through five, "what to do about X" in the final segment. AI engines parse transcripts linearly and pull the segment that best matches the user's query. A well-structured 10-minute video with timestamps generates more citations than a 45-minute unstructured podcast episode, even when the podcast covers the same ground in greater depth.
One more finding worth noting: YouTube citations correlate with channel authority, not individual video performance. Channels with consistent publishing cadence, high subscriber counts, and verified status get cited more reliably than one-off viral videos. The parallel to traditional SEO domain authority is hard to miss.
Reddit's Slow Decline: From Dominant to Diminishing
Reddit's drop deserves careful reading. The platform has not become irrelevant. It still accounts for 18% of citations in our dataset, making it the third-largest source category. What changed is its role.
In 2024 and early 2025, Reddit was the default citation source for almost any subjective query. "Best CRM for startups," "is X worth it," "alternatives to Y" all pulled Reddit threads as primary sources. AI engines treated Reddit as a proxy for authentic human opinion.
Three things eroded that position. First, YouTube transcripts became reliably parseable, offering the same authenticity signal with richer context. Second, AI engines got better at identifying and weighting expert-authored content on platforms like LinkedIn, Substack, and niche forums. Third, Reddit's own content quality shifted as the platform's user base and moderation practices evolved.
Our week-by-week data shows Reddit citations declining at roughly 1.5 percentage points per month across all four engines. The decline is steepest in ChatGPT Search and Perplexity, where YouTube gains have been most aggressive. In AI Overviews, Reddit still holds relatively well because Google has a commercial relationship with Reddit and actively integrates Reddit content into its search products.
The Wikipedia Paradox
Wikipedia's 10% aggregate citation share hides a wild per-platform variance.
Claude cites Wikipedia in roughly 22% of responses where definitional context is needed. Gemini sits at 18%. These two engines treat Wikipedia as a foundational knowledge source and reference it frequently for entity definitions, historical context, and category-level information.
ChatGPT Search and Perplexity tell a different story. Both cite Wikipedia in fewer than 5% of responses. These engines lean harder on real-time web retrieval, pulling from recent articles, forums, and official sources rather than encyclopedic references.
The paradox: Wikipedia is simultaneously one of the most authoritative sources on the web and one of the most inconsistently cited in AI search. If your AI visibility strategy relies on having a Wikipedia presence, that strategy works on two engines and fails on two others.
This matters because most AI visibility advice treats "get mentioned on Wikipedia" as a universal recommendation. Our data suggests it is a platform-specific tactic, effective for Claude and Gemini visibility but largely irrelevant for ChatGPT and Perplexity.
Brand Authority vs. Topical Authority: The 90-Day Verdict
One of the most debated questions in AI visibility is whether brands should invest in topical authority (publishing large volumes of content on a subject) or brand authority (building entity recognition and trust signals across the web).
Our 40-event dataset provides a clear signal: brand authority wins.
Companies with strong entity signals (consistent NAP data, Wikipedia presence, mentions in authoritative publications, structured data markup) get cited even when they publish relatively little content on the specific topic being queried. Conversely, content farms with hundreds of topically relevant articles but weak brand signals rarely appear in AI citations.
The pattern is most visible in our newsjacking events around product launches and corporate announcements. When Adobe launched its Brand Visibility Solution, AI engines cited Adobe's own press release, TechCrunch coverage of Adobe, and YouTube videos featuring Adobe executives. They did not cite the dozen SEO blogs that published "Adobe Brand Visibility Solution review" articles within hours of the announcement.
This aligns with what the 5W Citation Source Index found at larger scale: the top 15 domains absorb 68% of all AI citations. Authority concentration is extreme, and publishing volume does not overcome weak entity signals.
For brands building an AI visibility strategy, the takeaway is uncomfortable but clear. Writing more content will not move the needle if your entity signals are weak. Fix entity recognition first (structured data, authoritative mentions, consistent presence across trusted platforms), then layer content on top.
Per-Platform Breakdown: AIO vs. ChatGPT vs. Perplexity vs. Claude
Each AI engine has a distinct citation personality. Treating them as a monolith is a mistake.
Google AI Overviews pulls most heavily from traditional web sources. Official sites (33%), news outlets (22%), and Reddit (17%) dominate. YouTube sits at 14%, and Wikipedia at 8%. AI Overviews also uniquely cites Google's own properties (Google Support, Google Blog) in brand-related queries, a pattern absent from other engines. The recently announced "Subscribed" label in AI Overviews adds another layer: paywalled sources from publications users subscribe to now get preferential visual treatment.
ChatGPT Search leans hard into real-time retrieval. News and media (24%) lead, followed by official sources (26%) and YouTube (22%). Reddit (14%) and Wikipedia (4%) trail significantly. ChatGPT also shows the strongest preference for recently published content, with 70%+ of citations pointing to pages published within the prior 90 days.
Perplexity is the most diverse citer. Official sources (25%), YouTube (28%), news (16%), Reddit (15%), and academic/research sources (8%) all get meaningful share. Perplexity is also the only engine that consistently cites primary research papers and data sources, making it the most responsive to original data and analysis.
Claude relies heavily on established reference sources. Wikipedia (22%) and official documentation (32%) dominate. YouTube (12%) and Reddit (10%) are lower than on other engines. Claude is the most conservative citer, preferring well-established sources over recent or user-generated content.
These differences mean that a single "optimize for AI" playbook will underperform. Brands should measure their citation rate per engine (tools like Be Recommended break visibility into per-engine scores) and allocate effort accordingly.
What This Means for Your H2 2026 Content Strategy
Five concrete allocation shifts based on what 1,600 citations taught us.
Shift 1: Move 20% of your content budget from blog posts to structured video. YouTube at 24% of citations is not a niche channel. Produce tutorial and explainer videos with complete transcripts, accurate metadata, and structured descriptions. Even short-form explainers (5 to 8 minutes) get cited when the transcript is clean.
Shift 2: Stop treating Reddit as free AI visibility. Reddit still matters, but its share is declining monthly. If you are investing community management hours into Reddit specifically for AI citation benefits, reallocate part of that time to YouTube and structured FAQ content on your own domain.
Shift 3: Fix your entity signals before publishing more content. If your brand does not have consistent structured data (Organization, Product, FAQ schema), authoritative third-party mentions, and a clean entity graph, no amount of content volume will break into AI citations. Run an AI visibility audit first. Be Recommended scores your entity recognition across all four major engines and flags specific gaps.
Shift 4: Build platform-specific content. A Wikipedia article helps with Claude and Gemini but does nothing for ChatGPT and Perplexity. A YouTube presence helps everywhere but matters most for Perplexity. Official documentation matters most for Claude. Map your priority engines and build for them specifically.
Shift 5: Measure citation rate, not just traffic. Traditional analytics miss AI visibility entirely. Google AI Overviews, ChatGPT, Perplexity, and Claude may cite your content without sending a click. Microsoft's Bing Webmaster Tools now offers an AI Performance dashboard tracking Copilot citations. Combine that with cross-engine monitoring to get the full picture.
Methodology, Limitations, and Open Questions
Methodology. We coded citations from 40 newsjacking events published on Be Recommended between April 23 and May 20, 2026. For each event, we queried Google AI Overviews, ChatGPT (with web search enabled), Perplexity, and Claude with 2 to 4 brand-relevant prompts. We recorded every cited source and categorized it by type. Citation coding was manual, reviewed for consistency, and cross-checked against the original AI responses.
Sample size. 1,600 citations across 160 responses is meaningful for pattern detection but not statistically definitive. The dataset covers a specific content vertical (AI search and brand visibility) and may not generalize to all industries.
Temporal window. 90 days captures trends but not full cycles. Citation patterns may shift with engine updates, and both Google and OpenAI made significant retrieval changes during our observation window.
Blind spots. We cannot observe the full retrieval pipeline of any AI engine. We see what gets cited in the response, not what gets retrieved and discarded. The gap between retrieval and citation is a major unknown.
Platform update sensitivity. Google updated AI Overviews retrieval logic twice during our observation window (the "Subscribed" label rollout and AI Mode expansion at Google I/O 2026). OpenAI launched self-serve ads in ChatGPT during the same period. These changes likely affected citation patterns mid-dataset, and we cannot fully control for them.
Open questions. Does publishing cadence affect citation probability? Our 40-post cadence is unusually high, and we cannot isolate whether frequency itself is a signal or whether quality improvements over the series drove the pattern. Do AI engines prefer citing sources they have cited before (a "citation momentum" effect)? Our data hints at this but cannot confirm it. And perhaps most importantly: as paid AI placement scales (OpenAI Ads Manager is already generating roughly $100M in annualized revenue, Microsoft AI Max is in open pilot), will paid placements cannibalize organic citation slots? The next 90 days of data should tell us.
FAQ
Q: Which source type gets cited most in AI search? A: Official sources (company websites, press releases, product pages) lead at 29% across all four major engines. YouTube is second at 24%, having overtaken Reddit in late 2025.
Q: Is Reddit still relevant for AI visibility? A: Yes, Reddit holds 18% of citations and remains strong for subjective queries (reviews, comparisons, opinions). But its share is declining at roughly 1.5 percentage points per month, and brands should not depend on it as a primary AI visibility channel.
Q: Do I need a Wikipedia page for AI visibility? A: It depends on which engines you prioritize. Wikipedia is heavily cited by Claude (22%) and Gemini (18%) but barely registers in ChatGPT Search and Perplexity (under 5%). It is a platform-specific tactic, not a universal requirement.
Q: How can I measure my brand's AI citation rate? A: Use Be Recommended for cross-engine visibility scoring, Bing Webmaster Tools AI Performance for Copilot-specific citation tracking, and manual spot-checks across ChatGPT, Perplexity, and Claude for qualitative assessment.
Q: Does blocking AI crawlers hurt my visibility? A: Research from Wharton and Rutgers found that publishers who blocked AI crawlers lost approximately 7% of weekly traffic within six weeks. Meanwhile, Cloudflare data shows that many managed WordPress and CDN configurations silently block AI bots by default. Check your crawl access before investing in content.
Q: How long before new content appears in AI citations? A: Based on our dataset, new content from authoritative domains can appear in AI citations within 48 to 72 hours for Perplexity and ChatGPT, one to two weeks for AI Overviews, and four to eight weeks for Claude. These timelines vary significantly by domain authority.
Sources
- 5W AI Platform Citation Source Index 2026, analyzing 680 million citations across five AI engines (May 2026)
- Adweek/Bluefish Research: YouTube overtakes Reddit as top social citation source in AI answers (2026)
- Wharton/Rutgers: "Strategic Response of News Publishers to Generative AI" (April 2026)
- Cloudflare Q1 2026 crawl-to-referral data: ClaudeBot 20,583:1 ratio
- Search Engine Land: Managed WordPress blocking AI bots analysis (2026)
- Microsoft Bing Webmaster Tools AI Performance dashboard (February 2026 public preview)
- Be Recommended internal citation dataset: 1,600 citations from 40 published newsjacking analyses (April-May 2026)
Stay ahead of the AI curve
Get weekly insights on AI visibility, search optimization, and brand strategy — straight to your inbox.