
Your Brand Citation Might Be Wrong: Why Citation Count Is No Longer Enough in AI Search
By Jakub
TL;DR
Columbia's Tow Center tested eight AI search engines on 1,600 news queries. Collectively, they answered incorrectly over 60% of the time. ChatGPT Search was completely wrong on 57% of queries and never once declined to answer. Licensing deals with publishers did not improve accuracy. For brands, this means citation count alone is a dangerous metric. You need to measure citation accuracy, completeness, and reputational signal to understand whether AI search is helping or hurting your brand.
What Happened
On May 20, 2026, Bloomberg reported on updated findings from the Tow Center for Digital Journalism at Columbia University. The researchers had systematically tested eight generative search tools on 200 news excerpts from 20 publishers, asking each chatbot to identify the source article, publisher name, publication date, and URL.
The results were sobering. ChatGPT Search was completely accurate on just 28% of queries and completely wrong on 57%. Perplexity fared somewhat better at 37% completely incorrect. Grok 3 was the worst performer at 94% error rate. Across all eight engines, the collective inaccuracy rate exceeded 60%. The Tow Center had run a similar study 14 months earlier with comparable results, suggesting that despite model upgrades, AI search accuracy has not measurably improved.
Why It Matters for AI Visibility
The shift from "get cited" to "get cited correctly." Most GEO playbooks still focus on citation frequency: how often does your brand appear in AI-generated answers across a set of prompts? Tools like HubSpot's AEO Sensor and Profound measure this well. But the Tow Center data reveals a blind spot. A citation that attributes a competitor's claim to your brand, links to a syndicated copy instead of your canonical page, or fabricates a URL with your domain name is worse than no citation at all. The GEO conversation needs a second layer: citation quality.
Licensing deals don't fix the problem. One of the study's most striking findings was that publishers with content licensing agreements with OpenAI or Perplexity were not cited more accurately than those without deals. The San Francisco Chronicle, part of Hearst's strategic content partnership with OpenAI, was correctly identified in only one out of ten test queries. This matters for brands investing in AI search partnerships or structured data feeds. A licensing agreement or data pipeline solves your legal exposure, not your citation quality.
Confident false assertions are a systemic brand risk. ChatGPT incorrectly identified 134 articles in the study but only used hedging language 15 times out of 200 responses. It never declined to answer. When an AI search engine doesn't know the right source and fabricates one with high confidence, your brand can be attributed to claims you never published. The Tow Center documented cases where publishers were systematically misattributed across queries, with chatbots citing syndicated or plagiarized copies instead of canonical sources. For a deeper look at how to detect this, see our AI search citation audit guide. Scale this pattern across millions of daily AI search queries, and misattribution becomes a compounding reputational risk that most brands aren't tracking.
The problem is expanding beyond news into commerce. OpenAI's recent ChatGPT Shopping Research launch brings the same retrieval architecture into product recommendations, buyer guides, and brand comparisons. The Tow Center methodology was built for news, but the accuracy concerns extrapolate directly to any domain where AI search synthesizes information from multiple sources and attributes it with confidence.
What Brands Should Do Now
Audit your current AI search reporting for quality gaps. If your monitoring dashboard only shows citation count (how many times your brand appears across N prompts), you're seeing half the picture. Run a manual spot-check: sample 20 brand mentions across ChatGPT, Perplexity, and Gemini. For each, verify the brand name, URL, quoted claim, and surrounding context. Note every misattribution, fabricated URL, and out-of-context reference. That sample misattribution rate is your new baseline.
Set an escalation threshold. If your sample misattribution rate exceeds 15%, escalate it to your quarterly brand-risk review. A misattribution rate above 15% means AI search engines are actively associating your brand with claims you didn't make, at a frequency that compounds over time. Set up recurring monitoring, not just one-off audits.
Strengthen your canonical brand pages as defensive infrastructure. AI retrieval systems prefer structured, well-cited, schema-clean content over scraped HTML. Make sure your about page, product pages, and key comparison pages have clean JSON-LD markup (Organization, Product, FAQ), consistent NAP data, and authoritative internal and external links. This won't prevent every misattribution, but it gives the retrieval layer the best possible input for correct attribution.
Check your AI search presence after every major content publish. Don't assume that publishing a well-optimized blog post means AI search will represent it correctly. Within 48 hours of publishing, query your target keywords in ChatGPT, Perplexity, and Gemini. Verify that your brand is cited, the attribution is accurate, and the link points to your canonical URL. Document discrepancies and track them over time.
Key Takeaways
Citation count was GEO 1.0. Citation quality is GEO 2.0. Here's a four-layer measurement framework:
- Citation count measures how often your brand appears in AI answers across a prompt set. Existing tools (Profound, HubSpot AEO Sensor, Athena) cover this well. Necessary, but insufficient.
- Citation accuracy measures how many of those citations correctly attribute your brand name, URL, quoted content, and context. This is the layer the Tow Center study exposes as broken. Track your sample misattribution rate and false-positive citation rate.
- Citation completeness measures whether AI search surfaces your brand in the contexts where it belongs. Are you cited for your core competencies, or pulled out of context into tangential queries?
- Citation reputational signal measures the sentiment framing around your brand citation. Is the AI engine positioning you positively, neutrally, or negatively? Brand sentiment in AI-generated answers is the newest blind spot in reputation monitoring.
The Tow Center gave the industry a data anchor. The next step is operationalizing these findings into brand-side measurement. If your monitoring stack stops at citation count, misattributions accumulate in the AI search index months before you catch them.
FAQ
How accurate is ChatGPT Search at citing sources?
According to the Tow Center for Digital Journalism at Columbia University, ChatGPT Search was completely accurate on 28% of news queries and completely wrong on 57%. It rarely acknowledged uncertainty, using hedging language in only 15 out of 200 responses.
Do content licensing deals with OpenAI improve citation accuracy?
No. The Tow Center study found that publishers with licensing agreements were not cited more accurately than those without deals. Licensing addresses legal exposure, not retrieval quality.
What is the difference between citation count and citation accuracy?
Citation count measures how often a brand appears in AI-generated answers. Citation accuracy measures whether those citations correctly attribute the brand name, URL, and context. A high citation count with low accuracy means AI search may be misrepresenting your brand at scale.
How can brands monitor AI search citation quality?
Start with a manual spot-check: sample 20 brand mentions across ChatGPT, Perplexity, and Gemini. Verify brand name, URL, quoted claim, and context for each. Track your misattribution rate over time and escalate if it exceeds 15%.
Does AI search inaccuracy affect product recommendations too?
The Tow Center study focused on news, but the same retrieval architecture powers product recommendations and buyer guides in tools like ChatGPT Shopping Research. Accuracy concerns apply wherever AI synthesizes and attributes information from multiple sources.
Sources
- Tow Center for Digital Journalism: AI Search Has a Citation Problem (Columbia Journalism Review, March 2025)
- Tow Center: How ChatGPT Search (Mis)represents Publisher Content (Columbia Journalism Review, November 2024)
- AI search engines fail to produce accurate citations in over 60% of tests (Nieman Journalism Lab, March 2025)
Stay ahead of the AI curve
Get weekly insights on AI visibility, search optimization, and brand strategy — straight to your inbox.