AI Citations: 23 Signals to Boost Your Visibility
Last updated on September 8, 2026 at 12:09 PM.AI citation describes the mechanism by which ChatGPT, Perplexity or Google AI Overviews select and link to a webpage as a source within a generated answer. Pages cited in Google AI Overviews receive 120% more organic clicks per impression than non-cited pages, according to an analysis by Seer Interactive. Citability does not arise from good content alone – it requires technical accessibility, structural extractability and distributed authority across multiple platforms. This article presents the 23 evidence-based signals that AI systems use for source selection, explains the differences between platforms and provides a step-by-step guide for building citability in a B2B context.

Why position 1 on Google is no longer enough
Position-1 CTR on Google has dropped from around 27% to approximately 11%. Google AI Overviews appear for roughly 30% of all US search queries. ChatGPT processes several hundred million weekly active users by its own account. Perplexity has become the standard research tool for analysts, journalists and technical buyers. Anyone who does not appear as a source in these answers loses both traditional clicks and AI visibility.
The consequence for marketing decision-makers: citation optimisation – the deliberate optimisation for citability by AI systems – is no longer an experimental channel but a business-critical priority. The question is no longer "Do we rank on Google?" but "Are we selected as a source by ChatGPT, Perplexity and Google AI?" Google and AI answers use two different selection mechanisms. Visibility in ChatGPT, Perplexity and similar platforms follows different rules than classic ranking.
How to build visibility for both Google and AI answers simultaneously without creating two misaligned workstreams is demonstrated by Crispy Content through a data-driven and automated approach to Agentic SEO and Generative Engine Optimization.
Connected to this is a second question that gains weight in the AI era: How can a distinctive brand voice be defined in a way that an AI can reproduce it – rather than delivering the generic sound that remains interchangeable?
How a distinctive brand voice can be defined through voice profiles, corporate voice systems and styleguides is explained by Crispy Content in detail under Voice-Style-Engineering.
All of this only works, however, when backed by a strategy that fits the target audience and keeps measurable results in view – the concrete question of what growth costs and how you know it is working.
How competitive analyses, data-driven communication strategies and social media approaches can be developed into a methodology that targets sustainable growth for industrial companies, technology providers and service businesses alike is described by Crispy Content in its strategy portfolio.
What does AI citation mean? Core concepts and definitions
AI citation is the process by which a Large Language Model (LLM) selects and links to a webpage as evidence for a statement within a generated answer. Generative Engine Optimization (GEO) is the discipline that prepares content so that AI systems prefer it as a source. Retrieval-Augmented Generation (RAG) provides the technical foundation: the model retrieves external sources before generating an answer, reads them and selectively decides which ones to cite. Anyone who can distinguish these three terms understands the mechanism behind every AI source attribution.
Generative Engine Optimization (GEO) – Distinction from classic SEO
GEO optimises individual text passages for retrieval and citation in AI answers; SEO optimises entire pages for rankings and clicks. The Princeton study at KDD 2024 (Aggarwal et al.) showed: GEO techniques such as adding statistics, source references and direct quotes increase AI visibility by 30–40%, measured by Position-Adjusted Word Count. The difference lies in the optimisation unit: not the page as a whole counts, but the individual paragraph that can be extracted as an answer building block.
| Attribute | SEO | GEO |
|---|---|---|
| Optimisation goal | Ranking position in SERPs | Citation in AI answers |
| Success metric | Rankings, CTR, organic traffic | Citation rate, share of voice in LLM answers |
| Time horizon | Months to years | Weeks to months, continuous updating |
| Optimisation unit | Entire page (title, meta, body) | Individual paragraph or section |
| Primary lever | Backlinks, technical signals | Structural extractability, fact density |
Retrieval-Augmented Generation (RAG) – How AI systems select sources
RAG describes the process by which an LLM retrieves external sources before generating an answer, reads them and selectively cites them. ChatGPT retrieves multiple candidate pages but cites only a fraction – approximately 15% of retrieved pages, according to an analysis by Leapd (Azamfar, 2026). Perplexity conducts a real-time web search for every query and cites an average of around 22 sources per answer (Yext, 2025). The selection ratio is the decisive point: retrieval does not mean citation. Between the two lies the evaluation based on relevance, structure and trustworthiness.
Fan-out queries – Why a single keyword is not enough
AI systems decompose complex queries into multiple sub-queries, known as fan-out queries. The source that appears across as many of these sub-queries as possible gets cited. AirOps demonstrated a clear correlation between ranking for related queries and AI citations (referenced in Shepard, 2026). Anyone optimising for only one head keyword covers at best one of five or six sub-queries – and loses the citation to domains with broader topical coverage.
The core principle: How AI systems evaluate and select sources
AI systems evaluate sources in a three-step sequence: technical reachability, then semantic relevance, then trustworthiness. Only pages that pass all three stages get cited. Anyone who fails at the first stage – for example through a blocked robots.txt – never reaches the evaluation for relevance and authority.
Cyrus Shepard, former Head of Global SEO at Moz, analysed 54 studies and patents in 2026 and distilled 23 citation factors from them, weighted by replicability and evidence strength (Shepard, 2026). AI citation works like a peer-review process in academia: the model looks for the most reliable, best-structured and most easily extractable source.
"Most of the critical factors align with traditional SEO practices: win SEO, win AI citations – most of the time, with extra steps."
– Cyrus Shepard, former Head of Global SEO at Moz, May 2026
| Factor | Score | Key finding |
|---|---|---|
| URL accessibility | 9.5 | Page must be reachable by AI crawlers |
| Search rank | 9.4 | 38% of AI Overview citations come from the Google top 10 |
| Fan-out rank | 9.3 | Ranking for related sub-queries |
| Query-answer match | 9.2 | Semantic alignment with the question |
| Preview control | 9.2 | No nosnippet directives |
| Intent-format match | 9.0 | Appropriate format for search intent |
| Topic cluster rank | 8.9 | Deep topical coverage |
| Answer near page top | 8.8 | 44.2% of all LLM citations come from the first 30% of the text |
| AI-ready structure | 8.6 | Clear H2/H3 hierarchy, tables, lists |
| Factual specificity | 8.3 | Concrete numbers instead of vague statements |
Source: Shepard (2026), Zyppy Signal – AI Citation Ranking Factors.
The five signal groups of citability in detail
The 23 factors can be clustered into five groups: technical accessibility, relevance, structural extractability, content quality, and authority and trust. Each group addresses a different stage in the retrieval process. Ignoring one group means the other four can be served perfectly and the content still will not be cited – the system is built sequentially.
Technical accessibility – The non-negotiable foundation
GPTBot, OAI-SearchBot, PerplexityBot and Google-Extended need access. Activating Cloudflare AI scraper blocking means self-exclusion. Pages with a First Contentful Paint below 0.4 seconds are cited three times more frequently than pages with FCP above 1.13 seconds, according to Shepard (2026). Load time is a selection criterion in the retrieval process.
Worked example: A B2B company with 50 specialist articles accidentally blocks GPTBot in robots.txt. Result: zero AI citations despite Google top-10 rankings for 35 of those articles. Since URL accessibility is the first stage in the sequential evaluation model (score 9.5), the articles simply do not exist for the LLM.
Relevance – Semantic fit as the reason for citation
Page content must align semantically with the question asked. Intent-format match determines selection: "best" queries require comparison lists, "how-to" queries require step-by-step guides. Topic clusters beat individual pages – the more related questions a domain answers, the higher the citation probability across fan-out queries.
Structural extractability – How AI processes content
According to Shepard (2026), 44.2% of all LLM citations come from the first 30% of the text. FAQ sections with a direct answer in the first three sentences are cited three times more frequently. The optimal section length is 120–180 words between headings. Every section must be comprehensible in isolation – the LLM extracts fragments, not entire articles.
| Text position | Share of LLM citations |
|---|---|
| First 30% | 44.2% |
| Middle 30–70% | 31.1% |
| Last 30% | 24.7% |
Source: Shepard (2026), based on Dejan.ai analysis.
Content quality – Specificity beats generality
"Adults need a lot of protein" is not citable. "Experts recommend 0.8 g of protein per kg of body weight" is. The difference is functional: the LLM needs a statement it can embed as evidence in an answer without having to qualify it further. Hedging reduces citation probability. The Princeton GEO study (Aggarwal et al., KDD 2024) showed that adding statistics increases AI visibility by 22% and adding direct quotes by 37% – each measured on the Perplexity engine.
| Content signal | Effect on citation rate | Source |
|---|---|---|
| Statistics with source attribution | +22% | Aggarwal et al. (KDD 2024) |
| Direct expert quotes | +37% | Aggarwal et al. (KDD 2024) |
| FAQ schema with direct answer | +40% | SE Ranking (2025) |
| Visible year in title | +30% | SE Ranking (2025) |
| Structured data markup | +73% | SE Ranking (2025) |
Authority and trust – Building entity trust
According to the AI Platform Citation Source Index by 5WPR (2026), 89% of links cited by AI engines come from earned media. Brands are cited 6.5 times more frequently via third-party sources than via their own domains, according to the same analysis. Brand search volume is one of the strongest predictors of LLM citations – stronger than classic backlinks (Shepard, 2026). The logic behind this: when others write about a brand, the model interprets it as independent confirmation. Anyone publishing only on their own domain provides the LLM with no external trust anchor.
Platform differences: ChatGPT, Perplexity and Google AI Overviews cite differently
Only 11% of domains are cited by both ChatGPT and Perplexity. Only 2% of cited URLs appear across platforms (Yext, 2025). Each platform operates on its own citation logic. Treating all three as one channel means optimising for none.
| Attribute | ChatGPT | Google AI Overviews | Perplexity |
|---|---|---|---|
| Primary source | Wikipedia (7.8%) | Reddit (2.2%) | Reddit (6.6%) |
| Retrieval method | Bing + training data | Google index | Real-time web search |
| Citations per answer | ~7 | 6–14 | ~22 |
| Freshness preference | 29% from 2022 or older | Quarterly updates | 82% from the last 30 days |
| Brand citation rate | 0.59% | Medium | 13.05% |
Sources: Yext (2025), Azamfar/Leapd (2026), SE Ranking (2025).
ChatGPT – Encyclopaedic authority and Bing visibility
Commercial queries trigger a web search in ChatGPT in 53.5% of cases, informational queries only in 18.7% (Azamfar/Leapd, 2026). Pages with FAQ schema and inline citations receive a 40% higher citation weighting according to SE Ranking (2025). Domains with profiles on G2 or Capterra show a three times higher citation probability (Shepard, 2026). ChatGPT favours encyclopaedic sources with broad coverage – unless the query is highly specific.
Perplexity – Real-time retrieval and freshness as a lever
New content can be cited by Perplexity within hours of indexing. Visible year numbers in titles and headings improve the citation rate by 30% (SE Ranking, 2025). Perplexity traffic converts eleven times better than classic organic traffic according to Azamfar/Leapd (2026). The platform rewards recency more strongly than any other – 82% of citations come from the last 30 days. Anyone who wants to be visible here must publish like a newsroom.
Google AI Overviews – Semantic completeness over ranking position
Only 38% of AI Overview citations now come from the organic top 10 (Shepard, 2026). Multimodal content (text + image/video) shows a 156% higher selection rate according to SE Ranking (2025). Structured data markup improves the selection rate by 73% (SE Ranking, 2025). Google AI Overviews draw on sources that answer a question completely – even if they rank at position 15 or 20. Semantic completeness beats ranking position here.
First steps: Building citation optimisation systematically
Building citability follows a clear sequence: secure the technical foundation, then make content structurally extractable, then build authority via third-party sources. Reversing this order means investing in visibility that cannot be technically retrieved. The first two steps require no additional budget – they demand configuration and editorial work on existing content.
Step 1 – Audit technical accessibility (days 1–3)
- Check robots.txt: GPTBot, OAI-SearchBot, PerplexityBot and Google-Extended must have access. A single disallow line eliminates the entire citation opportunity for the affected platform.
- Adjust Cloudflare WAF rules: Deactivate AI scraper blocking or whitelist the relevant bots. The default setting blocks AI crawlers in many configurations.
- Optimise Core Web Vitals: Target FCP below 0.4 seconds. Every second above this threshold measurably reduces citation probability.
Step 2 – Make existing content extractable (weeks 1–2)
- Move the core statement to the top: The central answer of every article belongs in the first 300 words. The LLM reads sequentially and weights the beginning.
- Shorten sections: 120–180 words between headings. Split longer sections, merge shorter ones.
- Ensure isolated comprehensibility: Every key statement must work without surrounding context – the LLM extracts fragments.
- Add FAQ sections: Direct answer in the first sentence, elaboration afterwards.
Step 3 – Build topic clusters (months 1–3)
- Identify fan-out queries: What sub-questions does the AI ask about a main topic? Tools like AlsoAsked or the "People Also Ask" box provide starting points.
- Create standalone sections or articles: For each sub-question, produce a linked piece of content that answers the partial question completely.
- Strengthen internal linking: Cross-link cluster pages so the LLM recognises the domain's topical depth.
Step 4 – Build earned media and entity trust (ongoing)
- Place expert contributions: Industry media, specialist portals and guest articles create the external trust anchors that LLMs interpret as confirmation.
- Leverage review platforms: G2, Capterra and OMR Reviews are demonstrably relevant sources for ChatGPT.
- Create YouTube content: Brand mention in title and transcript. YouTube has a significant citation advantage over other video sources according to Shepard (2026).
Broad distribution increases AI citations by a multiple compared to owned publishing alone, according to 5WPR (2026).
Five common mistakes in AI citation strategy
Most companies do not fail because of missing content but because of structural errors that prevent AI systems from recognising existing content as a source. The following five mistakes are the most common – and the easiest to fix.
- Mistake 1 – Treating all AI platforms as one channel: Only 2% of URLs are cited across platforms (Yext, 2025). Each platform requires its own strategy based on its specific retrieval logic.
- Mistake 2 – Burying key statements in the bottom third of the page: 44.2% of citations come from the first 30% (Shepard, 2026). Place the direct answer in the first 300 words.
- Mistake 3 – Vague phrasing instead of specific facts: "Companies benefit from this" is not citable. Concrete numbers, named sources and clear positions are.
- Mistake 4 – Blocking AI crawlers: Blocking GPTBot or PerplexityBot in robots.txt eliminates every citation opportunity – regardless of content quality.
- Mistake 5 – Publishing only on your own domain: Brands are cited 6.5 times more frequently via third-party sources (5WPR, 2026). Build an earned media strategy with trade media, YouTube and review platforms.
| Attribute | Classic SEO logic | AI citation logic 2026 |
|---|---|---|
| Success criterion | Position 1–3 in SERPs | Citation in generated answer |
| Optimisation unit | Entire URL | Individual paragraph/section |
| Authority signal | Backlinks | Brand search volume + earned media |
| Freshness requirement | Evergreen is sufficient | Quarterly updates required |
| Platform strategy | One (Google) | Three+ (ChatGPT, Perplexity, AI Overviews) |
How AI citation will change by 2027
The citation logic of AI systems is not stable – it is changing
Gerrit Grunert
Gerrit Grunert is the founder and CEO of Crispy Content®. In 2019, he published his book "Methodical Content Marketing" published by Springer Gabler, as well as the series of online courses "Making Content." In his free time, Gerrit is a passionate guitar collector, likes reading books by Stefan Zweig, and listening to music from the day before yesterday.