Kojable research · AI citation source ecosystems
How Concentrated Are Source Domains in AI Citations?
A cross-platform analysis of domain reach, citation-occurrence concentration and source-set stability across ChatGPT, Gemini and Perplexity.
Key finding
AI citation domains show a steep head and a long tail. A relatively small number of domains account for a large share of captured citation occurrences, while thousands of other domain labels appear less frequently.
Qualification: Concentration differs across platforms but does not measure source authority, preference, quality, influence or market control.
- AI citation concentration
- Source-domain reach
- Cross-platform measurement
- 16 min read
Study overview
Executive summary
AI-generated answers often cite several sources at once. That makes a simple domain leaderboard difficult to interpret: one domain can appear once across many answers, while another can recur several times inside fewer answers. Reach and occurrence volume measure different dimensions of prominence.
This study examines more than 55,000 AI responses across ChatGPT, Gemini and Perplexity. More than 50,000 responses contained at least one captured citation, producing more than 450,000 captured citation occurrences and more than 300,000 distinct response-domain pairs across thousands of observed domain labels.
The distribution has a steep head and a broad long tail. The leading domain accounts for approximately 15% of captured citation occurrences, the top five approximately 34%, the top ten approximately 46% and the top fifty approximately 68%. Roughly a dozen domains reach half of occurrence volume, roughly 100 reach three-quarters and several hundred reach 90%.
These figures describe where citations accumulated in this cohort. They do not identify source quality, authority, trust, preference, influence or AI-search market share.
Direct answer
How concentrated are observed AI citations across source domains?
-
A concentrated head coexists with a very broad tail
A relatively small number of domain ranks account for much of the captured citation volume, but the observed distribution extends across thousands of source labels. ChatGPT has higher top-k concentration than Gemini and Perplexity in this cohort, yet those differences are descriptive rather than normative.
Related measurement
How this differs from source-composition research
Source composition asks what kinds of sources appear. It groups final citations into source families such as owned, comparable, community, official and other third-party evidence.
This article asks how concentrated final citations are across individual domain labels. A source family can be broad while a small number of domains dominate within it; an ecosystem can also contain thousands of domains while retaining a sharply concentrated head. Decision 4 establishes the source-composition result. This standalone publication establishes the concentration result.
Analytical units
What the study measures
Citation-positive responses
More than 50,000 responses contained at least one captured citation. This is the denominator for response-level domain reach.
Distinct response-domain pairs
A domain contributes at most one pair per response, even if it appears repeatedly inside that answer. The cohort contains more than 300,000 such pairs, or roughly 6–7 distinct observed domains per citation-positive response.
Citation occurrences
An occurrence is one captured citation record. A domain can contribute multiple occurrences in one answer. The cohort contains more than 450,000 occurrences across the three included platforms. Occurrences are not the same as unique pages, independent evidence, claims or user clicks.
Two measures of prominence
Response reach and citation-occurrence volume answer different questions
Response reach asks: In how many citation-positive responses did a domain appear at least once? Repeated citations to the same domain inside one answer count once.
Citation-occurrence volume asks: How many captured citation records were assigned to a domain? A domain can contribute several occurrences within one answer.
Occurrence concentration
A small source head accounts for much of the captured volume
| Ranked source set | Approximate occurrence share |
|---|---|
| Leading domain | ~15% |
| Top 5 domains | ~34% |
| Top 10 domains | ~46% |
| Top 50 domains | ~68% |
The leading observed domain reached well over half of citation-positive responses. Its occurrence share was lower because the reach and occurrence denominators differ. The occurrence result means about 15% of captured citation occurrences in this cohort—not 15% of AI search.
Recurrence identifies where observed citations accumulated; it does not explain why they accumulated there. Prompt composition, source availability, platform behaviour, citation rendering and the structure of the information environment can all contribute.
Distribution breadth
The long tail remains very broad
Roughly a dozen domains account for half of observed occurrence volume. Roughly 100 account for three-quarters. Several hundred are needed to reach 90%. The remaining distribution continues across thousands of observed domain labels.
The head is concentrated, but the tail remains very broad. A source strategy focused only on the first few ranks can miss topic-specific evidence environments; treating every low-frequency domain as equally important would ignore the steepness of the head.
Cross-platform comparison
Source concentration differs across AI platforms
| Platform | Top 5 | Top 10 |
|---|---|---|
| ChatGPT | ~40% | ~51% |
| Gemini | ~36% | ~48% |
| Perplexity | ~31% | ~45% |
ChatGPT is more concentrated on these particular top-k measures, Gemini is intermediate and Perplexity is less concentrated. These are descriptive differences in the observed citation distribution. Lower concentration is not inherently better, and the comparison does not rank evidence quality.
Gemini exposed the broadest observed domain inventory in this collection, while ChatGPT and Perplexity also drew from substantial source sets. Exact platform inventories are intentionally not published.
Segmentation
One aggregate leaderboard is insufficient
A domain can rank highly overall because it recurs across all three platforms, appears extremely often on one platform, or occupies a middle position on several. Those patterns imply different source portfolios and different monitoring priorities.
Useful analysis therefore separates platform, buyer-question type, source family and time rather than treating the pooled rank as a universal source score.
Multi-membership
Why response-reach percentages do not sum to 100%
One answer can cite many domains. Domain reach is a multi-membership measure. A single cited response can contribute to the reach numerator of several domains.
Occurrence share is different: each captured citation occurrence is assigned to one domain label, so occurrence shares can be added toward 100%. Calling a reach percentage “share of citations” changes both the denominator and the interpretation.
Longitudinal pattern
The leading position persisted while lower ranks moved
The leading anonymised source position remained stable across the observed periods, while lower-ranked domains moved more substantially. A one-time leaderboard can therefore combine a persistent leader, recurring middle ranks and period-sensitive sources.
The study cannot identify whether that movement reflects prompt composition, platform behavior, search changes, source availability, citation rendering or another mechanism. It does not identify a platform update, changing trust or deliberate preference.
Necessary evidence
The scientific finding does not require a named-domain leaderboard
The concentration result depends on ranks, cumulative shares, platform-level top-k comparisons and broad temporal direction—not public disclosure of the underlying source ecosystem. Rank-based reporting establishes the steep head, long tail and cross-platform differences without turning source names into a reconstructive fingerprint.
What recurrence means
What domain recurrence can—and cannot—tell us
It can describe the observed source environment
Recurrence can show whether citations are concentrated or diffuse, whether a domain appears broadly or repeatedly, how top-k structure differs by platform, how broad the tail is and whether prominent ranks persist across observation periods.
It cannot establish hidden qualities or mechanisms
A high recurrence rate does not establish trust, factual accuracy, authority, citation quality, independence, endorsement, influence, answer correctness, commercial impact or user behaviour. Final citations are survivors; the full candidate set that may have been discovered, ranked, filtered or rejected remains unobserved.
AEO, GEO and answer alignment
Implications for source strategy
-
Define the unit
Report whether a metric counts answers, occurrence rows, URLs, pages or domains. The denominator should travel with the metric.
-
Segment by platform
Pooled source ranks blend distinct ChatGPT, Gemini and Perplexity distributions.
-
Track persistence
Separate durable visibility from period-sensitive movement rather than relying on one snapshot.
-
Investigate before acting
Recurrence can prioritise analysis, but strategic fit still depends on relevance, credibility, accessibility and realistic publishing opportunities.
Operational framework
A better source-monitoring framework
For each platform and question set, preserve citation-positive response count, distinct domain count, response reach, occurrence volume, top-k concentration, long-tail thresholds, source-family classification and time-series persistence.
The useful progression is: how concentrated is the source environment; which source families account for that concentration; and which specific publishing environments matter for the relevant company, topic, platform and buyer question?
Methodology
How the study was measured
The observational collection spans ChatGPT, Gemini and Perplexity. Captured citation records were assigned to normalised domain labels. Subdomains can remain separate where they represent distinct cited domains, so domain aggregation is not equivalent to verified organisation ownership.
Response reach counts distinct citation-positive responses containing a domain at least once. Citation-occurrence share counts every captured occurrence assigned to a domain. Domains are ranked by occurrence volume, and cumulative shares describe the head and tail. Platform distributions are calculated separately.
Longitudinal comparisons report only broad persistence and movement. Exact dates, rank trajectories and period-level counts are not part of the public result.
Reporting design
Public-safe research details
The underlying observational dataset contains a distinctive source ecosystem, so this public version uses aggregate and rank-based reporting rather than a named domain leaderboard. It omits exact domain-level counts, exact dates, period denominators, raw URLs and precise rank trajectories.
The three publication figures were generated from conceptual marks or explicitly approved rounded values. They do not read private source tables or analytical chart geometry. This preserves the research conclusion while reducing contextual re-identification risk.
Study boundaries
Limitations
-
Final-output evidence
Captured citations do not reveal the full hidden retrieval and candidate-selection process.
-
Domain aggregation
Domain labels hide page-level differences; one organisation can control several domains and one domain can host many authors.
-
Repeated observations
Prompt variants and overlapping platform, topic and time distributions reduce statistical independence.
-
Cohort dependence
The collection is not a random sample of every possible AI query, and platform behaviour can change after observation.
-
No quality inference
Concentration does not measure factual support, source quality, trust, traffic, clicks, conversion or causal influence.
FAQ
Frequently asked questions
What is AI citation source concentration?
AI citation source concentration describes how much of the captured citation-occurrence volume accumulates among the leading source-domain ranks. It is a property of the observed citation distribution, not a measure of authority, quality or market share.
What is the difference between domain reach and citation-occurrence volume?
Domain reach counts the citation-positive responses in which a domain appeared at least once. Citation-occurrence volume counts every captured citation record assigned to that domain, including repeated occurrences within one answer.
How concentrated were the observed AI citation domains?
The leading domain accounted for about 15% of captured citation occurrences, the top five about 34%, the top ten about 46% and the top fifty about 68% in this observational cohort.
What does the long tail of AI citation sources mean?
A relatively small source head accounts for much of the captured volume, but the rest of the distribution extends across thousands of observed domain labels. Roughly a dozen domains reached half of occurrence volume, roughly 100 reached three-quarters and several hundred reached 90%.
Does citation concentration measure AI-search market share?
No. The denominator is captured citation occurrences in this observational cohort, not AI-search traffic, usage, revenue or market activity. A 15% occurrence share does not mean 15% of AI search.
Which platform had the highest top-k source concentration?
ChatGPT had the highest observed top-five and top-ten concentration in this cohort, followed by Gemini and then Perplexity. These are descriptive differences in the observed citation distribution.
Does lower source concentration mean better evidence quality?
No. Lower concentration can reflect a broader or more even source distribution, but it does not establish better evidence, greater independence, higher accuracy or stronger source quality.
Why are named source domains not published in this study?
The underlying observational dataset contains a distinctive source ecosystem, so this public version uses aggregate and rank-based reporting to reduce contextual re-identification risk. The scientific findings do not require a named domain leaderboard.
Did the leading source remain stable over time?
The leading anonymised source position remained stable across the observed periods, while lower-ranked positions moved more substantially. The study cannot identify which hidden mechanism caused that movement.
How should companies use source-concentration analysis?
Use concentration to prioritise deeper source analysis, then segment by platform, buyer question, source family and time. Investigate relevance, accessibility and publishing fit before treating a recurring domain as a strategic target.