Kojable research · Cross-Model Citation Study
Who Gets Cited by AI? Publishers, Authors and Evidence Visibility
An AI citation exposes more than a URL. It can make a publisher—and, where identity can be resolved, a named author—visible inside the answer's evidence environment.
Key finding
Named-author share ranged from 50.9% to 77.1%, author-resolution coverage from 92.3% to 98.6%, and publisher-resolution coverage from 90.5% to 100% across the four observed provider stacks.
Qualification: These are descriptive visibility and data-coverage measures from the same fixed study—not author or publisher quality scores, causal citation factors, a separate experiment or an independent replication. Identity review used AI-assisted web review under the recorded provenance method.
- 4 provider stacks
- 9 matched questions
- Publisher and author visibility
- 11 min read
Answer first
Which publishers and authors became visible?
Direct answer
The observed AI citations distributed visibility across identifiable publishers and, where page evidence supported it, named authors. Publisher concentration, publisher recurrence, author recurrence and identity resolution describe different parts of that visibility; none is a quality or causal score.
Claude had the highest observed named-author share at 77.1%. OpenAI was 69.3%, Perplexity 65.5% and Gemini 50.9%. Publisher identity was resolved for 90.5%–100% of cited URLs and author type for 92.3%–98.6%.
| Provider | Named-author share | Author resolution | Publisher resolution | Publisher HHI | Top-publisher share |
|---|---|---|---|---|---|
| Claude | 77.1% | 96.2% | 100.0% | 0.283 | 37.7% |
| Gemini | 50.9% | 92.3% | 90.5% | 0.219 | 30.4% |
| OpenAI | 69.3% | 98.6% | 98.6% | 0.201 | 25.5% |
| Perplexity | 65.5% | 95.2% | 98.4% | 0.270 | 33.0% |
Provider profiles use equal-weight question-macro summaries, so a question with many citations does not automatically dominate a provider's result.
Source identity
A citation exposes several identity layers
A visible citation URL can be normalised to a page and domain, linked to a publisher identity and, where the page supports it, connected to an author identity. Those layers should not be collapsed.
- URL and document
- The exact page selected as visible evidence, after conservative URL normalisation.
- Domain and publisher
- A domain is a network location; a publisher is the editorial or organisational identity behind the page. Multi-tenant platforms make these non-equivalent.
- Author
- A resolved named individual or another valid author type supported by page-level identity evidence.
This differs from the separate source-identity measurement study, which focuses on deciding whether two citation records identify the same exact page or document. Here the question is which publishers and authors become visible after that identity work.
Finding 1
Named authors were visible in roughly half to three-quarters of cited URLs

Named-author share reports the share of cited URLs assigned to a resolved named individual. It does not show that authorship caused selection, that the author was prominent in the generated answer or that the page was higher quality.
- Visibility is not preference
Claude's higher point estimate in this fixed panel does not establish a stable provider preference for named writers.
- Resolution changes the denominator story
High author-resolution coverage makes the observed named-author shares interpretable for this panel, but it remains a measurement result rather than provider performance.
Metric distinction
Named-author share and verified-author share are not nested metrics
Verified-author share is the share of source records with a qualifying upstream verification status under the method. A valid non-person author—such as an organisation or editorial desk—can have qualifying verification. Verified-author share is therefore not a subset of named-author share.
| Provider | Named-author share | Verified-author share | Author-resolution coverage |
|---|---|---|---|
| Claude | 77.1% | 80.8% | 96.2% |
| Gemini | 50.9% | 60.8% | 92.3% |
| OpenAI | 69.3% | 77.1% | 98.6% |
| Perplexity | 65.5% | 72.2% | 95.2% |
Finding 2
Publisher visibility was concentrated, but not equally

OpenAI had the lowest publisher HHI point estimate at 0.201 and Claude the highest at 0.283. Top-publisher shares ranged from 25.5% for OpenAI to 37.7% for Claude. Effective publisher count—the reciprocal of HHI—was approximately 4.15 for Claude, 5.09 for Gemini, 5.30 for OpenAI and 3.81 for Perplexity.
Concentration is not quality. A concentrated set can contain strong evidence; a broad set can contain weak evidence. Publisher and domain are also different units, particularly on multi-tenant platforms.
Broader recurrence view
Some publishers recurred across questions, providers or both
The recurrence table uses the broader ten-question successful-response benchmark. It retains publishers appearing in at least three questions, then sorts by questions, providers and citation events in descending order before showing the first ten.

| Publisher | Questions | Providers | Provider-question cells | URLs | Citation events |
|---|---|---|---|---|---|
| Search Engine Land | 7 | 3 | 11 | 19 | 50 |
| Growth Memo | 6 | 2 | 6 | 6 | 20 |
| Search Engine Journal | 4 | 3 | 6 | 4 | 10 |
| Semrush Blog | 3 | 4 | 6 | 3 | 12 |
| Omniscient Digital | 3 | 3 | 3 | 4 | 21 |
| Clutch | 3 | 2 | 4 | 1 | 30 |
| Authoritytech | 3 | 2 | 3 | 4 | 29 |
| Trakkr | 3 | 2 | 4 | 4 | 19 |
| Arxiv | 3 | 2 | 3 | 3 | 3 |
| G2 | 3 | 1 | 3 | 3 | 5 |
Search Engine Land had the widest question recurrence at seven questions. Semrush Blog appeared across all four provider stacks but only three questions. Those are different kinds of breadth, and neither establishes endorsement.
Resolved people
Named-author recurrence was narrower
The named-person view joins the author-recurrence output to the resolved author table, restricts the result to named_individual, and sorts by questions, providers and citation events. Non-person and malformed identity rows are excluded.
| Named author | Questions | Providers | Provider-question cells | URLs | Citation events |
|---|---|---|---|---|---|
| Kevin Indig | 6 | 2 | 6 | 4 | 15 |
| Cate Dombrowski | 3 | 3 | 3 | 3 | 20 |
| Casey Nifong | 3 | 3 | 4 | 2 | 13 |
| Jeanette Godreau | 3 | 2 | 4 | 1 | 30 |
| Jaxon Parrott | 3 | 2 | 3 | 3 | 19 |
| Toby Brissett | 3 | 2 | 3 | 3 | 3 |
| Greg Jarboe | 3 | 2 | 3 | 3 | 3 |
| Claire Taylor | 3 | 2 | 3 | 3 | 3 |
Kevin Indig appeared across six benchmark questions and two provider stacks. This is recurrence among resolved cited identities—not a ranking of authors, expertise or influence.
Eligibility gate
The author-overlap result did not qualify for a headline claim
Named-person coverage failed the prespecified headline eligibility gate for every provider pair: 0 of 9 matched questions qualified for each pair.
We can describe provider-level author visibility and author recurrence, but we do not have sufficient named-person coverage for a robust headline cross-provider author-overlap claim.
The overlap metrics remain part of the research record, but sparse eligible coverage means they should not be presented as evidence that provider stacks do or do not converge on the same people.
Method
How publisher and author identity were resolved
The pipeline uses conservative canonical URL, domain, publisher and author matching. It does not conduct broad personal crawling. Multi-tenant platforms are handled cautiously because a hosting domain may not identify the publisher of an individual page.
- Publisher resolution
Uses page, domain and available publisher signals to assign an organisational identity without assuming that every domain is one publisher.
- Author resolution
Uses available page-level identity evidence and records named individuals separately from valid non-person author types and unresolved cases.
- Review provenance
Ambiguous identity records were handled under the recorded
ai_assisted_web_reviewprovenance method; the outputs remain coverage-qualified.
- Conservative joins
- Normalised IDs and explicit joins are preferred to speculative name matching.
- Resolution denominator
- The share of cited URLs for which identity type can be resolved sufficiently under the method.
- Unknowns
- Missing or ambiguous identity evidence remains unresolved rather than being converted to a person or publisher.
Company implications
Treat evidence visibility as an ecosystem problem
A company can be described through its own pages, through external publishers and through the identifiable people who create useful evidence. The practical question is not simply “Are we cited?” but “Which evidence identities repeatedly shape how our category is explained?”
- Map publisher exposure
Track which publishers appear across material buyer questions and provider stacks, while separating breadth from event volume.
- Make authorship explicit
Use clear, accurate bylines and author context where a named person is genuinely responsible for a page.
- Strengthen corroboratable evidence
Improve claims, primary evidence and third-party coverage without treating recurrence as endorsement.
These are content-governance and measurement practices, not a claim that adding an author name or targeting a recurring publisher causes AI citation.
Study design
The same fixed Cross-Model Citation Study
This companion is not a separate experiment or independent replication. The designed benchmark contained ten B2B buyer questions across Claude, Gemini, OpenAI and Perplexity: 40 expected cells, 39 successful responses and one missing Gemini response. The primary provider profile uses the nine questions completed by all four provider stacks, with one observed run per provider-question cell.
- Provider profile
- Nine matched questions; equal-weight question-macro summaries.
- Recurrence views
- The broader ten-question successful-response benchmark, with explicit deterministic filters and ordering.
- Observed runs
- One run per provider-question cell; repeated runs are required to assess output stability.
Evidence boundaries
Limitations
- Fixed observational panel
The results describe the recorded provider stacks and questions, not universal or stable provider traits.
- Identity evidence varies
Coverage depends on available page, publisher and byline signals; unresolved cases remain unknown.
- No causal effect
The study does not estimate whether named authorship, publisher recurrence or publisher concentration causes retrieval or citation.
- No quality ranking
Resolution, recurrence, verified status, concentration and citation volume do not establish expertise, accuracy, authority or endorsement.
- Publisher is not domain
Multi-tenant and shared-hosting environments require conservative organisational matching.
- Author-overlap coverage gate
No provider pair had sufficient named-person coverage for a headline overlap claim.
- Different analysis scopes
Provider profiles use nine matched questions; recurrence views use successful responses across the ten-question design.
- Review provenance
Ambiguous identity work used AI-assisted web review and remains subject to the recorded evidence and matching rules.
Research record
Research and reproducibility
The public package contains the study methodology, author and publisher identity outputs, recurrence tables, provider profiles, coverage diagnostics, author-overlap eligibility results and publication manifests.
View the Cross-Model Citation Study publication package on GitHub.
FAQ
Frequently asked questions
Which provider had the highest named-author share?
Claude had the highest observed named-author share in the nine-question matched panel at 77.1%. OpenAI was 69.3%, Perplexity 65.5% and Gemini 50.9%.
Does a named author make a page more likely to be cited?
This study cannot establish that. Named-author share describes cited pages; it does not isolate authorship as a causal citation feature.
What does author resolution mean?
It is the share of cited URLs for which the author type could be resolved sufficiently under the extraction method. It is a data-coverage measure, not a provider quality score.
What does verified-author share mean?
It is the share of source records with qualifying upstream verification status under the method. It is not limited only to named individuals, so it should not be interpreted as a subset of named-author share.
Which publisher recurred most broadly?
Search Engine Land appeared across seven of the ten benchmark questions and three provider stacks. Semrush Blog appeared across all four provider stacks but across three questions. These describe different kinds of recurrence.
Which named author recurred across the most questions in the illustrative recurrence table?
Kevin Indig appeared across six benchmark questions and two provider stacks in the resolved author-recurrence data.
Did the four AI systems cite the same authors?
The study calculated author-overlap metrics, but named-person coverage did not pass the headline eligibility gate for any provider pair. We therefore do not make a headline cross-provider author-overlap claim.
Is publisher concentration good or bad?
Neither inherently. Concentration describes how citation activity is distributed among publishers. A concentrated evidence set can contain strong sources; a broad one can contain weak sources, and vice versa.
From benchmark to company evidence
See what this looks like for your company
Research shows how AI systems behave across a broader sample. Kojable helps you measure how those systems describe, cite and compare your company across the buyer questions that matter.
