Kojable research · Cross-Model Citation Study
What Do AI-Cited Pages Look Like? Freshness, Length, Structure and Schema
The pages cited in this fixed benchmark were often substantial, structured pieces of web content—but the study does not show that those characteristics caused the citations.
Key finding
Question-level median cited-page word counts averaged roughly 2.1k–2.6k words, list presence averaged 74%–92%, and table presence averaged 33%–47%. Dated-source medians ranged from approximately 76–120 days.
Non-causal qualification: This companion profiles pages that were cited in the same fixed Cross-Model Citation Study. It does not estimate whether word count, headings, images, lists, tables, schema or freshness caused retrieval or citation.
- 4 provider stacks
- 9 matched questions
- Page features and source age
- 11 min read
Answer first
What did the cited pages look like?
Direct answer
The cited pages in this fixed benchmark were generally substantial and structurally rich, but those characteristics do not establish why the pages were cited.
Typical question-level median page length was roughly 2.1k–2.6k words. Lists appeared on roughly 74%–92% of feature-evaluable cited pages, tables on roughly 33%–47%, and dated-source medians ranged from roughly 76–120 days.
| Provider | Median source age | Date coverage | Median words | Median headings | Median images |
|---|---|---|---|---|---|
| Claude | 76 days | 88.8% | 2,228 | 21.6 | 15.6 |
| Gemini | 81 days | 64.8% | 2,147 | 20.8 | 4.9 |
| OpenAI | 88 days | 70.3% | 2,216 | 19.3 | 6.6 |
| Perplexity | 120 days | 97.2% | 2,605 | 26.8 | 8.3 |
These are equal-weight question-macro summaries. Continuous page characteristics are calculated from question-level cited-page medians before provider-level averaging, so high-citation questions do not automatically dominate.
Finding 1
How fresh were the cited sources?
Source age is measured relative to the timestamp of the AI response that cited the page—not relative to this article's publication date.

| Provider | Median source age | Benchmark-panel interval | Publication-date coverage |
|---|---|---|---|
| Claude | 76 days | 44–112 days | 88.8% |
| Gemini | 81 days | 48–119 days | 64.8% |
| OpenAI | 88 days | 77–98 days | 70.3% |
| Perplexity | 120 days | 86–157 days | 97.2% |
Claude had the lowest source-age point estimate, but one observed run per provider-question cell and unequal date denominators do not establish a stable provider preference for fresher sources.
Coverage boundary
Freshness must be read with date coverage
Missing, ambiguous or unparseable publication dates remain unknown. They are not assigned zero age, treated as old or assigned the response date.
| Provider | Published within 90 days | Published within 30 days |
|---|---|---|
| Claude | 54.7% | 23.2% |
| Gemini | 53.8% | 34.0% |
| OpenAI | 54.8% | 15.0% |
| Perplexity | 45.7% | 10.8% |
Finding 2
The cited pages were usually substantial

| Provider | Median word count | Median heading count | Median image count |
|---|---|---|---|
| Claude | 2,228 | 21.6 | 15.6 |
| Gemini | 2,147 | 20.8 | 4.9 |
| OpenAI | 2,216 | 19.3 | 6.6 |
| Perplexity | 2,605 | 26.8 | 8.3 |
Longer pages may reflect research-heavy B2B questions, source format, publisher templates or retrieval and selection processes. The study does not identify an ideal length or show that adding headings or images increases citation probability.
Finding 3
Lists, tables and schema appeared on cited pages

| Provider | List present | Table present | FAQ schema | HowTo schema | Extraction coverage |
|---|---|---|---|---|---|
| Claude | 73.9% | 32.9% | 16.4% | 16.4% | 100.0% |
| Gemini | 74.4% | 38.5% | 32.3% | 29.4% | 89.6% |
| OpenAI | 88.1% | 40.3% | 26.2% | 26.2% | 98.1% |
| Perplexity | 92.2% | 47.0% | 34.7% | 32.9% | 93.1% |
Lists and tables can improve human readability, while schema can make page semantics explicit. Their prevalence here is descriptive; it does not establish that any of these features increased retrieval or citation probability.
Metric governance
Why generic structured-data share is not headlined
The locked table contains a generic structured_data_share field, but question-level evaluability is too limited and uneven for a credible four-provider comparison.
- Claude
- 0 evaluable questions
- Gemini
- 3 evaluable questions
- OpenAI
- 1 evaluable question
- Perplexity
- 2 evaluable questions
A point estimate of 1.0 in a sparse evaluable subset does not mean that 100% of a provider's cited pages use structured data. The field remains technical context rather than a headline result and is excluded from Figure 3.
Method
What the page-feature extraction measures
The pipeline creates one cached feature record per canonical URL and extracts observable HTML/page properties. Full page bodies are not stored.
- Accepted targets
- Appropriate public HTTP(S) HTML/XHTML targets with successful responses.
- Unit of reuse
- One feature record per canonical URL; the same page's characteristics may be reused when its canonical URL recurs.
- Failure handling
- Fetch failure means unknown, not feature absence.
- Scale signals
Word count, heading count and image count.
- Structure signals
List presence, table presence and FAQ/HowTo schema signals.
- Identity and provenance signals
Visible author/date signals, structured-data signals and canonical tags.
Interpretation boundary
Why this is not an SEO recipe
2,500 words + 25 headings + lists + tables + FAQ schema = AI citation.
The data does not support that formula. It describes pages that were already cited and does not isolate the effect of any individual feature.
- Topic
Research-heavy B2B questions may naturally lead to longer, more structured pages.
- Source format
Editorial-style pages were common in this benchmark and can partly explain scale and structure.
- Publisher templates
Site-wide design systems can add headings, images, lists or schema independently of content quality.
- Retrieval and selection
The observed final page profile can be shaped by unobserved candidate pools, retrieval systems and source selection.
A causal recommendation requires an intervention or otherwise comparable candidate-level design.
Company implications
Useful diagnostics without a causal checklist
The goal is not to imitate an average cited page. It is to make the information buyers and AI systems encounter accurate, current, explicit and easy to verify.
- Make important evidence easy to locate
Do not bury material claims in vague narrative.
- Make comparisons and criteria explicit
Lists and tables can improve human clarity without implying a citation effect.
- Keep important information current
Outdated pricing, capabilities or positioning create a representation risk distinct from simple publication age.
- Structure content around real buyer questions
Clear sections improve usefulness even before considering AI retrieval.
- Make important claims corroboratable
A well-structured page can still fail to provide convincing evidence.
The diagnostic chain is representation → source → claim → evidence → page clarity → verification.
Study design
The same fixed Cross-Model Citation Study
This companion is not a separate experiment or independent replication.
- Designed benchmark
- 10 B2B buyer questions across 4 provider stacks and 40 expected cells.
- Observed responses
- 39 successful responses, with one missing Gemini response.
- Primary comparison
- 9 matched questions and one observed run per provider-question cell.
- Provider summaries
- Question-macro averaging; continuous features begin with question-level medians.
Benchmark-panel intervals show sensitivity to the composition of this fixed question panel. They do not measure provider run-to-run stability.
Evidence boundaries
Limitations
- Source-age coverage
Publication dates are unavailable for some cited pages, so freshness remains coverage-qualified.
- Extraction dependence
Failed or unsupported page fetches remain unknown; coverage differs across provider sets.
- Recurring canonical URLs
Page features may be reused for the same canonical page, while source age remains response-specific because it uses response-generation time.
- No causal feature effects
The study does not estimate the effect of words, headings, images, lists, tables, FAQ schema, HowTo schema or freshness on retrieval or citation probability.
- Unequal candidate observability
Claude and OpenAI candidate pools are treated as complete under the study contract; Gemini and Perplexity expose citation-biased pools. This article therefore centres comparable final cited-page profiles.
Research record
Research and reproducibility
The public package contains temporal methodology, page-feature methodology, source-date provenance, extraction coverage, provider summaries, benchmark-panel uncertainty outputs, reconciliation checks and publication manifests.
View the Cross-Model Citation Study publication package on GitHub.
FAQ
Frequently asked questions
How old were the pages cited by the AI systems?
Question-macro median cited-source age ranged from approximately 76 days for Claude to 120 days for Perplexity. Publication-date coverage differed by provider and must be considered alongside those estimates.
Which provider cited the freshest sources?
Claude had the lowest median source-age point estimate at about 76 days, followed by Gemini at 81 days, OpenAI at 88 days and Perplexity at 120 days. The study does not establish a stable provider preference for fresher sources.
How long were cited pages?
Question-macro median cited-page word-count summaries ranged from roughly 2,147 words for Gemini to 2,605 for Perplexity.
Did most cited pages contain lists?
Yes in this benchmark. List presence ranged from approximately 73.9% for Claude to 92.2% for Perplexity among feature-evaluable cited pages.
How common were tables?
Table presence ranged from approximately 32.9% for Claude to 47.0% for Perplexity.
Was FAQ schema common?
Observed FAQ-schema share ranged from about 16.4% for Claude to 34.7% for Perplexity. This is descriptive, not evidence that FAQ schema causes citation.
Does adding schema make a page more likely to be cited?
This study cannot answer that. A controlled or candidate-level intervention analysis would be needed to estimate a causal effect.
Does publishing newer content make AI systems cite it more?
Not established here. The study describes the ages of pages that were cited; it does not isolate publication recency as a causal factor.
Should every page be 2,000+ words?
No. The observed word counts describe cited pages in this particular research benchmark. Appropriate length depends on the question, evidence and usefulness of the page.
From benchmark to company evidence
See what this looks like for your company
Research shows how AI systems behave across a broader sample. Kojable helps you measure how those systems describe, cite and compare your company across the buyer questions that matter.
