Kojable research · AI citation measurement · Source identity

A Named Source Is Not a Verified Citation: Why Source Identity Matters in AI Citation Measurement

Published By Piush Vaish

A companion analysis of 496 finance AI responses showing why nearly complete source collection can coexist with almost no exact page identity—and why page-level claims require deeper source resolution than a recognizable label.

Key finding

492 of 496 prompt runs contained at least one identifiable named source, yet only 2 of 4,660 named-source objects had canonical page identities in the collected records. Response coverage was almost complete; exact document identity was not.

Qualification: A declared canonical URL is one strong identity signal, but the broader requirement is stable, auditable page or document identity. Final resolved URLs, redirects, normalized URLs, content fingerprints, duplication and syndication relationships may also contribute.

  • Source identity
  • Citation resolution
  • Measurement design
  • 17 min read
Measurement-readiness chart showing 6,884 raw source objects, 2,224 artifacts, 4,660 named-source objects, only two canonical page identities in the collected records and 3,912 objects with unknown source type.
496prompt runs requested
492responses with identifiable named sources
4,660named-source objects after artifact removal
2 / 4,660objects with canonical page identities in the collected recordsPage-identity limitThis is an identity-coverage observation, not a generic citation-accuracy score.

Study overview

Executive summary

An AI response names a source. Has the cited webpage actually been identified? Not necessarily.

In 496 target-oriented finance prompt runs, the system returned 6,884 raw source objects. After 2,224 retrieval artifacts were removed, 4,660 identifiable named-source objects remained. Yet only two had canonical page identities in the collected records. The other 4,658—99.96% of the named-source set—depended on title-based fallback identity.

A recognizable label can support response-, target-domain- or publisher-level analysis without supporting exact-page attribution. It may show that a response had a named source, a target-owned domain appeared or an inferred publisher recurred. It may still be unable to establish the exact article, page persistence, document duplication, page features or the claim a source supported.

The governing principle is: the resolution of the source identity sets the maximum resolution of the claim. You cannot make page-level claims from publisher-level evidence.

Answer first

Direct answer

A named AI source is not automatically a verified page-level citation. A source may be identifiable enough for response-, target-domain- or publisher-level analysis while remaining unresolved at the exact-page or document level.

Every published result should state both its denominator and its deepest reliable identity level. Publisher-level evidence should not produce page-level conclusions or prescriptions.

Companion analysis

Research lineage

This is the fourth companion publication using the same underlying 496-response Finance GEO dataset. It is not an independent experiment.

Research 01

What 496 Grounded AI Responses Reveal About Finance GEO Visibility reports the full empirical source landscape.

Research 02

Branded AI Visibility Is Not Market Visibility separates branded discoverability from general market visibility.

Research 03

The External Publisher Ecosystem Behind Finance AI Answers examines external-source fragmentation and distinguishes incidence from authority.

Research 04

This paper asks when a source record is resolved well enough to support a page-level citation claim.

Measurement boundary

The measurement paradox: complete responses, incomplete identities

Collection quality was strong at response level: 495 of 496 runs succeeded, 494 contained at least one raw source object and 492 contained an identifiable named source. Document identity was almost entirely unresolved: only 2 of 4,660 named-source objects had canonical page identities in the collected records.

The 4,658 title-fallback identities represent 99.96% of the named-source set. Separately, 3,912 named-source objects—83.9%—could not be assigned a reliable non-unknown source type. Collection completeness, identity coverage and source-type completeness are different dimensions and should not be collapsed into one quality score.

Research snapshot

What each observed count can support

Observed values and the analytical level each can support.
MeasurementObserved valueWhat it supports
Requested prompt runs496Study scope
Successful responses495Response-level completeness
Responses with raw source objects494Raw source coverage
Responses with named sources492Named-source coverage
Raw source objects6,884Returned object volume
Retrieval artifacts removed2,224Source-layer cleaning
Named-source objects4,660Named-source analysis
Canonical page identities in collected records2Very limited page-level attribution
Title-fallback identities4,658Publisher/title-level fallback analysis
Unknown source types3,912Weak full-channel classification
Non-unknown source types748Partial source-type analysis

Native methodological framework

The source-identity resolution chain

AI citation measurement is an identity-resolution process, not a single citation/no-citation field. Each stage supports a different analytical output.

  1. Raw source object

    Whatever source-related record the system returned. Strongest output: returned source-object volume.

  2. Artifact filtered

    Interface elements and non-source retrieval artifacts are removed before analytical counting.

  3. Usable named source

    A recognizable name or label supports named-source coverage, but not necessarily page attribution.

  4. Publisher / target-domain identity

    Association with an organization, platform or company domain supports publisher response incidence and target-owned presence.

  5. Resolved page identity

    Sufficient URL and retrieval data identify a specific page location, enabling page incidence and persistence analysis.

  6. Canonical / deduplicated document identity

    Redirects, URL variants, duplicates and syndication are reconciled for unique-document analysis and page-feature joins.

  7. Claim-level evidence mapping

    A verified document is mapped to the answer statement it supports, contradicts or qualifies.

Figure 1

Measurement readiness of the source records

Measurement-readiness chart showing 2,224 artifacts among 6,884 raw objects, only two canonical page identities among 4,660 named-source objects, and 3,912 objects with unknown source type, or 83.9 percent.
Figure 1. The dataset is highly observable at the named-source level but almost entirely unresolved at the canonical-page level. Of 4,660 named-source objects, only two had canonical page identities in the collected records, while 3,912 remained unknown by source type.
Open full-resolution figure

The funnel contains three different losses of resolution: 2,224 of 6,884 raw objects—32.3%—were artifacts; document-level coverage nearly disappears between named-source and canonical-page identity; and 83.9% of named-source objects remain unknown by source type. These gaps affect object-volume, exact-document and channel-composition claims respectively.

Identity levels

A complete response is not a complete citation record

A complete record answers separate questions: did the system return a source-related object; was it analytically usable; did it have a recognizable label; could it be associated with a publisher or target domain; could the exact page be resolved; could variants and copies be deduplicated; and could the document be mapped to a specific answer claim?

The finance study is strong at the earlier response, named-source and domain levels and very limited at the exact-page level. That boundary should also be the boundary of its published claims.

Fallback limitation

Why title-based identity is unstable

  • One document, several labels

    Truncation, punctuation, subtitles, capitalization and display-label changes can make one page appear to be several sources.

  • Several documents, one label

    Generic names such as Pricing, About Us, Products, Press Release or Help Center can collapse distinct pages.

  • Syndication creates false multiplicity

    Originals, distribution copies, partner sites and mirrors may represent one underlying document.

  • Platforms hide the document

    YouTube or Reddit may identify a platform without resolving the channel, creator, video, subreddit, thread or post.

  • Titles change over time

    Headline, slug and metadata changes can create apparent churn while the underlying content remains materially the same.

These are general risks, not diagnosed failure modes for every unresolved record. Fallback identity should remain visibly labeled as fallback identity.

Technical position

Canonical URL is valuable, but verified page identity is the requirement

A declared canonical URL is an important signal, not the entire identity system. A robust pipeline may combine the original returned URL, fetched URL, final URL after redirects, redirect chain, HTTP status, declared canonical URL, normalized URL, publisher domain, normalized title, publication and modification dates, content fingerprint, duplicate relationship and syndication relationship.

The goal is a stable, auditable page or document identity. A page without a declared canonical can sometimes still be resolved reliably; a declared canonical alone does not solve every duplication, syndication or provenance problem.

Failure modes

Five analytical errors caused by weak source identity

  1. Inflated or compressed counts

    Title variants can count one page several times, while generic labels can collapse several pages into one.

  2. False page-performance rankings

    “Most cited article” requires verified page identity; a title cannot reliably separate an article, homepage, product page, help document or syndicated copy.

  3. Invalid page-feature joins

    Backlinks, schema, authorship, word count and page type cannot be attached reliably without page identity. Even then, causal or predictive work also needs eligible-but-uncited comparison pages.

  4. Misclassified source channels

    One domain can contain editorial, sponsored, press-release, product, forum and research content. Publisher identity does not define page type.

  5. Unreliable longitudinal tracking

    Title changes, URL changes, redirects and syndication can create artificial gains, losses or churn without stable document identity.

Reusable framework

Match the claim to the identity level

The minimum identity level required for common AI citation questions and support in this dataset.
Research questionMinimum reliable identity levelSupport
Successful response?Prompt runStrong
Any source-related object?Raw source object / responseStrong
Usable named source?Named source / responseStrong
Target-owned source present?Target-domain identityStrong within target-oriented prompts
Which external publishers recur?Publisher identity / responseDescriptively useful
Which publishers co-appear?Publisher pair / responseHypothesis-generating
Which exact page appears most often?Resolved / deduplicated pageUnsupported
Which page persists?Stable document identityUnsupported
Which page features predict citation?Verified cited + eligible uncited pagesUnsupported
Exact source-type mix?Reliable page/type classificationWeak
Which claim does each source support?Page identity + claim mappingUnsupported

Target-owned sources appearing in 450 of 496 target-oriented prompt runs is a defensible target-domain statement; it does not identify the company page. PR Newswire appearing in 67 responses is publisher incidence, not 67 verified unique releases.

Denominator discipline

“Citation count” is not one denominator

Raw source-object count

6,884 returned records, including artifacts and objects that do not map one-to-one to documents.

Named-source-object count

4,660 usable objects with recognizable identity, not necessarily verified pages.

Publisher response incidence

Responses containing an inferred publisher at least once; this is not a unique-article count.

Target-domain response incidence

Responses containing the target company domain; this does not resolve the exact owned page.

Resolved page incidence

Responses containing one verified page; broadly unsupported by the current records.

Unique-document incidence

Frequency after redirect, duplicate, URL-variant and syndication resolution.

Claim-level evidence incidence

Frequency with which a verified document supports a specific answer claim.

Every metric should state what is counted, its denominator, the identity level and the deduplication rule.

Decision boundary

Source identity determines the level of action

Publisher-level evidence supports publisher-level investigation: recurrence, topic context and representation accuracy. Target-domain evidence supports first-party visibility diagnosis. Stable page identity enables page persistence, factual consistency and update analysis. Claim mapping enables support, contradiction, qualification and staleness diagnosis.

Buyer and research checklist

How to audit an AI citation measurement product

A polished dashboard can label rows “top sources” or “top pages” while relying on very different identity systems. Ask:

  • What does each row represent?

    Raw object, title label, publisher, normalized URL, resolved page or deduplicated document?

  • Is the raw evidence preserved?

    Can researchers inspect the original source object, label and URL?

  • How are URLs resolved?

    Are redirects followed, final URLs stored and declared canonicals captured where available?

  • How is normalization disclosed?

    Are URL variants, duplicates and syndicated documents identified transparently?

  • Can the page be audited?

    Can the underlying page be opened, and does its identity remain stable across dates?

  • Are failure states visible?

    Unresolved sources should remain visibly unresolved rather than receiving synthesized certainty.

  • Are identity levels separated?

    Does the product distinguish publisher, page and deduplicated document identity?

  • Can sources be mapped to claims?

    Can the product show the answer passage and source relationship?

  • Is classification auditable?

    Do source-type labels include confidence, method and provenance?

Proposed standard

A minimum source-identity standard

  1. Response provenance

    Model, version, interface, retrieval mode, prompt, locale, timestamp, response ID and source order.

  2. Raw source representation

    Complete returned object, original label, title, URL, object type and stable object hash.

  3. Resolved page identity

    Fetched and final URL, redirects, status, canonical signal, domain, title and dates where available.

  4. Document normalization and deduplication

    Normalized URL, fingerprint, duplicate and syndication relationships, document family and resolution confidence.

  5. Analytical classification

    Owned/external status, publisher, page, source type, method, confidence, provenance and review status.

  6. Claim-level mapping

    Answer passage, claim identifier and support, contradiction or qualification relationship where observable.

  7. Explicit failure states

    Unresolved URLs, inaccessible pages, ambiguous publishers, suspected artifacts, duplicates, missing canonical signals, unknown types and uncertain mappings.

Unknown values should remain unknown. Filling missing identity with synthetic certainty makes tables look complete while weakening the result.

Future capabilities

What stronger identity would unlock

  • Page persistence

    Separate publisher persistence, page persistence and underlying document persistence.

  • Syndication analysis

    Identify apparent sources that are copies of one release, article, announcement or report.

  • Cited versus eligible-but-uncited analysis

    Build the negative panel required for stronger page-selection claims.

  • Content-update and page-feature analysis

    Attach changes, authorship, type, structure, relevance and link features to the correct page.

  • Claim-support analysis

    Test whether a source supports, contradicts, qualifies or only tangentially relates to the answer claim.

  • Reproducibility

    Allow another researcher to open and inspect the recorded evidence.

These are capabilities enabled by stronger future identity resolution, not analyses completed in the current study.

Supported conclusions

What this evidence supports

  • Highly complete response and named-source collection

    495 of 496 runs succeeded and 492 contained an identifiable named source.

  • Measurable target-domain presence

    Owned-source presence can be measured at response and domain level within the target-oriented design.

  • Descriptively useful publisher incidence

    Stable inferred publisher identities can be compared at response level; co-appearance can generate hypotheses.

  • A material source-identity limit

    The near-absence of canonical page identities prevents reliable exact-page analysis.

  • A material source-type limit

    With 83.9% unknown by source type, precise full-dataset channel composition is unsupported.

Claim boundary

What this evidence does not support

  • Exact-page and unique-document rankings

    Title fallbacks are not verified unique documents and cannot support a reliable page leaderboard.

  • Page-feature or authority effects

    The data do not support page-feature joins, causal authority effects or exact page-level optimization prescriptions.

  • Precise channel composition

    Unknown source types prevent a precise full-dataset source-channel mix.

  • Persistence and syndication

    Stable document identity is required to track pages or reconcile copies over time.

  • Claim attribution or quality

    Source presence alone does not show which claim was supported, evidence quality or publisher authority.

  • Causal attribution

    Cited-source observations alone cannot establish why a source appeared or what caused an answer outcome.

Next study design

Research agenda

A stronger pipeline should preserve raw objects; resolve URLs and redirects; store final URLs and canonical signals; normalize variants transparently; fingerprint documents; detect duplicates and syndication; classify with provenance; map sources to passages; build eligible-but-uncited comparison sets; repeat across models, markets and time; and report every outcome at its correct identity level.

The objective is not to force every source into a clean identity. It is to distinguish resolved evidence from unresolved evidence honestly enough that each conclusion matches what the data know.

Conclusion

Identity resolution is part of the measurement

The finance dataset is highly complete at response level and almost entirely unresolved at exact-document level. Both are true: 495 of 496 prompt runs succeeded and 492 contained a named source, while only 2 of 4,660 named-source objects had canonical page identities and 83.9% remained unknown by source type.

The dataset can credibly measure response completeness, named-source coverage, target-owned presence and inferred publisher recurrence. It cannot credibly rank exact pages, measure page persistence or feature effects, estimate precise channel composition or attribute claims to documents.

The wider lesson is not merely “collect canonical URLs.” It is to build and report an explicit source-identity chain. Raw objects, named sources, publisher identities, resolved pages, deduplicated documents and claim-level evidence are distinct analytical levels.

The resolution of the source identity sets the maximum resolution of the claim. Until page identity is resolved, page-level precision is an appearance—not a result.

Frequently asked questions

Frequently asked questions

What is the source-identity problem in AI citation measurement?

It is the gap between receiving a recognizable source label and resolving the exact page or document that label represents. A source can be identifiable at publisher level while remaining unresolved at page level.

Does a named AI source count as a verified citation?

It can count as a named-source observation, but not automatically as a verified page-level citation. Page-level claims require a reliably resolved page or document identity.

Why are canonical URLs important?

Canonical URLs can provide a strong page-identity signal and help normalize duplicates or URL variants. But verified document identity can also use final resolved URLs, redirect history, normalized URLs, content fingerprints and other provenance.

How complete was source collection in the finance study?

Of 496 requested prompt runs, 495 completed successfully, 494 contained raw source objects and 492 contained at least one identifiable named source.

How many named sources had canonical page identities?

Only 2 of 4,660 named-source objects had canonical page identities in the collected records. The other 4,658 relied on title-based fallback identities.

Can the dataset rank the most-cited individual pages?

Not reliably. Exact page ranking requires stable page identity across the source records.

Can publisher incidence still be measured?

Yes, descriptively, when publisher identity can be inferred consistently. The dataset can count responses containing an inferred publisher while remaining unable to identify the exact article behind every occurrence.

Why does source identity matter for longitudinal tracking?

Without stable document identity, title changes, URL changes, redirects and syndicated copies can look like gains, losses or churn even when the underlying evidence has not materially changed.

What should an AI citation dashboard expose?

It should distinguish raw objects, publisher identities and verified page identities; preserve original metadata; expose resolved URLs and failure states; and keep unresolved records visibly unresolved.

What is the main methodological rule from this study?

The resolution of the source identity sets the maximum resolution of the claim. Publisher-level evidence should not be used to make page-level conclusions.

Piush Vaish, Founder of Kojable

Author

About the author

Piush VaishFounder and CEO of Kojable

Piush Vaish is the founder and CEO of Kojable, a repeat founder and data scientist with more than 10 years of experience building and productising AI, machine-learning and data products. His experience spans high-growth technology companies and large enterprise environments. He combines technical depth with customer discovery, creative problem-solving and a strong bias towards shipping useful products. He writes about AI search, AEO, GEO, agentic discovery and AI product strategy.

Read more about Piush Vaish