Kojable research · Finance GEO · 496 grounded AI responses
What 496 Grounded AI Responses Reveal About Finance GEO Visibility
An analysis of 496 target-oriented finance prompt runs across 49 target domains, measuring first-party source presence, external publisher incidence, source identity and response-level co-citation.
Key finding
Target-owned sources appeared in 450 of 496 prompt runs, an observed prompt-weighted rate of 90.7%.
Qualification: These were target-oriented branded and comparison finance prompts. The result does not measure unaided company discovery, general finance-market visibility, share of voice or comparative company performance.
- Finance GEO
- AI citation measurement
- Publisher co-citation
- 18 min read
Study overview
Executive summary
Finance companies increasingly want to know whether generative AI systems can find, use and cite their information. The obvious temptation is to turn every set of AI answers into a ranking: which company is most visible, which publisher has the most authority, and which content strategy wins. The evidence supports a narrower and more useful conclusion.
Across 496 target-oriented prompt runs covering 49 finance domains, target-owned sources appeared in 450 responses, or 90.7%. This is a strong descriptive signal that first-party information was commonly discoverable when questions explicitly concerned a company or compared it with alternatives. The prompt-level 95% Wilson interval was 87.9% to 93.0%, but that interval should not be treated as fully accounting for clustering by target, prompt design or the single model snapshot.
The same responses contained 3,028 external named-source objects associated with 1,197 inferred external publisher identities. AI answers were not grounded in first-party material alone; they drew from a broad and fragmented outside information environment spanning distribution services, video, communities, reviews and specialist finance sources.
The study also exposes a measurement boundary that matters to anyone buying or reporting GEO research. Response-level source presence is measurable here. Page-level performance, exact channel mix, causal authority effects and general market visibility are not. The source records usually identify a publisher by name but almost never resolve the exact page URL.
The result is a detailed picture of how sources assemble around target-oriented finance questions—and a clear standard for what the evidence does and does not justify.
Answer first
Direct answer
When a finance company was already part of a target-oriented question, a source attributable to that company appeared in 450 of 496 prompt runs: 90.7%.
That result is evidence of common first-party source presence under the tested prompts. It is not evidence that the companies collectively held 90.7% of finance AI visibility, would be independently discovered across unbranded demand, or can be ranked reliably against one another.
Study design
Research snapshot
- Input files
50 files; no exact duplicate input files were detected.
- Target domains
49 finance-related target domains.
- Prompt runs
496 requested prompt-response observations.
- Successful responses
495 successful responses and one failed response.
- Recorded model
gemini-2.5-flash-lite.- Language and market
US English.
- Collection window
14–15 January 2026.
- Raw grounding source objects
6,884 returned objects.
- Named source objects
4,660 after artifact removal.
- Retrieval artifacts removed
2,224 objects, or 32.3% of all returned objects.
- Unit of analysis
Prompt response and response-level source presence; a publisher contributes at most once to a response-level incidence count.
- Prompt design
Target-oriented branded, comparison and related finance prompts.
- Inference boundary
Descriptive and exploratory; not a balanced market benchmark, causal design or unbranded share-of-voice study.
Finding 1
The response collection is complete enough for presence analysis
The first analytical question is not which source won. It is whether the responses are sufficiently complete to support any conclusion at all.
Of 496 requested prompt runs, 495 completed successfully, a 99.8% success rate. At least one raw source object appeared in 494 responses, and at least one identifiable named source appeared in 492. Four responses lacked a named source, including the failed response.
High missingness could create a false visibility pattern. If source-bearing answers represented only a small or selectively successful subset, the observed citation rate might say more about data loss than source use. Here, raw-source coverage reached 99.6% and named-source coverage reached 99.2%. The central response-level presence measures therefore rest on nearly the full collection.
The conclusion should remain narrow. The study can describe whether a named publisher or target domain appeared in a response. Completeness alone does not supply canonical URLs, balanced experimental exposure or causal controls.
Finding 2
Target-owned citations were common under the tested prompts
A target-owned source appeared in 450 of the 496 prompt runs. The observed prompt-weighted rate is 90.7%, with a descriptive 95% Wilson interval of 87.9% to 93.0%.
The plain-language interpretation is that when a prompt explicitly focused on a finance company or placed it in a comparison, the AI response usually included material attributable to that company’s own domain. First-party publishing was not merely background material in these target-oriented queries; it frequently became part of the visible evidence set.
The result establishes a baseline capability for the tested question type. If the owned domain is absent even when the company is named, that can be a useful discoverability signal for follow-up. The result does not show whether the company would appear in a generic category question, whether the owned source was cited first, what share of citations it received, or whether the final answer recommended the company.
Finding 3
Target-level variation is descriptive, not a ranking
The median target-level owned-source rate was 90%. Nineteen of the 49 targets appeared through an owned source in every tested response. Eleven targets fell below 90%, and the lowest observed target-level rate was 60%.
The observed spread suggests differences worth investigating. It does not establish a company league table. Each target received only 6 to 20 prompt runs, with most around 10, and the prompt mix was not a balanced randomised exposure design. A result of 100% may simply represent 10 successes from 10 tested prompts. One or two responses can materially change a target percentage when the denominator is small.
The correct portfolio-level claim is that target-owned information was commonly surfaced under branded and comparison-oriented finance questions, with descriptive variation across companies. The data do not establish which company has the strongest overall GEO performance.
Finding 4
External grounding is broad and fragmented
High first-party presence did not remove the use of outside sources. The study identified 3,028 external named-source objects spanning 1,197 inferred external publisher identities. This is a long-tailed information environment rather than one in which a small number of publishers consistently arbitrate every answer.
PR Newswire appeared in 67 prompt runs, YouTube in 62, Reddit in 50, Business Wire in 48 and G2 in 42. Specialist and analyst-style sources also recurred: FF News appeared in 31 responses; FinTech Futures and Gartner in 26 each; NerdWallet in 23; and Global FinTech Series, fintech.global and CanvasBusinessModel.com in 21 each.
The important finding is the coexistence of several directional evidence channels:
- Press and distribution
Company-distributed announcements appeared through services such as PR Newswire and Business Wire.
- Video and community
YouTube and Reddit recurred as distinct forms of public explanation and discussion.
- Reviews and comparison
G2 and related sources can frame product discovery and comparison questions.
- Specialist, analyst and educational sources
Finance and fintech media, analyst-style publishers and reference sources supplied additional context.
This mixture suggests that target-oriented AI answers assemble evidence from sources serving different functions. A press release may supply a launch fact; a review platform may frame comparison; a specialist publication may provide industry context; and community or video sources may add explanation or experience. This interpretation is directional because 83.9% of named source objects lacked a governed source-type label.
The evidence supports a diversified visibility thesis: first-party content is foundational, but the surrounding external information environment still matters. It does not support a precise claim that any channel contributes a specific share of finance AI visibility.
Finding 5
Source identity sets a hard ceiling on page-level conclusions
AI citation analysis can mistakenly treat every returned source object as a verified webpage. This dataset shows why that assumption is risky.
Of 6,884 returned source objects, 2,224—32.3%—were retrieval artifacts rather than usable named sources. Among the remaining 4,660 named sources, only two contained a canonical page identity. The other 4,658 relied on title-based fallback identity. In addition, 3,912 named source objects, or 83.9%, could not be assigned a reliable source type.
This creates a clear hierarchy of claims:
- Supported at the observed grain
Whether a response contained a named source; whether a target-owned source was present; how often an inferred publisher identity recurred; and which inferred publishers co-appeared.
- Weakly supported or unavailable
Which exact article or landing page appeared; whether similarly named labels resolve to the same page; exact source-channel shares; and which page features caused citation.
A publisher name can support an ecosystem analysis, but page optimisation requires page identity. Without a resolved URL, researchers cannot reliably join citation outcomes to page text, publication date, schema, backlinks, author information or technical performance. “The model named this publisher” and “the model cited this exact page” are different findings.
Finding 6
Co-citation reveals recurring information ecosystems
Individual publisher counts show prevalence. Co-citation asks which publishers repeatedly appear in the same response. In this study, co-citation means two inferred publishers appearing within the same AI response.
- Support
The number of prompt runs containing both publishers.
- Jaccard similarity
The share of either publisher’s combined response appearances in which both appeared.
- Lift
The observed joint response rate divided by the rate expected if the publishers appeared independently.
The normalised graph contained 923 publisher pairs with support in at least two prompt runs. Each publisher and pair was counted at most once per response, avoiding repeated source-object inflation.
Business Wire and Unit21 appeared together in 13 prompt runs, with Jaccard similarity of 0.25 and lift of 7.84. Business Wire and FinTech Futures appeared together in 12, with lift of 4.73. FF News and FinTech Futures appeared together in 10, with lift of 6.10. Business Wire and FF News also appeared together in 10, with lift of 3.31.
Reddit and YouTube appeared together in 10 responses, but their lower lift of 1.59 indicates that part of the overlap reflects how common both publishers were individually. G2 and Gartner appeared together in eight responses, with lift of 3.60.
Lower-support pairs produced some of the highest lift values. New Frontier Funding and Pipe co-appeared eight times with lift of 43.73; SAP and SAP Fioneer co-appeared eight times with lift of 39.36; and IBS Intelligence and Plumery co-appeared seven times with lift of 31.31. High lift can arise from rare marginal counts, so these are hypothesis-generating patterns for follow-up rather than proof of authority, influence or a real-world ecosystem relationship.
The defensible conclusion is that responses can draw from recurring source bundles. Co-citation can generate questions about which specialist outlets cluster with distribution services, which comparison sources cluster with analyst sources, and which topic contexts recur. It cannot establish that one publisher links to, endorses, partners with or influences another.
Supported claims
What the evidence supports
- Operationally reliable collection
495 of 496 runs completed successfully, and 492 contained at least one named source.
- Common first-party presence under tested prompts
Target-owned sources appeared in 450 target-oriented prompt runs, an observed prompt-weighted rate of 90.7%.
- Broad and fragmented external grounding
The collection contained 3,028 external named-source objects across 1,197 inferred external publishers.
- Recurring publisher and channel hypotheses
Distribution, video, community, review, specialist and analyst-style publishers recurred, although channel conclusions remain directional.
- Useful response-level co-citation
Support, Jaccard similarity and lift identify source/topic patterns that can be tested in later research.
- A measurement ceiling set by identity quality
Publisher-level presence can be studied, but exact page-level recommendations require resolved page identity.
Unsupported claims
What the evidence does not support
- Unbranded or general market visibility
The prompts were target-oriented, so 90.7% is not finance share of voice or unaided discovery.
- A stable ranking of 49 companies
Unequal samples of 6 to 20 runs, mostly around 10, cannot support a precision league table.
- Causal publisher authority or answer impact
The study observes source presence without a matched panel of eligible uncited pages or a causal intervention.
- Exact page-level optimisation
Only two canonical page identities were available among 4,660 named source objects.
- Precise channel shares or source-quality equivalence
Source type was unknown for 83.9% of named source objects, and incidence does not measure quality or endorsement.
- Publisher partnerships or relationships
Co-citation records joint response appearance, not links, syndication, partnership, endorsement, influence or causation.
- Longitudinal, cross-model or international generality
The collection used one model snapshot, US English and a single 14–15 January 2026 window.
Evidence-led strategy
Implications for finance GEO and AI answer alignment
First-party discoverability is a baseline capability
Owned sources appeared in more than nine out of ten target-oriented responses. Finance companies need clear, accessible first-party information capable of answering branded and comparison questions. If the owned domain is absent even when the company is named, that is a useful discoverability gap to investigate.
AI answers draw from a wider information ecosystem
External sources remained extensive despite high owned presence. The recurring publishers span press distribution, video, communities, reviews, specialist media and analyst-style information. Finance GEO is therefore not reducible to publishing more pages on the corporate domain.
External channels may serve different evidence functions
The mix of external sources and co-appearance patterns suggests that answers combine factual announcement sources, explanatory sources, comparison sources and contextual reporting. This is an inference from observed publisher types and co-citation—not a causal test—and should be used to design follow-up questions rather than prescribe an exact channel mix.
Measurement resolution sets the ceiling of the recommendation
This dataset supports response-level and inferred publisher-level conclusions. It does not support exact page optimisation because canonical page identities are almost entirely absent. Page-level improvement requires page-level source identity.
Scientific boundaries
Methodology and limitations
- Response and source grains
A source object is an individual grounding item returned by the model. A named source is a non-artifact object from which the pipeline inferred a publisher. An owned-citation run is a response containing at least one named source attributed to the target domain. It does not measure citation position, citation share, recommendation sentiment or whether the answer favoured the target.
- Target-oriented prompt design
The prompts already concerned a target company or compared it with alternatives. The design measures conditioned source presence, not which companies an AI system would independently discover or recommend across general finance demand.
- Unequal and small target samples
Targets received between 6 and 20 runs, with most around 10. A 100% observed rate can be 10 successes from 10 prompts. Small denominator changes and unequal exposure prevent a stable target ranking.
- Prompt-weighted headline
The 450/496 headline gives each prompt run equal weight. It is not a perfectly balanced target-level benchmark and should not be converted into a market-share claim.
- Descriptive confidence interval
The 87.9% to 93.0% Wilson interval describes prompt-level binomial uncertainty. Runs are clustered by target, prompt design and a single model snapshot; a target-cluster bootstrap or hierarchical model would be preferable for stronger external inference.
- Source identity
Only 2 of 4,660 named source objects had canonical page identities. The other 4,658 used title fallback identity, commonly because a raw grounding redirect did not resolve to a final landing page.
- Source taxonomy
3,912 of 4,660 named source objects—83.9%—remained unknown source type. Channel interpretations are directional, not precise source-share estimates.
- Co-citation interpretation
Co-citation is response-level co-appearance. High lift with low support can be driven by rare marginals and should generate hypotheses, not claims of authority, partnership, influence or causation.
- Time, model and market scope
The analysis records one model,
gemini-2.5-flash-lite, in US English during 14–15 January 2026. It does not establish longitudinal stability, cross-model generality or international generality.
Research roadmap
Recommended next research phase
- Improve source identity at collection time
Capture final landing URL, redirect chain, canonical URL, HTTP status, retrieval timestamp and resolver/version lineage while the grounding response is fresh.
- Separate branded from unbranded discovery
Create distinct strata for branded questions, comparisons, unbranded category discovery, problem-led questions and high-intent selection questions.
- Balance and repeat the experiment
Run matched prompt structures across targets, models, markets and collection waves with balanced samples and multiple replicates.
- Expand outcomes without collapsing them
Measure owned citation presence, owned citation share, citation position, recommendation direction, answer sentiment, source quality and grounding coverage as separate outcomes rather than one arbitrary score.
- Govern publisher classification
Introduce a versioned source taxonomy with rule provenance, confidence and manual review for high-incidence domains.
- Use cluster-aware inference
For portfolio benchmarks, use a target-cluster bootstrap or hierarchical model rather than treating every prompt run as independent.
- Require repeated graph evidence
Compare support, Jaccard and lift across multiple collection waves before interpreting high-lift co-citation edges.
Method and governance
Research and reproducibility details
The reviewed publication package in this repository contains the primary article, the methodological review, the five published figures and the derived article metrics used as the numerical source of truth for this page. The source folders are retained for provenance and are not published as duplicate resources.
The underlying collection and executable analysis pipeline is maintained separately from the website in the form referenced by the analytical review. This website repository does not contain a verified public package that independently reproduces the study from raw responses. For that reason, this publication does not present repository commands or claim independent reproducibility from the website alone.
The page reports the reviewed derived metrics and methodological boundaries supplied with the publication package. Publication review and sign-off are by Piush Vaish; this is an author review, not independent external peer review or replication.
FAQ
Frequently asked questions
What does the 90.7% figure actually measure?
It is the prompt-weighted share of the 496 target-oriented finance prompt runs that contained at least one named source attributable to the target company: 450 of 496 runs.
Does this mean these companies have 90.7% AI visibility?
No. The prompts already concerned a target company or compared it with alternatives. The result does not measure unaided discovery, unbranded finance demand, market share of voice or comparative company performance.
Can the study identify which exact pages AI systems preferred?
No. Only 2 of 4,660 named source objects had canonical page identities; the other 4,658 relied on title fallback identity. The data support publisher-presence analysis, not exact page-level optimisation.
Which external sources appeared most often?
By response-level incidence, PR Newswire appeared in 67 prompt runs, YouTube in 62, Reddit in 50, Business Wire in 48 and G2 in 42. These counts describe recurrence, not endorsement, authority or causal influence.
What does co-citation mean in this study?
Co-citation means that two inferred publishers appeared within the same AI response. It does not mean one linked to the other, or that a partnership, endorsement, syndication path, influence mechanism or causal relationship exists.
Can the 49 companies be ranked from these results?
No. Targets received between 6 and 20 runs, most had about 10, and the prompt mix was not a balanced randomised exposure design. Target percentages are descriptive and do not support a stable league table.