Kojable research · GEO experimental design · Finance AI

Why GEO Rankings, Authority Claims and Causal Conclusions Require Stronger Experiments

Published By Piush Vaish

A capstone analysis of 496 finance AI responses showing why descriptive GEO scores cannot automatically become market rankings, authority effects or causal conclusions—and what stronger experiments would require.

Key finding

A descriptive GEO result can be mathematically correct while the stronger interpretation is unsupported. Rankings require comparable opportunity and uncertainty; authority associations require cited and eligible-but-uncited candidate pages; causal claims require a design capable of isolating cause.

Qualification: The current 496-response dataset is strong for bounded descriptive analysis. It was not designed as a balanced market-ranking experiment or a causal page-authority study.

  • Experimental design
  • GEO rankings
  • Authority and causality
  • 17 min read
Distribution of observed target-owned source-presence rates from 60 to 100 percent across 49 finance targets. Targets received unequal exposure of 6 to 20 prompt runs, so the chart is descriptive rather than a precision market ranking.
496target-oriented finance prompt runs
49target domainsDesign boundaryTargets did not receive equal prompt exposure.
6–20prompt runs per targetUnequal exposureRaw percentages therefore have different precision.
2 / 4,660named-source objects with canonical page identitiesPage-identity limitThis is not a citation-accuracy score; it limits page-level authority analysis.

Capstone synthesis

Executive summary

GEO measurements are easy to sort. A company appears in 100% of tested answers, another in 80%, and a cited domain has a high SEO authority metric. A descending table can look like market leadership or a causal visibility factor. But sorting observations does not create an experiment, and describing cited winners does not establish why they were selected.

The six-paper Finance GEO series supports bounded descriptive claims: target-owned information was common in target-oriented prompts, external sources were fragmented, some identities recurred together and exact page identity was severely limited. It does not support a stable company market ranking or a causal authority effect.

Evidence requirements rise with the claim. Descriptive statements require correct observations, denominators, identity and scope. Comparative rankings require comparable opportunity, eligibility, repetition and uncertainty. Market claims require a population and sampling frame. Authority associations require cited and eligible-but-uncited pages. Causal claims require a design capable of separating cause from correlated alternatives.

The experiment must match the sentence in the headline.

Answer first

Direct answer

The existing study supports descriptive observations within its tested scope. A defensible market ranking needs an explicit estimand, eligible candidate population, representative or deliberately weighted query universe, comparable opportunity, repeated observations and uncertainty.

A defensible authority association needs verified cited and eligible-but-uncited pages. A causal authority claim requires an intervention or credible quasi-experimental design.

Sixth companion paper

Research lineage

This capstone synthesizes the same underlying 496-response Finance GEO dataset. It is not another independent experiment.

Research 01 · Evidence landscape

What 496 Grounded AI Responses Reveal About Finance GEO Visibility asks what the responses contained.

Research 02 · Measurement meaning

Branded AI Visibility Is Not Market Visibility separates branded discoverability from general market visibility.

Research 03 · External ecosystem

The External Publisher Ecosystem Behind Finance AI Answers identifies recurring external source identities.

Research 04 · Measurement resolution

A Named Source Is Not a Verified Citation defines the page-identity boundary.

Research 05 · Co-citation

Co-Citation Maps Co-Presence, Not Influence examines joint source-identity appearance.

Research 06 · Experimental design

This paper defines what stronger comparisons, market claims, authority associations and causal conclusions require.

Research snapshot

The descriptive evidence and its boundaries

Observed Finance GEO results and the level of inference each supports.
MeasurementObserved valueWhat it supports
Prompt runs / successful responses496 / 495Descriptive response universe and high completeness
Responses with named sources492Strong named-source coverage
Target-owned source responses450 / 496 · 90.7%Target-oriented owned-source presence
Target distribution49 targets · median 90% · range 60%–100%Sample-specific descriptive variation
Runs per target6–20Unequal exposure and precision
Environmentgemini-2.5-flash-lite · US English · 14–15 Jan 2026One model, locale and short wave
Named / canonical page identities4,660 / 2Named-source analysis; severe page-level limitation

Methodological hierarchy

Three levels of claim require three levels of evidence

1. Descriptive claim

“Target-owned sources appeared in 450 of 496 target-oriented runs.”

Requires observed outcomes, correct denominator, correct identity level and explicit scope. The current dataset supports this.

2. Comparative or generalisation claim

“Company A is more visible than Company B.”

Requires comparable opportunity, eligibility, balanced or predeclared weighting, repeats, uncertainty and a defined environment.

3. Causal claim

“Higher page authority causes AI citation.”

Requires verified pages, cited and eligible-but-uncited candidates, temporal order, confounder handling and experimental or credible quasi-experimental identification.

Estimand-first framework

Define the estimand before building the ranking

An estimand is the exact quantity the study intends to estimate. “Which brand wins GEO?” does not define one.

  1. Population

    Which companies, products, buyers and markets are in scope?

  2. Query universe

    Which documented demand and intent categories should the prompts represent?

  3. Eligibility

    Which companies could plausibly answer each query?

  4. Outcome

    Mention, shortlist, recommendation, position, citation, accuracy or another explicit measure?

  5. Environment and time

    Which models, interfaces, markets, locales, retrieval modes and collection window?

  6. Repetition

    How many repeated observations are needed for the desired precision?

  7. Weighting

    How do query strata contribute to the overall result?

  8. Uncertainty

    How will clustered variation, ties and rank stability be reported?

Example estimand: Among eligible US finance companies, what proportion of repeated, unbranded category prompts mentions each company over a defined 30-day period across specified AI systems?

Outcome discipline

“AI visibility” is not one outcome

Mention, shortlist inclusion, recommendation, ordered position, target-owned source presence, external-source presence, sentiment, factual accuracy, positioning accuracy, citation persistence and representation consistency can move independently.

A company can have strong branded discoverability but weak unbranded selection, frequent mentions but poor accuracy, or strong source participation without recommendation. A composite must publish its formula, weights and sensitivity rather than concealing these distinctions.

Five design boundaries

Why descriptive target variation is not a market ranking

  • 1. Unequal target exposure

    Targets received 6–20 runs. A 6/6 result sorts above 19/20, but the smaller estimate does not contain stronger evidence.

  • 2. Unbalanced prompt opportunity

    Different intent, comparison and difficulty mixes provide different opportunities for an owned source to appear.

  • 3. Limited repetition

    One generative realization is not persistent performance; repeated trials are needed to estimate variability.

  • 4. Target-oriented prompts

    Companies were already in scope. Research 02 explains why this is not unbranded market discovery.

  • 5. One environment

    One model, US English and a short collection window do not establish broad rank stability.

Figure 1

A sortable distribution is not automatically a precision ranking

Observed target-owned source-presence rates range from 60 to 100 percent across 49 finance targets. Unequal target exposure of 6 to 20 prompt runs means the distribution is descriptive and should not be treated as a stable precision ranking.
Figure 1. Observed target-owned source presence ranged from 60% to 100% across 49 targets. Targets received between 6 and 20 prompt runs, so the sortable descriptive distribution does not by itself support a stable 1–49 market ranking.
Open full-resolution figure

Uncertainty

A raw ordering is not a statistically separable ranking

Two observed rates such as 90% and 85% can be forced into first and second position even when the available precision cannot distinguish them. Responsible rankings may need numerators, denominators, effective sample sizes, uncertainty intervals, exposure thresholds, weighting sensitivity, model/date stability, ties and rank bands.

Every target need not occupy a unique rank merely because software can sort the values.

Opportunity definition

Absence is uninterpretable without defining opportunity

Company ranking

Which companies were eligible for this query? Geography, product, customer type, use case, regulation and commercial fit can make absence irrelevant rather than a visibility failure.

Authority / source selection

Which pages plausibly could have supported the query? The candidate panel needs cited and relevant eligible-but-uncited pages from the same opportunity.

Variable definition

“Authority” must be defined as a metric

Authority is not one natural variable. A study may mean Domain Authority, Domain Rating, Page Authority, backlink count, referring-domain count, publisher prominence or another explicitly defined measure.

A credible study specifies the exact metric, whether it is page- or domain-level, when it was measured and how it relates to the unit selected. A later domain metric cannot automatically explain an unresolved page observation.

Selection problem

Cited-only data describe the selected winners

Collecting authority metrics for cited sources describes pages that already experienced the outcome. It cannot estimate a source-selection effect without relevant pages that could plausibly have been cited but were not.

A football analysis of goal scorers cannot identify which attributes increase scoring probability without considering players who had comparable scoring opportunities but did not score. Likewise, arbitrary unrelated uncited pages are not a control group. The key contrast is cited versus eligible-but-uncited within the same opportunity.

Research 04 consistency

Page-level features require verified page/document identity

The study had 4,660 named-source objects but only two canonical page identities in the collected records. A page-level study must verify identity using original and fetched URLs, final redirected URL, redirect chain, canonical signal, normalized URL, content fingerprint, duplicate relationship and syndication relationship.

A canonical URL is one signal; stable, auditable document identity is the broader requirement. Without it, backlinks, Page Authority, authorship, freshness and content structure cannot reliably be joined to the page that produced the observation. See Research 04.

Inference boundary

Association would still not prove causation

A future candidate-panel study might support: higher authority metric values are conditionally associated with citation probability after controlling for defined covariates.

That is weaker than “higher authority causes citation.” Authority can correlate with relevance, brand prominence, age, publishing activity, content completeness, freshness and retrieval accessibility. Confounder adjustment can help, but causal language requires credible identification through matched designs, within-prompt models, longitudinal natural experiments, controlled publication interventions or suitable quasi-experimental variation.

Figure 2 · Native framework

Claim–Evidence Ladder

Evidence requirements increase as the sentence moves from bounded description to market generalisation, association and causality.
ClaimMinimum evidence required
“Owned sources appeared in 90.7% of these target-oriented runs.”Complete outcomes, correct denominator and target attribution.
“Target A had a higher observed rate than Target B in this sample.”Target numerators/denominators and a comparable metric definition.
“Target A is more visible than Target B.”Comparable opportunity, balanced/weighted prompts, repeated trials and uncertainty.
“Target A leads the general finance market.”Defined population, representative unbranded query universe, eligibility and market scope.
“Publisher X recurred frequently.”Response-level source incidence.
“High-authority pages are associated with citation.”Verified cited and eligible uncited candidates, contemporaneous features and confounder adjustment.
“Authority causes citation.”Credible causal design establishing temporal order and isolating the authority effect.
Figure 2. Claim–Evidence Ladder. The stronger the sentence, the stronger the experimental design required.

Figure 3 · Native comparison

Two different experiments for two different claims

Market Visibility Ranking Study

Purpose: Estimate comparative company visibility.

  1. Explicit estimand
  2. Eligible company universe
  3. Representative unbranded queries
  4. Comparable opportunity
  5. Repeated runs
  6. Multiple dates
  7. Defined models, markets and locales
  8. Separate outcomes
  9. Cluster-aware analysis
  10. Uncertainty, ties and rank stability

Output: Defensible comparative visibility estimate—not raw sorted target-oriented percentages.

Authority / Source-Selection Study

Purpose: Estimate whether a defined authority metric is associated with source selection.

  1. Precisely defined metric
  2. Reproducible page candidates
  3. Cited pages
  4. Eligible-but-uncited pages
  5. Verified document identity
  6. Contemporaneous features
  7. Within-prompt comparison
  8. Confounder adjustment
  9. Out-of-sample validation
  10. Stronger design before causal language

Output: Association or causal estimate appropriate to the design—not cited-winner properties alone.

Figure 3. A ranking experiment and an authority-effect experiment require different candidate universes, outcomes and controls.

Proposed protocol

Designing a defensible market-ranking study

  1. 1. Define the estimand and population

    Specify the outcome, eligible companies, product categories, markets and buyers.

  2. 2. Build a representative query universe

    Use neutral unbranded prompts stratified by intent with weights declared before results.

  3. 3. Balance opportunity

    Use eligibility, matched prompt families, balanced sampling or predeclared weighting.

  4. 4. Repeat and collect across time

    Choose repeats for desired precision and use multiple waves to estimate persistence.

  5. 5. Define environments

    Specify models, interfaces, locales, languages and retrieval modes before pooling.

  6. 6. Separate outcomes

    Measure mention, shortlist, recommendation, position, citations, sentiment and accuracy separately.

  7. 7. Model clustering

    Respect repeats within prompts, prompts within intent families, companies, dates and models.

  8. 8. Report uncertainty and ties

    Publish sample sizes, intervals, rank stability and performance bands without forcing unique order.

Different proposed protocol

Designing a defensible authority-association study

  1. 1. Define the authority metric

    Name the exact page/domain measure and timestamp.

  2. 2. Build a candidate panel

    Identify reproducibly the cited and eligible-but-uncited pages for each prompt.

  3. 3. Verify document identity

    Resolve URLs, redirects, canonical signals, fingerprints, duplicates and syndication.

  4. 4. Measure features at the correct time

    Timestamp authority and content variables at or before the citation opportunity.

  5. 5. Compare within opportunity

    Compare pages relevant to the same prompt to reduce topic/context differences.

  6. 6. Adjust plausible confounders

    Include relevance, prominence, freshness, completeness, page type, activity and accessibility.

  7. 7. Validate out of sample

    Test on new prompts, targets, dates or models.

  8. 8. Report association honestly

    Use conditional-association language unless a causal identification strategy exists.

Counterfactual requirement

What would justify causal language?

The causal question is: what would happen to citation probability if the proposed authority-related exposure changed while relevant alternatives remained comparable?

Causal conclusions need more than correlation among selected sources. Depending on the construct, credible approaches may include matched designs, within-prompt candidate models, controlled publishing treatments, longitudinal natural experiments or quasi-experimental timing. The current dataset did not conduct these analyses.

Supported conclusions

What the current evidence supports

  • High collection completeness

    Response and named-source coverage were operationally strong.

  • Target-oriented owned-source presence

    450 of 496 runs contained a target-owned source; observed target rates varied descriptively.

  • External and co-citation patterns

    External incidence was fragmented and 923 source-identity pairs recurred at least twice.

  • Severe page-identity limitation

    The source records do not broadly support exact-page analysis.

  • Hypotheses and a roadmap

    The series identifies measurement gaps and concrete next-stage experiments.

Claim boundary

What the current evidence does not support

  • Stable company or market rankings

    No stable 1–49 order, market-wide share of voice, unaided leadership or cross-model ranking is established.

  • One universal visibility score

    Distinct outcomes cannot be collapsed without a transparent decision-specific formula.

  • Page-authority or feature effects

    Cited-only unresolved source data cannot estimate page-level selection associations.

  • Causal authority claims

    The dataset does not identify causal Domain Authority, Page Authority or other page-feature effects.

  • Broad market generality

    One short target-oriented model snapshot does not represent every market or system.

Decision use

Questions before trusting a ranking or authority claim

Before trusting a GEO ranking
  1. What outcome and estimand?
  2. Branded or unbranded prompts?
  3. Which companies are eligible?
  4. Comparable opportunity?
  5. How many repeats?
  6. Which models, markets and dates?
  7. How are prompts weighted?
  8. Is uncertainty shown?
  9. Are ties or bands allowed?
  10. Sample description or market generalisation?
Before trusting an authority claim
  1. Which exact metric?
  2. Page or domain level?
  3. Verified documents?
  4. Eligible uncited comparison pages?
  5. How was the candidate universe built?
  6. Features measured before exposure?
  7. Confounders addressed?
  8. Within-opportunity comparison?
  9. Out-of-sample validation?
  10. Association or causation?
  11. If causal, what identification design?

Next generation

Research agenda

Market-visibility work should build representative unbranded query universes, explicit eligibility, repeated waves, defined environments and uncertainty-aware ranks. Source-selection work should build verified cited and eligible-uncited candidate panels with contemporaneous features and within-prompt comparisons. Causal work needs predeclared hypotheses, temporal ordering, credible interventions or quasi-experimental variation, replication and out-of-sample validation.

Every published GEO claim should state its estimand, numerator, denominator, identity level, eligibility rule, environment, time window, repetition, uncertainty and whether the conclusion is descriptive, associative or causal.

Conclusion

The experiment must match the sentence in the headline

The Finance GEO series contains strong descriptive evidence. It should be used for the questions it answers: target-oriented owned-source presence, descriptive target variation, fragmented external incidence, recurring source co-presence and source-identity limits.

A descriptive rate does not become a market ranking because it can be sorted. A source feature does not become a citation cause because it is common among cited winners.

A defensible ranking begins with an estimand, eligible population, representative opportunity, repeats and uncertainty. A defensible authority study begins with verified documents and a cited-versus-eligible-but-uncited contrast. A causal claim requires a design capable of identifying cause.

The experiment must match the sentence in the headline.

Frequently asked questions

Frequently asked questions

Can the 49 finance companies be ranked from this dataset?

They can be sorted by observed target-owned source rate, but that ordering is not a stable market ranking. Exposure was unequal, prompt mixes were unbalanced, repetitions were limited and prompts were target-oriented.

What does 90.7% measure?

It is the prompt-weighted share of 496 target-oriented runs containing at least one target-owned source. It measures owned-source presence after the company was already in scope.

Why is a raw ranking not enough?

A sorted order forces small or uncertain differences into unique positions. Defensible ranking requires comparable opportunity, repeated observations and uncertainty-aware reporting.

What is an estimand?

It is the exact quantity a study intends to estimate, including population, query universe, eligibility, outcome, environment, time period and weighting.

Why do unbranded prompts matter for market rankings?

Target-oriented prompts already place a company in scope. General market visibility requires questions where the model must first decide which eligible companies to include.

Why cannot cited pages alone establish an authority effect?

Every observed page has already experienced citation. Estimating an association with citation probability also requires relevant pages that were eligible to be cited but were not.

What is an eligible-but-uncited page?

It is a page that plausibly could have supported the same prompt but did not appear. Such pages provide the comparison required for candidate-selection analysis.

Does a high authority metric prove causality?

No. Authority can correlate with relevance, prominence, age, publishing activity, backlinks and accessibility. Association does not identify cause.

Why does page identity matter?

Page-level authority and content features belong to specific documents. Without verified page or document identity, those features cannot be reliably attached to the selected source.

Is a canonical URL mandatory?

A canonical URL is an important identity signal, but the broader requirement is stable verified page or document identity using resolved URLs, redirects, normalization, fingerprints and duplication relationships.

Should GEO rankings show ties?

Yes, when available uncertainty does not support a unique ordering. Rank bands or ties can be more honest than forcing every company into a separate position.

What makes a market-visibility ranking stronger?

A representative unbranded query universe, explicit eligibility, comparable exposure, repeated runs, multiple dates, defined environments, separate outcomes and cluster-aware uncertainty.

What makes an authority study stronger?

Verified pages, a reproducible cited and eligible-uncited candidate panel, contemporaneous metrics, within-prompt comparison, confounder adjustment and out-of-sample validation.

When is causal language justified?

When the design credibly estimates what would happen under a change in the proposed cause, generally through an intervention or strong quasi-experimental identification rather than cited-source correlation.

What is the main methodological rule?

The experiment must match the sentence in the headline. Stronger claims require stronger experimental design.

Piush Vaish, Founder of Kojable

Author

About the author

Piush VaishFounder and CEO of Kojable

Piush Vaish is the founder and CEO of Kojable, a repeat founder and data scientist with more than 10 years of experience building and productising AI, machine-learning and data products. His experience spans high-growth technology companies and large enterprise environments. He combines technical depth with customer discovery, creative problem-solving and a strong bias towards shipping useful products. He writes about AI search, AEO, GEO, agentic discovery and AI product strategy.

Read more about Piush Vaish