Kojable research · 1,500 finance-oriented prompts

When Do Personas Really Change AI Responses?

Published Updated By Piush Vaish

Lessons from a 1,500-prompt similarity study of 12 finance personas, three topic groups and four intent types.

By Piush Vaish, founder and CEO of Kojable.

Key finding

After adjustment, responses associated with the same persona retained a modest similarity signal of +0.0142.

Qualification: Prompt construction explained approximately 76% of the raw response-similarity gap, and the historical experiment cannot establish that personas independently caused the remaining association.

  • Persona prompting
  • AI response similarity
  • Grounding queries
  • 12 min read
1,500prompts
1,494usable responses
12finance personas
+0.0142adjusted response gap Qualification Residual association, not a causal estimate.

Study overview

Executive summary

Personas are a standard ingredient in AI prompting. A model may be asked to respond as if it were advising a CFO, a compliance officer, a founder or a finance analyst. The assumption is intuitive: if the persona changes, the answer should change in a useful and predictable way.

But detecting a difference is not the same as proving that the persona caused it. Persona prompts usually change several things at once: the job title, objective, risk tolerance, vocabulary, constraints, time horizon and success metrics. If two outputs look different, the difference may simply reflect those altered words. Conversely, highly similar answers may indicate either desirable factual consistency or ineffective personalisation.

We analysed 1,500 finance-oriented prompts and their grounded AI responses. The reanalysis asked how much persona-related signal remained in responses and grounding queries after accounting for prompt construction and experimental design.

The result is nuanced. Personas leave a measurable signature, but most of the raw similarity effect is explained by the prompts themselves. The remaining signal is modest, and the historical experiment cannot establish that personas independently caused it.

Related research: See the separate 750-response matched AEO study of persona-conditioned search, pipeline, AEO and SEO language.

Answer first

Direct answer

Within this synthetic prompt generator and single collection run, responses associated with the same persona retained a modest similarity signal after adjustment for observed prompt semantics and design factors.

The raw within-persona response-similarity gap was +0.0586. After controlling for base query, template identifier, prompt length and local prompt-semantic components, it fell to +0.0142. Adjustment therefore removed approximately 76% of the raw gap. This is an exploratory association, not evidence that personas independently caused the remaining response differences.

Comparison of raw and adjusted within-persona similarity effects for prompts, AI responses and grounding queries.
Figure 1. Raw and adjusted same-persona similarity effects. The response gap fell from +0.0586 to +0.0142 after adjustment, while the grounding-query gap fell from +0.0890 to +0.0375.
Open full-resolution figure

Study design

The experiment

The prompt population covered 12 finance personas:

  • CFO
  • FP&A lead
  • treasury manager
  • compliance officer
  • finance operations manager
  • payments operations lead
  • accounts-receivable manager
  • accounts-payable manager
  • founder
  • revenue-operations lead
  • finance analyst
  • internal auditor

The prompts spanned three topic groups—cash flow, payment processing and fraud detection—and four intent types: informational, commercial, transactional and educational. Each persona was represented by exactly 125 prompts, and each topic group contained 500 prompts.

Every persona carried a structured profile. A CFO, for example, was associated with runway protection, board-ready planning and a risk-averse posture. A compliance officer was associated with control evidence, audit findings and very low risk tolerance. These attributes were inserted into hand-written prompt templates alongside topic, industry, geography, company size, metric and integration information.

The prompts were sent to a grounded generative model, which produced a response, a set of web-search queries and grounding-source metadata. Of the 1,500 calls, 1,494 produced usable responses.

Analytical control

Why the original analysis needed a more controlled approach

A simple analysis could compare every response with every other response, label each pair “same persona” or “different persona,” and test whether the first group has a higher average similarity. That creates two problems.

First, the observations are not independent. With 1,494 successful responses, there are more than 1.1 million response pairs, but every response appears in many pairs. Treating those pairs as independent observations creates extremely narrow confidence intervals and overstated statistical significance.

Second, same-persona prompts share repeated vocabulary and profile attributes by construction. Higher within-persona similarity may therefore be a prompt-generation effect rather than evidence of independent model behaviour.

The improved analysis:

  1. reconstructed and retained full prompt metadata, including template identifiers and sampled dimensions;
  2. audited treatment coverage, missingness, query hygiene and runtime characteristics;
  3. replaced persona rankings with uncertainty intervals;
  4. adjusted operational outcomes for base query, template and prompt length;
  5. used deterministic local TF–IDF and latent semantic analysis rather than additional paid embedding calls;
  6. residualised response and query representations on prompt semantics and design factors; and
  7. used template-cluster bootstraps that avoid counting a duplicated original response as a perfect self-pair.

This does not turn the historical run into a randomised controlled experiment, but it produces a more defensible exploratory assessment.

Results

Detailed findings

The generator clearly inserts a persona signature

The strongest persona effect appears before the model is called. The raw within-persona minus cross-persona prompt-similarity difference was +0.0815, with a 95% template-cluster bootstrap interval of +0.0632 to +0.1078.

This is useful as a manipulation check: the prompt generator successfully produces text that is more similar within a persona than across personas. It should not be interpreted as evidence that the model independently learned or applied a persona.

The prompt audit also showed uneven treatment intensity: 98.8% of prompts contained either a persona label or persona-profile phrase, 31.0% contained the literal persona label, and 92.7% explicitly contained the assigned base query. The remaining prompts weaken treatment consistency because “persona” is not a uniform experimental treatment.

Most response similarity is explained by prompt construction

Within-persona minus cross-persona similarity differences and 95% template-cluster bootstrap intervals.
RepresentationRaw differenceRaw 95% intervalAdjusted differenceAdjusted 95% intervalAttenuation
Responses+0.0586+0.0470 to +0.0748+0.0142+0.0100 to +0.0237Approximately 76%
Grounding queries+0.0890+0.0720 to +0.1123+0.0375+0.0296 to +0.0552Approximately 58%

The remaining response association is positive, suggesting that persona-conditioned prompts may leave a residual signature beyond the observed prompt-design variables. Its magnitude is much smaller than the unadjusted result, and unobserved prompt differences may still explain part of it.

Query similarity also attenuates after adjustment

The adjusted query signal was larger than the adjusted response signal. One interpretation is that persona wording affects how the grounded model frames its information search even when overall search volume remains stable. Another is that residual prompt features not fully captured by the local semantic model continue to influence the generated queries. The experiment cannot distinguish between these explanations.

Persona does not explain grounding volume

The descriptive dashboard showed grounding coverage ranging from 71.0% for finance-operations managers to 84.0% for founders, but the uncertainty intervals overlap substantially.

Grounding coverage estimates with uncertainty intervals for the 12 finance personas.
Figure 2. Grounding coverage by persona with uncertainty intervals. The observed range from 71.0% to 84.0% does not by itself establish a stable persona hierarchy because the intervals overlap substantially.
Open full-resolution figure
Persona omnibus tests after adjustment for base query, template and prompt length.
OutcomePersona omnibus p-valueInterpretation
Grounding coverage0.900No persona contribution detected
Sources per response0.490No persona contribution detected
Clean queries per responseApproximately 1.000No persona contribution detected
Response length0.00184Overall persona association detected

Intent was more informative than persona for grounding coverage. Transactional prompts had coverage of approximately 69.6%, while commercial prompts achieved approximately 82.3%. A dashboard can always identify the highest and lowest category, but extrema are not automatically evidence of a stable effect.

Response length changes modestly

Response length was the only operational outcome with a clear overall persona association after adjustment. Adjusted means ranged from approximately 626.6 words for finance-operations managers to 649.3 words for compliance officers: about 22.7 words, or 3.6%.

This difference is statistically detectable, but its business importance is uncertain. A longer answer is not necessarily more relevant, accurate or useful. Length is better treated as a diagnostic or cost metric than as a measure of personalisation quality.

Adjusted estimates for grounding coverage, source volume, clean-query volume and response length across personas.
Figure 3. Adjusted operational outcomes by persona. Grounding coverage, sources per response and clean queries per response showed no detected persona contribution; response length showed a modest overall association.
Open full-resolution figure

Some persona pairings are plausible, but exploratory

After removing variation associated with observed prompt and design factors, several response-centroid pairings stood out: compliance officer and internal auditor at +0.248, CFO and founder at +0.147, and FP&A lead and finance analyst at +0.136. CFO and founder was the clearest positive query-centroid pairing at +0.104.

These pairings make intuitive sense, but they are hypotheses for future validation. The persona taxonomy was defined by the same profiles used to generate the prompts, so the analysis cannot independently validate that taxonomy.

Heatmaps comparing residual response-centroid and query-centroid similarity across the 12 finance personas.
Figure 4. Residual persona-centroid similarity for responses and grounding queries. The strongest response pairing was compliance officer with internal auditor (+0.248); the query-side structure was weaker.
Open full-resolution figure

Operational evidence

Data quality matters as much as modelling

  • All six failed responses were transient 503 errors.
  • The failures occurred in one contiguous payment-processing collection window, at indices 812–841.
  • Five of the six failures were transactional prompts.
  • The saved records contain 297 blank query entries.
  • The run accumulated approximately 9.41 hours of recorded model-processing time.
  • One call recorded more than 1,349 seconds of processing time.
  • The source metadata contains 3,797 source records but only 1,738 distinct titles.
Data-quality audit summarising prompt treatment coverage, response failures, blank query entries and runtime characteristics.
Figure 5. Data-quality audit for the historical collection. Six calls failed, 297 blank query entries were retained, and persona treatment intensity varied across prompt templates.
Open full-resolution figure

Because prompts were collected in topic blocks, model availability, search-index changes and time are partially confounded with topic. Future runs should randomise or interleave prompt conditions and record precise request timestamps, configuration, retry history and model-version metadata.

Source analysis also needs care. The saved URIs are grounding redirect URLs rather than canonical destination URLs, while exact source titles are unstable identifiers. Source relevance, authority, diversity, ownership and citation support are more meaningful outcomes than raw title uniqueness.

Evaluation framework

What persona similarity should measure

Semantic similarity is not itself a business outcome. For facts, policies and compliance requirements, similarity across personas may be desirable because the underlying truth should not change. For recommendations and explanations, some divergence may be useful: a CFO may need risk, capital-allocation and board-level implications, while an operations lead may need implementation detail and process controls.

A strong evaluation should distinguish between:

  • invariant content, which should remain consistent;
  • persona-relevant framing, which should change appropriately;
  • unsupported divergence, where personalisation changes facts or introduces risk;
  • superficial variation, where wording changes without improving usefulness.

This suggests a better primary metric: persona-relevance lift over a matched generic response, subject to factuality and safety guardrails.

Next study

The confirmatory experiment

The next study should use a matched, blocked design. For every stable prompt skeleton:

  1. Hold the topic, intent, industry, geography, company size, metric, wording and integration constant.
  2. Render the same skeleton under every persona.
  3. Include a generic or no-persona condition.
  4. Include a shuffled or deliberately mismatched persona as a placebo control.
  5. Repeat each condition across multiple seeds or collection periods.
  6. Randomise and interleave the execution order.

The prompt skeleton—not the response pair—should be the primary inferential unit. Hierarchical models or block-restricted permutations can then estimate controlled persona contrasts without treating millions of dependent pairs as independent evidence.

The primary endpoint should be decided in advance. A practical evaluation could combine blinded human ratings of persona relevance; factual and citation support; topic fidelity; source authority and diversity; finance and compliance safety; latency, token use and cost; and stability across repeated generations.

Embedding similarity can remain a useful diagnostic, but it should be validated against human judgements before it becomes a decision KPI.

Method and governance

Methodology and reproducibility

Design and sample

The study used an existing artifact of 1,500 finance-oriented prompts and grounded-model calls. It included 12 personas with 125 prompts each, three topic groups with 500 prompts each, and four intent types. There were 1,494 usable responses.

Representations and adjustment

Deterministic local TF–IDF and latent semantic analysis provided the similarity proxy. Response and query representations were residualised on observed prompt semantics and design factors. Operational outcomes were adjusted for base query, template and prompt length; the query-similarity analysis also adjusted for the number of clean queries.

Uncertainty and dependence

Template-cluster bootstraps were used for uncertainty intervals and avoided counting a duplicated original response as a perfect self-pair. The analysis is exploratory rather than confirmatory.

Recorded gaps

The source package does not record the collection dates, language, geography, exact model name or version, precise request timestamps, configuration or retry history. The grounded generative model is therefore not attributed to a specific model release or product environment.

Reproducibility and sign-off

No public analysis script, notebook or reproducibility package accompanies this publication. The published article and figures document the reported design, findings, interpretation and limitations.

The final article, figures, interpretation and methodological limitations were reviewed and approved for publication by Piush Vaish, founder and CEO of Kojable, on 5 August 2026. This was an author review and publication sign-off rather than an independent external review or replication.

Study boundaries

Limitations

  • The historical run was not a matched randomised experiment, so it cannot establish that personas caused the residual response or query signals.
  • Persona prompts changed several profile and wording attributes at once. Unobserved prompt differences may remain after adjustment.
  • Prompt templates and persona profiles defined the same taxonomy later examined by the analysis, so the study cannot independently validate that taxonomy.
  • Repeated response pairs are dependent. The cluster bootstrap improves inference but does not create independent observations.
  • Prompts were collected in topic blocks, partially confounding topic with time, model availability and search-index conditions.
  • Treatment intensity was uneven: some prompts used a literal role, others only profile phrases, and a small share contained no meaningful persona cue.
  • The deterministic local semantic representation is a proxy. It should be validated against blinded domain-rated outcomes.
  • Similarity, response length and grounding volume do not directly measure relevance, factuality, usefulness, safety or business value.

The practical boundary is clear: this study identifies a modest residual association and a better evaluation design. It does not prove that persona prompting independently improves AI behaviour.

FAQ

Frequently asked questions

Do personas independently cause AI responses to change?

This study cannot establish that. It found a modest residual association after adjustment, but the historical design did not hold all non-persona prompt features constant.

How much of the raw response-similarity gap did adjustment remove?

Approximately 76%. The same-persona response-similarity gap fell from +0.0586 before adjustment to +0.0142 after adjustment.

Did persona explain grounding coverage or search volume?

No persona contribution was detected for grounding coverage, sources per response or clean queries per response after adjustment.

How should a confirmatory persona study be designed?

Use matched prompt skeletons across every persona, include generic and placebo controls, randomise and interleave execution, repeat conditions, and evaluate persona relevance alongside factuality and safety.

Piush Vaish, Founder of Kojable

Author

About the author

Piush VaishFounder and CEO of Kojable

By Piush Vaish, founder and CEO of Kojable.

Piush Vaish is the founder and CEO of Kojable, a repeat founder and data scientist with more than 10 years of experience building and productising AI, machine-learning and data products. He writes about AI search, AEO, GEO, agentic discovery and AI product strategy.

Read more about Piush Vaish