Kojable research · 1,500 finance-oriented prompts
When Do Personas Really Change AI Responses?
Lessons from a 1,500-prompt similarity study of 12 finance personas, three topic groups and four intent types.
By Piush Vaish, founder and CEO of Kojable.
Key finding
After adjustment, responses associated with the same persona retained a modest similarity signal of +0.0142.
Qualification: Prompt construction explained approximately 76% of the raw response-similarity gap, and the historical experiment cannot establish that personas independently caused the remaining association.
- Persona prompting
- AI response similarity
- Grounding queries
- 12 min read
Study overview
Executive summary
Personas are a standard ingredient in AI prompting. A model may be asked to respond as if it were advising a CFO, a compliance officer, a founder or a finance analyst. The assumption is intuitive: if the persona changes, the answer should change in a useful and predictable way.
But detecting a difference is not the same as proving that the persona caused it. Persona prompts usually change several things at once: the job title, objective, risk tolerance, vocabulary, constraints, time horizon and success metrics. If two outputs look different, the difference may simply reflect those altered words. Conversely, highly similar answers may indicate either desirable factual consistency or ineffective personalisation.
We analysed 1,500 finance-oriented prompts and their grounded AI responses. The reanalysis asked how much persona-related signal remained in responses and grounding queries after accounting for prompt construction and experimental design.
The result is nuanced. Personas leave a measurable signature, but most of the raw similarity effect is explained by the prompts themselves. The remaining signal is modest, and the historical experiment cannot establish that personas independently caused it.
Related research: See the separate 750-response matched AEO study of persona-conditioned search, pipeline, AEO and SEO language.
Answer first
Direct answer
Within this synthetic prompt generator and single collection run, responses associated with the same persona retained a modest similarity signal after adjustment for observed prompt semantics and design factors.
The raw within-persona response-similarity gap was +0.0586. After controlling for base query, template identifier, prompt length and local prompt-semantic components, it fell to +0.0142. Adjustment therefore removed approximately 76% of the raw gap. This is an exploratory association, not evidence that personas independently caused the remaining response differences.
Study design
The experiment
The prompt population covered 12 finance personas:
- CFO
- FP&A lead
- treasury manager
- compliance officer
- finance operations manager
- payments operations lead
- accounts-receivable manager
- accounts-payable manager
- founder
- revenue-operations lead
- finance analyst
- internal auditor
The prompts spanned three topic groups—cash flow, payment processing and fraud detection—and four intent types: informational, commercial, transactional and educational. Each persona was represented by exactly 125 prompts, and each topic group contained 500 prompts.
Every persona carried a structured profile. A CFO, for example, was associated with runway protection, board-ready planning and a risk-averse posture. A compliance officer was associated with control evidence, audit findings and very low risk tolerance. These attributes were inserted into hand-written prompt templates alongside topic, industry, geography, company size, metric and integration information.
The prompts were sent to a grounded generative model, which produced a response, a set of web-search queries and grounding-source metadata. Of the 1,500 calls, 1,494 produced usable responses.
Analytical control
Why the original analysis needed a more controlled approach
A simple analysis could compare every response with every other response, label each pair “same persona” or “different persona,” and test whether the first group has a higher average similarity. That creates two problems.
First, the observations are not independent. With 1,494 successful responses, there are more than 1.1 million response pairs, but every response appears in many pairs. Treating those pairs as independent observations creates extremely narrow confidence intervals and overstated statistical significance.
Second, same-persona prompts share repeated vocabulary and profile attributes by construction. Higher within-persona similarity may therefore be a prompt-generation effect rather than evidence of independent model behaviour.
The improved analysis:
- reconstructed and retained full prompt metadata, including template identifiers and sampled dimensions;
- audited treatment coverage, missingness, query hygiene and runtime characteristics;
- replaced persona rankings with uncertainty intervals;
- adjusted operational outcomes for base query, template and prompt length;
- used deterministic local TF–IDF and latent semantic analysis rather than additional paid embedding calls;
- residualised response and query representations on prompt semantics and design factors; and
- used template-cluster bootstraps that avoid counting a duplicated original response as a perfect self-pair.
This does not turn the historical run into a randomised controlled experiment, but it produces a more defensible exploratory assessment.
Results
Detailed findings
The generator clearly inserts a persona signature
The strongest persona effect appears before the model is called. The raw within-persona minus cross-persona prompt-similarity difference was +0.0815, with a 95% template-cluster bootstrap interval of +0.0632 to +0.1078.
This is useful as a manipulation check: the prompt generator successfully produces text that is more similar within a persona than across personas. It should not be interpreted as evidence that the model independently learned or applied a persona.
The prompt audit also showed uneven treatment intensity: 98.8% of prompts contained either a persona label or persona-profile phrase, 31.0% contained the literal persona label, and 92.7% explicitly contained the assigned base query. The remaining prompts weaken treatment consistency because “persona” is not a uniform experimental treatment.
Most response similarity is explained by prompt construction
| Representation | Raw difference | Raw 95% interval | Adjusted difference | Adjusted 95% interval | Attenuation |
|---|---|---|---|---|---|
| Responses | +0.0586 | +0.0470 to +0.0748 | +0.0142 | +0.0100 to +0.0237 | Approximately 76% |
| Grounding queries | +0.0890 | +0.0720 to +0.1123 | +0.0375 | +0.0296 to +0.0552 | Approximately 58% |
The remaining response association is positive, suggesting that persona-conditioned prompts may leave a residual signature beyond the observed prompt-design variables. Its magnitude is much smaller than the unadjusted result, and unobserved prompt differences may still explain part of it.
Query similarity also attenuates after adjustment
The adjusted query signal was larger than the adjusted response signal. One interpretation is that persona wording affects how the grounded model frames its information search even when overall search volume remains stable. Another is that residual prompt features not fully captured by the local semantic model continue to influence the generated queries. The experiment cannot distinguish between these explanations.
Persona does not explain grounding volume
The descriptive dashboard showed grounding coverage ranging from 71.0% for finance-operations managers to 84.0% for founders, but the uncertainty intervals overlap substantially.
| Outcome | Persona omnibus p-value | Interpretation |
|---|---|---|
| Grounding coverage | 0.900 | No persona contribution detected |
| Sources per response | 0.490 | No persona contribution detected |
| Clean queries per response | Approximately 1.000 | No persona contribution detected |
| Response length | 0.00184 | Overall persona association detected |
Intent was more informative than persona for grounding coverage. Transactional prompts had coverage of approximately 69.6%, while commercial prompts achieved approximately 82.3%. A dashboard can always identify the highest and lowest category, but extrema are not automatically evidence of a stable effect.
Response length changes modestly
Response length was the only operational outcome with a clear overall persona association after adjustment. Adjusted means ranged from approximately 626.6 words for finance-operations managers to 649.3 words for compliance officers: about 22.7 words, or 3.6%.
This difference is statistically detectable, but its business importance is uncertain. A longer answer is not necessarily more relevant, accurate or useful. Length is better treated as a diagnostic or cost metric than as a measure of personalisation quality.
Some persona pairings are plausible, but exploratory
After removing variation associated with observed prompt and design factors, several response-centroid pairings stood out: compliance officer and internal auditor at +0.248, CFO and founder at +0.147, and FP&A lead and finance analyst at +0.136. CFO and founder was the clearest positive query-centroid pairing at +0.104.
These pairings make intuitive sense, but they are hypotheses for future validation. The persona taxonomy was defined by the same profiles used to generate the prompts, so the analysis cannot independently validate that taxonomy.
Operational evidence
Data quality matters as much as modelling
- All six failed responses were transient 503 errors.
- The failures occurred in one contiguous payment-processing collection window, at indices 812–841.
- Five of the six failures were transactional prompts.
- The saved records contain 297 blank query entries.
- The run accumulated approximately 9.41 hours of recorded model-processing time.
- One call recorded more than 1,349 seconds of processing time.
- The source metadata contains 3,797 source records but only 1,738 distinct titles.
Because prompts were collected in topic blocks, model availability, search-index changes and time are partially confounded with topic. Future runs should randomise or interleave prompt conditions and record precise request timestamps, configuration, retry history and model-version metadata.
Source analysis also needs care. The saved URIs are grounding redirect URLs rather than canonical destination URLs, while exact source titles are unstable identifiers. Source relevance, authority, diversity, ownership and citation support are more meaningful outcomes than raw title uniqueness.
Evaluation framework
What persona similarity should measure
Semantic similarity is not itself a business outcome. For facts, policies and compliance requirements, similarity across personas may be desirable because the underlying truth should not change. For recommendations and explanations, some divergence may be useful: a CFO may need risk, capital-allocation and board-level implications, while an operations lead may need implementation detail and process controls.
A strong evaluation should distinguish between:
- invariant content, which should remain consistent;
- persona-relevant framing, which should change appropriately;
- unsupported divergence, where personalisation changes facts or introduces risk;
- superficial variation, where wording changes without improving usefulness.
This suggests a better primary metric: persona-relevance lift over a matched generic response, subject to factuality and safety guardrails.
Next study
The confirmatory experiment
The next study should use a matched, blocked design. For every stable prompt skeleton:
- Hold the topic, intent, industry, geography, company size, metric, wording and integration constant.
- Render the same skeleton under every persona.
- Include a generic or no-persona condition.
- Include a shuffled or deliberately mismatched persona as a placebo control.
- Repeat each condition across multiple seeds or collection periods.
- Randomise and interleave the execution order.
The prompt skeleton—not the response pair—should be the primary inferential unit. Hierarchical models or block-restricted permutations can then estimate controlled persona contrasts without treating millions of dependent pairs as independent evidence.
The primary endpoint should be decided in advance. A practical evaluation could combine blinded human ratings of persona relevance; factual and citation support; topic fidelity; source authority and diversity; finance and compliance safety; latency, token use and cost; and stability across repeated generations.
Embedding similarity can remain a useful diagnostic, but it should be validated against human judgements before it becomes a decision KPI.
Method and governance
Methodology and reproducibility
Design and sample
The study used an existing artifact of 1,500 finance-oriented prompts and grounded-model calls. It included 12 personas with 125 prompts each, three topic groups with 500 prompts each, and four intent types. There were 1,494 usable responses.
Representations and adjustment
Deterministic local TF–IDF and latent semantic analysis provided the similarity proxy. Response and query representations were residualised on observed prompt semantics and design factors. Operational outcomes were adjusted for base query, template and prompt length; the query-similarity analysis also adjusted for the number of clean queries.
Uncertainty and dependence
Template-cluster bootstraps were used for uncertainty intervals and avoided counting a duplicated original response as a perfect self-pair. The analysis is exploratory rather than confirmatory.
Recorded gaps
The source package does not record the collection dates, language, geography, exact model name or version, precise request timestamps, configuration or retry history. The grounded generative model is therefore not attributed to a specific model release or product environment.
Reproducibility and sign-off
No public analysis script, notebook or reproducibility package accompanies this publication. The published article and figures document the reported design, findings, interpretation and limitations.
The final article, figures, interpretation and methodological limitations were reviewed and approved for publication by Piush Vaish, founder and CEO of Kojable, on 5 August 2026. This was an author review and publication sign-off rather than an independent external review or replication.
Study boundaries
Limitations
- The historical run was not a matched randomised experiment, so it cannot establish that personas caused the residual response or query signals.
- Persona prompts changed several profile and wording attributes at once. Unobserved prompt differences may remain after adjustment.
- Prompt templates and persona profiles defined the same taxonomy later examined by the analysis, so the study cannot independently validate that taxonomy.
- Repeated response pairs are dependent. The cluster bootstrap improves inference but does not create independent observations.
- Prompts were collected in topic blocks, partially confounding topic with time, model availability and search-index conditions.
- Treatment intensity was uneven: some prompts used a literal role, others only profile phrases, and a small share contained no meaningful persona cue.
- The deterministic local semantic representation is a proxy. It should be validated against blinded domain-rated outcomes.
- Similarity, response length and grounding volume do not directly measure relevance, factuality, usefulness, safety or business value.
The practical boundary is clear: this study identifies a modest residual association and a better evaluation design. It does not prove that persona prompting independently improves AI behaviour.
FAQ
Frequently asked questions
Do personas independently cause AI responses to change?
This study cannot establish that. It found a modest residual association after adjustment, but the historical design did not hold all non-persona prompt features constant.
How much of the raw response-similarity gap did adjustment remove?
Approximately 76%. The same-persona response-similarity gap fell from +0.0586 before adjustment to +0.0142 after adjustment.
Did persona explain grounding coverage or search volume?
No persona contribution was detected for grounding coverage, sources per response or clean queries per response after adjustment.
How should a confirmatory persona study be designed?
Use matched prompt skeletons across every persona, include generic and placebo controls, randomise and interleave execution, repeat conditions, and evaluate persona relevance alongside factuality and safety.




