Kojable research · Follow-on persona analysis · 1,494 responses

Beyond Persona Labels: Do AI Personas Still Differ After Explicit Cues Are Removed?

Published By Piush Vaish

A follow-on analysis of 1,494 finance-persona responses testing whether residual persona differentiation survives removal of exact role and profile language—and where that remaining signal appears most strongly.

By Piush Vaish, founder and CEO of Kojable.

Key finding

After exact persona and profile phrases were removed, approximately 76% of the adjusted response separation and 60% of the adjusted query separation remained.

Qualification: This is a post-generation sensitivity analysis, not a counterfactual intervention. The remaining separation should not be interpreted as a causal or cue-independent persona effect.

  • Persona prompting
  • Prompt echo
  • AI response framing
  • Search-query framing
  • 17 min read
1,500attempted prompts
1,494usable responses
12finance personas
76%of the adjusted response point estimate remained Qualification Proportion of an adjusted point estimate, not a causal effect or performance score.

Study overview

Executive summary

Kojable’s earlier persona-similarity research established an important boundary: outputs associated with the same persona were more similar to one another than outputs associated with different personas, but much of the raw separation was associated with prompt construction rather than something that could be cleanly attributed to the persona itself.

That result left a more specific question unresolved. When a residual persona signal remains after accounting for prompt design, what is that signal actually made of? Does the model merely repeat obvious persona language such as a role name, stated goal, KPI, risk tolerance or operating constraint? Or does some broader difference in vocabulary, emphasis and information-seeking remain after those exact cues are stripped away?

This follow-on analysis uses the same underlying finance-persona dataset. It removes exact persona names and exact profile phrases from the already generated responses and search queries, then recomputes the adjusted persona-separation diagnostic.

Literal persona cues explain a meaningful share of the remaining differentiation, particularly in generated search queries, but they do not explain all of it. After exact cue removal, approximately 76% of the adjusted response separation and 60% of the adjusted query separation remain.

The strongest residual signal appears upstream in the model’s generated search language rather than in its final answers. At the same time, persona assignment shows no detectable relationship with whether a response contains a source or with how many sources it contains.

Research lineage

How this extends the previous persona-similarity study

This is a continuation of Kojable’s earlier persona-similarity research, not a separate collection run.

The previous study asked whether persona-associated outputs still showed measurable semantic separation after accounting for prompt construction and observed design factors. This analysis begins where that study ended. It asks how much of that adjusted separation survives when explicit persona labels and profile phrases are removed from the generated text itself.

The distinction matters. The earlier question asks whether a residual signal exists. This question asks whether that signal is largely literal prompt echo or whether broader persona-associated framing remains after exact cues are removed.

The same dataset is useful for both questions because this analysis acts as a sensitivity test on the first. It does not create a new causal experiment and should not be interpreted as one.

Answer first

Direct answer

Answer

Yes, some persona-associated differentiation remains after exact persona and profile phrases are removed from the generated outputs. Literal persona echo is insufficient to explain the full adjusted pattern observed in this corpus.

For final responses, the design-adjusted within-persona minus cross-persona similarity difference was 0.0153. After exact cue removal, it fell to 0.0117. The estimated contribution associated with literal persona/profile phrases was 0.0036, with a 95% bootstrap interval of 0.0024 to 0.0059.

For generated search queries, adjusted separation was 0.0413 and fell to 0.0248 after exact cue removal. The estimated literal-cue contribution was 0.0165, with a 95% interval of 0.0112 to 0.0274.

Exact persona wording therefore explains part of the remaining signal, but not all of it. Removing words after generation cannot tell us what the model would have produced if those cues had never been present in the prompt.

Study design

One dataset, three related text diagnostics

The underlying collection attempted 1,500 prompts across 12 finance personas:

  • AP manager
  • AR manager
  • CFO
  • FP&A lead
  • compliance officer
  • finance analyst
  • finance operations manager
  • founder
  • internal auditor
  • payments operations lead
  • revenue operations lead
  • treasury manager

The prompt generator attempted 125 prompts per persona. Six attempts failed or returned blank responses, leaving 1,494 successful nonblank responses for analysis, with either 124 or 125 observations per persona. Each response was matched to the exact seeded prompt that generated it, so persona assignment came from prompt-generation metadata rather than keywords inferred from the response.

  1. Raw separation

    How much more similar texts from the same persona are than texts from different personas.

  2. Design-adjusted separation

    The same comparison after accounting for prompt template, prompt length, prompt-text features and non-persona factors rendered into the prompt. Query models also account for the number of generated queries.

  3. Literal-cue-ablated separation

    The adjusted comparison after removing exact persona names and exact profile phrases from the saved responses and queries.

The cue dictionary includes exact persona/profile language representing the assigned role, stated goal, constraint, risk tolerance, KPI and time horizon.

Text similarity was represented with deterministic local TF-IDF and latent semantic analysis. The principal diagnostic is within-persona minus cross-persona cosine similarity. A positive value means texts associated with the same persona are, on average, more similar than texts associated with different personas.

Uncertainty intervals use 1,999 paired bootstrap resamples at the prompt-template level. These are semantic diagnostics, not business-calibrated outcome measures. There is no established threshold at which a particular cosine-similarity difference becomes commercially important.

Analytical bridge

The residual signal that motivated this study

The raw response separation was 0.0587, with a 95% bootstrap interval of 0.0472 to 0.0748. Generated queries showed a larger raw difference of 0.0890, with a 95% interval of 0.0719 to 0.1128.

The prompts themselves were also separated by persona: the raw prompt-text difference was 0.0815, with a 95% interval of 0.0626 to 0.1078. The model did not receive interchangeable inputs across personas.

After adjustment for observed design factors, response separation fell to 0.0153, with a 95% interval of 0.0110 to 0.0252. Query separation fell to 0.0413, with a 95% interval of 0.0335 to 0.0599. Relative to the raw point estimates, these were descriptive attenuations of approximately 74% and 54%, respectively—not a causal decomposition.

Line chart comparing response and query persona-separation estimates across raw text, design-adjusted text and exact-cue-ablated adjusted text; both series decline, with query separation higher throughout.
Figure 1. From raw separation to cue-ablated separation. Adjustment removes variation associated with observed prompt design; the follow-on step removes exact persona and profile phrases from the generated text. Query separation remains higher than response separation throughout.
Open full-resolution figure

The bridge into this paper is deliberately narrow: a smaller positive signal remained after adjustment. The cue-ablation analysis asks what happens when one additional source of apparent differentiation—literal persona/profile language in the generated output—is removed.

Finding 1

Explicit cues explain part—but not all—of the remaining differentiation

The central test removes exact persona names and exact profile phrases from the saved outputs before recomputing the adjusted diagnostic.

Adjusted separation before and after exact cue removal.
TextDesign adjustedAfter cue removalLiteral-cue contributionContribution 95% intervalAdjusted estimate remaining
Responses0.01530.01170.00360.0024 to 0.0059Approximately 76%
Generated queries0.04130.02480.01650.0112 to 0.0274Approximately 60%

For responses, approximately 24% of the adjusted point estimate was associated with exact persona/profile wording, while approximately 76% remained. For queries, approximately 40% was associated with exact wording and approximately 60% remained.

Two labelled horizontal stacked bars show that 76 percent of adjusted response separation and 60 percent of adjusted query separation remain after exact persona and profile cues are removed; the remaining shares are shown separately from literal-cue contributions.
Figure 2. Literal-cue sensitivity of the adjusted persona-separation diagnostic. Exact persona and profile language explains a meaningful share of the remaining pattern, especially in queries, but does not account for all observed differentiation.
Open full-resolution figure

This is evidence against the simplest version of the prompt-echo hypothesis. If persona differentiation were only a matter of repeating a job title or exact profile phrase, the separation should collapse once those strings are removed. It does not.

The defensible conclusion is that exact persona/profile language is one contributor to persona-associated differentiation, but it is not sufficient to explain the full adjusted pattern in these responses and queries.

Finding 2

The stronger residual signal appears in information-seeking language

Across every stage of the analysis, generated queries are more separated by persona than final responses. After design adjustment, response separation was 0.0153 and query separation was 0.0413. After exact cue removal, the corresponding estimates were 0.0117 and 0.0248.

The difference survives after obvious persona language has been stripped out. This suggests that persona conditioning may be expressed more strongly in how the system frames an information need than in the wording of the final answer. A finance persona can influence which concepts are foregrounded, which relationships are investigated and how a question is decomposed into searches, even when the final synthesis converges towards common financial language and answer structures.

That interpretation is relevant for systems using generated fan-out queries, retrieval planning or tool calls before producing a final response. Persona-associated differences may be more visible upstream, where the system decides what to investigate, than downstream as a stylistic rewrite.

Finding 3

Different framing did not produce detectable differences in basic evidence volume

Of the 1,494 analysed responses, 1,171 contained at least one recorded source, corresponding to an overall grounding-coverage rate of approximately 78.4%.

After adjusting the grounding model for base query, prompt template and prompt length, the omnibus persona test produced p=0.8998. The data therefore provide no evidence that the probability of including at least one source differed across the 12 personas.

Responses contained an average of 2.541 sources. The adjusted negative-binomial source-count model produced an omnibus persona test of p=0.4900, again providing no evidence of a persona-level difference.

Two labelled forest plots show adjusted grounding probability and source-count estimates for all 12 personas; intervals overlap broadly and the omnibus persona tests are p equals 0.8998 and p equals 0.4900.
Figure 3. Adjusted grounding coverage and source quantity by persona. The semantic-framing results are not accompanied by a detectable persona-level difference in whether sources appear or how many source objects are returned.
Open full-resolution figure

A secondary output-volume analysis likewise found no persona effect on the number of generated queries, with an omnibus p-value of approximately 0.9999.

Response length did vary across personas in an adjusted omnibus test, with p approximately 0.0018, but length is not a validated measure of quality, usefulness or efficiency and should not be used to rank personas.

Contrast result

Persona-associated query content differed, but persona assignment did not detectably change basic query quantity, grounding presence or source volume.

Source presence and source count remain weak proxies for evidence quality. This analysis did not adjudicate source authority, factual correctness, citation entailment or claim-level coverage. The correct conclusion is “no detected persona effect on source presence or quantity”, not “all personas are equally well grounded”.

Interpretation

What might the residual signal represent?

Broader semantic framing

Removing “CFO” or an exact KPI phrase does not remove every concept statistically or semantically associated with capital allocation, liquidity, board reporting, runway or financial risk.

Topic emphasis

Two personas can answer the same broad problem while prioritising different parts of it, such as controls and auditability, cash impact or workflow execution.

Information-seeking structure

The stronger query result suggests persona conditioning may affect how the system decomposes a task into information needs, even when final answers are more standardised.

Remaining design correlation

Because the historical prompt system was not a fully crossed matched experiment, broader persona-associated prompt differences remain a plausible alternative explanation.

These mechanisms are not mutually exclusive. The present analysis cannot determine how much of the residual belongs to each one.

Claim boundary

What this analysis does not prove

  • A causal persona effect

    The analysis removes text after generation. It does not regenerate the same task under a prompt in which persona cues were never supplied.

  • A useful residual

    A persistent semantic signature could reflect useful role-relevant framing, superficial paraphrase, correlated prompt structure or a combination.

  • A “best” persona

    The diagnostics do not show which persona is more accurate, helpful, persuasive, efficient or commercially valuable.

  • Better retrieval

    More distinctive query language does not mean the resulting searches retrieve better evidence.

  • Grounding quality

    Grounding presence and source count do not measure source authority, citation correctness or claim support.

The central distinction is between being detectably different and being demonstrably better.

Design implications

Implications for persona design

  • Do not evaluate personas by surface wording alone

    An evaluation that only checks for an assigned role or repeated KPI will overstate persona differentiation. Literal cue repetition is a measurable part of the signal.

  • Evaluate the information-seeking layer separately

    For systems using retrieval, agents or fan-out search, inspect query formulation and evidence selection as well as final-response text.

  • Separate useful framing from stylistic variation

    A persona is valuable only if its differences improve a pre-specified outcome such as prioritisation, task completion, retrieval relevance, factuality, risk handling or decision support.

  • Treat persona definitions as system controls

    The observed pattern is more consistent with framing controls than a simple voice layer. Teams should specify which aspects should remain invariant, which should change and how those changes will be tested.

Research agenda

What the next experiment should test

A true counterfactual study should move beyond post-hoc cue removal.

  1. Use matched prompt skeletons

    Hold the underlying task, topic, constraints and wording constant while varying only the persona treatment.

  2. Include generic and no-persona controls

    Generate a neutral condition alongside the intended persona condition for each task.

  3. Include swapped-persona controls

    Apply a deliberately mismatched persona to distinguish role-relevant steering from generic prompt elaboration.

  4. Replicate each condition

    Generate multiple responses per task-persona cell so within-condition variability can be separated from treatment differences.

  5. Evaluate retrieval outcomes directly

    Measure document relevance, source authority, diversity, coverage and whether retrieved evidence supports the eventual answer.

  6. Use blinded outcome assessment

    Domain experts should rate outputs without seeing the persona label, using outcomes such as factual accuracy, task completion, prioritisation, actionability, risk handling and citation support.

Only then can the analysis move from “persona-associated text remains different” to “persona conditioning improves a defined outcome”.

Method and reproducibility

Methodology and reproducibility

Sample

All 1,494 successful nonblank responses from 1,500 attempted prompts were analysed. The collection covers 12 finance personas, with 124 or 125 usable observations per persona.

Semantic representation

Semantic features were derived with deterministic TF-IDF and truncated latent semantic analysis. The primary measure is within-persona minus cross-persona cosine similarity.

Adjustment

Adjusted text features accounted for prompt template, prompt length, prompt semantic features and non-persona design factors rendered in the prompt. Query diagnostics additionally adjusted for query count.

Cue ablation

Exact-cue ablation removed exact persona names and exact profile phrases representing role, goal, constraint, risk tolerance, KPI and time horizon from saved outputs before recomputing the adjusted comparison.

Uncertainty

Confidence intervals used 1,999 paired bootstrap resamples of prompt templates, reflecting shared structure introduced by the prompt design.

Operational models

Grounding coverage used a binomial model. Source and query counts used negative-binomial models; the source-count model used a fixed dispersion parameter of one. Reported persona tests are omnibus tests unless noted.

Study boundaries

Limitations

  • Single synthetic collection

    The study evaluates one synthetic prompt collection and one saved generation per prompt.

  • Incomplete crossing

    Prompt templates are not fully crossed with every persona.

  • Missing controls

    There is no matched generic, swapped-persona or no-persona control.

  • Semantic proxy

    Local latent semantic analysis is reproducible but is not a validated finance-domain quality metric.

  • Post-generation ablation

    Exact cue removal is a sensitivity analysis, not a counterfactual intervention, and it cannot remove every semantic association related to a persona.

  • Residual design differences

    Residual prompt differences may remain after adjustment.

  • Count-model assumption

    The source-count model uses a fixed negative-binomial dispersion parameter of one.

  • Evidence-quality boundary

    Source presence and quantity do not establish authority, correctness or claim support.

  • No business outcomes

    The study contains no user, revenue, conversion, decision-quality, latency or cost outcomes.

A positive residual means only that exact persona/profile strings do not explain all observed same-persona semantic structure. It should not be relabelled as a causal or intrinsic persona effect.

FAQ

Frequently asked questions

Is the remaining AI persona signal just prompt echo?

Exact persona and profile wording explains part of the adjusted separation, but not all of it. After exact phrases were removed, positive response and query separation remained. Because the ablation happened after generation, that residual is not a proven cue-independent or causal persona effect.

How much adjusted response separation remained after explicit cue removal?

Approximately 76% of the adjusted response point estimate remained: separation fell from 0.0153 to 0.0117. This percentage is a descriptive share of the adjusted point estimate, not a causal effect or performance score.

Why were generated search queries more persona-sensitive than final answers?

The study suggests persona-associated framing may be more visible when the system formulates information needs than when it synthesises a final answer. It does not isolate the mechanism, and broader prompt wording, correlated topics and unmeasured design differences remain possible explanations.

Does this study show that persona prompting improves retrieval or grounding?

No. Different query language does not establish better retrieval. The analysis found no detectable persona-level difference in grounding presence or source quantity, and it did not assess source authority, factual correctness, citation support or business outcomes.

Piush Vaish, Founder of Kojable

Author

About the author

Piush VaishFounder and CEO of Kojable

By Piush Vaish, founder and CEO of Kojable.

Piush Vaish is the founder and CEO of Kojable, a repeat founder and data scientist with more than 10 years of experience building and productising AI, machine-learning and data products. He writes about AI search, AEO, GEO, agentic discovery and AI product strategy.

Read more about Piush Vaish