Kojable research · Persona conditioning · 750 grounded responses

Persona Prompts Change More Than Tone: Evidence from 750 Grounded AI Responses

Published By Piush Vaish

Across 150 matched business scenarios, five persona conditions produced persistent differences in grounded search formulation and final answers—even after direct persona language was removed and across five different task types.

Key finding

When the same underlying business problem was presented to Gemini under five different persona conditions, the resulting differences extended beyond explicit persona wording. The answers remained strongly distinguishable after direct persona language was removed, the model formulated substantially different search queries, and the differentiation remained high across diagnostic, measurement, comparison, action-planning and risk-control tasks.

Qualification: These results show persistent persona-associated differences in this historical grounded-AI run. They do not establish that persona instructions caused the differences independently of collection order, that different publishers were retrieved, that real people in these roles would behave similarly, or that more diverse answers are necessarily better answers.

  • Persona conditioning
  • Grounded AI
  • Search-query differentiation
  • 18 min read
150neutral matched scenarios
5persona conditions
749substantive responses
Exploratorymatched model-behaviour study Major limitation Persona condition was perfectly confounded with within-scenario collection position.

At a glance

Research snapshot

Study design, cohort and primary interpretation boundary.
Study elementDetail
Research questionHow persistent are persona-associated differences across grounded search formulation, final answers and different task types?
ModelGemini 3.6 Flash with Google Search grounding
Neutral scenarios150
Persona conditions5
Historical responses750
Substantive responses after quality screening749
Primary matched scenarios149
Task typesDiagnostic, measurement, comparison, action plan, risk control
Topic clustersSearch visibility, content authority, conversion/pipeline, AI optimisation, Answer Alignment
Primary statusExploratory matched model-behaviour study
Major limitationPersona condition was perfectly confounded with within-scenario collection position

Study overview

Executive summary

Persona prompting is often treated as a superficial prompt-engineering choice.

A researcher may ask an AI system to “act as a CMO,” “act as an SEO Manager” or “act as a Founder” and receive noticeably different responses. But that observation alone tells us very little.

The obvious alternative explanation is that the system simply repeats the language supplied in the persona prompt.

We designed a matched experiment to examine that problem more deeply.

The study began with 150 neutral business scenarios. Each scenario was rendered under five complete persona-conditioning bundles:

  • CMO
  • VP of Product Marketing
  • SEO Manager
  • Demand Generation Director
  • Founder/CEO

That produced 750 historical grounded responses from Gemini.

Rather than analysing the responses once and stopping, we subjected the experiment to four successive analytical tests.

  1. First

    We audited whether the historical run was sufficiently stable to analyse at all.

  2. Second

    We removed direct persona language from the generated answers to test whether apparent persona differences were mostly prompt echo.

  3. Third

    We inspected the search queries generated during grounded answering to determine whether persona-associated differentiation appeared before the final prose.

  4. Fourth

    We tested whether the differentiation was concentrated in certain tasks or remained broadly present across different forms of business work.

The result is a more nuanced picture than either “personas do not matter” or “personas behave like real executives.”

The five persona conditions were associated with persistent differences in both answer construction and search formulation. Those differences survived direct persona-language removal and appeared across all five task families.

But the study also defines important boundaries.

We cannot yet establish whether the personas produced materially different factual claims, recommendations or decisions. We cannot identify publisher-level source differences from the stored grounding data. And because persona order was fixed during the historical run, we cannot isolate persona from collection-position effects.

The strongest conclusion is therefore narrower:

Persona conditioning should be treated as a meaningful experimental variable in grounded-AI research, not merely as a change in tone.

Direct answer

Do persona prompts actually change AI answers?

Answer

In this historical matched Gemini experiment, persona-conditioned answers were strongly and persistently different.

Those differences remained after direct persona wording was removed and generalized across topic clusters, task families and time blocks.

The same differentiation was also visible in the search queries Gemini formulated during grounded answering.

Because persona and collection position were perfectly confounded, the study cannot establish that the persona instruction alone caused those differences.

Research question

Why we ran this study

Persona prompts are increasingly used in AI research, content analysis, customer research and decision support.

A common approach is to ask the same AI system to respond from different roles:

“Answer as a CMO.”

“Answer as an SEO Manager.”

“Answer as a Founder.”

The assumption is that these prompts expose different perspectives.

But there are at least three ways this could be misleading.

  • Prompt vocabulary

    First, the system could simply repeat vocabulary from the persona prompt.

  • Underlying process

    Second, the final answer could sound different while relying on essentially the same underlying information-gathering process.

  • Task sensitivity

    Third, persona differentiation could appear only for certain kinds of questions and disappear for others.

We wanted to separate these possibilities.

Instead of asking whether persona answers merely look different, the study progressively tested:

  1. Historical stability

    Whether the historical experiment was trustworthy enough to analyse.

  2. Direct persona language

    Whether differences remained after direct persona language was removed.

  3. Grounded search queries

    Whether differences were visible in the model's grounded search queries.

  4. Task families

    Whether those differences persisted across different kinds of tasks.

The resulting analysis turns persona prompting from a stylistic question into an experimental-design question.

For earlier Kojable work on role framing and persona effects, see How AI Role Prompts Change the Decision Lens, AI Persona Effects Beyond Prompt Echo, Persona Conditioning in AEO, and Persona Similarity in AI Responses.

Study design

Study design

The core experiment used 150 neutral scenarios.

Every neutral scenario was presented under all five persona-conditioning bundles, giving five answers to the same underlying business problem.

The personas represented different decision contexts rather than simply different job titles.

The CMO condition emphasized strategic priority, brand impact and resource trade-offs.

The VP of Product Marketing emphasized positioning, buyer understanding, launch impact and validation.

The SEO Manager emphasized query coverage, entity signals, discoverability and measurement.

The Demand Generation Director emphasized audience quality, funnel progression, conversion and pipeline.

The Founder/CEO emphasized category position, company risk, growth and resource concentration.

The 150 scenarios were also balanced across five task families:

Five task families used across the neutral scenarios.
Task familyWhat the scenario asked the model to do
DiagnosticIdentify causes, evidence and diagnostic sequences
MeasurementDefine baselines, measurement plans and validation
ComparisonCompare approaches, criteria and trade-offs
Action planRecommend actions, experiments and next steps
Risk controlIdentify risks, guardrails, controls and escalation

The scenarios spanned five topic clusters, giving six neutral scenarios in every cluster × task cell.

Gemini could use Google Search grounding while answering.

This matters because it allowed us to examine not only the final response, but also the search queries reported during grounded generation.

Finding 1

The historical run was usable—but only for exploratory claims

Before studying persona effects, we audited whether the underlying run was trustworthy.

Of the 750 historical responses, 749 passed the substantive-response quality screen.

One response—N093::SEO Manager—contained raw tool-call syntax rather than a substantive answer and was excluded from the new substantive cohort.

The important question was whether removing that response materially changed the existing matched results.

It did not.

A major matched effect-size statistic for the first historical lexical analysis changed from:

0.4705 → 0.4738

and the second from:

0.4534 → 0.4518.

No effect direction or corrected pairwise decision changed.

That means the main historical persona-separation patterns were not being produced by the malformed response.

We also examined the collection window for evidence of substantial drift.

The full run lasted approximately 16.6 hours. Retries and high-latency responses became somewhat more common later in the run, but the predefined within-cluster drift screens did not reach the level required to classify the experiment as unstable.

The final quality decision was therefore:

Suitable with limitations.

The most important limitation is structural.

The five personas were collected in the same order for every scenario:

CMO → VP of Product Marketing → SEO Manager → Demand Generation Director → Founder/CEO

Persona condition and within-scenario collection position are therefore perfectly confounded.

If the third answer in every block differs systematically from the first answer, the historical data cannot tell us whether that occurred because of the persona instruction, the collection position, or some combination of the two.

No regression can recover that missing counterfactual after the fact.

For that reason, this report uses language such as:

“persona-conditioned outputs were associated with…”

rather than:

“the persona instruction caused…”

The experiment supports exploratory matched comparisons, not causal persona claims.

Finding 2

The differences were not simply the model repeating the persona prompt

The strongest challenge to a persona experiment is prompt echo.

If we tell an AI system:

“As an SEO Manager responsible for organic and AI-search performance, focus on query coverage, entity signals and content discoverability…”

then an SEO-oriented answer is not surprising.

Indeed, the responses contained substantial direct persona-language carryover.

The exact persona title appeared in 63% of substantive responses.

Across the persona-conditioning fields, responses reused much of the supplied vocabulary:

Mean vocabulary coverage across persona-conditioning fields.
Persona fieldMean vocabulary coverage
Persona context74.3%
Persona focus99.7%
Persona deliverable95.0%

The model clearly followed the persona instructions.

The question was whether those injected words explained the observed separation.

To test this, we created progressively more aggressive response views.

The raw answers were compared with versions in which:

  1. Persona labels

    Explicit persona labels were masked.

  2. Persona profile

    Direct phrases from persona context, focus and deliverable instructions were masked.

  3. Residual representation

    Additional persona-exclusive initialisms were masked in a residual representation.

If persona separation were mainly prompt parroting, the answers should become substantially more similar after that language was removed.

They did not.

Mean cosine similarity across progressively masked response representations.
Response representationMean cosine similarity
Raw response0.14317
Persona-label masked0.14344
Persona-profile masked0.14636
Residual0.14552

Measured as 1 − cosine similarity, direct persona-language masking removed only a very small fraction of the observed separation.

The residual representation removed approximately 0.27% of the measured separation relative to the raw answers.

That does not mean the persona instructions ceased influencing the answer. Masking cannot remove implicit instruction-following, paraphrases or changes in structure.

But it does show that the observed differences were not primarily the consequence of copying a few persona-specific phrases.

Comparison of persona-language carryover and raw versus masked response similarity in the matched persona study.
Figure 1. Direct persona-language masking changed average cross-persona similarity only slightly. The figure should be described as a lexical sensitivity analysis, not evidence of a causal persona effect.
Open full-resolution figure

Could we still identify the persona after masking?

We tested this independently using text classification.

There were five persona classes, giving a nominal balanced chance level of 20%.

A classifier trained on residual response text achieved:

Residual macro-F1 across four validation designs.
Validation designResidual macro-F1
Scenario-grouped validation0.996
Leave-one-cluster-out0.933
Leave-response-family-out0.987
Time-block holdout0.939

The important result is not simply that the classifier worked.

It is that persona identity remained highly predictable even when:

  • The same neutral scenario could not leak between train and test.

  • An entire topic cluster was unseen during training.

  • An entire response family was held out.

  • Later or earlier portions of the historical run were separated.

Residual differentiation also appeared in interpretable output dimensions.

The strongest remaining differences included:

  • Governance framing
  • Execution markers
  • Commercial vocabulary
  • Owner/action markers
  • Technical vocabulary

For example, the SEO Manager condition remained the most technically oriented, Demand Generation remained the most commercially oriented, and CMO responses contained more governance-oriented language.

These are model-output associations, not claims about how real people in those professions think.

The core result is:

Persona-conditioned answers remained strongly distinguishable after direct persona-profile language was removed.

For related cue-ablation research on a separate finance-persona corpus, see AI Persona Effects Beyond Prompt Echo.

Persona classification macro-F1 across grouped and held-out validation designs after persona-language masking.
Figure 2. Persona identity remained highly predictable from residual response text under grouped, held-out-cluster, held-out-family and time-block validation.
Open full-resolution figure

Finding 3

Persona differentiation appeared before the final answer—in the search queries

If persona-associated differences existed only in the final wording, the grounded search process might still look essentially identical.

That is not what we observed.

Across the 749 substantive responses, Gemini reported 1,937 usable search-query entries.

There were no within-response duplicate queries.

Only nine normalized queries were shared across persona conditions within matched scenarios.

Literal query sets were therefore almost entirely different.

The exact-query divergence metric was:

0.9991

on a scale where 1 represents completely different normalized query sets.

That number is so close to the ceiling that it is not useful for measuring how much two searches differ.

We therefore also examined graded representations.

Mean divergence across exact, token and lexical query representations.
Query comparisonMean divergence
Exact normalized query sets0.9991
Token sets0.6822
TF-IDF lexical representation0.8870

The token and lexical measures show that the searches differed not only in exact wording, but also substantially in their vocabulary and formulation.

Heatmap or summary of grounded search-query divergence between persona pairs.
Figure 3. Matched persona conditions produced highly differentiated grounded search-query formulations. Exact divergence is near ceiling, so token and lexical measures are more informative about degree.
Open full-resolution figure

Persona signatures were visible in search behavior

The search-query language broadly echoed the output patterns observed in the final answers.

SEO Manager queries contained substantially more technical language.

Demand Generation queries contained substantially more commercial language.

CMO queries contained much more governance-oriented language.

For example:

Technical, commercial and governance density in persona-conditioned search queries.
Persona conditionTechnical query density / 1k wordsCommercial densityGovernance density
CMO9.0930.2835.90
Demand Generation Director6.2275.643.52
Founder/CEO9.5434.9810.02
SEO Manager74.9217.771.06
VP of Product Marketing8.8333.164.32

Again, these are historical persona-condition/position associations, not measurements of how real executives search.

Do more different searches correspond to more different answers?

Exact-query divergence was almost saturated, leaving too little variation to explain differences in answer distance.

Its within-scenario association with residual answer distance was essentially zero:

r = −0.012

But the graded query measures showed a different pattern.

Within matched scenarios:

  • Token divergence vs residual answer distance: r = 0.514
  • Lexical query divergence vs residual answer distance: r = 0.495

Across the ten persona-pair profiles, the descriptive rank correspondence was also high:

  • Token query divergence vs residual answer distance: Spearman 0.939
  • Lexical divergence vs residual answer distance: Spearman 0.818

These results are consistent with a relationship between how differently two persona conditions formulate searches and how differently their final answers are expressed.

They do not establish causal mediation.

Search and answer generation occurred inside the same grounded model call, and the historical experiment did not independently randomize search behavior.

The publication-safe conclusion is:

Persona-associated differentiation was visible in both search-query formulation and final answer construction.

For broader Kojable research on grounded query formulation, see How Predictable Are Gemini’s Fan-Out Queries?.

Comparison of persona-pair query divergence and residual answer divergence.
Figure 4. Persona pairs with more lexically differentiated queries also tended to have more differentiated residual answers. This is an exploratory association, not causal mediation.
Open full-resolution figure

Evidence boundary

What we could not establish about sources

The grounded run also stored 9,799 source records.

At first glance, that appears to offer an opportunity to compare which publishers different persona conditions relied on.

It did not.

Every stored source URI was an unresolved grounding-wrapper URL rather than a reliably attributable publisher URL.

We therefore predefined a source-quality gate.

Required reliable publisher coverage was 95%.

Observed reliable publisher coverage was:

0%.

The source-analysis gate failed.

As a result, this study does not report:

  • Publisher overlap
  • Source-domain divergence
  • Persona-specific publisher preference
  • Query-to-source associations
  • Source-to-answer associations

This is an important distinction.

The study establishes query differentiation.

It does not establish source-selection differentiation.

Failure to identify publisher differences should not be interpreted as evidence that the personas used the same sources.

The source question remains unresolved.

Finding 4

Persona differentiation persisted across all five task types

A persona effect that appears only in one narrow task would have limited general relevance.

We therefore examined whether answer differentiation changed materially across:

  • Action plans
  • Comparisons
  • Diagnostics
  • Measurement
  • Risk control

Using the persona-residual response representation from the previous analysis, mean answer diversity was:

Residual answer diversity across five task families.
TaskResidual answer diversity
Action plan0.8603
Comparison0.8463
Diagnostic0.8570
Measurement0.8478
Risk control0.8614

Risk-control scenarios showed the highest observed answer differentiation.

Comparison scenarios showed the lowest.

But the difference between the highest and lowest task means was only:

0.0150

or approximately:

1.76% of the overall level of task differentiation.

In other words, the difference between task types was small compared with the overall level of persona-associated separation.

That led to the predefined conclusion:

Broadly persistent across tasks.

Persona-conditioned answer differentiation was not confined to one type of work.

Residual answer and lexical query differentiation across action plan, comparison, diagnostic, measurement and risk-control tasks.
Figure 5. Residual answer differentiation remained high across all five task families; the total range between highest and lowest task means was small relative to overall separation.
Open full-resolution figure

The same broad pattern appeared in search queries

Lexical search-query divergence was:

Lexical search-query divergence across five task families.
TaskLexical query divergence
Action plan0.8829
Comparison0.8825
Diagnostic0.8928
Measurement0.8710
Risk control0.9053

Risk control again produced the highest observed query differentiation.

Measurement produced the lowest.

The range in search divergence was somewhat wider than the range in answer diversity, but differentiation remained high for all five task types.

Which persona was most distinctive?

SEO Manager had the highest mean answer distinctiveness for every task family.

CMO or Founder/CEO was consistently among the least distinctive.

The nearest persona pair in every task was:

CMO / Founder/CEO

The most distant pairs frequently involved SEO Manager.

That aligns with both the residual-answer analysis and the grounded-query analysis.

But “distinctive” should not be confused with “better.”

The study has not established that SEO Manager answers contain more correct information, better recommendations or greater business value.

It establishes that the outputs occupy a more distinct region of the observed response space.

Persona distinctiveness profiles across five task families.
Figure 6. SEO Manager was the most distinctive answer condition across all five task families, while CMO and Founder/CEO were consistently closer. Distinctiveness is not quality.
Open full-resolution figure

Panel-size analysis

How much diversity is lost when the persona panel gets smaller?

A practical implication follows from the matched five-persona design.

If the objective is simply to capture the range of answer representations generated under the five persona conditions, can we reduce the panel?

We tested every fixed:

  • Two-persona subset
  • Three-persona subset
  • Four-persona subset

For each panel, we measured how well the selected responses covered the five-persona response-space geometry.

Average normalized retention was approximately:

Approximate retention of the observed five-persona answer-space diversity.
Panel sizeApproximate retention of observed five-persona answer-space diversity
2 personas41%
3 personas61%
4 personas81%
5 personas100% reference panel

No three-persona panel reached the predefined 80% preservation threshold.

Four-persona panels exceeded 80%, but remained below the separate 95% near-full-coverage criterion.

The resulting panel decision was:

Five-persona panel materially additive.

This does not mean five personas are required for every AI research project.

Nor does it mean a three-persona panel preserves only 61% of useful business information.

The metric describes response-space geometry, not facts, strategic value or decision quality.

A persona that appears lexically redundant could still contribute a uniquely important recommendation.

That distinction motivated the next phase of the research.

Observed five-persona response-space coverage retained by smaller persona panels.
Figure 7. Two-, three- and four-persona panels retained approximately 41%, 61% and 81% of the observed five-persona response-space geometry. Coverage is representational, not substantive utility.
Open full-resolution figure

Evidence summary

What the study establishes

Taken together, the analyses support four main conclusions.

First, the historical matched run is sufficiently robust for exploratory analysis after removing one malformed response, although its fixed collection order prevents causal persona attribution.

Second, persona-associated answer differentiation survives removal of directly injected persona language.

Third, the differentiation is visible in grounded search-query formulation as well as in the final answers.

Fourth, it remains broadly persistent across five different kinds of business task.

These findings support the idea that persona choice should be treated as a meaningful experimental variable when designing grounded-AI research.

They do not establish that the persona conditions accurately simulate real executives.

Evidence boundaries

What the study does not establish

Several tempting conclusions go beyond the evidence.

It does not establish causal persona effects

Persona identity was perfectly aliased with within-scenario collection position.

A new randomized experiment is required to isolate persona from order.

It does not establish real-human role behavior

The experiment studies one AI system responding to prompt bundles.

It does not measure CMOs, founders, SEO managers or demand-generation leaders.

It does not establish different publisher selection

The source records could not be reliably resolved to publishers.

Search-query differences are measurable; publisher differences are not.

It does not establish better answers

The analyses through this stage measure representational differentiation.

They do not establish:

  • Factual correctness
  • Recommendation quality
  • Evidence quality
  • Usefulness
  • Business outcomes

It does not establish substantive decision divergence

Two responses can be lexically far apart while making the same recommendation.

Conversely, two responses can sound similar while recommending different actions.

That distinction is the subject of the next phase.

Substantive status

The unresolved question: does different mean substantively different?

After establishing persistent representational differences, we attempted to measure a harder layer:

Do the persona conditions make materially different factual claims, recommendations, decisions and risk assessments?

The first deterministic extraction pipeline failed its measurement-validity check.

It produced candidate text but accepted no substantive units under its original rules.

We therefore did not interpret that result as evidence that the persona answers were substantively equivalent.

Instead, we created a separate measurement-validation stage.

A frozen 50-response calibration set has now been prepared across all five personas and all five task types.

The next step requires:

  1. Two independent human reviewers
  2. Separate adjudication
  3. Validation of factual-claim extraction
  4. Validation of recommendation extraction
  5. Validation of risk extraction
  6. Validation of decision classification
  7. Validation of proposition-equivalence matching

Until that calibration is complete, the correct substantive conclusion is:

Not established.

This matters because it separates two questions that are often collapsed in AI evaluation:

Do outputs differ?

and:

Do those differences matter?

This study answers the first much more clearly than the second.

Interpretation

Implications for AI research

The practical implication is not that every company should immediately create five personas.

It is that persona choice should not be treated as decorative prompt wording.

If the same underlying scenario produces different grounded searches and persistently different answer structures under different persona conditions, then evaluating an AI system through only one persona can under-sample the range of responses the system is capable of producing.

That matters for research designs involving:

  • AI visibility
  • Answer Alignment
  • Buyer-question testing
  • AI search evaluation
  • Competitive research
  • Decision-support systems
  • Prompt-panel design

A single generic prompt may answer:

“What does the model say?”

A persona panel begins to answer:

“How does the answer space change when the decision context changes?”

That is a different research problem.

The results also suggest that persona-conditioned research should be evaluated at more than the final prose layer.

Search formulation itself changed substantially.

For grounded AI systems, that means evaluation can potentially examine:

prompt → query formulation → evidence retrieval → answer

rather than treating the answer as a black box.

This study could measure the first, second and fourth parts of that pathway.

Publisher-level evidence retrieval remains an unresolved layer.

Implications for AI Answer Alignment

For companies monitoring how AI systems represent them, the findings raise an important methodological question.

A company may appear correctly represented under one broad query but differently represented when the same underlying issue is approached from:

  • Strategic risk
  • Technical implementation
  • Revenue generation
  • Positioning
  • Executive prioritisation

If persona-conditioned contexts systematically alter both search formulation and answer construction, a single generic question may not fully represent the decision environments in which buyers or stakeholders encounter the company.

This does not mean research programmes should create arbitrary personas simply to generate more outputs.

It means persona design should be deliberate, tested and tied to distinct decision contexts.

The next question is whether those contexts produce genuinely different substantive conclusions.

That is where the current research programme is heading.

Method and reproducibility

Methodology

The study uses a matched design.

Each of 150 neutral scenarios was rendered under five persona-conditioning bundles.

The neutral scenario is the primary comparison block.

This avoids treating 750 responses as if they were 750 independent business questions.

The quality audit identified one malformed response and created a substantive cohort of 749 responses.

Analyses requiring all five personas used 149 complete scenarios and 745 responses.

Where statistical uncertainty was estimated, resampling preserved the neutral-scenario structure rather than independently resampling individual persona pairs.

Multiple pairwise tests used family-wise correction where specified.

The prompt-parroting analysis used a deterministic, uncapped TF-IDF representation for lexical similarity and trained classifiers only on training-fold data.

The retrieval analysis normalized recorded search queries offline.

No live search, source resolution or new model generation was used during the completed analytical layers.

The task analysis reused frozen answer and query distances from earlier stages rather than recomputing the representations.

All major analysis stages were designed to replay deterministically from the frozen historical response corpus.

Study boundaries

Limitations

The most important limitation is the fixed historical collection order.

Every scenario used the same persona sequence.

Persona and position are therefore inseparable in the historical data.

A randomized replicated experiment is required for causal claims.

There was also only one model response per persona/scenario cell.

The study therefore cannot estimate within-condition generation variability.

The five task families used different scenario templates, so differences between task families should be interpreted descriptively rather than as randomized task effects.

The source metadata did not contain resolvable publisher identities.

Source differentiation therefore remains unidentified.

Finally, the similarity and diversity metrics used here measure lexical and structural representation.

They do not establish semantic equivalence, factual correctness or business utility.

Next phase

The next experiment

There are two clear next steps.

The first is already underway:

Validate substantive divergence through independent human calibration.

That will test whether the five persona conditions contribute genuinely different:

  • Facts
  • Recommendations
  • Decisions
  • Risks

The second is experimental redesign.

A confirmatory collection should randomize persona order, record generation controls and collect repeated outputs per persona/scenario.

That would allow the research programme to distinguish:

  • Persona-condition effects
  • Generation variability
  • Collection-position effects
  • True within-persona repeatability

Only then should stronger causal or equivalence claims be considered.

FAQ

Frequently asked questions

Is this just the AI repeating the persona prompt?

The model did repeat substantial persona-related vocabulary, but removing direct persona labels and profile phrases eliminated only a very small fraction of the measured answer separation. Persona identity also remained highly predictable from the residual responses.

Did different personas make Gemini search differently?

Yes at the query level. Exact search-query overlap was extremely low, and graded token and lexical representations also showed substantial differentiation.

The study cannot establish that different publishers were retrieved because the stored source URLs could not be reliably resolved.

Which persona was most different?

SEO Manager was the most distinctive response condition across all five task families in the observed response-space analysis.

This does not mean SEO Manager produced the best or most useful answers.

Were some tasks more persona-sensitive than others?

Risk-control scenarios had the highest observed answer and query differentiation, but differences between task families were small relative to the overall level of persona separation.

The main conclusion was that differentiation was broadly persistent across tasks.

Do we need five personas?

Five personas preserved the full reference response space by definition. The best two-, three- and four-persona subsets retained approximately 41%, 61% and 81% respectively under the study's geometric coverage measure.

That does not yet show that five personas preserve more useful facts or decisions. Substantive coverage is still being validated.

Does this prove that real CMOs and SEO Managers think differently?

No.

The experiment measures AI outputs under persona-conditioning prompts. It does not observe real people in those roles and should not be interpreted as organizational psychology research.

Does more differentiation mean better answers?

No.

Different answers may still be equally correct, equally useful or substantively equivalent.

Measuring substantive factual, recommendation, decision and risk divergence is the next stage of the research.

Conclusion

Conclusion

Persona prompting is often discussed as if it were mainly a question of tone.

This experiment suggests a more consequential interpretation.

Across 150 matched scenarios, the five persona conditions were associated with persistent differences in the model's final answers.

Those differences survived direct persona-language removal.

They were also visible in the grounded search queries the model formulated.

And they remained broadly present across five different task families.

At the same time, the study draws a clear boundary around what has not yet been established.

We cannot isolate persona from historical collection position.

We cannot identify publisher-level source differences.

And we cannot yet say whether the observed response diversity corresponds to materially different facts, recommendations or decisions.

The strongest current conclusion is therefore not:

Personas make AI answers better.

It is:

Persona conditioning changes the observed answer space enough that it should be treated as an experimental design variable—not merely as a stylistic prompt choice.

The next phase is to determine whether that representational diversity translates into substantive decision value.

Piush Vaish, Founder of Kojable

Author

About the author

Piush VaishFounder and CEO of Kojable

By Piush Vaish, founder and CEO of Kojable.

Piush Vaish is the founder and CEO of Kojable, a repeat founder and data scientist with more than 10 years of experience building and productising AI, machine-learning and data products. He writes about AI search, AEO, GEO, agentic discovery and AI product strategy.

Read more about Piush Vaish