Kojable research · Persona conditioning · 750 grounded responses
Persona Prompts Change More Than Tone: Evidence from 750 Grounded AI Responses
Across 150 matched business scenarios, five persona conditions produced persistent differences in grounded search formulation and final answers—even after direct persona language was removed and across five different task types.
Key finding
When the same underlying business problem was presented to Gemini under five different persona conditions, the resulting differences extended beyond explicit persona wording. The answers remained strongly distinguishable after direct persona language was removed, the model formulated substantially different search queries, and the differentiation remained high across diagnostic, measurement, comparison, action-planning and risk-control tasks.
Qualification: These results show persistent persona-associated differences in this historical grounded-AI run. They do not establish that persona instructions caused the differences independently of collection order, that different publishers were retrieved, that real people in these roles would behave similarly, or that more diverse answers are necessarily better answers.
- Persona conditioning
- Grounded AI
- Search-query differentiation
- 18 min read
At a glance
Research snapshot
| Study element | Detail |
|---|---|
| Research question | How persistent are persona-associated differences across grounded search formulation, final answers and different task types? |
| Model | Gemini 3.6 Flash with Google Search grounding |
| Neutral scenarios | 150 |
| Persona conditions | 5 |
| Historical responses | 750 |
| Substantive responses after quality screening | 749 |
| Primary matched scenarios | 149 |
| Task types | Diagnostic, measurement, comparison, action plan, risk control |
| Topic clusters | Search visibility, content authority, conversion/pipeline, AI optimisation, Answer Alignment |
| Primary status | Exploratory matched model-behaviour study |
| Major limitation | Persona condition was perfectly confounded with within-scenario collection position |
Study overview
Executive summary
Persona prompting is often treated as a superficial prompt-engineering choice.
A researcher may ask an AI system to “act as a CMO,” “act as an SEO Manager” or “act as a Founder” and receive noticeably different responses. But that observation alone tells us very little.
The obvious alternative explanation is that the system simply repeats the language supplied in the persona prompt.
We designed a matched experiment to examine that problem more deeply.
The study began with 150 neutral business scenarios. Each scenario was rendered under five complete persona-conditioning bundles:
- CMO
- VP of Product Marketing
- SEO Manager
- Demand Generation Director
- Founder/CEO
That produced 750 historical grounded responses from Gemini.
Rather than analysing the responses once and stopping, we subjected the experiment to four successive analytical tests.
- First
We audited whether the historical run was sufficiently stable to analyse at all.
- Second
We removed direct persona language from the generated answers to test whether apparent persona differences were mostly prompt echo.
- Third
We inspected the search queries generated during grounded answering to determine whether persona-associated differentiation appeared before the final prose.
- Fourth
We tested whether the differentiation was concentrated in certain tasks or remained broadly present across different forms of business work.
The result is a more nuanced picture than either “personas do not matter” or “personas behave like real executives.”
The five persona conditions were associated with persistent differences in both answer construction and search formulation. Those differences survived direct persona-language removal and appeared across all five task families.
But the study also defines important boundaries.
We cannot yet establish whether the personas produced materially different factual claims, recommendations or decisions. We cannot identify publisher-level source differences from the stored grounding data. And because persona order was fixed during the historical run, we cannot isolate persona from collection-position effects.
The strongest conclusion is therefore narrower:
Persona conditioning should be treated as a meaningful experimental variable in grounded-AI research, not merely as a change in tone.
Direct answer
Do persona prompts actually change AI answers?
Answer
In this historical matched Gemini experiment, persona-conditioned answers were strongly and persistently different.
Those differences remained after direct persona wording was removed and generalized across topic clusters, task families and time blocks.
The same differentiation was also visible in the search queries Gemini formulated during grounded answering.
Because persona and collection position were perfectly confounded, the study cannot establish that the persona instruction alone caused those differences.
Research question
Why we ran this study
Persona prompts are increasingly used in AI research, content analysis, customer research and decision support.
A common approach is to ask the same AI system to respond from different roles:
“Answer as a CMO.”
“Answer as an SEO Manager.”
“Answer as a Founder.”
The assumption is that these prompts expose different perspectives.
But there are at least three ways this could be misleading.
- Prompt vocabulary
First, the system could simply repeat vocabulary from the persona prompt.
- Underlying process
Second, the final answer could sound different while relying on essentially the same underlying information-gathering process.
- Task sensitivity
Third, persona differentiation could appear only for certain kinds of questions and disappear for others.
We wanted to separate these possibilities.
Instead of asking whether persona answers merely look different, the study progressively tested:
- Historical stability
Whether the historical experiment was trustworthy enough to analyse.
- Direct persona language
Whether differences remained after direct persona language was removed.
- Grounded search queries
Whether differences were visible in the model's grounded search queries.
- Task families
Whether those differences persisted across different kinds of tasks.
The resulting analysis turns persona prompting from a stylistic question into an experimental-design question.
For earlier Kojable work on role framing and persona effects, see How AI Role Prompts Change the Decision Lens, AI Persona Effects Beyond Prompt Echo, Persona Conditioning in AEO, and Persona Similarity in AI Responses.
Study design
Study design
The core experiment used 150 neutral scenarios.
Every neutral scenario was presented under all five persona-conditioning bundles, giving five answers to the same underlying business problem.
The personas represented different decision contexts rather than simply different job titles.
The CMO condition emphasized strategic priority, brand impact and resource trade-offs.
The VP of Product Marketing emphasized positioning, buyer understanding, launch impact and validation.
The SEO Manager emphasized query coverage, entity signals, discoverability and measurement.
The Demand Generation Director emphasized audience quality, funnel progression, conversion and pipeline.
The Founder/CEO emphasized category position, company risk, growth and resource concentration.
The 150 scenarios were also balanced across five task families:
| Task family | What the scenario asked the model to do |
|---|---|
| Diagnostic | Identify causes, evidence and diagnostic sequences |
| Measurement | Define baselines, measurement plans and validation |
| Comparison | Compare approaches, criteria and trade-offs |
| Action plan | Recommend actions, experiments and next steps |
| Risk control | Identify risks, guardrails, controls and escalation |
The scenarios spanned five topic clusters, giving six neutral scenarios in every cluster × task cell.
Gemini could use Google Search grounding while answering.
This matters because it allowed us to examine not only the final response, but also the search queries reported during grounded generation.
Finding 1
The historical run was usable—but only for exploratory claims
Before studying persona effects, we audited whether the underlying run was trustworthy.
Of the 750 historical responses, 749 passed the substantive-response quality screen.
One response—N093::SEO Manager—contained raw tool-call syntax rather than a substantive answer and was excluded from the new substantive cohort.
The important question was whether removing that response materially changed the existing matched results.
It did not.
A major matched effect-size statistic for the first historical lexical analysis changed from:
0.4705 → 0.4738
and the second from:
0.4534 → 0.4518.
No effect direction or corrected pairwise decision changed.
That means the main historical persona-separation patterns were not being produced by the malformed response.
We also examined the collection window for evidence of substantial drift.
The full run lasted approximately 16.6 hours. Retries and high-latency responses became somewhat more common later in the run, but the predefined within-cluster drift screens did not reach the level required to classify the experiment as unstable.
The final quality decision was therefore:
Suitable with limitations.
The most important limitation is structural.
The five personas were collected in the same order for every scenario:
CMO → VP of Product Marketing → SEO Manager → Demand Generation Director → Founder/CEO
Persona condition and within-scenario collection position are therefore perfectly confounded.
If the third answer in every block differs systematically from the first answer, the historical data cannot tell us whether that occurred because of the persona instruction, the collection position, or some combination of the two.
No regression can recover that missing counterfactual after the fact.
For that reason, this report uses language such as:
“persona-conditioned outputs were associated with…”
rather than:
“the persona instruction caused…”
The experiment supports exploratory matched comparisons, not causal persona claims.
Finding 2
The differences were not simply the model repeating the persona prompt
The strongest challenge to a persona experiment is prompt echo.
If we tell an AI system:
“As an SEO Manager responsible for organic and AI-search performance, focus on query coverage, entity signals and content discoverability…”
then an SEO-oriented answer is not surprising.
Indeed, the responses contained substantial direct persona-language carryover.
The exact persona title appeared in 63% of substantive responses.
Across the persona-conditioning fields, responses reused much of the supplied vocabulary:
| Persona field | Mean vocabulary coverage |
|---|---|
| Persona context | 74.3% |
| Persona focus | 99.7% |
| Persona deliverable | 95.0% |
The model clearly followed the persona instructions.
The question was whether those injected words explained the observed separation.
To test this, we created progressively more aggressive response views.
The raw answers were compared with versions in which:
- Persona labels
Explicit persona labels were masked.
- Persona profile
Direct phrases from persona context, focus and deliverable instructions were masked.
- Residual representation
Additional persona-exclusive initialisms were masked in a residual representation.
If persona separation were mainly prompt parroting, the answers should become substantially more similar after that language was removed.
They did not.
| Response representation | Mean cosine similarity |
|---|---|
| Raw response | 0.14317 |
| Persona-label masked | 0.14344 |
| Persona-profile masked | 0.14636 |
| Residual | 0.14552 |
Measured as 1 − cosine similarity, direct persona-language masking removed only a very small fraction of the observed separation.
The residual representation removed approximately 0.27% of the measured separation relative to the raw answers.
That does not mean the persona instructions ceased influencing the answer. Masking cannot remove implicit instruction-following, paraphrases or changes in structure.
But it does show that the observed differences were not primarily the consequence of copying a few persona-specific phrases.
Could we still identify the persona after masking?
We tested this independently using text classification.
There were five persona classes, giving a nominal balanced chance level of 20%.
A classifier trained on residual response text achieved:
| Validation design | Residual macro-F1 |
|---|---|
| Scenario-grouped validation | 0.996 |
| Leave-one-cluster-out | 0.933 |
| Leave-response-family-out | 0.987 |
| Time-block holdout | 0.939 |
The important result is not simply that the classifier worked.
It is that persona identity remained highly predictable even when:
The same neutral scenario could not leak between train and test.
An entire topic cluster was unseen during training.
An entire response family was held out.
Later or earlier portions of the historical run were separated.
Residual differentiation also appeared in interpretable output dimensions.
The strongest remaining differences included:
- Governance framing
- Execution markers
- Commercial vocabulary
- Owner/action markers
- Technical vocabulary
For example, the SEO Manager condition remained the most technically oriented, Demand Generation remained the most commercially oriented, and CMO responses contained more governance-oriented language.
These are model-output associations, not claims about how real people in those professions think.
The core result is:
Persona-conditioned answers remained strongly distinguishable after direct persona-profile language was removed.
For related cue-ablation research on a separate finance-persona corpus, see AI Persona Effects Beyond Prompt Echo.
Finding 3
Persona differentiation appeared before the final answer—in the search queries
If persona-associated differences existed only in the final wording, the grounded search process might still look essentially identical.
That is not what we observed.
Across the 749 substantive responses, Gemini reported 1,937 usable search-query entries.
There were no within-response duplicate queries.
Only nine normalized queries were shared across persona conditions within matched scenarios.
Literal query sets were therefore almost entirely different.
The exact-query divergence metric was:
0.9991
on a scale where 1 represents completely different normalized query sets.
That number is so close to the ceiling that it is not useful for measuring how much two searches differ.
We therefore also examined graded representations.
| Query comparison | Mean divergence |
|---|---|
| Exact normalized query sets | 0.9991 |
| Token sets | 0.6822 |
| TF-IDF lexical representation | 0.8870 |
The token and lexical measures show that the searches differed not only in exact wording, but also substantially in their vocabulary and formulation.
Persona signatures were visible in search behavior
The search-query language broadly echoed the output patterns observed in the final answers.
SEO Manager queries contained substantially more technical language.
Demand Generation queries contained substantially more commercial language.
CMO queries contained much more governance-oriented language.
For example:
| Persona condition | Technical query density / 1k words | Commercial density | Governance density |
|---|---|---|---|
| CMO | 9.09 | 30.28 | 35.90 |
| Demand Generation Director | 6.22 | 75.64 | 3.52 |
| Founder/CEO | 9.54 | 34.98 | 10.02 |
| SEO Manager | 74.92 | 17.77 | 1.06 |
| VP of Product Marketing | 8.83 | 33.16 | 4.32 |
Again, these are historical persona-condition/position associations, not measurements of how real executives search.
Do more different searches correspond to more different answers?
Exact-query divergence was almost saturated, leaving too little variation to explain differences in answer distance.
Its within-scenario association with residual answer distance was essentially zero:
r = −0.012
But the graded query measures showed a different pattern.
Within matched scenarios:
- Token divergence vs residual answer distance: r = 0.514
- Lexical query divergence vs residual answer distance: r = 0.495
Across the ten persona-pair profiles, the descriptive rank correspondence was also high:
- Token query divergence vs residual answer distance: Spearman 0.939
- Lexical divergence vs residual answer distance: Spearman 0.818
These results are consistent with a relationship between how differently two persona conditions formulate searches and how differently their final answers are expressed.
They do not establish causal mediation.
Search and answer generation occurred inside the same grounded model call, and the historical experiment did not independently randomize search behavior.
The publication-safe conclusion is:
Persona-associated differentiation was visible in both search-query formulation and final answer construction.
For broader Kojable research on grounded query formulation, see How Predictable Are Gemini’s Fan-Out Queries?.
Evidence boundary
What we could not establish about sources
The grounded run also stored 9,799 source records.
At first glance, that appears to offer an opportunity to compare which publishers different persona conditions relied on.
It did not.
Every stored source URI was an unresolved grounding-wrapper URL rather than a reliably attributable publisher URL.
We therefore predefined a source-quality gate.
Required reliable publisher coverage was 95%.
Observed reliable publisher coverage was:
0%.
The source-analysis gate failed.
As a result, this study does not report:
- Publisher overlap
- Source-domain divergence
- Persona-specific publisher preference
- Query-to-source associations
- Source-to-answer associations
This is an important distinction.
The study establishes query differentiation.
It does not establish source-selection differentiation.
Failure to identify publisher differences should not be interpreted as evidence that the personas used the same sources.
The source question remains unresolved.
Finding 4
Persona differentiation persisted across all five task types
A persona effect that appears only in one narrow task would have limited general relevance.
We therefore examined whether answer differentiation changed materially across:
- Action plans
- Comparisons
- Diagnostics
- Measurement
- Risk control
Using the persona-residual response representation from the previous analysis, mean answer diversity was:
| Task | Residual answer diversity |
|---|---|
| Action plan | 0.8603 |
| Comparison | 0.8463 |
| Diagnostic | 0.8570 |
| Measurement | 0.8478 |
| Risk control | 0.8614 |
Risk-control scenarios showed the highest observed answer differentiation.
Comparison scenarios showed the lowest.
But the difference between the highest and lowest task means was only:
0.0150
or approximately:
1.76% of the overall level of task differentiation.
In other words, the difference between task types was small compared with the overall level of persona-associated separation.
That led to the predefined conclusion:
Broadly persistent across tasks.
Persona-conditioned answer differentiation was not confined to one type of work.
The same broad pattern appeared in search queries
Lexical search-query divergence was:
| Task | Lexical query divergence |
|---|---|
| Action plan | 0.8829 |
| Comparison | 0.8825 |
| Diagnostic | 0.8928 |
| Measurement | 0.8710 |
| Risk control | 0.9053 |
Risk control again produced the highest observed query differentiation.
Measurement produced the lowest.
The range in search divergence was somewhat wider than the range in answer diversity, but differentiation remained high for all five task types.
Which persona was most distinctive?
SEO Manager had the highest mean answer distinctiveness for every task family.
CMO or Founder/CEO was consistently among the least distinctive.
The nearest persona pair in every task was:
CMO / Founder/CEO
The most distant pairs frequently involved SEO Manager.
That aligns with both the residual-answer analysis and the grounded-query analysis.
But “distinctive” should not be confused with “better.”
The study has not established that SEO Manager answers contain more correct information, better recommendations or greater business value.
It establishes that the outputs occupy a more distinct region of the observed response space.
Panel-size analysis
How much diversity is lost when the persona panel gets smaller?
A practical implication follows from the matched five-persona design.
If the objective is simply to capture the range of answer representations generated under the five persona conditions, can we reduce the panel?
We tested every fixed:
- Two-persona subset
- Three-persona subset
- Four-persona subset
For each panel, we measured how well the selected responses covered the five-persona response-space geometry.
Average normalized retention was approximately:
| Panel size | Approximate retention of observed five-persona answer-space diversity |
|---|---|
| 2 personas | 41% |
| 3 personas | 61% |
| 4 personas | 81% |
| 5 personas | 100% reference panel |
No three-persona panel reached the predefined 80% preservation threshold.
Four-persona panels exceeded 80%, but remained below the separate 95% near-full-coverage criterion.
The resulting panel decision was:
Five-persona panel materially additive.
This does not mean five personas are required for every AI research project.
Nor does it mean a three-persona panel preserves only 61% of useful business information.
The metric describes response-space geometry, not facts, strategic value or decision quality.
A persona that appears lexically redundant could still contribute a uniquely important recommendation.
That distinction motivated the next phase of the research.
Evidence summary
What the study establishes
Taken together, the analyses support four main conclusions.
First, the historical matched run is sufficiently robust for exploratory analysis after removing one malformed response, although its fixed collection order prevents causal persona attribution.
Second, persona-associated answer differentiation survives removal of directly injected persona language.
Third, the differentiation is visible in grounded search-query formulation as well as in the final answers.
Fourth, it remains broadly persistent across five different kinds of business task.
These findings support the idea that persona choice should be treated as a meaningful experimental variable when designing grounded-AI research.
They do not establish that the persona conditions accurately simulate real executives.
Evidence boundaries
What the study does not establish
Several tempting conclusions go beyond the evidence.
It does not establish causal persona effects
Persona identity was perfectly aliased with within-scenario collection position.
A new randomized experiment is required to isolate persona from order.
It does not establish real-human role behavior
The experiment studies one AI system responding to prompt bundles.
It does not measure CMOs, founders, SEO managers or demand-generation leaders.
It does not establish different publisher selection
The source records could not be reliably resolved to publishers.
Search-query differences are measurable; publisher differences are not.
It does not establish better answers
The analyses through this stage measure representational differentiation.
They do not establish:
- Factual correctness
- Recommendation quality
- Evidence quality
- Usefulness
- Business outcomes
It does not establish substantive decision divergence
Two responses can be lexically far apart while making the same recommendation.
Conversely, two responses can sound similar while recommending different actions.
That distinction is the subject of the next phase.
Substantive status
The unresolved question: does different mean substantively different?
After establishing persistent representational differences, we attempted to measure a harder layer:
Do the persona conditions make materially different factual claims, recommendations, decisions and risk assessments?
The first deterministic extraction pipeline failed its measurement-validity check.
It produced candidate text but accepted no substantive units under its original rules.
We therefore did not interpret that result as evidence that the persona answers were substantively equivalent.
Instead, we created a separate measurement-validation stage.
A frozen 50-response calibration set has now been prepared across all five personas and all five task types.
The next step requires:
- Two independent human reviewers
- Separate adjudication
- Validation of factual-claim extraction
- Validation of recommendation extraction
- Validation of risk extraction
- Validation of decision classification
- Validation of proposition-equivalence matching
Until that calibration is complete, the correct substantive conclusion is:
Not established.
This matters because it separates two questions that are often collapsed in AI evaluation:
Do outputs differ?
and:
Do those differences matter?
This study answers the first much more clearly than the second.
Interpretation
Implications for AI research
The practical implication is not that every company should immediately create five personas.
It is that persona choice should not be treated as decorative prompt wording.
If the same underlying scenario produces different grounded searches and persistently different answer structures under different persona conditions, then evaluating an AI system through only one persona can under-sample the range of responses the system is capable of producing.
That matters for research designs involving:
- AI visibility
- Answer Alignment
- Buyer-question testing
- AI search evaluation
- Competitive research
- Decision-support systems
- Prompt-panel design
A single generic prompt may answer:
“What does the model say?”
A persona panel begins to answer:
“How does the answer space change when the decision context changes?”
That is a different research problem.
The results also suggest that persona-conditioned research should be evaluated at more than the final prose layer.
Search formulation itself changed substantially.
For grounded AI systems, that means evaluation can potentially examine:
prompt → query formulation → evidence retrieval → answer
rather than treating the answer as a black box.
This study could measure the first, second and fourth parts of that pathway.
Publisher-level evidence retrieval remains an unresolved layer.
Implications for AI Answer Alignment
For companies monitoring how AI systems represent them, the findings raise an important methodological question.
A company may appear correctly represented under one broad query but differently represented when the same underlying issue is approached from:
- Strategic risk
- Technical implementation
- Revenue generation
- Positioning
- Executive prioritisation
If persona-conditioned contexts systematically alter both search formulation and answer construction, a single generic question may not fully represent the decision environments in which buyers or stakeholders encounter the company.
This does not mean research programmes should create arbitrary personas simply to generate more outputs.
It means persona design should be deliberate, tested and tied to distinct decision contexts.
The next question is whether those contexts produce genuinely different substantive conclusions.
That is where the current research programme is heading.
Method and reproducibility
Methodology
The study uses a matched design.
Each of 150 neutral scenarios was rendered under five persona-conditioning bundles.
The neutral scenario is the primary comparison block.
This avoids treating 750 responses as if they were 750 independent business questions.
The quality audit identified one malformed response and created a substantive cohort of 749 responses.
Analyses requiring all five personas used 149 complete scenarios and 745 responses.
Where statistical uncertainty was estimated, resampling preserved the neutral-scenario structure rather than independently resampling individual persona pairs.
Multiple pairwise tests used family-wise correction where specified.
The prompt-parroting analysis used a deterministic, uncapped TF-IDF representation for lexical similarity and trained classifiers only on training-fold data.
The retrieval analysis normalized recorded search queries offline.
No live search, source resolution or new model generation was used during the completed analytical layers.
The task analysis reused frozen answer and query distances from earlier stages rather than recomputing the representations.
All major analysis stages were designed to replay deterministically from the frozen historical response corpus.
Study boundaries
Limitations
The most important limitation is the fixed historical collection order.
Every scenario used the same persona sequence.
Persona and position are therefore inseparable in the historical data.
A randomized replicated experiment is required for causal claims.
There was also only one model response per persona/scenario cell.
The study therefore cannot estimate within-condition generation variability.
The five task families used different scenario templates, so differences between task families should be interpreted descriptively rather than as randomized task effects.
The source metadata did not contain resolvable publisher identities.
Source differentiation therefore remains unidentified.
Finally, the similarity and diversity metrics used here measure lexical and structural representation.
They do not establish semantic equivalence, factual correctness or business utility.
Next phase
The next experiment
There are two clear next steps.
The first is already underway:
Validate substantive divergence through independent human calibration.
That will test whether the five persona conditions contribute genuinely different:
- Facts
- Recommendations
- Decisions
- Risks
The second is experimental redesign.
A confirmatory collection should randomize persona order, record generation controls and collect repeated outputs per persona/scenario.
That would allow the research programme to distinguish:
- Persona-condition effects
- Generation variability
- Collection-position effects
- True within-persona repeatability
Only then should stronger causal or equivalence claims be considered.
FAQ
Frequently asked questions
Is this just the AI repeating the persona prompt?
The model did repeat substantial persona-related vocabulary, but removing direct persona labels and profile phrases eliminated only a very small fraction of the measured answer separation. Persona identity also remained highly predictable from the residual responses.
Did different personas make Gemini search differently?
Yes at the query level. Exact search-query overlap was extremely low, and graded token and lexical representations also showed substantial differentiation.
The study cannot establish that different publishers were retrieved because the stored source URLs could not be reliably resolved.
Which persona was most different?
SEO Manager was the most distinctive response condition across all five task families in the observed response-space analysis.
This does not mean SEO Manager produced the best or most useful answers.
Were some tasks more persona-sensitive than others?
Risk-control scenarios had the highest observed answer and query differentiation, but differences between task families were small relative to the overall level of persona separation.
The main conclusion was that differentiation was broadly persistent across tasks.
Do we need five personas?
Five personas preserved the full reference response space by definition. The best two-, three- and four-persona subsets retained approximately 41%, 61% and 81% respectively under the study's geometric coverage measure.
That does not yet show that five personas preserve more useful facts or decisions. Substantive coverage is still being validated.
Does this prove that real CMOs and SEO Managers think differently?
No.
The experiment measures AI outputs under persona-conditioning prompts. It does not observe real people in those roles and should not be interpreted as organizational psychology research.
Does more differentiation mean better answers?
No.
Different answers may still be equally correct, equally useful or substantively equivalent.
Measuring substantive factual, recommendation, decision and risk divergence is the next stage of the research.
Conclusion
Conclusion
Persona prompting is often discussed as if it were mainly a question of tone.
This experiment suggests a more consequential interpretation.
Across 150 matched scenarios, the five persona conditions were associated with persistent differences in the model's final answers.
Those differences survived direct persona-language removal.
They were also visible in the grounded search queries the model formulated.
And they remained broadly present across five different task families.
At the same time, the study draws a clear boundary around what has not yet been established.
We cannot isolate persona from historical collection position.
We cannot identify publisher-level source differences.
And we cannot yet say whether the observed response diversity corresponds to materially different facts, recommendations or decisions.
The strongest current conclusion is therefore not:
Personas make AI answers better.
It is:
Persona conditioning changes the observed answer space enough that it should be treated as an experimental design variable—not merely as a stylistic prompt choice.
The next phase is to determine whether that representational diversity translates into substantive decision value.






