Kojable research · Persona conditioning · Task persistence
Do AI Persona Differences Depend on the Task? Evidence Across Five Business Task Types
Across five business task types, persona-associated answer differentiation stayed consistently high. The range between the highest- and lowest-diversity task means was only about 1.76% of the overall level.
Key finding
Persona-associated answer differentiation remained high across all five task families in this matched grounded-AI experiment.
Mean residual answer diversity ranged from 0.846 to 0.861 across action planning, comparison, diagnosis, measurement and risk control.
The full range between the highest- and lowest-diversity task means was only 0.015, or about 1.76% of the overall level of task differentiation.
The study therefore classified persona differentiation as:
Broadly persistent across tasks.
Qualification: The task families used different scenario content and templates, and persona condition was historically confounded with collection position. These results describe task-associated patterns in one frozen model run; they do not establish causal task effects or real-role behavior.
- Persona conditioning
- Five task families
- 149 complete scenarios
- 11 min read
At a glance
Research snapshot
| Study element | Detail |
|---|---|
| Research question | Are persona-conditioned AI answers more different for some business tasks than others? |
| Parent study | 150 matched scenarios × 5 persona conditions |
| Primary complete scenarios | 149 |
| Task families | Action plan, comparison, diagnostic, measurement, risk control |
| Personas | CMO, VP Product Marketing, SEO Manager, Demand Generation Director, Founder/CEO |
| Primary answer metric | Residual answer distance |
| Lowest task mean | Comparison: 0.8463 |
| Highest task mean | Risk control: 0.8614 |
| Range | 0.0150 |
| Relative range | 1.76% |
| Task decision | Broadly persistent across tasks |
| Major limitation | Task, scenario content and template were not independently randomized |
Direct answer
Do AI persona differences depend strongly on the task?
Answer
Not in this experiment.
Persona-associated answer differentiation was high across all five task families.
Risk-control scenarios had the highest observed residual answer diversity at 0.8614. Comparison scenarios had the lowest at 0.8463.
But the gap between them was small compared with the overall level of separation.
The study's predefined conclusion was therefore:
Persona differentiation was broadly persistent across task types rather than strongly task-specific.
That does not mean task framing had no effect. It means the overall level of persona-associated separation remained high regardless of whether the model was diagnosing a problem, defining measurement, comparing options, planning actions or controlling risk.
Research question
Why this question matters
Persona prompting is often used as if the role itself is the main source of variation.
But tasks differ too.
A model asked to diagnose a problem may naturally produce a different kind of answer from one asked to compare options or define a risk-control plan.
That creates an important design question.
If persona differentiation appears only in certain task types, researchers may be overgeneralizing from narrow prompt formats.
For example:
- Personas might matter for strategic comparison but not measurement
- They might matter for risk control but disappear in diagnosis
- Task structure might overwhelm persona framing entirely
Our broader V2 study had already found that persona-conditioned outputs remained strongly distinguishable after direct persona language was removed.
This analysis asks:
Does that differentiation hold across different kinds of work?
Task design
The five task families
The experiment used five response types.
1. Action plan
The model was asked to recommend concrete steps, experiments or next actions.
2. Comparison
The model compared approaches, options, trade-offs or criteria.
3. Diagnostic
The model identified likely causes, evidence and diagnostic sequences.
4. Measurement
The model defined baselines, metrics, validation plans or monitoring approaches.
5. Risk control
The model identified risks, guardrails, escalation rules or control mechanisms.
The original design contained 30 scenarios in each task family.
After one malformed response was excluded by the quality gate, the primary analysis retained:
- 30 action-plan scenarios
- 30 comparison scenarios
- 30 diagnostic scenarios
- 29 measurement scenarios
- 30 risk-control scenarios
Each complete scenario contained all five persona conditions.
Finding 1
Answer differentiation stayed high across every task family
Using the residual answer representation from the persona cue-masking analysis, mean persona-pair answer diversity was:
| Task | Residual answer diversity | 95% CI |
|---|---|---|
| Action plan | 0.8603 | 0.8558–0.8650 |
| Comparison | 0.8463 | 0.8379–0.8541 |
| Diagnostic | 0.8570 | 0.8498–0.8640 |
| Measurement | 0.8478 | 0.8423–0.8529 |
| Risk control | 0.8614 | 0.8548–0.8680 |
Every task mean remained above 0.84.
That matters because persona differentiation did not collapse in the more structured task types.
Measurement still showed residual answer diversity of 0.8478. Comparison, the lowest task, still measured 0.8463.
The practical result is not:
“Personas matter for one special kind of question.”
It is:
The observed persona-associated differentiation remained high across all five task families.
Finding 2
Risk control was highest, comparison lowest—but the spread was small
Risk control had the highest observed answer diversity: 0.8614.
Comparison had the lowest: 0.8463.
Difference: 0.0150.
That difference is about 1.76% of the overall level of task differentiation.
The study prospectively defined a larger task-specificity signal as requiring an answer-diversity range of at least 0.03, together with corrected contrast evidence, cross-cluster direction stability and supporting interaction evidence.
The observed range did not reach that threshold.
The primary task conclusion was therefore:
Broadly persistent across tasks
rather than:
Task-conditioned differentiation
Risk-control prompts did produce somewhat more differentiated answers, but not enough to support the claim that persona conditioning matters dramatically more when the task is about risk.
Finding 3
Grounded search-query differentiation was also high across all five tasks
The task analysis did not examine final answers alone.
It reused the lexical search-query divergence from the grounded-query study.
Mean lexical query divergence was:
| Task | Lexical query divergence | 95% CI |
|---|---|---|
| Action plan | 0.8829 | 0.8701–0.8940 |
| Comparison | 0.8825 | 0.8696–0.8949 |
| Diagnostic | 0.8928 | 0.8786–0.9058 |
| Measurement | 0.8710 | 0.8608–0.8813 |
| Risk control | 0.9053 | 0.8907–0.9188 |
Risk control again ranked highest. Measurement ranked lowest.
The contrast between risk control and measurement was 0.0343, with a 95% CI of 0.0160–0.0514 and corrected p-value 0.040. The direction agreed across all five topic clusters.
So query formulation showed somewhat more task variation than final answer diversity.
But the absolute level of query differentiation remained high in every task.
Finding 4
SEO Manager was the most distinctive answer condition in every task
For each task, we measured a persona's average answer distance from the other four conditions.
SEO Manager ranked highest in all five task families.
| Task | Highest answer distinctiveness | Lowest answer distinctiveness |
|---|---|---|
| Action plan | SEO Manager: 0.874 | Founder/CEO: 0.854 |
| Comparison | SEO Manager: 0.864 | CMO: 0.837 |
| Diagnostic | SEO Manager: 0.874 | Founder/CEO: 0.849 |
| Measurement | SEO Manager: 0.862 | CMO: 0.838 |
| Risk control | SEO Manager: 0.880 | CMO: 0.852 |
The consistency is notable.
SEO Manager was not merely distinctive for technical measurement tasks. It was also most distinctive for comparison, diagnosis, action planning and risk control.
But distinctive is not better.
The analysis does not establish that SEO Manager answers were more accurate, complete, useful or better researched.
Finding 5
CMO and Founder/CEO were the nearest pair across every task
Across all five task families, the closest observed pair was:
CMO / Founder/CEO
| Task | CMO / Founder distance |
|---|---|
| Action plan | 0.828 |
| Comparison | 0.819 |
| Diagnostic | 0.828 |
| Measurement | 0.816 |
| Risk control | 0.829 |
That does not mean the two conditions were interchangeable.
A distance around 0.82 is still substantial, and lexical proximity does not establish that they produce the same facts, recommendations, priorities or risks.
Interaction evidence
Did persona-task interactions matter?
A strongly task-conditioned persona pattern would imply that the relative distinctiveness of personas changes meaningfully by task.
For example, SEO might be highly distinctive in measurement but ordinary in comparison, or CMO might separate strongly in risk control but converge elsewhere.
That pattern was not strong enough to satisfy the study's predefined task-conditioning rule.
Persona profiles were relatively stable across tasks.
The analysis therefore did not identify a clear persona × task interaction large enough to overturn the broad-persistence conclusion.
This matters for study design.
If the role structure changed dramatically by task, a research team might need entirely different persona panels for different question types.
The historical data did not support that conclusion.
Robustness
The pattern also held across topic clusters
Each task family appeared across five topic clusters:
- AI optimisation
- Answer Alignment
- Content authority
- Conversion/pipeline
- Search visibility
Across task × cluster cells, answer differentiation ranged approximately 0.825–0.884.
Lexical query differentiation ranged approximately 0.836–0.927.
These cells contained only five or six scenarios each, so they should not be treated as standalone findings.
But they provide a useful robustness check: the broad task-persistence result was not obviously driven by one topic cluster.
Study design
What this means for AI research design
The practical implication is not that task framing can be ignored.
Task design still changes what the model is asked to do, what evidence may be relevant, the structure of the answer and what success looks like.
But if a research programme is using personas to sample different decision contexts, this experiment suggests you should not assume persona differentiation disappears just because the task becomes more structured.
In this study, persona-associated differences remained high when the model was asked to:
- Diagnose
- Measure
- Compare
- Plan
- Control risk
That supports a more deliberate matrix design:
persona × task
rather than choosing either personas or tasks as the only source of variation.
Interpretation
Implications for AI Answer Alignment
For Answer Alignment research, different task forms can represent different moments in a buyer or stakeholder decision.
A company may be evaluated through questions such as:
- What is going wrong?
- How should we measure it?
- Which option is better?
- What should we do next?
- What could go wrong?
If persona-associated differences remain high across all of those task forms, then testing only one task style may under-sample the ways an AI system can represent the same company or category.
That does not mean every research programme needs all five personas × all five tasks.
It means the dimensions are not obviously interchangeable.
Persona variation remained visible even after the task changed.
Study boundaries
What this study does not establish
- It does not establish causal task effects
The task families used different scenario content and templates. They were not randomized versions of the exact same scenario.
- It does not prove risk control is the best task
Risk control had the highest observed answer and query differentiation. Higher differentiation is not higher quality.
- It does not prove SEO Manager is the best persona
SEO Manager was most distinctive. Distinctiveness is not usefulness.
- It does not establish semantic or business-value differences
The metrics measure lexical and structural response geometry. They do not establish different factual claims or better decisions.
- It does not remove the persona/order confound
Persona conditions were collected in a fixed within-scenario order in the historical run.
Open question
The unresolved question: does task-persistent diversity remain substantive?
The main V2 programme now shows that persona-associated differentiation:
- Survives direct cue masking
- Appears in grounded search-query formulation
- Persists across five task types
- Remains materially additive in a five-persona panel
But those results are still primarily representational.
The next question is whether the task-persistent differences correspond to different factual propositions, recommendations, decisions and risks.
That substantive layer is currently awaiting independent human calibration.
Until that validation is complete, the correct conclusion is:
Persona-associated representational diversity persists across tasks. Whether substantive decision diversity persists to the same degree is not yet established.
Method and reproducibility
Methodology
The task analysis reused frozen outputs from earlier stages of the V2 persona programme.
Primary analyses used 149 complete scenarios.
Each complete scenario contributed five persona-conditioned answers and ten unordered persona pairs.
Residual answer distance was 1 − saved residual cosine similarity from the persona cue-masking analysis.
No response vectorization was refit for the task analysis.
Task-level headlines averaged the five topic-cluster means equally.
Uncertainty intervals resampled whole scenarios within cluster × task cells.
The grounded-query task comparison reused the saved lexical query divergence from the earlier retrieval-pathway analysis.
Task contrasts used cluster-preserving scenario-label permutations with family-wise correction.
No new model generation, live search or source resolution was performed.
Limitations
Limitations
- Task content and template structure
The five task families differ in scenario content and template structure. This means task comparisons are descriptive associations rather than clean randomized task effects.
- One generation per cell
There was one response generation per persona/scenario cell, so within-condition variability is not estimated.
- Small task × cluster cells
Task × cluster cells contain only five or six complete scenarios.
- Lexical and structural representation
The answer-distance representation is lexical/structural, not a semantic-equivalence or correctness measure.
- Publisher layer unresolved
Publisher-level source divergence remained unmeasurable in the underlying grounded corpus.
- Persona/order confounding
Persona condition remained confounded with historical collection position.
FAQ
Frequently asked questions
Which task showed the biggest persona differences?
Risk control had the highest observed residual answer diversity at 0.8614.
Which task showed the smallest?
Comparison had the lowest at 0.8463.
Was that difference large?
No. The full range was 0.0150, about 1.76% of the overall level of task differentiation.
Did persona differences disappear for measurement tasks?
No. Measurement still had residual answer diversity of 0.8478.
Which persona was most distinctive across tasks?
SEO Manager was the most distinctive answer condition in all five task families.
Which pair was consistently closest?
CMO and Founder/CEO were the nearest pair across all five task families.
Does this prove risk prompts are more useful?
No. Higher differentiation does not establish higher quality or usefulness.
Should I use different personas for different task types?
The data did not show a strong enough persona × task interaction to justify a universal task-specific persona-panel rule. Persona profiles were relatively stable across tasks in this experiment.
Conclusion
Conclusion
Persona-conditioned AI differences did not disappear when the task changed.
Across action planning, comparison, diagnosis, measurement and risk control, residual answer differentiation remained consistently high.
Risk control produced the highest observed answer diversity. Comparison produced the lowest.
But the overall spread was small: 0.015, or about 1.76% of the overall differentiation level.
Search-query differentiation showed somewhat more task variation, but remained high across all five task types.
The strongest conclusion is therefore:
In this matched grounded-AI experiment, persona-associated answer differentiation was broadly persistent across business task types rather than confined to one narrow prompt format.
That makes task design an important companion variable—not a substitute for persona design.
For the broader programme, see Persona Prompts Change More Than Tone: Evidence from 750 Grounded AI Responses.