Kojable research · Persona conditioning · Task persistence

Do AI Persona Differences Depend on the Task? Evidence Across Five Business Task Types

Published By Piush Vaish

Across five business task types, persona-associated answer differentiation stayed consistently high. The range between the highest- and lowest-diversity task means was only about 1.76% of the overall level.

Key finding

Persona-associated answer differentiation remained high across all five task families in this matched grounded-AI experiment.

Mean residual answer diversity ranged from 0.846 to 0.861 across action planning, comparison, diagnosis, measurement and risk control.

The full range between the highest- and lowest-diversity task means was only 0.015, or about 1.76% of the overall level of task differentiation.

The study therefore classified persona differentiation as:

Broadly persistent across tasks.

Qualification: The task families used different scenario content and templates, and persona condition was historically confounded with collection position. These results describe task-associated patterns in one frozen model run; they do not establish causal task effects or real-role behavior.

  • Persona conditioning
  • Five task families
  • 149 complete scenarios
  • 11 min read
149complete matched scenarios
5business task types
0.0150answer-diversity range
1.76%relative rangeDecisionBroadly persistent across tasks.

At a glance

Research snapshot

Study design, task-level answer spread and primary interpretation boundary.
Study elementDetail
Research questionAre persona-conditioned AI answers more different for some business tasks than others?
Parent study150 matched scenarios × 5 persona conditions
Primary complete scenarios149
Task familiesAction plan, comparison, diagnostic, measurement, risk control
PersonasCMO, VP Product Marketing, SEO Manager, Demand Generation Director, Founder/CEO
Primary answer metricResidual answer distance
Lowest task meanComparison: 0.8463
Highest task meanRisk control: 0.8614
Range0.0150
Relative range1.76%
Task decisionBroadly persistent across tasks
Major limitationTask, scenario content and template were not independently randomized

Direct answer

Do AI persona differences depend strongly on the task?

Answer

Not in this experiment.

Persona-associated answer differentiation was high across all five task families.

Risk-control scenarios had the highest observed residual answer diversity at 0.8614. Comparison scenarios had the lowest at 0.8463.

But the gap between them was small compared with the overall level of separation.

The study's predefined conclusion was therefore:

Persona differentiation was broadly persistent across task types rather than strongly task-specific.

That does not mean task framing had no effect. It means the overall level of persona-associated separation remained high regardless of whether the model was diagnosing a problem, defining measurement, comparing options, planning actions or controlling risk.

Research question

Why this question matters

Persona prompting is often used as if the role itself is the main source of variation.

But tasks differ too.

A model asked to diagnose a problem may naturally produce a different kind of answer from one asked to compare options or define a risk-control plan.

That creates an important design question.

If persona differentiation appears only in certain task types, researchers may be overgeneralizing from narrow prompt formats.

For example:

  • Personas might matter for strategic comparison but not measurement
  • They might matter for risk control but disappear in diagnosis
  • Task structure might overwhelm persona framing entirely

Our broader V2 study had already found that persona-conditioned outputs remained strongly distinguishable after direct persona language was removed.

This analysis asks:

Does that differentiation hold across different kinds of work?

Task design

The five task families

The experiment used five response types.

1. Action plan

The model was asked to recommend concrete steps, experiments or next actions.

2. Comparison

The model compared approaches, options, trade-offs or criteria.

3. Diagnostic

The model identified likely causes, evidence and diagnostic sequences.

4. Measurement

The model defined baselines, metrics, validation plans or monitoring approaches.

5. Risk control

The model identified risks, guardrails, escalation rules or control mechanisms.

The original design contained 30 scenarios in each task family.

After one malformed response was excluded by the quality gate, the primary analysis retained:

  • 30 action-plan scenarios
  • 30 comparison scenarios
  • 30 diagnostic scenarios
  • 29 measurement scenarios
  • 30 risk-control scenarios

Each complete scenario contained all five persona conditions.

Finding 1

Answer differentiation stayed high across every task family

Using the residual answer representation from the persona cue-masking analysis, mean persona-pair answer diversity was:

Residual answer diversity and scenario-bootstrap intervals by task family.
TaskResidual answer diversity95% CI
Action plan0.86030.8558–0.8650
Comparison0.84630.8379–0.8541
Diagnostic0.85700.8498–0.8640
Measurement0.84780.8423–0.8529
Risk control0.86140.8548–0.8680

Every task mean remained above 0.84.

That matters because persona differentiation did not collapse in the more structured task types.

Measurement still showed residual answer diversity of 0.8478. Comparison, the lowest task, still measured 0.8463.

The practical result is not:

“Personas matter for one special kind of question.”

It is:

The observed persona-associated differentiation remained high across all five task families.

Finding 2

Risk control was highest, comparison lowest—but the spread was small

Risk control had the highest observed answer diversity: 0.8614.

Comparison had the lowest: 0.8463.

Difference: 0.0150.

That difference is about 1.76% of the overall level of task differentiation.

The study prospectively defined a larger task-specificity signal as requiring an answer-diversity range of at least 0.03, together with corrected contrast evidence, cross-cluster direction stability and supporting interaction evidence.

The observed range did not reach that threshold.

The primary task conclusion was therefore:

Broadly persistent across tasks

rather than:

Task-conditioned differentiation

Risk-control prompts did produce somewhat more differentiated answers, but not enough to support the claim that persona conditioning matters dramatically more when the task is about risk.

Finding 3

Grounded search-query differentiation was also high across all five tasks

The task analysis did not examine final answers alone.

It reused the lexical search-query divergence from the grounded-query study.

Mean lexical query divergence was:

Lexical grounded-query divergence and scenario-bootstrap intervals by task family.
TaskLexical query divergence95% CI
Action plan0.88290.8701–0.8940
Comparison0.88250.8696–0.8949
Diagnostic0.89280.8786–0.9058
Measurement0.87100.8608–0.8813
Risk control0.90530.8907–0.9188

Risk control again ranked highest. Measurement ranked lowest.

The contrast between risk control and measurement was 0.0343, with a 95% CI of 0.0160–0.0514 and corrected p-value 0.040. The direction agreed across all five topic clusters.

So query formulation showed somewhat more task variation than final answer diversity.

But the absolute level of query differentiation remained high in every task.

Residual answer and lexical query differentiation across five task families.
Residual answer differentiation remained high across all five task families; the spread between task means was small relative to overall separation.
Open full-resolution figure

Finding 4

SEO Manager was the most distinctive answer condition in every task

For each task, we measured a persona's average answer distance from the other four conditions.

SEO Manager ranked highest in all five task families.

Highest and lowest observed answer distinctiveness by task family.
TaskHighest answer distinctivenessLowest answer distinctiveness
Action planSEO Manager: 0.874Founder/CEO: 0.854
ComparisonSEO Manager: 0.864CMO: 0.837
DiagnosticSEO Manager: 0.874Founder/CEO: 0.849
MeasurementSEO Manager: 0.862CMO: 0.838
Risk controlSEO Manager: 0.880CMO: 0.852

The consistency is notable.

SEO Manager was not merely distinctive for technical measurement tasks. It was also most distinctive for comparison, diagnosis, action planning and risk control.

But distinctive is not better.

The analysis does not establish that SEO Manager answers were more accurate, complete, useful or better researched.

Persona answer distinctiveness profiles across five task families.
SEO Manager was the most distinctive answer condition across all five tasks; CMO and Founder/CEO were consistently closer. Distinctiveness is not quality.
Open full-resolution figure

Finding 5

CMO and Founder/CEO were the nearest pair across every task

Across all five task families, the closest observed pair was:

CMO / Founder/CEO

CMO / Founder/CEO residual answer distance by task family.
TaskCMO / Founder distance
Action plan0.828
Comparison0.819
Diagnostic0.828
Measurement0.816
Risk control0.829

That does not mean the two conditions were interchangeable.

A distance around 0.82 is still substantial, and lexical proximity does not establish that they produce the same facts, recommendations, priorities or risks.

Interaction evidence

Did persona-task interactions matter?

A strongly task-conditioned persona pattern would imply that the relative distinctiveness of personas changes meaningfully by task.

For example, SEO might be highly distinctive in measurement but ordinary in comparison, or CMO might separate strongly in risk control but converge elsewhere.

That pattern was not strong enough to satisfy the study's predefined task-conditioning rule.

Persona profiles were relatively stable across tasks.

The analysis therefore did not identify a clear persona × task interaction large enough to overturn the broad-persistence conclusion.

This matters for study design.

If the role structure changed dramatically by task, a research team might need entirely different persona panels for different question types.

The historical data did not support that conclusion.

Robustness

The pattern also held across topic clusters

Each task family appeared across five topic clusters:

  • AI optimisation
  • Answer Alignment
  • Content authority
  • Conversion/pipeline
  • Search visibility

Across task × cluster cells, answer differentiation ranged approximately 0.825–0.884.

Lexical query differentiation ranged approximately 0.836–0.927.

These cells contained only five or six scenarios each, so they should not be treated as standalone findings.

But they provide a useful robustness check: the broad task-persistence result was not obviously driven by one topic cluster.

Study design

What this means for AI research design

The practical implication is not that task framing can be ignored.

Task design still changes what the model is asked to do, what evidence may be relevant, the structure of the answer and what success looks like.

But if a research programme is using personas to sample different decision contexts, this experiment suggests you should not assume persona differentiation disappears just because the task becomes more structured.

In this study, persona-associated differences remained high when the model was asked to:

  • Diagnose
  • Measure
  • Compare
  • Plan
  • Control risk

That supports a more deliberate matrix design:

persona × task

rather than choosing either personas or tasks as the only source of variation.

Interpretation

Implications for AI Answer Alignment

For Answer Alignment research, different task forms can represent different moments in a buyer or stakeholder decision.

A company may be evaluated through questions such as:

  • What is going wrong?
  • How should we measure it?
  • Which option is better?
  • What should we do next?
  • What could go wrong?

If persona-associated differences remain high across all of those task forms, then testing only one task style may under-sample the ways an AI system can represent the same company or category.

That does not mean every research programme needs all five personas × all five tasks.

It means the dimensions are not obviously interchangeable.

Persona variation remained visible even after the task changed.

Study boundaries

What this study does not establish

  • It does not establish causal task effects

    The task families used different scenario content and templates. They were not randomized versions of the exact same scenario.

  • It does not prove risk control is the best task

    Risk control had the highest observed answer and query differentiation. Higher differentiation is not higher quality.

  • It does not prove SEO Manager is the best persona

    SEO Manager was most distinctive. Distinctiveness is not usefulness.

  • It does not establish semantic or business-value differences

    The metrics measure lexical and structural response geometry. They do not establish different factual claims or better decisions.

  • It does not remove the persona/order confound

    Persona conditions were collected in a fixed within-scenario order in the historical run.

Open question

The unresolved question: does task-persistent diversity remain substantive?

The main V2 programme now shows that persona-associated differentiation:

  • Survives direct cue masking
  • Appears in grounded search-query formulation
  • Persists across five task types
  • Remains materially additive in a five-persona panel

But those results are still primarily representational.

The next question is whether the task-persistent differences correspond to different factual propositions, recommendations, decisions and risks.

That substantive layer is currently awaiting independent human calibration.

Until that validation is complete, the correct conclusion is:

Persona-associated representational diversity persists across tasks. Whether substantive decision diversity persists to the same degree is not yet established.

Method and reproducibility

Methodology

The task analysis reused frozen outputs from earlier stages of the V2 persona programme.

Primary analyses used 149 complete scenarios.

Each complete scenario contributed five persona-conditioned answers and ten unordered persona pairs.

Residual answer distance was 1 − saved residual cosine similarity from the persona cue-masking analysis.

No response vectorization was refit for the task analysis.

Task-level headlines averaged the five topic-cluster means equally.

Uncertainty intervals resampled whole scenarios within cluster × task cells.

The grounded-query task comparison reused the saved lexical query divergence from the earlier retrieval-pathway analysis.

Task contrasts used cluster-preserving scenario-label permutations with family-wise correction.

No new model generation, live search or source resolution was performed.

Limitations

Limitations

  • Task content and template structure

    The five task families differ in scenario content and template structure. This means task comparisons are descriptive associations rather than clean randomized task effects.

  • One generation per cell

    There was one response generation per persona/scenario cell, so within-condition variability is not estimated.

  • Small task × cluster cells

    Task × cluster cells contain only five or six complete scenarios.

  • Lexical and structural representation

    The answer-distance representation is lexical/structural, not a semantic-equivalence or correctness measure.

  • Publisher layer unresolved

    Publisher-level source divergence remained unmeasurable in the underlying grounded corpus.

  • Persona/order confounding

    Persona condition remained confounded with historical collection position.

FAQ

Frequently asked questions

Which task showed the biggest persona differences?

Risk control had the highest observed residual answer diversity at 0.8614.

Which task showed the smallest?

Comparison had the lowest at 0.8463.

Was that difference large?

No. The full range was 0.0150, about 1.76% of the overall level of task differentiation.

Did persona differences disappear for measurement tasks?

No. Measurement still had residual answer diversity of 0.8478.

Which persona was most distinctive across tasks?

SEO Manager was the most distinctive answer condition in all five task families.

Which pair was consistently closest?

CMO and Founder/CEO were the nearest pair across all five task families.

Does this prove risk prompts are more useful?

No. Higher differentiation does not establish higher quality or usefulness.

Should I use different personas for different task types?

The data did not show a strong enough persona × task interaction to justify a universal task-specific persona-panel rule. Persona profiles were relatively stable across tasks in this experiment.

Conclusion

Conclusion

Persona-conditioned AI differences did not disappear when the task changed.

Across action planning, comparison, diagnosis, measurement and risk control, residual answer differentiation remained consistently high.

Risk control produced the highest observed answer diversity. Comparison produced the lowest.

But the overall spread was small: 0.015, or about 1.76% of the overall differentiation level.

Search-query differentiation showed somewhat more task variation, but remained high across all five task types.

The strongest conclusion is therefore:

In this matched grounded-AI experiment, persona-associated answer differentiation was broadly persistent across business task types rather than confined to one narrow prompt format.

That makes task design an important companion variable—not a substitute for persona design.

For the broader programme, see Persona Prompts Change More Than Tone: Evidence from 750 Grounded AI Responses.

Piush Vaish, Founder of Kojable

Author

About the author

Piush VaishFounder and CEO of Kojable

By Piush Vaish, founder and CEO of Kojable.

Piush Vaish is the founder and CEO of Kojable, a repeat founder and data scientist with more than 10 years of experience building and productising AI, machine-learning and data products. He writes about AI search, AEO, GEO, agentic discovery and AI product strategy.

Read more about Piush Vaish